Image Processing Method, Apparatus, Device, and Storage Medium
By performing multi-level downsampling of images and calculating the attention map processing features, the feature continuity loss problem caused by image chunking is solved, and more accurate image processing results are achieved.
Patent Information
- Application Number
- CN202111601437.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-07-31
- Filing Date
- 2021-12-24
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2041-12-24
AI Technical Summary
The prior art performs chunking processing on images in image processing, resulting in the continuity of edge features of image blocks being destroyed, affecting the accuracy of feature extraction and processing results.
By performing multi-level downsampling processing on the initial features of the target image, an attention map is obtained, and the initial features and downsampling features are processed at each level to avoid image segmentation and extract complete features.
Improve the accuracy of image processing tasks, ensure the completeness of feature extraction and the accuracy of processing results.
Smart Images

Figure CN114332574B_ABST
Abstract
Description
[0001] This application claims the priority of a Chinese patent application with the application number 202110877152.X and the invention title "Image Processing Method, Device, Equipment and Storage Medium", which was filed on July 31, 2021, and the entire content of which is incorporated herein by reference. Technical Field
[0002] This application relates to the field of artificial intelligence technology, and particularly relates to an image processing method, device, equipment and storage medium. Background Art
[0003] Image processing tasks such as image segmentation and image classification have relatively wide applications in the field of computer vision (CV).
[0004] In the related art, when performing an image processing task, an input target image can be divided into multiple image blocks, and the multiple image blocks are sequentially input into a Transformer encoder for encoding processing, so as to extract the image features of the target image, and then the image processing result can be obtained according to the image features obtained by encoding.
[0005] However, the above solution needs to divide the image into blocks, which may cause the feature continuity at the edges of the image blocks to be destroyed, thereby affecting the accuracy of feature extraction and further affecting the accuracy of the processing result of the target image. Summary of the Invention
[0006] Embodiments of this application provide an image processing method, device, equipment and storage medium, which can improve the accuracy of the processing result of an image. The technical solution is as follows.
[0007] On the one hand, an image processing method is provided, and the method includes:
[0008] Performing n-level downsampling processing on the initial features of the target image to obtain n downsampled features; n is a positive integer;
[0009] Obtaining an attention map of the target image based on the target downsampled feature among the n downsampled features; the target downsampled feature is obtained by the last-level downsampling processing of the n-level downsampling processing;
[0010] Based on the attention map, processing the initial features and the n downsampled features respectively to obtain attention-processed features corresponding to the initial features and the n downsampled features respectively;
[0011] Based on the initial features and the attention-processed features corresponding to the n downsampled features respectively, obtaining a processing result of performing a specified image processing task on the target image.
[0012] On the other hand, an image processing apparatus is provided, the apparatus comprising:
[0013] A downsampling module for performing n-level downsampling processing on the initial features of the target image to obtain n downsampled features;
[0014] An attention map acquisition module for acquiring an attention map of the target image based on a target downsampled feature among the n downsampled features; the target downsampled feature is obtained by the last level of downsampling processing in the n-level downsampling processing;
[0015] An attention processing module for processing the initial features and the n downsampled features respectively based on the attention map to obtain attention processing features corresponding to the initial features and the n downsampled features respectively;
[0016] A result acquisition module for obtaining a processing result of performing a specified image processing task on the target image based on the attention processing features corresponding to the initial features and the n downsampled features respectively.
[0017] In a possible implementation manner, the attention processing module 703 includes:
[0018] An upsampling unit for upsampling the attention map to obtain a first attention map with a first scale; the first scale is the scale of a first feature, and the first feature is any one of the initial features and the n downsampled features;
[0019] A processing unit for performing matrix multiplication processing on the first attention map and the first feature to obtain an attention processing feature corresponding to the first feature.
[0020] In a possible implementation manner, the attention map acquisition module is used for,
[0021] acquiring the attention map based on a query dimension feature and a key dimension feature of the target downsampled feature;
[0022] The processing unit is used for performing matrix multiplication processing on the value dimension feature of the first attention map and the first feature to obtain an attention processing feature corresponding to the first feature.
[0023] In a possible implementation manner, the result acquisition module includes:
[0024] A fusion unit for fusing the attention processing features corresponding to the initial features and the n downsampled features respectively to obtain image features of the target image;
[0025] A result acquisition unit, configured to acquire the processing result based on the image feature.
[0026] In a possible implementation, the fusion unit is configured to
[0027] Upsample the attention processing features respectively corresponding to the n downsampled features to obtain n upsampled features; the scale of the upsampled features is the same as the scale of the initial feature;
[0028] Concatenate the initial feature and the n upsampled features to obtain the image feature.
[0029] In a possible implementation, the downsampling module is configured to
[0030] Perform downsampling processing on a second feature to obtain an intermediate sampled feature; the second feature is any one of the initial feature and the n downsampled features except the target downsampled feature;
[0031] Perform convolution processing on the intermediate sampled feature to obtain the next-level downsampled feature of the second feature.
[0032] In a possible implementation, the downsampling module is configured to
[0033] Perform convolution processing on a third feature to obtain an intermediate convolution feature; the third feature is any one of the initial feature and the n downsampled features except the target downsampled feature;
[0034] Perform downsampling processing on the intermediate convolution feature to obtain the next-level downsampled feature of the third feature.
[0035] On the other hand, a computer device is provided, which includes a processor and a memory. At least one computer instruction is stored in the memory, and the at least one computer instruction is loaded and executed by the processor to implement the above-mentioned image processing method.
[0036] On another hand, a computer-readable storage medium is provided, in which at least one computer instruction is stored, and the at least one computer instruction is loaded and executed by a processor to implement the above-mentioned image processing method.
[0037] On yet another hand, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the above-mentioned image processing method.
[0038] The beneficial effects brought by the technical solutions provided in the embodiments of the present application at least include:
[0039] By performing multi-level downsampling on the initial features of the target image, calculating the attention map through the features obtained by the last-level downsampling, and then processing the initial features and the downsampling features at each level through the attention map to extract the features of the target image. During this process, it is not necessary to segment the target image, avoiding the loss of feature continuity caused by image segmentation. At the same time, through the processing of the initial features and the downsampling features at each level by the attention map, the complete features of the target image can be extracted, thus ensuring the accuracy of image feature extraction and further improving the accuracy of the processing results of image processing tasks.
[0040] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] The drawings here are incorporated into the specification and form a part of this specification, showing the embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.
[0042] Figure 1 is a system composition diagram of an image processing system related to various embodiments of the present application;
[0043] Figure 2 is a schematic flowchart of an image processing method shown according to an exemplary embodiment;
[0044] Figure 3 is an image processing framework diagram shown according to an exemplary embodiment;
[0045] Figure 4 is a schematic flowchart of an image processing method shown according to an exemplary embodiment;
[0046] Figure 5 is Figure 4 a schematic framework diagram of the DAB structure related to the shown embodiment;
[0047] Figure 6 is Figure 4 a framework diagram of an image processing model related to the shown embodiment;
[0048] Figure 7 is a structural block diagram of an image processing device shown according to an exemplary embodiment;
[0049] Figure 8 is a schematic structural diagram of a computer device shown according to an exemplary embodiment. Detailed implementation manners
[0050] Here, exemplary embodiments will be described in detail, and examples thereof are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0051] Before describing the various embodiments shown in the present application, several concepts related to the present application will be introduced first.
[0052] Please refer to Figure 1 , which shows a system configuration diagram of an image processing system related to the various embodiments of the present application. As Figure 1 shown, the system includes an image acquisition device 120, a terminal 140, and a server 160; optionally, the system may further include a database 180.
[0053] The image acquisition device 120 may be a camera device or a camera head device for acquiring images. For example, in the medical field, the images acquired by the image acquisition device 120 may be medical images including blood vessels or biological tissues, such as fundus images (including blood vessels under the retina), gastroscope images, colonoscope images, oral internal images, and the like. In addition to the medical field, the solutions in the embodiments of the present application can also be applied to other fields, such as autonomous driving, information search, and other fields.
[0054] The image acquisition device 120 may include an image output interface, such as a Universal Serial Bus (USB) interface, a High Definition Multimedia Interface (HDMI) interface, or an Ethernet interface, etc.; or, the above image output interface may also be a wireless interface, such as a Wireless Local Area Network (WLAN) interface, a Bluetooth interface, etc.
[0055] Correspondingly, according to the different types of the above image output interfaces, there are also various ways for an operator to export the images captured by the image acquisition device 120. For example, the images can be imported into the terminal 140 through a wired or short-distance wireless method, or the images can also be imported into the terminal 140 or the server 160 through a local area network or the Internet.
[0056] The terminal 140 can be a terminal device with certain processing capabilities and interface display functions. For example, the terminal 140 can be a mobile phone, a tablet computer, an e-book reader, smart glasses, a laptop computer, a desktop computer, and so on.
[0057] The terminal 140 can include terminals used by developers or users. For example, in the medical field, the terminal 140 can be a terminal used by medical staff.
[0058] When the terminal 140 is implemented as a terminal used by a developer, the developer can develop a machine learning model for performing specified image processing tasks on images through the terminal 140 and deploy the machine learning model to the server 160 or the terminal used by the user.
[0059] When the terminal 140 is implemented as a terminal used by a user (such as medical staff), an application for obtaining and presenting the processing results of the image can be installed in the terminal 140. After the terminal 140 obtains the image collected by the image acquisition device 120, it can obtain the processing results obtained by performing specified image processing tasks on the image through the above application and present the processing results.
[0060] In Figure 1 In the system shown, the terminal 140 and the image acquisition device 120 are physically separate entity devices. Optionally, in another possible implementation, when the terminal 140 is implemented as a terminal used by a user, the terminal 140 and the image acquisition device 120 can also be integrated into a single entity device; for example, the terminal 140 can be a terminal device with an image acquisition function.
[0061] Among them, the server 160 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms.
[0062] For example, when the solution shown in the present application is applied to the medical field, it can be implemented as a part of the medical cloud. Among them, the medical cloud refers to a medical health service cloud platform created by using "cloud computing" on the basis of new technologies such as cloud computing, mobile technology, multimedia, wireless communication, big data, and the Internet of Things, combined with medical technology, realizing the sharing of medical resources and the expansion of the medical scope. Due to the application and combination of cloud computing technology, the medical cloud improves the efficiency of medical institutions and facilitates residents to seek medical treatment. For example, the current hospital appointment registration, medical insurance, etc. are all the products of the combination of cloud computing and the medical field. The medical cloud also has the advantages of data security, information sharing, dynamic expansion, and overall layout.
[0063] Among them, the above-mentioned server 160 can be a server that provides background services for the application programs installed in the terminal 140. This background server can perform version management of application programs, perform background processing on the images obtained by the application programs and return the processing results, perform background training on the machine learning models developed by developers, and so on.
[0064] The above-mentioned database 180 can be a Redis database, or it can also be other types of databases. Among them, the database 180 is used to store various types of data.
[0065] Optionally, the terminal 140 and the server 160 are connected through a communication network. Optionally, the image acquisition device 120 and the server 160 are connected through a communication network. Optionally, this communication network is a wired network or a wireless network.
[0066] Optionally, the system can also include a management device ( Figure 1 not shown), and this management device is connected to the server 160 through a communication network. Optionally, the communication network is a wired network or a wireless network.
[0067] Optionally, the above-mentioned wireless network or wired network uses standard communication technologies and / or protocols. The network is usually the Internet, but can also be any network, including but not limited to any combination of LAN (Local Area Network), MAN (Metropolitan Area Network), WAN (Wide Area Network), mobile, wired or wireless network, private network or virtual private network. In some embodiments, technologies and / or formats including HTML (Hyper Text Mark-up Language), XML (Extensible Markup Language), etc. are used to represent data exchanged through the network. In addition, conventional encryption technologies such as SSL (Secure Socket Layer), TLS (Transport Layer Security), VPN (Virtual Private Network), IPsec (Internet Protocol Security), etc. can be used to encrypt all or some of the links. In other embodiments, customized and / or dedicated data communication technologies can be used to replace or supplement the above data communication technologies.
[0068] Figure 2 is a schematic flowchart of an image processing method shown according to an exemplary embodiment. This method can be executed by a computer device. For example, the computer device can be a server, or the computer device can also be a terminal, or the computer device can include a server and a terminal. Among them, the server can be the server 160 in the above Figure 1 shown embodiment, and the terminal can be the terminal 140 in the above Figure 1 shown embodiment. As Figure 2 shown, the image processing method can include the following steps.
[0069] Step 201, perform n-level downsampling processing on the initial features of the target image to obtain n downsampled features; n is a positive integer.
[0070] Among them, the above-mentioned n-level downsampling refers to performing downsampling processing on the basis of the features of the upper layer.
[0071] For example, the first-level downsampling refers to performing downsampling on the initial features to obtain the downsampled features corresponding to the first-level downsampling; the second-level downsampling refers to performing downsampling processing on the downsampled features corresponding to the first-level downsampling to obtain the downsampled features corresponding to the second-level downsampling, and so on.
[0072] Step 202: Obtain an attention map of the target image based on the target downsampled feature among the n downsampled features; the target downsampled feature is obtained through the last-level downsampling process of the n-level downsampling process.
[0073] In the embodiments of the present application, the computer device can calculate the attention map of the target image through the downsampled feature obtained by the last-level downsampling (i.e., the above-mentioned target downsampled feature), thereby reducing the computational complexity of calculating the attention map of the entire target image and improving the computational efficiency.
[0074] Step 203: Process the initial feature and the n downsampled features respectively based on the attention map to obtain the attention-processed features corresponding to the initial feature and the n downsampled features respectively.
[0075] In the embodiments of the present application, since the attention map is calculated through the downsampled feature obtained by the last-level downsampling, in order to extract the features of the entire target image as completely as possible, the initial feature and the downsampled features at each level are processed through the attention map respectively, thereby reducing the feature loss during the feature extraction process.
[0076] Step 204: Obtain the processing result of performing a specified image processing task on the target image based on the attention-processed features corresponding to the initial feature and the n downsampled features respectively.
[0077] In the embodiments of the present application, after the computer device obtains the attention-processed features corresponding to the initial feature and the n downsampled features respectively, it can perform a specified image processing task on the target image based on the attention-processed features corresponding to the initial feature and the n downsampled features respectively to obtain the processing result.
[0078] In summary, the solution shown in the embodiments of the present application downsamples the initial feature of the target image at multiple levels, calculates the attention map through the feature obtained by the last-level downsampling, and then processes the initial feature and the downsampled features at each level through the attention map to achieve the extraction of the features of the target image. During this process, it is not necessary to segment the target image, avoiding the loss of feature continuity caused by image segmentation. At the same time, through the processing of the initial feature and the downsampled features at each level by the attention map, the complete features of the target image can be extracted, thereby ensuring the accuracy of the image feature extraction and further improving the accuracy of the processing result of the image processing task.
[0079] In a possible implementation manner, the solutions shown in the embodiments of the present application can be implemented based on AI (Artificial Intelligence) technology and can implement any type of image processing task in the field of computer vision. That is, the above Figure 2Each step in the illustrated embodiment can be implemented by a pre-trained image processing model to achieve a specified image processing task. For example, the above-specified image processing task can include, but is not limited to, tasks such as image classification, image segmentation, object detection, etc.
[0080] Among them, AI is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning, and decision-making. Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level technologies and software-level technologies. Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technologies, operation / interaction systems, mechatronics, etc. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0081] Computer vision is a science that studies how to make machines "see". Further, it refers to using cameras and computers to replace human eyes for tasks such as target recognition and measurement in machine vision, and further performing graphic processing to make the computer-processed images more suitable for human eyes to observe or be transmitted to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies, and attempts to build an artificial intelligence system that can obtain information from images or multi-dimensional data. Computer vision technology usually includes technologies such as image processing, image recognition, image semantic understanding, image retrieval, optical character recognition (OCR), video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, virtual reality, augmented reality, and map construction.
[0082] Machine learning is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills, and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learning from demonstration.
[0083] For example, taking the image classification task in the medical field (e.g., identifying whether the tissues and organs in a medical image are normal) as an example, please refer to Figure 3 , which shows a processing framework diagram of an image provided by an exemplary embodiment of the present application. As Figure 3 shown, the computer device can extract the initial features of the medical image and input them into the image processing model 30. The downsampling branch 31 in the image processing model 30 performs n-level downsampling to obtain the downsampling features corresponding to each level of downsampling. Among them, the initial features and the downsampling features of each level are input into the attention branch 32 in the image processing model 30. The attention branch extracts the attention map based on the last-level downsampling feature, and processes the initial features and the downsampling features of each level through the attention map to obtain the attention processing features corresponding to the initial features and the downsampling features of each level respectively. The attention processing features are input into the classification branch 33 in the image processing model 30, and the classification branch 33 outputs the classification result of the medical image. For example, it outputs the probability of whether the tissues and organs in the medical image are normal.
[0084] Among them, steps 201 to 203 in the above Figure 2 shown embodiment can be implemented by a neural network. In the embodiment of the present application, this neural network can be called a Dense Attention Block (DAB) structure of a Dense Convolutional Network (DenseNet) based on the Attention mechanism.
[0085] Figure 4 is a schematic flowchart of an image processing method shown according to an exemplary embodiment. This method can be executed by a computer device. For example, this computer device can be a server, or this computer device can also be a terminal, or this computer device can include a server and a terminal. Among them, the server can be the server 160 in the above Figure 1 shown embodiment, and the terminal can be the terminal 140 in the above Figure 1 shown embodiment. As Figure 4 shown, this image processing method can include the following steps.
[0086] Step 401, obtain the initial features of the target image.
[0087] In an exemplary solution of the embodiment of the present application, the computer device can perform convolution processing on the target image through the convolutional layer in the image processing model to obtain the initial features of the target image.
[0088] For example, please refer to Figure 5, which shows a schematic framework diagram of the DAB structure involved in the embodiments of the present application. Among them, the input of the above DAB structure can be any three-dimensional or four-dimensional matrix, representing a two-dimensional or three-dimensional image or feature with any number of channels (N*M*C1 or N*M*L*C1, where N, M, L represent scales, and C1 represents the number of input channels). The output of the above DAB structure is a feature with the same scale as the input (N*M*C2 or N*M*L*C2, and C2 represents the number of output channels). In the subsequent embodiments of the present application, the processing of two-dimensional images is taken as an example for illustration.
[0089] As Figure 5 shown, for the target image, the computer device can process it through the convolutional layer in the DAB structure in the image processing model to obtain the initial feature 51 of the target image.
[0090] For example, for an input multi-channel two-dimensional image or feature, Figure 5 the DAB structure shown first extracts features with the same scale as the input (the number of channels may be different) through a Convolutional Neural Network (CNN) layer. Here, the features with the same scale as the input can be the above initial features.
[0091] Step 402: Perform n-level downsampling processing on the initial feature of the target image to obtain n downsampled features; n is a positive integer.
[0092] In a possible implementation manner, the process of performing n-level downsampling processing on the initial feature to obtain n downsampled features may include:
[0093] Perform downsampling processing on the second feature to obtain an intermediate sampled feature; the second feature is any one of the initial feature and the n downsampled features except the target downsampled feature;
[0094] Perform convolutional processing on the intermediate sampled feature to obtain the next-level downsampled feature of the second feature.
[0095] That is to say, in a possible implementation manner of the embodiments of the present application, the above downsampling processing may include a downsampling process and a convolutional processing process, and in the downsampling processing, the computer device can first perform downsampling on the feature to be downsampled (i.e., the above second feature), and then perform a convolutional operation on the downsampling result to obtain the next-level downsampled feature.
[0096] In another possible implementation manner, the process of performing n-level downsampling processing on the initial feature to obtain n downsampled features may include:
[0097] Perform a convolution operation on the third feature to obtain an intermediate convolution feature; the third feature is any one of the initial feature and the n downsampled features except the target downsampled feature.
[0098] Perform a downsampling operation on the intermediate convolution feature to obtain the next-level downsampled feature of the third feature.
[0099] That is to say, in the downsampling operation, the computer device can also first perform a convolution operation on the feature to be downsampled (i.e., the above-mentioned third feature), and then perform downsampling on the convolution result to obtain the next-level downsampled feature.
[0100] For example, in Figure 5 In the DAB structure shown, after the computer device downsamples the initial feature 51, it can continue to extract the next-level downsampled feature 52 through another convolutional layer. This operation can be repeated arbitrarily many times before the image scale is reduced to 1. Among them, Figure 5 In
[0101] Or, in Figure 5 In the DAB structure shown, the computer device can also perform a convolution operation on the initial feature and then perform downsampling to obtain the next-level downsampled feature. This operation can also be repeated 3 times, and the length and width of each layer of downsampled feature are half of the previous-level feature.
[0102] Step 403, obtain the attention map of the target image based on the target downsampled feature among the n downsampled features; the target downsampled feature is obtained by the last-level downsampling process of the n-level downsampling process.
[0103] In the embodiments of the present application, the computer device can obtain the attention map based on the query dimension feature and the key dimension feature of the target downsampled feature.
[0104] As Figure 5 shown, after obtaining the last-level downsampled feature 52, the attention map 53 (Attention Map) can be calculated based on the last-level downsampled feature 52. In the embodiments of the present application, the attention calculation method in Transformer can be borrowed. That is to say, for a given last-level downsampled feature y ∈ R N×M×C , the query dimension feature q, the key dimension feature k, and the value dimension feature v are obtained through three parallel convolutional neural networks. For the convenience of calculation, q, k, and v can all be converted into two-dimensional matrices, where each row in the two-dimensional matrix corresponds to a pixel point in the downsampled feature y, that is, Then, in the attention map 53, the attention of pixel point j to pixel point i is as follows:
[0105]
[0106] where Softmax represents the normalization function, T represents the matrix transpose, and · represents matrix multiplication; c q 、c k 、c v respectively represent the number of channels of q, k, and v.
[0107] Step 404: Process the initial feature and the n downsampled features respectively based on the attention map to obtain the attention-processed features corresponding to the initial feature and the n downsampled features respectively.
[0108] In a possible implementation manner, the process of processing the initial feature and the n downsampled features respectively based on the attention map to obtain the attention-processed features corresponding to the initial feature and the n downsampled features respectively may include:
[0109] Upsample the attention map to obtain a first attention map with a scale of the first scale; the first scale is the scale of the first feature, and the first feature is any one of the initial feature and the n downsampled features;
[0110] Perform matrix multiplication processing on the first attention map and the first feature to obtain the attention-processed feature corresponding to the first feature.
[0111] In the embodiments of the present application, the attention map is calculated through the last-level downsampled feature. Therefore, the scale of the attention map is the same as that of the last-level downsampled feature, while the scales of the initial feature to the (n - 1)-th level downsampled features are all larger than the scale of the attention map. To ensure the correctness of the attention mechanism, it is necessary to first upsample the attention map so that the scale of the attention map is the same as the scale of the feature to be matrix-multiplied (i.e., the above-mentioned first feature), so as to correctly perform the attention mechanism operation.
[0112] Among them, for the last-level downsampled feature, the kernel scale of the upsampling of the corresponding attention map can be 1*1, that is, directly perform attention calculation on the attention map and the last-level downsampled feature.
[0113] In a possible implementation manner, the process of performing matrix multiplication processing on the first attention map and the first feature to obtain the attention-processed feature corresponding to the first feature may include:
[0114] Perform matrix multiplication processing on the value dimension features of the first attention map and the first feature to obtain the attention-processed feature corresponding to the first feature.
[0115] In the embodiment of the present application, after obtaining the attention map A, applying it to the value dimension feature v can obtain the feature SA after attention processing:
[0116] SA = A · v
[0117] That is to say, the above-mentioned target downsampled feature is the feature of the last level of downsampling and has undergone multiple levels of downsampling. In order to provide features with sufficient spatial accuracy for subsequent tasks, the embodiment of the present application upsamples the attention map A to the same scale as the features of the previous several levels and multiplies them with the matrix of the features of the previous several levels respectively to obtain the initial feature and the attention processing features corresponding to each level of downsampled features.
[0118] For example, in Figure 5 after the computer device calculates the attention map 53 through the last-level downsampled feature 52, for the last-level downsampled feature 52, directly multiply the attention map 53 with the value dimension feature v of the last-level downsampled feature 52 to obtain the attention processing feature 54 corresponding to the last-level downsampled feature 52; for the initial feature 51 and the first two levels of downsampled features 52, after upsampling the attention map 53 to the same scale respectively, multiply it with the value dimension feature v of the corresponding feature to obtain the attention processing features 54 corresponding to the initial feature 51 and the first two levels of downsampled features 52 respectively.
[0119] For instance, in Figure 5In this case, since the attention map 53 is calculated from the last-stage downsampled feature 52, the scale of the attention map 53 is the same as that of the last-stage downsampled feature 52 (for example, both scales are 100*100). When performing attention calculation on the last-stage downsampled feature 52, there is no need to upsample the attention map 53, and the matrix multiplication is directly performed between the value dimension feature of the last-stage downsampled feature 52 and the attention map 53. Since the last-stage downsampled feature 52 is obtained by performing three-stage downsampling on the initial feature 51, the scale of the attention map 53 is smaller than that of the initial feature 51 and the first two-stage downsampled features 52. Therefore, the attention map 53 needs to be upsampled to the same scale as the initial feature 51 and the first two-stage downsampled features 52 respectively, and then the matrix multiplication is performed with the initial feature 51 and the first two-stage downsampled features 52 respectively. For example, taking the kernel scale of the three-stage downsampling as 2*2, when performing attention calculation on the second-stage downsampled feature 52, the scale of the attention map 53 is upsampled to 200*200 and then the attention calculation is performed with the second-stage downsampled feature 52 (such as performing attention calculation on the value dimension feature of the second-stage downsampled feature 52); when performing attention calculation on the first-stage downsampled feature 52, the scale of the attention map 53 is upsampled to 400*400 and then the attention calculation is performed with the first-stage downsampled feature 52; when performing attention calculation on the initial feature 51, the scale of the attention map 53 is upsampled to 800*800 and then the attention calculation is performed with the initial feature 51.
[0120] Among them, the upsampling of the attention map 53 can be directly performed with the corresponding scale ratio based on the attention map 53. For example, taking the kernel scale of the downsampling as 2*2, for the second-stage downsampled feature 52, the computer device performs upsampling with a kernel scale of 2*2 on the basis of the attention map 53 to obtain the upsampled attention map corresponding to the second-stage downsampled feature 52; for the first-stage downsampled feature 52, the computer device performs upsampling with a kernel scale of 4*4 on the basis of the attention map 53 to obtain the upsampled attention map corresponding to the first-stage downsampled feature 52; for the initial feature, the computer device performs upsampling with a kernel scale of 8*8 on the basis of the attention map 53 to obtain the upsampled attention map corresponding to the initial feature 51.
[0121] Alternatively, the upsampling of the attention map 53 described above can also be performed on the attention map 53 step by step. For example, taking the downsampling kernel scale of 2*2 as an example, for the second-level downsampled feature 52, the computer device performs upsampling with a kernel scale of 2*2 on the basis of the attention map 53 to obtain the upsampled attention map corresponding to the second-level downsampled feature 52; for the first-level downsampled feature 52, the computer device performs upsampling with a kernel scale of 2*2 on the basis of the upsampled attention map corresponding to the second-level downsampled feature 52 to obtain the upsampled attention map corresponding to the first-level downsampled feature 52; for the initial feature, the computer device performs upsampling with a kernel scale of 2*2 on the basis of the upsampled attention map corresponding to the first-level downsampled feature 52 to obtain the upsampled attention map corresponding to the initial feature 51.
[0122] In the embodiments of the present application, since the scale of the finally calculated attention map A is NM×NM, it is easy to cause insufficient video memory after upsampling. To reduce the requirement for video memory, the method of axis attention can be introduced to calculate the attention map in the embodiments of the present application. Its basic principle is to split the attention map A into two parts, horizontal and vertical, and calculate them separately.
[0123] Taking the horizontal attention map as an example, after obtaining the query-dimensional feature q and the key-dimensional feature k of the target downsampled feature, the computer device can first average the q and k of different pixel points in each column of the target downsampled feature, and then calculate the attention value of each column in the attention map from the obtained average values of q and k. Subsequently, when performing the operation of the attention mechanism, the attention value of each column in the obtained attention map is applied to the value-dimensional feature of each pixel in each column of the corresponding target downsampled feature. The same processing is also performed on the initial feature and other downsampled features. The calculation of the vertical attention map can be analogized in the same way. In this way, the scale of the attention map can be reduced from NM×NM to N×N + M×M, thus greatly reducing the requirement for video memory.
[0124] In addition, the attention mechanism part in the embodiments of the present application can be implemented by using the multi-head attention mechanism. Among them, the multi-head attention mechanism is one of the cores of Transformer, and this mechanism can also be applied to the DAB framework in the embodiments of the present application. For example, the same feature can be input into multiple parallel blocks (blocks), and the features output by each block are concatenated together as the input of the next layer.
[0125] That is to say, the DAB framework in the embodiments of the present application is a model framework constructed based on the multi-head attention mechanism. Among them, the multi-head attention mechanism contains multiple groups of (Q, K, V) matrices. One group of (Q, K, V) matrices represents the operation of one attention mechanism. After splicing these multiple matrices and multiplying by a projection matrix, the output of the final multi-head attention mechanism can be obtained, that is, the attention processing feature in the embodiments of the present application.
[0126] After obtaining the above attention processing features, the computer device can obtain the processing result of performing a specified image processing task on the target image based on the initial feature and the attention processing features corresponding to the n downsampled features respectively. This process can refer to the subsequent steps.
[0127] Step 405: Fuse the initial feature and the attention processing features corresponding to the n downsampled features respectively to obtain the image feature of the target image.
[0128] In the embodiments of the present application, the DAB framework can concatenate the initial feature and the attention processing features corresponding to the n downsampled features respectively as the final output of the DAB framework (i.e., the image feature of the target image).
[0129] In a possible implementation manner, the process of fusing the initial feature and the attention processing features corresponding to the n downsampled features respectively to obtain the image feature of the target image may include:
[0130] Upsample the attention processing features corresponding to the n downsampled features respectively to obtain n upsampled features; the scale of the upsampled features is the same as the scale of the initial feature;
[0131] Concatenate the initial feature and the n upsampled features to obtain the image feature.
[0132] In the embodiments of the present application, since the scales of the initial feature and the attention processing features corresponding to the n downsampled features are different, in order to realize the fusion of multiple attention processing features, the computer device can upsample the attention processing features corresponding to the n downsampled features respectively to obtain n upsampled features with the same scale as the initial feature, and then concatenate the n upsampled features with the initial feature to obtain the image feature.
[0133] For example, in Figure 5Among them, the scales of the attention processing features 54 corresponding to the n downsampled features in the DAB framework are smaller than the scale of the attention processing feature 54 corresponding to the initial feature. Therefore, for the attention processing features 54 corresponding to the n downsampled features, the DAB framework can perform upsampling processing through an upsampling layer to obtain n upsampled features. After the n upsampled features are cascaded with the attention processing feature 54 corresponding to the initial feature, the image feature 55 can be obtained, and the image feature 55 can be used as the output of the DAB framework.
[0134] For example, in Figure 5 Among them, the purpose of upsampling the scale of the attention processing feature 54 is to make the scales of the upsampled features obtained by upsampling each attention processing feature 54 the same, so that they can be cascaded subsequently. Among them, the scale of the attention processing feature 54 corresponding to the initial feature is the highest. At this time, the attention processing feature 54 corresponding to the initial feature does not need to be upsampled, and only the attention processing features 54 corresponding to the three-level downsampled features 52 need to be upsampled so that the scales of the upsampled features 54 corresponding to the three-level downsampled features 52 are all the same as the scale of the attention processing feature 54 corresponding to the initial feature 51. For example, taking the scale of the initial feature 51 as 800*800 and the kernel scales of downsampling as 2*2, the scale of the attention processing feature 54 corresponding to the initial feature is 800*800, and the scales of the attention processing features 54 corresponding to the three-level downsampled features 52 are 400*400, 200*200, and 100*100 respectively; at this time, the computer device can upsample the scales of the attention processing features 54 corresponding to the three-level downsampled features 52 to 800*800 respectively, and then splice them with the attention processing feature 54 corresponding to the initial feature 51 to obtain the image feature 55.
[0135] Step 406, obtaining a processing result of performing a specified image processing task on the target image based on the image feature.
[0136] After obtaining the above image feature, the computer device can execute a specified image processing task according to the above image feature to obtain a processing result.
[0137] In the embodiments of the present application, a single DAB framework can be used to extract image features, or multiple DAB frameworks can be cascaded to extract the final image features, where the output of each DAB framework is used as the input of the next layer.
[0138] For example, please refer to Figure 6 , which shows the framework diagram of the image processing model involved in the embodiments of the present application. As Figure 6As shown, taking the case where the image processing model is used for image classification, the image processing model includes 3 DAB frameworks and a Multilayer Perceptron (MLP) network. Among them, the 3 DAB frameworks and the MLP network are cascaded. The MLP can be composed of two fully connected layers. The features extracted by the DAB can be flattened and input into the MLP to obtain the final classification result. After the medical image is input into the first DAB framework 61, the DAB framework 61 performs the processing shown in the above steps 401 to 405 on the medical image and outputs image features. The image features output by this DAB framework 61 are used as the input of the DAB framework 62 (for example, as the initial features of the DAB framework 62), and so on. The image features output by the DAB framework 63 will be input into the MLP network 64 to output the image classification result (for example, output the probability of whether the tissues and organs in the image are normal).
[0139] Among them, the above image processing model can be trained by image samples pre-annotated according to the task type. For example, taking image classification as an example, in the model training stage, the model training device can obtain sample images labeled with normal samples and abnormal samples, and input the sample images into the image processing model. The image processing model outputs the predicted classification results of the sample images. Then, combining the predicted classification results and the annotation information of the sample images, the loss function is calculated, and then the parameters of the image processing model are adjusted through the loss function. For example, adjust the parameters of the 3 DAB frameworks and the MLP network in Figure 6 . After multiple rounds of iterative training, a converged image processing model is obtained. This converged image processing model can execute the solution shown in the above embodiments of the present application.
[0140] Among them, the above loss function can adopt the cross entropy loss function. For the results of manual annotation, the label of the normal sample is 0, and the label of the abnormal sample is 1. The loss function can be defined as:
[0141] L ce =-(ylog(p)+(1 - y)log(1 - p))
[0142] Among them, y is the sample label of the sample image, and p is the probability that the image processing model outputs that the sample image is a positive sample. According to this loss function, the present application uses the gradient descent method based on Adam to update the parameters of the image processing model. For example, in a certain image classification task, betas in Adam=(0.95, 0.9995). The initial learning rate can be 0.001, and it is reduced to one-fifth every 20 epochs (all samples entering the model for calculation once is called 1 epoch). It can be trained for a total of 100 epochs, and the batch size can be 50.
[0143] When the solution shown in this application is implemented as an image classification network, it can be used for various medical images and even the classification problems of other images.
[0144] In the embodiments of this application, DAB involved as a basic module of a neural network, its usage is similar to that of Transformer and CNN modules, and it can be applied to feature extraction in most image processing tasks. It is also possible to stack multiple modules to generate a deeper network, extract higher-dimensional features, and handle complex tasks. For any image recognition, segmentation, or detection tasks, replace the CNN or Transformer module with this module to extract features, and then input them into the classification head, segmentation head, or detection head. The training process and inference method can both use the solutions of the corresponding tasks themselves.
[0145] In the DAB involved in the embodiments of this application, the number of layers of the CNN network, kernel scale, number of feature channels, and kernel scale of downsampling can all be freely adjusted according to task requirements.
[0146] Among them, in the embodiments of this application, the feature extraction layer can be replaced by other networks in addition to the basic CNN implementation, such as Residual Network (ResNet), Dilated CNN, etc.
[0147] In summary, the solution shown in the embodiments of this application performs multi-level downsampling on the initial features of the target image, calculates the attention map through the features obtained by the last-level downsampling, and then processes the initial features and the downsampling features at all levels through the attention map to achieve the extraction of the features of the target image. In this process, there is no need to segment the target image, avoiding the loss of feature continuity caused by image segmentation. At the same time, through the processing of the initial features and the downsampling features at all levels by the attention map, the complete features of the target image can be extracted, thus ensuring the accuracy of image feature extraction and further improving the accuracy of the processing results of image processing tasks.
[0148] Among them, the solution shown in the above embodiments of this application can be implemented or executed in combination with the blockchain. For example, some or all of the steps in the above embodiments can be executed in the blockchain system; or, the data required for the execution of each step in the above embodiments or the generated data can be stored in the blockchain system; for example, the training samples used in the above model training, and the model input data such as the target image in the model application process can be obtained by the computer device from the blockchain system; for another example, the parameters of the model obtained after the above model training (including the parameters of the DAB architecture and the parameters of other parts in the image processing model) can be stored in the blockchain system.
[0149] It is understandable that in the specific embodiments of the present application, when it comes to relevant data such as medical data and medical images, which are user information, when the above embodiments of the present application are applied to specific products or technologies, user permission or consent needs to be obtained, and the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards of relevant countries and regions.
[0150] Figure 7 It is a structural block diagram of an image processing device shown according to an exemplary embodiment. This device can implement Figure 2 or Figure 4 all or part of the steps in the method provided by the shown embodiment. This image processing device includes:
[0151] A downsampling module 701, configured to perform n-level downsampling processing on the initial features of the target image to obtain n downsampled features;
[0152] An attention map acquisition module 702, configured to obtain an attention map of the target image based on a target downsampled feature among the n downsampled features; the target downsampled feature is obtained by the last-level downsampling processing of the n-level downsampling processing;
[0153] An attention processing module 703, configured to process the initial features and the n downsampled features respectively based on the attention map to obtain attention processing features corresponding to the initial features and the n downsampled features respectively;
[0154] A result acquisition module 704, configured to obtain a processing result of performing a specified image processing task on the target image based on the attention processing features corresponding to the initial features and the n downsampled features respectively.
[0155] In a possible implementation manner, the attention processing module 703 includes:
[0156] An upsampling unit, configured to upsample the attention map to obtain a first attention map with a first scale; the first scale is the scale of a first feature, and the first feature is any one of the initial features and the n downsampled features;
[0157] A processing unit, configured to perform matrix multiplication processing on the first attention map and the first feature to obtain an attention processing feature corresponding to the first feature.
[0158] In a possible implementation manner, the attention map acquisition module 702 is configured to,
[0159] obtain the attention map based on the query dimension feature and the key dimension feature of the target downsampled feature;
[0160] The processing unit is configured to perform matrix multiplication on the first attention map and the value dimension feature of the first feature to obtain an attention processing feature corresponding to the first feature.
[0161] In a possible implementation manner, the result acquisition module 704 includes:
[0162] A fusion unit configured to fuse the initial feature and the attention processing features respectively corresponding to the n downsampled features to obtain an image feature of the target image;
[0163] A result acquisition unit configured to obtain the processing result based on the image feature.
[0164] In a possible implementation manner, the fusion unit is configured to
[0165] Upsample the attention processing features respectively corresponding to the n downsampled features to obtain n upsampled features; the scale of the upsampled features is the same as the scale of the initial feature;
[0166] Concatenate the initial feature and the n upsampled features to obtain the image feature.
[0167] In a possible implementation manner, the downsampling module 701 is configured to
[0168] Perform downsampling on a second feature to obtain an intermediate sampled feature; the second feature is any one of the initial feature and the n downsampled features except the target downsampled feature;
[0169] Perform convolution on the intermediate sampled feature to obtain a next-level downsampled feature of the second feature.
[0170] In a possible implementation manner, the downsampling module 701 is configured to
[0171] Perform convolution on a third feature to obtain an intermediate convolutional feature; the third feature is any one of the initial feature and the n downsampled features except the target downsampled feature;
[0172] Perform downsampling on the intermediate convolutional feature to obtain a next-level downsampled feature of the third feature.
[0173] In summary, the solution shown in the embodiments of the present application downsamples the initial features of the target image at multiple levels, calculates the attention map through the features obtained by the last-level downsampling, and then processes the initial features and the downsampled features at each level through the attention map to extract the features of the target image. During this process, there is no need to segment the target image, avoiding the loss of feature continuity caused by image segmentation. At the same time, through the processing of the initial features and the downsampled features at each level by the attention map, the complete features of the target image can be extracted, thus ensuring the accuracy of image feature extraction and further improving the accuracy of the processing results of image processing tasks.
[0174] Figure 8 FIG. 4 is a schematic structural diagram of a computer device according to an exemplary embodiment. The computer device 800 includes a central processing unit (CPU, Central Processing Unit) 801, a system memory 804 including a random access memory (Random Access Memory, RAM) 802 and a read-only memory (Read-Only Memory, ROM) 803, and a system bus 805 connecting the system memory 804 and the central processing unit 801. The computer device 800 further includes a basic input / output system 806 for facilitating the transfer of information between various components within the computer, and a mass storage device 807 for storing an operating system 813, application programs 814, and other program modules 815.
[0175] The mass storage device 807 is connected to the central processing unit 801 through a mass storage controller (not shown) connected to the system bus 805. The mass storage device 807 and its associated computer-readable medium provide non-volatile storage for the computer device 800. That is to say, the mass storage device 807 may include a computer-readable medium (not shown) such as a hard disk or a compact disc read-only memory (Compact Disc Read-Only Memory, CD-ROM) drive.
[0176] Without loss of generality, the computer-readable medium may include a computer storage medium and a communication medium. The computer storage medium includes volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. The computer storage medium includes RAM, ROM, flash memory or other solid-state storage technologies, CD-ROM, or other optical storage, magnetic tape cartridges, tapes, disk storage, or other magnetic storage devices. Of course, those skilled in the art will appreciate that the computer storage medium is not limited to the above several types. The above-mentioned system memory 804 and mass storage device 807 may be collectively referred to as a memory.
[0177] The computer device 800 can be connected to the Internet or other network devices through the network interface unit 811 connected to the system bus 805.
[0178] The memory further includes one or more programs stored in the memory, and the central processing unit 801 implements Figure 2 or Figure 4 all or part of the steps of any of the illustrated methods.
[0179] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory including a computer program (instructions), and the above program (instructions) can be executed by a processor of a computer device to complete the methods illustrated in various embodiments of the present application. For example, the non-transitory computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), magnetic tape, floppy disk, and optical data storage device, etc.
[0180] In an exemplary embodiment, a computer program product or a computer program is also provided, and the computer program product or the computer program includes computer instructions stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods illustrated in the above various embodiments.
[0181] Those skilled in the art will readily conceive of other embodiments of the present application after considering the specification and practicing the invention disclosed herein. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include common general knowledge or conventional technical means in the technical field not disclosed in the present application. The specification and embodiments are only to be regarded as exemplary, and the true scope and spirit of the present application are pointed out by the claims.
[0182] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.
Claims
1. An image processing method, characterized in that, The method includes: Performing n-level downsampling processing on the initial features of the target image to obtain n downsampled features; n is a positive integer; Obtaining an attention map of the target image based on a target downsampled feature among the n downsampled features; the target downsampled feature is obtained by the last-level downsampling processing of the n-level downsampling processing; Upsampling the attention map to obtain a first attention map with a first scale; the first scale is the scale of a first feature, and the first feature is any one of the initial feature and the n downsampled features; Performing matrix multiplication on the first attention map and the first feature to obtain an attention-processed feature corresponding to the first feature; Based on the initial feature and the attention-processed features respectively corresponding to the n downsampled features, obtaining a processing result for performing a specified image processing task on the target image.
2. The method according to claim 1, wherein The obtaining an attention map of the target image based on a target downsampled feature among the n downsampled features includes: Obtaining the attention map based on the query dimension feature and the key dimension feature of the target downsampled feature; The performing matrix multiplication on the first attention map and the first feature to obtain an attention-processed feature corresponding to the first feature includes: Performing matrix multiplication on the first attention map and the value dimension feature of the first feature to obtain an attention-processed feature corresponding to the first feature.
3. The method according to claim 1, wherein The obtaining a processing result for performing a specified image processing task on the target image based on the initial feature and the attention-processed features respectively corresponding to the n downsampled features includes: Fusing the initial feature and the attention-processed features respectively corresponding to the n downsampled features to obtain an image feature of the target image; Obtaining the processing result based on the image feature.
4. The method according to claim 3, characterized in that, The fusing the initial feature and the attention-processed features respectively corresponding to the n downsampled features to obtain an image feature of the target image includes: Upsampling the attention-processed features respectively corresponding to the n downsampled features to obtain n upsampled features; the scale of the upsampled features is the same as the scale of the initial feature; Cascading the initial feature and the n upsampled features to obtain the image feature.
5. The method according to claim 1, characterized in that, The performing n-level downsampling processing on the initial feature to obtain n downsampled features includes: Performing downsampling processing on a second feature to obtain an intermediate downsampled feature; the second feature is any one of the initial feature and the n downsampled features except the target downsampled feature; Performing convolution processing on the intermediate downsampled feature to obtain the next-level downsampled feature of the second feature.
6. The method according to claim 1, wherein The performing n-level downsampling processing on the initial feature to obtain n downsampled features includes: Performing convolution processing on a third feature to obtain an intermediate convolution feature; the third feature is any one of the initial feature and the n downsampled features except the target downsampled feature; Perform downsampling processing on the intermediate convolutional feature to obtain the next-level downsampled feature of the third feature.
7. An image processing apparatus, characterized in that, The device includes: A downsampling module, configured to perform n-level downsampling processing on the initial feature of the target image to obtain n downsampled features; An attention map acquisition module, configured to obtain an attention map of the target image based on a target downsampled feature among the n downsampled features; the target downsampled feature is obtained by the last-level downsampling processing of the n-level downsampling processing; An attention processing module, including an upsampling unit and a processing unit; The upsampling unit is configured to upsample the attention map to obtain a first attention map with a first scale; the first scale is the scale of a first feature, and the first feature is any one of the initial feature and the n downsampled features; The processing unit is configured to perform matrix multiplication processing on the first attention map and the first feature to obtain an attention processing feature corresponding to the first feature; A result acquisition module, configured to obtain a processing result of performing a specified image processing task on the target image based on the initial feature and the attention processing features respectively corresponding to the n downsampled features.
8. A computer device, characterized in that, The computer device includes a processor and a memory, and at least one computer instruction is stored in the memory, and the at least one computer instruction is loaded and executed by the processor to implement the image processing method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, At least one computer instruction is stored in the storage medium, and the at least one computer instruction is loaded and executed by a processor to implement the image processing method according to any one of claims 1 to 6.
10. A computer program product, characterized in that, The computer program product includes computer instructions, and the computer instructions are executed by a processor of a computer device, so that the computer device executes the image processing method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Traffic identifier detection method based on multi-scale circulation attention network
CN108647585A
Medical image segmentation system based on attention routing
CN113129310A