Classification System and Method for Information in Images
The proposed classification system addresses the complexity and overfitting issues in multi-task medical imaging by using a convolutional neural network with attention networks to generate and fuse attention maps, enhancing diagnostic accuracy and reducing model size.
Patent Information
- Application Number
- CN202110282440.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-03-16
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2041-03-16
AI Technical Summary
Existing medical imaging AI-assisted diagnostic tools are difficult to perform multiple diagnostic tasks at the same time, resulting in the model being too large or unable to effectively share convolutional layer features, resulting in overfitting or inaccurate diagnosis.
The architecture of a convolutional neural network and attention network is adopted to combine the convolutional neural network and attention network. Through the shared feature map, attention circuits and fusion circuits are used for multi-task learning, selecting the characteristics of attention for different classification tasks, and generating classification results through classifiers.
It realizes the commonality and accuracy of the model in multiple diagnostic tasks, reduces the complexity of the model and data volume, and improves the accuracy of the classification task, especially in medical imaging diagnosis, which can handle multiple symptoms simultaneously.
Smart Images

Figure CN115147688B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to artificial intelligence and machine learning, and in particular to a classification system and method for information in images. Background Art
[0002] Artificial Intelligence (AI) has shown satisfactory results in learning from data to simulate human brain understanding, reasoning, planning, communication, and perception. Therefore, many studies have proposed using AI-assisted diagnostic tools for medical images. The medical images include: X-rays, Computed Tomography (CT), Magnetic Resonance Imaging (MRI), and color images, etc. However, most AI-assisted diagnostic tools using medical images can only diagnose a single symptom.
[0003] To perform multiple diagnostic tasks, one approach is to design and train multiple models. For a system that performs multiple diagnostic tasks, the model complexity or parameter size will be proportional to the number of diagnostic tasks that the system needs to perform. Therefore, the models generated by this approach are often too large and have many parameters, and are easily overfitted due to a small amount of training data.
[0004] To perform multiple diagnostic tasks, another approach is hard parameter sharing, which shares all convolutional layers in the model and then uses different classifiers or regressors at the end of the neural network. However, the models generated by this approach are too general, so that all the detailed features required to judge various symptoms cannot be found in the shared convolutional layers. Summary of the Invention
[0005] In view of this, the present invention proposes a classification system and method for information in images. In order to make the artificial intelligence model more general and utilize the relationships between related symptoms, the architecture corresponding to the classification system proposed by the present invention can simultaneously learn all related symptoms and is also specific enough to evaluate each individual task.
[0006] A method for classifying information in an image according to an embodiment of the present invention includes: a convolutional neural network receiving an input image and generating a plurality of shared feature maps; an attention network generating a plurality of attention maps based on these shared feature maps; a fusion circuit selecting at least two of these attention maps to perform a fusion operation to generate a fusion map; and a classifier generating a classification result based on the fusion map.
[0007] A method for classifying information in an image according to an embodiment of the present invention includes: a first convolutional layer receiving an input image; the first convolutional layer performing a convolutional operation according to the input image to generate a first feature map; a first attention circuit performing an attention operation according to the first feature map to generate a first attention map; a second convolutional layer performing a convolutional operation according to the first feature map to generate a second feature map; a second attention circuit performing another attention operation according to the second feature map and the first attention map to generate a second attention map; a third convolutional layer performing a convolutional operation according to the second feature map to generate a third feature map; a third attention circuit performing another attention operation according to the third feature map and the second attention map to generate a third attention map; a fusion circuit selecting at least two of the first attention map, the second attention map, and the third attention map to perform a fusion operation to generate a fusion map; and a classifier generating a classification result according to the fusion map.
[0008] A system for classifying information in an image according to an embodiment of the present invention includes a first convolutional layer, a second convolutional layer, a third convolutional layer, a first attention circuit, a second attention circuit, a third attention circuit, a fusion circuit, and a classifier. The first convolutional layer receives an input image and performs a convolutional operation according to the input image to generate a first feature map. The second convolutional layer is communicatively connected to the first convolutional layer and performs a convolutional operation according to the first feature map to generate a second feature map. The third convolutional layer is communicatively connected to the second convolutional layer and performs a convolutional operation according to the second feature map to generate a third feature map. The first attention circuit is communicatively connected to the first convolutional layer and performs an attention operation according to the first feature map to generate a first attention map. The second attention circuit is communicatively connected to the second convolutional layer and the first attention circuit and performs another attention operation according to the second feature map and the first attention map to generate a second attention map. The third attention circuit is communicatively connected to the third convolutional layer and the second attention circuit and performs another attention operation according to the third feature map and the second attention map to generate a third attention map. The fusion circuit is communicatively connected to at least two of the first attention circuit, the second attention circuit, and the third attention circuit and performs a fusion operation according to at least two of the first attention map, the second attention map, and the third attention map to generate a fusion map. The classifier is communicatively connected to the fusion circuit and generates a classification result according to the fusion map.
[0009] The above description of the content of this application and the following description of the embodiments are used to demonstrate and explain the spirit and principle of the present invention, and provide a further explanation of the scope of the patent application of the present invention. Description of the Drawings
[0010] Figure 1 is an architecture diagram of a system for classifying information in an image according to an embodiment of the present invention;
[0011] Figure 2It is a flowchart of a method for classifying information in an image according to an embodiment of the present invention;
[0012] Figure 3 Illustrates Figure 1 A detailed architecture diagram of a sub-architecture of as an example;
[0013] Figure 4 It is a flowchart of a method for classifying information in an image according to an embodiment of the present invention;
[0014] Figure 5 Illustrated as an execution flowchart of an attention operation;
[0015] Figure 6 Illustrated as an execution flowchart of a dimension adjustment operation; and
[0016] Figure 7 It is a schematic diagram of an embodiment of a fusion operation.
[0017] Symbol description
[0018] CNN Convolutional Neural Network
[0019] L1, L2, L3, L4 First convolutional layer, second convolutional layer, third convolutional layer, fourth convolutional layer
[0020] N1, N2 First attention network, second attention network
[0021] A1~A4 First attention circuit, second attention circuit, third attention circuit, fourth attention circuit
[0022] B1~B4 Attention circuit
[0023] F1, F2 First fusion circuit, second fusion circuit
[0024] C1, C2 First classifier, second classifier
[0025] S0, S12~S14, S22~S24, S40~S50 Steps
[0026] A11, A21, A31, A41 Mask generation circuit
[0027] K1, K2, K3, K4 Attention mask
[0028] A12, A22, A32, A42 Bit multiplication circuit
[0029] A13, A23, A33, A43 Dimension adjustment circuit
[0030] M1, M2, M3, M4 Attention map Detailed implementation manners
[0031] The detailed features and characteristics of the present invention are described in detail in the embodiments below. The content is sufficient for any person skilled in the relevant art to understand the technical content of the present invention and implement it accordingly. Based on the content disclosed in this specification, the scope of the patent application, and the drawings, any person skilled in the relevant art can easily understand the related concepts and characteristics of the present invention. The following embodiments further illustrate the viewpoints of the present invention in detail, but do not limit the scope of the present invention in any way.
[0032] The classification system and method for information in an image proposed by the present invention can be used to classify a single medical image for multiple tasks. For example, when the present invention is applied to mammograms, in addition to diagnosing whether a tumor is malignant or benign, it can also simultaneously classify and output multiple tasks listed in Table 1 below.
[0033] Table 1
[0034]
[0035]
[0036] Figure 1 is a framework diagram of the classification system for information in an image according to an embodiment of the present invention. This framework is applicable to two classification tasks. Figure 2 is a flowchart of the classification method for information in an image according to an embodiment of the present invention. This flowchart corresponds to Figure 1 the framework diagram of.
[0037] Figure 1 Illustrates a convolutional neural network including multiple convolutional layers L1 - L4, a first attention network N1 including multiple attention circuits A1 - A4, a second attention network N2 including multiple attention circuits B1 - B4, first and second fusion circuits F1 - F2, and first and second classifiers C1 - C2. The arrows in each block in the figure represent the data flow direction of the output of this block.
[0038] The sub - architecture composed of the convolutional neural network CNN, the first attention network N1, the first fusion circuit F1, and the first classifier C1 can handle the first classification task, and the sub - architecture composed of the convolutional neural network CNN, the second attention network N2, the second fusion circuit, and the second classifier C2 can handle the second classification task. The classification task can be binary classification, such as "malignant" or "benign" as shown in Table 1. The classification task can also be multi - class classification, such as the multiple output classifications corresponding to the "symptom" task as shown in Table 1. The present invention uses attention networks in the above two sub - architectures respectively, so it can utilize the shared features of different classification tasks and find the differences between classification tasks.
[0039] Please refer to Figure 2 . Step S0 is that the convolutional neural network CNN receives an input image and generates multiple shared feature maps. Specifically, the convolutional neural network CNN receives an input image from the outside through the convolutional layer L1, such as an X-ray image of mammography. The convolutional neural network CNN can extract multiple features of the input image through the convolutional layers L1 to L4. Since it is common knowledge to use the convolutional neural network CNN for image feature extraction, the detailed steps of feature extraction will not be elaborated here.
[0040] Step S12 is that the first attention network N1 generates multiple first attention maps based on these shared feature maps. For example, the hidden layer L1 generates a shared feature map and transmits it to the attention circuit A1, and the attention circuit A1 generates a first attention map based on this shared feature map and transmits it to the attention circuit A2. The data paths between the remaining hidden layers L2 to L4 and the attention circuits A2 to A4 can be analogized as above.
[0041] It should be noted that there is a one-to-one correspondence between the attention circuits A1 to A4 and the convolutional layers L1 to L4. Similarly, there is a one-to-one correspondence between the attention circuits B1 to B4 and the convolutional layers L1 to L4. Through the above data structure, the attention circuits A1 to A4 or B1 to B4 can adaptively adjust the parts that need to be concerned about for the shared feature maps respectively generated by each convolutional layer L1 to L4.
[0042] Step S13 is that the first fusion circuit F1 selects at least two of these first attention maps to perform a first fusion operation to generate a first fusion map. Figure 1 The illustrated example is that the first fusion circuit F1 selects the three attention circuits A2, A3, and A4 to perform the fusion operation.
[0043] Step S14 is that the first classifier C1 generates a first classification result based on the first fusion map. The first classifier C1 is implemented in the form of a fully-connected layer, for example.
[0044] The processes of steps S22 to S24 are roughly the same as the processes of steps S12 to S14, except that: at least one of the multiple second attention maps generated in step S22 is different from the multiple first attention maps generated in step S21. In other words, for different classification tasks, the places where the attention maps focus are also different.
[0045] In addition, in step S23, the second fusion circuit F2 selects three of the attention circuits B1, B3, and B4, etc., to perform the fusion operation. When the classification tasks are different, in order to improve the classification accuracy, the attention maps to be obtained are also different. The fusion circuits F1 to F2 proposed by the present invention simulate the comprehensive consideration of the macroscopic and microscopic aspects in human perception by fusing images at different levels. In practice, the number of selected attention maps is a hyper-parameter pre-determined by the user, and at least two attention maps need to be selected. Different layers in the convolutional neural network CNN contain different information. The low-level features have more color, edge, and spatial information, while the high-level features have more semantic information. The fusion map retains multi-level information by fusing multiple attention maps, thereby expanding the receptive field. A larger receptive field, for example, includes long-distance relationships between pixels and is helpful for classification tasks. For example, the wound size in medical images can range from very small to very large. Although fusing more layers can include more features, this behavior is not always beneficial to the classification task. Since the discriminative power of low-level features is weak, fusing too many layers may reduce the classification accuracy. Therefore, the number of fused layers (i.e., the number of selected attention maps) can be a hyper-parameter for adjusting the model.
[0046] Similarly, which attention maps output by the attention circuits are selected to perform the fusion operation, the number of convolutional layers, and the number of fully connected layers in the classifier are all hyper-parameters that can be pre-configured by the user.
[0047] In addition, it should be noted that although Figure 1 shows the classification system architecture diagrams for two classification tasks. However, the present invention does not limit the number of tasks applicable to the classification system. For example, according to the connection form between the first attention network A1 and the convolutional neural network, the user can add a third attention network to communicate with the convolutional neural network CNN. The third attention network shares multiple shared feature maps generated by the convolutional neural network CNN with the first attention network N1. Therefore, compared with training three independent classification models respectively, the architecture adopted by the present invention is undoubtedly more flexible and reduces the overall data volume of the classification system.
[0048] To clearly illustrate the internal structure of the attention network, Figure 3 shows Figure 1 a detailed architecture diagram of a sub-architecture as an example. This sub-architecture consists of a convolutional neural network CNN, a first attention network N1, a first fusion circuit F1, and a first classifier C1. The reader should be able to understand that the details of the sub-architecture composed of the convolutional neural network CNN, the second attention network N2, the second fusion circuit F2, and the second classifier C2 are basically the same as Figure 3 the same.
[0049] AsFigure 3 As shown in Figure 3 , a classification system for information in an image according to an embodiment of the present invention includes: a first convolutional layer L1, a second convolutional layer L2, a third convolutional layer L3, a fourth convolutional layer L4, a first attention circuit A1, a second attention circuit A2, a third attention circuit A4, a first fusion circuit F1, and a first classifier C1.
[0050] Figure 3 Cooperate with Figure 1 Draw. In Figure 1 In order to clearly show that the first fusion circuit F1 and the second fusion circuit F2 correspond to different convolutional layers respectively, a classification system with four convolutional layers L1 to L4 is taken as an example. The present invention does not limit the upper limit number of convolutional layers. The present invention can also adopt more than five convolutional layers, depending on the user's hyperparameter configuration, but at least three convolutional layers are required. A person of ordinary skill in the art can easily derive the architecture diagrams with different numbers of convolutional layers according to Figure 1 And Figure 3 The architecture diagrams shown.
[0051] Figure 4 Is a flowchart of a method for classifying information in an image according to an embodiment of the present invention, and this flowchart corresponds to Figure 3 The architecture diagram of.
[0052] As Figure 3 And steps S40 and S41 shown, the first convolutional layer L1 receives the input image and performs a convolutional operation according to the input image to generate a first feature map.
[0053] As Figure 3 And steps S43 shown, the second convolutional layer L2 is communicatively connected to the first convolutional layer L1 and performs a convolutional operation according to the first feature map to generate a second feature map.
[0054] As Figure 3 And steps S45 shown, the third convolutional layer L3 is communicatively connected to the second convolutional layer L2 and performs a convolutional operation according to the second feature map to generate a third feature map.
[0055] As Figure 3 And steps S47 shown, the fourth convolutional layer L4 is communicatively connected to the third convolutional layer L3 and performs a convolutional operation according to the third feature map to generate a fourth feature map.
[0056] The main difference in the convolutional operations performed by the first to fourth convolutional layers L1 to L4 lies in the size of the input image and the output image of each convolutional layer. The first to fourth feature maps are not shown in Figure 3 .
[0057] As Figure 3As shown in step S42, the first attention circuit A1 is communicatively connected to the first convolutional layer L1 and performs an attention operation based on the first feature map to generate a first attention map M1. In an embodiment, each time the attention circuit processes a layer, the size of the attention map becomes smaller, which means that there are fewer parts to be concerned about in the image, and the features generated by the higher-level attention circuits are more important as classification judgment indicators.
[0058] Figure 5 The execution flowchart of the attention operation is shown. The masking generation circuit A11 in the first attention circuit A1 performs a 1×1 convolution operation based on the first feature map, then performs batch normalization on the result of the 1×1 convolution operation, and then inputs the result of the batch normalization into the sigmoid function to generate an attention mask K1. The values of the attention mask can be real numbers or {0, 1}. An example of a 3×3 attention mask is shown in Table II below. The larger the value in the attention mask, the higher the importance of the corresponding feature map pixel at that position.
[0059] Table II
[0060]
[0061]
[0062] As Figure 3 shown, the bit multiplication circuit A12 in the first attention circuit A1 performs a pixel multiplication operation based on the attention mask K1 and the first feature map. Performing the pixel multiplication operation is equivalent to magnifying the important pixels in the feature map by the attention mask and ignoring the unimportant pixels.
[0063] Figure 6 The execution flowchart of the dimension adjustment operation is shown. As Figure 3As shown, the dimension adjustment circuit A13 in the first attention circuit A1 performs a 3×3 convolution operation on the output result of the bit multiplication circuit A12, then performs batch normalization on the result of the 3×3 convolution operation, and then inputs the result of the batch normalization into a rectified linear unit (ReLU) to generate a first attention map. In an embodiment, the dimension adjustment circuit is used to match the number of channels between adjacent two layers (such as A1 and A2, or A2 and A3). For example, assume that the size of the first feature map output by L1 is 200×300×16, and the size of the second feature map output by L2 is 100×150×64; then the size of M1 must be 3×3×64, where 64 is the number of channels of L2. In another embodiment of the present invention, each of the dimension adjustment circuits A13 to A43 is further connected to a pooling layer to adjust the length and width dimensions of the attention maps M1 to M4. For example, the maximum-pooling method can be used to perform down-sampling on the length and width dimensions.
[0064] As Figure 3 shown in and step S44, the second attention circuit A2 is communicatively connected to the second convolutional layer L2 and the first attention circuit A1, and performs another attention operation based on the second feature map and the first attention map M1 to generate a second attention map M2.
[0065] As Figure 3 shown in and step S46, the third attention circuit A3 is communicatively connected to the third convolutional layer L3 and the second attention circuit A2, and performs another attention operation based on the third feature map and the second attention map M2 to generate a third attention map M3.
[0066] As Figure 3 shown in and step S48, the fourth attention circuit A4 is communicatively connected to the third convolutional layer L4 and the third attention circuit A3, and performs another attention operation based on the fourth feature map and the third attention map M3 to generate a fourth attention map M4.
[0067] The detailed architectures of the second attention circuit A2, the third attention circuit A3, and the fourth attention circuit A4 are basically the same. Here, the second attention circuit A2 is taken as an example and described as follows. The said another attention operation is basically the same as Figure 5 the execution process of the said attention operation: the masking generation circuit A21 in the second attention circuit A2 sequentially performs a 1×1 convolution operation, batch normalization, and an S function based on the second feature map and the first attention map M1 to generate an attention mask K2. The bit multiplication circuit A22 in the second attention circuit A2 performs a pixel multiplication operation based on the attention mask K2 and the first feature map.
[0068] AsFigure 6 As shown, the dimension adjustment circuit A23 in the second attention circuit A2 performs a 3×3 convolution operation on the output result of the bit multiplication circuit A22, then performs batch normalization on the result of the 3×3 convolution operation, and then inputs the result of the batch normalization into the rectified linear unit to generate the second attention map M2.
[0069] Overall, for each of the attention circuits A1 to A4, first, attention masks K1 to K4 are generated based on the shared feature maps output by the convolutional layers L1 to L4 and the attention maps of the previous-level attention circuit (if any). Then, the attention masks are subjected to a masking operation (pixel multiplication operation) with the shared feature maps, and the result of the above masking operation is further subjected to dimension adjustment to finally generate an attention map.
[0070] As Figure 3 shown, the fusion circuit F1 is communicatively connected to at least two of the first attention circuit A1, the second attention circuit A2, the third attention circuit A3, and the fourth attention circuit A4. In Figure 3 the example, the fusion circuit F1 is communicatively connected to the second attention circuit A2, the third attention circuit A3, and the fourth attention circuit A4. As shown in step S49, the fusion circuit F1 selects at least two of the first attention map M1, the second attention map M2, the third attention map M3, and the fourth attention map M4 to perform a fusion operation to generate a fusion map. In Figure 3 the example, the fusion circuit F1 performs a fusion operation based on the second attention map M2, the third attention map M3, and the fourth attention map M4.
[0071] Figure 7 is a schematic diagram of an embodiment of the fusion operation. Taking Figure 3 the example, the second attention map M2 is an attention map of a lower layer, so its size is larger and the number of channels is smaller. The fourth attention map M4 is an attention map of a higher layer, with a smaller size and a larger number of channels. Therefore, before the fusion circuit fuses the three attention maps M2 to M4, an image normalization operation needs to be performed to generate M2' to M4', and then an image mixing operation is carried out. The image normalization operation includes size adjustment and number-of-channels adjustment. The size adjustment is based on the attention map M2 with the largest size, and the attention maps M3 and M4 with smaller sizes are subjected to up-sampling operations to make the sizes of these two attention maps M3 and M4 the same as the size of the attention map M2. The number-of-channels adjustment is based on the attention map M4 with the largest number of channels, and 1×1 convolution operations are performed on the attention maps M2 and M3 with smaller numbers of channels by providing a larger number of convolutional kernels to make the numbers of channels of these two attention maps M2 and M3 the same as the number of channels of the attention map M4.
[0072] Another embodiment of the channel number adjustment is to adjust the channel numbers of all attention maps M2 to M4 downward to the same value. This value can be less than the minimum channel number in the attention maps M2 to M4. The specific implementation of the downward adjustment is also through 1×1 convolution operation, and the dimensionality reduction operation of all attention maps M2 to M4 is achieved by providing a smaller number of convolution kernels.
[0073] Two embodiments of the mixing operation include bit-scale addition operation, or concatenation operation.
[0074] As described above, there are two implementation methods in the image normalization operation (size up adjustment and channel number up adjustment, size up adjustment and channel number down adjustment), and there are two implementation methods in the mixing operation (superposition operation, concatenation operation). Therefore, the fusion operation described in step S50 can combine the two implementation methods of the image normalization operation and the two implementation methods of the mixing operation, and has four implementation methods.
[0075] As Figure 3 shown in step S50, the classifier C1 is communicatively connected to the fusion circuit F1 and generates a classification result based on the fusion map. The classifier C1 can adopt the form of a fully connected layer of "512→64→2" for example, and finally generate a binary classification result.
[0076] It should be noted that Figure 4 the flowchart of can also first complete the feature map extraction process of the convolutional neural network CNN, and then continue to execute the attention map generation process of the attention network, as shown in the following execution process: S40→S41→S43→S45→S47→S42→S44→S46→S48→S49→S50.
[0077] The following supplements the execution process of only using three convolutional layers and three attention circuits in the image information classification method of the present invention: The first convolutional layer receives the input image; the first convolutional layer performs a convolution operation based on the input image to generate a first feature map; the first attention circuit performs an attention operation based on the first feature map to generate a first attention map; the second convolutional layer performs a convolution operation based on the first feature map to generate a second feature map; the second attention circuit performs another attention operation based on the second feature map and the first attention map to generate a second attention map; the third convolutional layer performs a convolution operation based on the second feature map to generate a third feature map; the third attention circuit performs another attention operation based on the third feature map and the second attention map to generate a third attention map; the fusion circuit selects at least two from at least the first attention map, the second attention map, and the third attention map to perform a fusion operation to generate a fusion map; and the classifier generates a classification result based on the fusion map.
[0078] In summary, the information classification system for images proposed by the present invention is equivalent to a multi-task learning system. The present invention extracts features shared by convolutional layers through the adoption of an attention mechanism and utilizes an image hierarchy to perform multiple classification tasks.
[0079] The information classification system for images proposed by the present invention can store the trained model and the attention network on a remote server. Users working locally can capture physiological images using a smartphone or webcam and then upload them to the classification system on the server for symptom judgment. Another implementation is to store a lightweight attention network locally and store multiple feature maps generated by the convolutional neural network on a remote server, and interact through a mobile application and the network service provided by the server, thereby achieving the diagnosis of medical images assisted by artificial intelligence in the form of edge computing.
[0080] The present invention adopts multi-task training and shares model features among different tasks. In order to make the system proposed by the present invention general enough to utilize the shared features of different tasks and subtle enough to find the differences between tasks, the present invention adopts the architecture of an attention network. The effects of the present invention include: the proposed architecture is general enough to utilize the shared features of different tasks and can also find the differences between tasks. The architecture proposed by the present invention takes into account the multi-level structure of image data and can find not only rough image features but also fine image features.
[0081] Although the present invention is disclosed as above in the foregoing embodiments, it is not intended to limit the present invention. Any changes and modifications made without departing from the spirit and scope of the present invention fall within the scope of patent protection of the present invention. For the scope of protection defined by the present invention, please refer to the appended patent application scope.
Claims
1. A method for classifying information in an image, characterized in that, Comprising: Receiving an input image by a convolutional neural network and generating a plurality of shared feature maps; Generating a plurality of first attention maps by a first attention network according to the plurality of shared feature maps; Selecting at least two of the plurality of first attention maps by a first fusion circuit to perform a first fusion operation to generate a first fusion map; and Generating a first classification result by a first classifier according to the first fusion map; Performing a 1×1 convolution operation, batch normalization, and S function in sequence by a second attention circuit according to a second feature map and the first attention map to generate an attention mask; And performing at least one pixel multiplication operation by the second attention circuit according to the attention mask and the second feature map to generate a second attention map.
2. The method for classifying information in an image according to claim 1, further comprising: Generating a plurality of second attention maps by a second attention network according to the plurality of shared feature maps; Selecting at least two of the plurality of second attention maps by a second fusion circuit to perform a second fusion operation to generate a second fusion map; and Generating a second classification result by a second classifier according to the second fusion map; wherein At least one of the plurality of second attention maps is different from the plurality of first attention maps.
3. A method for classifying information in an image, characterized in that, Comprising: Receiving an input image by a first convolutional layer; Performing a convolution operation by a first convolutional layer according to the input image to generate a first feature map; Performing an attention operation by a first attention circuit according to the first feature map to generate a first attention map; Performing the convolution operation by a second convolutional layer according to the first feature map to generate a second feature map; Performing another attention operation by a second attention circuit according to the second feature map and the first attention map to generate a second attention map; performing the convolution operation by a third convolutional layer according to the second feature map to generate a third feature map; Performing the another attention operation by a third attention circuit according to the third feature map and the second attention map to generate a third attention map; Selecting at least two of at least the first attention map, the second attention map, and the third attention map by a fusion circuit to perform a fusion operation to generate a fusion map; and Generating a classification result by a classifier according to the fusion map; Performing another attention operation by a second attention circuit according to the second feature map and the first attention map to generate a second attention map includes: performing a 1×1 convolution operation, batch normalization, and S function in sequence by the second attention circuit according to the second feature map and the first attention map to generate an attention mask; And performing at least one pixel multiplication operation by the second attention circuit according to the attention mask and the second feature map to generate the second attention map.
4. The method for classifying information in an image according to claim 3, wherein performing the attention operation by the first attention circuit according to the first feature map to generate the first attention map includes: The first attention circuit performs 1×1 convolution operations, batch normalization, and the S function in sequence based on the first feature map order to generate an attention mask; and The first attention circuit performs at least one pixel multiplication operation based on the attention mask and the first feature map to generate the first attention map.
5. The method for classifying information in an image according to claim 3, wherein the fusion circuit selects at least two of the first attention map, the second attention map, and the third attention map to perform the fusion operation to generate the fusion map, including: Adjusting the sizes of the at least two attention maps to be the same; Adjusting the channel number dimensions of the at least two attention maps to be the same; and Performing a mixing operation on the at least two attention maps after adjusting the size and channel number.
6. A classification system for information in an image, characterized in that, including: A first convolutional layer, receiving an input image, and performing a convolution operation based on the input image to generate a first feature map; A second convolutional layer, communicatively connected to the first convolutional layer, and performing the convolution operation based on the first feature map to generate a second feature map; A third convolutional layer, communicatively connected to the second convolutional layer, and performing the convolution operation based on the second feature map to generate a third feature map; A first attention circuit, communicatively connected to the first convolutional layer, and performing an attention operation based on the first feature map to generate a first attention map; A second attention circuit, communicatively connected to the second convolutional layer and the first attention circuit, and performing another attention operation based on the second feature map and the first attention map to generate a second attention map; A third attention circuit, communicatively connected to the third convolutional layer and the second attention circuit, and performing the another attention operation based on the third feature map and the second attention map to generate a third attention map; A fusion circuit, communicatively connected to at least two of the first attention circuit, the second attention circuit, and the third attention circuit, and performing a fusion operation based on at least two of the first attention map, the second attention map, and the third attention map to generate a fusion map; and A classifier, communicatively connected to the fusion circuit, and generating a classification result based on the fusion map; The second attention circuit is used for: the second attention circuit performs 1×1 convolution operations, batch normalization, and the S function in sequence based on the second feature map and the first attention map to generate an attention mask; and the second attention circuit performs at least one pixel multiplication operation based on the attention mask and the second feature map to generate the second attention map.
Citation Information
Patent Citations
Image semantic segmentation method and device, electronic equipment and storage medium
CN112465828A