Colorectal cancer grading method and device for perceiving attention based on semantic token

By using semantic token perceive attention network in colorectal cancer grading, obtaining semantic tokens, channel tokens and spatial semantic token feature maps, the problem of limited computing resources in colorectal cancer grading is solved, the image feature representation ability and local semantic modeling ability are improved, and more accurate colorectal cancer grading is achieved.

CN120107669APending Publication Date: 2025-06-06THE THIRD AFFILIATED HOSPITAL OF GUANGZHOU MEDICAL UNIVERSITY (GUANGZHOU SEVERE MATERNAL TREATMENT CENTER GUANGZHOU ROUJI HOSPITAL)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510171531.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-17
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

At this stage, the colorectal cancer grading method based on deep learning cannot be applied to clinical practice on a large scale due to limited computing resources, resulting in the inability to effectively represent the information of tag histological images. Image segmentation leads to distortion of feature details, unable to contain the microstructure of the entire tissue, and loss of important data information.

Method used

A colorectal cancer grading method based on semantic token perception attention is adopted. A semantic token, channel token and spatial semantic token feature map is obtained through the semantic token perception attention network, and a feedforward network is combined to output the prediction level.

Benefits of technology

It effectively solves the problem of limited computing resources in colorectal cancer grading, improves image feature representation ability, enhances local semantic modeling ability, and better integrates the microscopic regional information of histological images, thereby improving the accuracy of colorectal cancer grading.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107669A_ABST
    Figure CN120107669A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of medical image processing, in particular to a colorectal cancer grading method and device for perceiving attention based on semantic tokens, and the method comprises the steps: obtaining a histological image of a target region, segmenting the histological image into small image patches with preset sizes, and extracting a feature map; inputting the feature map into a semantic token perception attention network to obtain an output feature map with semantic token perception attention, and inputting the output feature map into a feedforward network to obtain a colorectal cancer prediction level of the target area; according to the method, global information of a histological image is learned through the class token semantic perception attention module, and fine local pixel-level details are modeled through the channel token semantic perception attention module to compensate feature details of the whole image, so that a global dependency relationship between channels is established; the spatial token semantic perception attention module captures dual attention spatial semantic information to fuse microscopic region information of the histological image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical image processing, and in particular to a method and device for grading colorectal cancer based on semantic token-aware attention. Background Art

[0002] Colorectal cancer is one of the common malignant tumors caused by multiple factors. In clinical practice, doctors usually assess the grade of colorectal cancer based on the degree of tumor cell infiltration in the intestinal wall.

[0003] In clinical diagnosis, experienced pathologists grade colorectal cancer by analyzing the abnormalities of individual cancer cells and the spatial organization of aberrant glandular structures. Deep learning has also been applied to colorectal cancer grading because of its outstanding performance in signal detection, image recognition and other fields.

[0004] Although deep learning-based methods have performed well in colorectal cancer grading, clinical practice is still greatly limited due to the following issues:

[0005] Due to limited computing resources, these methods always tend to perform representation learning on small image blocks (image patches) in high-resolution digital histology images. However, firstly, class tokens based on small image blocks can only represent local information, but cannot represent the information of the labeled histology images, such as Figure 1 As shown in area A in ; secondly, image segmentation always leads to distortion of the details of the marked features in the entire image, such as Figure 1 As shown in B, this reduces the ability of local modeling semantics, further leading to the weakening of pixel-level channel dependence in colorectal cancer cells. Thirdly, image patch segmentation strategies often fail to include the entire tissue microstructure, thereby losing important data information and regions of interest, such as Figure 1 As shown in area B. Summary of the invention

[0006] In view of this, the purpose of an embodiment of the present invention is to provide a method and device for colorectal cancer grading based on semantic token-aware attention, so as to solve the problem that the current colorectal cancer grading method based on deep learning cannot be applied on a large scale in clinical practice.

[0007] In a first aspect, an embodiment of the present invention provides a method for grading colorectal cancer based on semantic token-aware attention, comprising:

[0008] Acquire a histological image of a target area, segment the histological image into small image patches of a preset size and extract a feature map;

[0009] The feature map is input into a semantic token-aware attention network to obtain an output feature map with semantic token-aware attention, wherein the output feature map includes a class semantic token feature map, a channel token feature map and a space semantic token map; wherein the semantic token-aware attention network includes at least a class token semantic-aware attention module, a channel token semantic-aware attention module and a space token semantic-aware attention module, wherein the class token semantic-aware attention module is used to acquire class tokens of different categories to learn the global information of the feature map feature map, and output the class semantic token feature map; the channel token semantic-aware attention module is used to model local pixel-level details to compensate for feature information, and output the channel token feature map; the space token semantic-aware attention module is used to capture dual attention space semantic information to fuse the microscopic region information of the histological image, and output the space semantic token map;

[0010] The output feature map is input into a feed-forward network to obtain the predicted grade of colorectal cancer in the target area.

[0011] Optionally, segmenting the histological image into small image patches of a preset size and extracting a feature map comprises:

[0012] The histological image is segmented into small image patches of a preset size, and feature maps are extracted from the small image patches based on DenseNet161.

[0013] Optionally, acquiring class tokens of different categories to learn global information of the feature map includes:

[0014] calculating class tokens of different classifications for the feature map;

[0015] Using self-attention to learn attention weights for all categories, processing the feature map to obtain an attention weight map;

[0016] A class semantic token feature map of different classification information is obtained based on the attention weight map and the class token.

[0017] Optionally, calculating class tokens of different classifications for the feature map comprises:

[0018] Class tokens for different categories are calculated for the feature map based on the following formula:

[0019]

[0020] Where X is the feature map, N is the length of the feature map, i is the index of the feature map, and CToken(X) represents the class token of the feature map X.

[0021] Optionally, the calculation formula of the attention weight map is:

[0022]

[0023] Among them, Att(Q,K,V) is the attention weight map, Q, K and V represent query, key and value respectively, and d is the length of Q. Att(Q,K,V) represents the attention weight map generated based on query Q, key K and value V, and Softmax() represents the Softmax activation function.

[0024] Optionally, the class semantic token feature map for obtaining different classification information based on the attention weight map and the class token includes:

[0025] Adding the attention weight map to the class token to obtain an integrated feature map Z;

[0026] The integrated feature map Z is processed as follows to obtain a class semantic token:

[0027] ResMLP(Z)=Z+MLP(Norm(Z));

[0028] Among them, MLP represents multi-layer perceptron, Norm represents normalization, ResMLP represents residual multi-layer perceptron, and ResMLP(Z) represents class semantic token map.

[0029] Optionally, modeling local pixel-level details to compensate for feature information includes:

[0030] Feeding the feature map into a convolutional layer to obtain channel tokens;

[0031] The channel tokens are flattened in the channel dimension and passed into the self-attention mechanism to generate channel token weights, the channel token weights are concatenated with the feature map, and then passed into the feed-forward network to obtain the channel token attention map;

[0032] Multiplying the channel token attention map by the channel token and transmitting the result to the feedforward network to generate a channel token feature map;

[0033] The calculation formula of the channel token feature map is:

[0034] CHTSAA(X)=FFN{FFN{C[Att(CHToken(X),X)]}};

[0035] Among them, CHToken, Att and FFN represent channel token, self-attention and feed-forward network respectively, X is a feature map, CHToken(X) represents the channel token of feature map X, and CHTSAA(X) represents the channel token feature map.

[0036] Optionally, capturing dual attention spatial semantic information to fuse microscopic region information of the histological image includes:

[0037] Flattening the feature map and feeding it to a first linear layer to generate a first query, a first key, and a first value;

[0038] feeding the feature map to a second linear layer, generating a second query, and generating an attention weight for the feature map based on the second query and the first key and the first value;

[0039] The feature map with attention weight added is input into a third linear layer to generate a second key and a second value; based on the first query, the second key and the second value generate a spatial semantic token map.

[0040] Optionally, generating a spatial semantic token graph based on the first query, the second key and the second value comprises:

[0041] Transmitting a multiplication result of the first query, the second key and the second value to a forward transmission network to obtain a spatial semantic token graph;

[0042] The calculation formula of the spatial semantic token map is:

[0043] SPTSAA(X)=FFN(Att(X),Att T (X));

[0044] Among them, FFN and Att represent the feedforward network and attention mechanism respectively, and SPTSAA(X) represents the spatial semantic token map.

[0045] In a second aspect, an embodiment of the present invention provides a colorectal cancer grading device based on semantic token-aware attention, the device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, the processor executing the computer program running in a module of the following device:

[0046] A feature extraction module, used for acquiring a histological image of a target area, segmenting the histological image into small image patches of a preset size and extracting a feature map;

[0047] An attention generation module is used to input the feature map into a semantic token-aware attention network to obtain an output feature map with semantic token-aware attention, wherein the output feature map includes a class semantic token feature map, a channel token feature map and a space semantic token map; wherein the semantic token-aware attention network includes at least a class token semantic-aware attention module, a channel token semantic-aware attention module and a space token semantic-aware attention module, wherein the class token semantic-aware attention module is used to obtain class tokens of different categories to learn the global information of the feature map, and output the class semantic token feature map; the channel token semantic-aware attention module is used to model local pixel-level details to compensate for feature information, and output the channel token feature map; the space token semantic-aware attention module is used to capture dual attention space semantic information to fuse the microscopic area information of the histological image, and output the space semantic token map;

[0048] A graded prediction module is used to input the output feature map into a feedforward network to obtain a predicted grade of colorectal cancer in the target area.

[0049] The embodiments of the present invention include the following beneficial effects: the embodiments of the present invention acquire different categories of class tokens based on the class token semantic-aware attention module to learn the global information of the histological image; and model fine local pixel-level details based on the channel token semantic-aware attention module to compensate for the feature details of the entire image to establish the global dependency relationship between channels; and at the same time, capture the dual attention space semantic information based on the space token semantic-aware attention module to fuse the microscopic regional information of the histological image, thereby solving the restriction problem of large-scale application of the current deep learning-based colorectal cancer grading method in clinical practice. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0051] Figure 1 Examples of image patches for high-resolution digital histology images;

[0052] Figure 2 is a schematic diagram of the architecture of a semantic token-aware attention network proposed in some embodiments of the present invention;

[0053] Figure 3 is a flow chart of a method for colorectal cancer grading based on semantic token-aware attention provided by an embodiment of the present invention;

[0054] Figure 4 is a schematic diagram of the structure of a token-like semantic-aware attention module in some embodiments of the present invention;

[0055] Figure 5 is a schematic diagram of the structure of a channel token semantic-aware attention module in some embodiments of the present invention;

[0056] Figure 6 is a schematic diagram of the structure of a spatial token semantic-aware attention module in some embodiments of the present invention;

[0057] Figure 7a is the confusion matrix of the semantic token-aware attention network tested on the CRC dataset;

[0058] Figure 7b is the confusion matrix of the semantic token-aware attention network tested on the ECRC dataset;

[0059] Figure 8 is a schematic diagram of a colorectal cancer grading device based on semantic token-aware attention according to some embodiments of the present invention. DETAILED DESCRIPTION

[0060] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0061] It should be noted that, although the functional modules are divided in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart. The terms "first", "second", etc. in the specification, claims and the above drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence.

[0062] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0063] In addition, described feature, structure or characteristic can be combined in one or more embodiments in any suitable manner. In the following description, many specific details are provided to provide a full understanding of the embodiments of the present application. However, those skilled in the art will appreciate that the technical scheme of the present application can be put into practice without one or more of the specific details, or other methods, components, devices, steps, etc. can be adopted. In other cases, known methods, devices, realizations or operations are not shown or described in detail to avoid blurring the various aspects of the application.

[0064] The block diagrams shown in the accompanying drawings are only functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software form, or in one or more hardware charging modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0065] The flowcharts shown in the accompanying drawings are only exemplary and do not necessarily include all the contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps can be decomposed, and some operations / steps can be combined or partially combined, so the actual execution order may change according to actual conditions.

[0066] Colorectal cancer is one of the common malignant tumors caused by multiple factors. In clinical practice, doctors usually assess the grade of colorectal cancer based on the degree of tumor cell infiltration in the intestinal wall.

[0067] In clinical diagnosis, experienced pathologists grade colorectal cancer by analyzing the abnormalities of individual cancer cells and the spatial organization of aberrant glandular structures. Deep learning has also been applied to colorectal cancer grading because of its outstanding performance in signal detection, image recognition and other fields.

[0068] Although deep learning methods have performed well in the field of colorectal cancer grading, large-scale clinical practice remains challenging due to the following issues: Due to limited computing resources, these methods always tend to perform representation learning in small image blocks (image patches) in high-resolution digital histology images. Figure 1 is an example of an image patch of a high-resolution digital histology image. First, the class tokens generated by small image patches can only represent local information, but cannot represent the information of the labeled histology image, such as Figure 1 As shown in A; secondly, image patch segmentation always leads to distortion of the details of the marked features in the entire image (such as Figure 1B), which reduces the ability of local modeling semantics, further leading to the weakening of pixel-level channel dependence in colorectal cancer cells; thirdly, the segmentation strategy that generates small image patches often cannot contain the entire tissue microstructure, resulting in the loss of important data information and areas of interest, such as Figure 1 As shown in B.

[0069] In response to the defects of the above-mentioned existing deep learning-based colorectal cancer grading methods, an embodiment of the present invention proposes a semantic token-aware attention network, which includes at least a class token semantic-aware attention module, a channel token semantic-aware attention module and a space token semantic-aware attention module. The class token semantic-aware attention module is used to obtain class tokens of different categories to learn the global information of the feature map, the channel token semantic-aware attention module is used to model local pixel-level details to compensate for feature information, and the space token semantic-aware attention module is used to capture dual attention space semantic information to fuse microscopic regional information of histological images. In some embodiments of the present invention, the network structure of the semantic token-aware attention network is as follows: Figure 2 shown.

[0070] The present invention also proposes a colorectal cancer grading method based on the perceptual semantic token perceptual attention network. Figure 3 As shown, the method comprises the following steps:

[0071] S310, acquiring a histological image of a target area, segmenting the histological image into small image patches of a preset size and extracting a feature map.

[0072] Specifically, a histological image of the target area is first obtained. In one embodiment of the present invention, the resolution of the histological image is 1792×1792 pixels, and then the histological image is segmented into image patches with a resolution of 224×224. The final image patches are fed into a backbone network based on DenseNet161 to extract features, and the resolution of the generated feature map is 7x7.

[0073] S320, inputting the feature map into a semantic token-aware attention network to obtain a feature map with semantic token-aware attention, wherein the semantic token-aware attention network includes at least a class token semantic-aware attention module, a channel token semantic-aware attention module and a space token semantic-aware attention module, the class token semantic-aware attention module is used to obtain class tokens of different categories to learn the global information of the feature map, the channel token semantic-aware attention module is used to model local pixel-level details to compensate for feature information, and the space token semantic-aware attention module is used to capture dual attention space semantic information to fuse the microscopic area information of the histological image.

[0074] Digital histology images are suitable for analysis by convolutional neural networks. Due to limited computational resources, convolutional neural networks are usually used for representation learning based on image patches (e.g., 224×224) of digital histology images. However, the class tokens generated by the image patches can only represent local information, resulting in the inability of the image patches to represent the label information of the histology images. To solve this problem, the present invention designs a class token semantic-aware attention module for obtaining class tokens of different classification information, in which self-attention is used to learn the attention weight categories of all categories, defined as:

[0075] CLTSAA(X)=ResMLP{CToken(X)+Att[Norm(X)]};

[0076] Among them, CLTSAA(X) represents the class token semantic-aware attention module acting on the feature map X, X represents the input feature map, ResMLP, CToken, Att, Norm and MLP represent residual multi-layer perceptron, class labeling, self-attention, normalization and multi-layer perceptron respectively.

[0077] Specific:

[0078] First, the class token of the feature map is calculated by the following formula:

[0079]

[0080] Where X is the feature map, N is the length of the feature map, i is the index of the feature map, and CToken(X) represents the class token of the feature map X.

[0081] The feature map X is normalized and input into the self-attention to obtain the attention weight map, defined as:

[0082]

[0083] Among them, Att(Q,K,V) is the attention weight map, Q, K and V represent query, key and value, d is the length of Q, and Softmax() represents the Softmax activation function.

[0084] Finally, the attention weight map Att(Q, K, V) is added to the class token CToken(X) to obtain an integrated feature map Z;

[0085] The integrated feature map Z is processed as follows to obtain a semantic token-like feature map:

[0086] ResMLP(Z)=Z+MLP(Norm(Z));

[0087] Among them, MLP, Norm and ResMLP represent multi-layer perceptron, normalized and residual multi-layer perceptron respectively, and ResMLP(Z) represents the semantic token-like feature map obtained by applying the residual multi-layer perceptron to the input feature map Z.

[0088] In some embodiments of the present invention, the structure of the token-like semantic-aware attention module is as follows: Figure 4 .

[0089] Image segmentation always leads to distortion of the details of labeled features in the whole image, which reduces the ability of local semantic modeling and further leads to weakened pixel-level channel dependencies in colorectal cancer cells. To solve this problem, the present invention designs a channel token semantic-aware attention module, which aims to learn fine local pixel-level details to establish global dependencies between channels, defined as:

[0090] CHTSAA(X)=FFN{FFN{C[Att(CHToken(X),X)]}};

[0091] Among them, CHToken, Att and FFN represent channel token, self-attention and feedforward network respectively, X is a feature map, CHToken(X) means obtaining the channel token of X, and CHTSAA(X) means that the channel token semantic-aware attention module acts on the input feature map X to generate a channel token feature map.

[0092] Specific:

[0093] First, the input feature map is fed into a convolutional layer to obtain channel tokens;

[0094] Second, channel tokens are flattened in the channel dimension and transformed into a self-attention mechanism to produce channel token weights;

[0095] Third, the channel token weights are concatenated with the feature map and then fed into the forward transmission network after softmax to obtain the channel token attention map;

[0096] Fourth, the channel token attention map is multiplied with the channel token and transmitted to the final feed-forward network to generate the channel token feature map.

[0097] In some embodiments of the present invention, the structure of the channel token semantic-aware attention module is as follows: Figure 5 .

[0098] The common patch segmentation strategy is to segment into small image blocks, which usually fails to integrate the microstructure of the entire tissue and loses important data information and areas of interest. In order to obtain effective spatial token information, the present invention develops a spatial token semantic-aware attention module to capture the dual attention weight feature map from the input. The calculation formula of the spatial semantic token map is:

[0099] SPTSAA(X)=FFN(Att(X),Att T (X));

[0100] Among them, FFN and Att represent the feedforward network and attention mechanism respectively, the superscript T of Att represents the transpose of the matrix; SPTSAA(X) represents the spatial semantic token map obtained by the spatial token semantic-aware attention module acting on the input feature map X.

[0101] Specific:

[0102] First, the input feature map is fed into two linear layers to generate six feature spaces Q s , K s 、V s , Q r , K r and V r .

[0103] Second, query Q r Multiply by the softmax transpose K s , then multiply by V s To generate attention weights.

[0104] Third, the attention weight map features with the input added will be fed into the linear layer to generate another transposed key V r Value Q r .

[0105] Fourth, Q s , Q r , and V r The multiplication result is transmitted to the forward transmission network to obtain the spatial semantic tag.

[0106] In some embodiments of the present invention, the structure of the spatial token semantic-aware attention module is as follows: Figure 6 .

[0107] S330, inputting the output feature map into a feed-forward network to obtain the predicted grade of colorectal cancer in the target area.

[0108] Specifically, the class semantic token feature map, the channel token feature map and the space semantic token map are spliced ​​and input into the feedforward network to obtain the colorectal cancer prediction grade of the target area.

[0109] One embodiment of the present invention is Figure 3 The method described in S310-S330 (abbreviated as TSAA-Net) was compared with nine state-of-the-art colorectal cancer grading methods on two public datasets (CRC dataset and ECRC dataset). The CRC dataset is a colorectal cancer dataset, and the ECRC dataset is an extended colorectal cancer dataset. The comparison results are as follows:

[0110] Table 1: Results of different methods on CRC and ECRC datasets

[0111]

[0112]

[0113] Comparison on the CRC dataset:

[0114] As can be seen from Table 1, the semantic token-aware attention network proposed in the present invention shows the best performance on the CRC dataset, achieving 99.28%, 99.05%, 99.54% and 99.28% accuracy, precision, recall and 1-score respectively, proving the superiority of the semantic token-aware attention network for cancer color grading. In addition, it can be seen from Table 1 that the proposed semantic token-aware attention network outperforms the classic deep learning models, namely ResNet50, MobileNet, InceptionV3 and Xception. This shows that the mechanism-based module can further help the network extract effective features from the original input image. In addition, some CNN methods combined with traditional machine learning, such as CNN-SVM, CNN-LR and CNN-LSTM, are also used as basic comparisons, and the semantic token-aware attention network proposed in the present invention can also achieve better performance, indicating that it can focus more feature details for colorectal cancer grading. Finally, the graph-based method CGC-Net and the attention-based method CAC-Net were tested on the dataset. We can find that the proposed semantic token-aware attention network is more accurate than these two methods, indicating that it is more suitable for clinical diagnosis.

[0115] Comparison on the ECRC dataset:

[0116] From Table 1, it can be observed that the semantic token-aware attention network performs better than other methods in color cancer grading. Specifically, compared with the second-best CAC-Net method, the proposed semantic token-aware attention network can achieve significant improvements of 10.00%, 10.83%, 12.00%, and 11.43% in accuracy, precision, recall, and 1-score, respectively, demonstrating the effectiveness of the proposed semantic token-aware attention network for colorectal cancer grading. In addition, the semantic token-aware attention network outperforms other CNN-based methods in colorectal cancer grading, providing more confidence for pathologists to use the proposed semantic token-aware attention network for clinical diagnosis.

[0117] Analysis of the confusion matrix:

[0118] Figure 7a is the confusion matrix of the semantic token-aware attention network tested on the CRC dataset, Figure 7b is the confusion matrix of the semantic token-aware attention network tested on the ECRC dataset, where all test results of cross-validation are statistically analyzed. Figure 7a and Figure 7b It can be clearly observed that the false negative and false positive results of other categories are relatively small, which indicates that the proposed semantic token-aware attention network can serve as an effective tool for cancer grading in colorectal cancer.

[0119] Figure 8 A colorectal cancer grading device based on semantic token-aware attention is shown according to some embodiments of the present invention. Figure 8 As shown, the colorectal cancer grading device 800 based on semantic token-aware attention includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to run in the modules of the following devices:

[0120] A feature extraction module 810 is used to obtain a histological image of a target area, segment the histological image into small image patches of a preset size and extract a feature map;

[0121] an attention generation module 820, configured to input the feature map into a semantic token-aware attention network to obtain a feature map with semantic token-aware attention, wherein the semantic token-aware attention network at least includes a class token semantic-aware attention module, a channel token semantic-aware attention module, and a space token semantic-aware attention module, wherein the class token semantic-aware attention module is configured to obtain class tokens of different categories to learn the global information of the feature map, the channel token semantic-aware attention module is configured to model local pixel-level details to compensate for feature information, and the space token semantic-aware attention module is configured to capture dual attention space semantic information to fuse microscopic region information of the histological image;

[0122] The grade prediction module 830 is used to input the output feature map into a feed-forward network to obtain the predicted grade of colorectal cancer in the target area.

[0123] In summary, the semantic token-aware attention-based colorectal cancer grading method and device provided by the embodiments of the present invention learn the global information of the histological image by acquiring different categories of class tokens based on the class token semantic-aware attention module; and modeling fine local pixel-level details based on the channel token semantic-aware attention module to compensate for the feature details of the entire image to establish a global dependency relationship between channels; at the same time, the spatial token semantic-aware attention module captures the dual attention spatial semantic information to fuse the microscopic regional information of the histological image, thereby solving the current restriction problem of large-scale application of colorectal cancer grading methods based on deep learning in clinical practice.

[0124] The device embodiments described above are merely illustrative, and the units described as separate components may or may not be physically separated, that is, they may be located in one place or distributed on multiple network units. Some or all of the charging modules may be selected according to actual needs to achieve the purpose of the present embodiment.

[0125] Those skilled in the art will appreciate that all or some of the steps in the methods disclosed above, and the functional charging modules / units in the systems and devices may be implemented as software, firmware, hardware, or appropriate combinations thereof.

[0126] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0127] It should be understood that in the present application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the objects associated before and after are in an "or" relationship. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0128] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0129] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0130] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0131] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including multiple instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (Read-Only Memory, referred to as ROM), random access memory (Random Access Memory, referred to as RAM), disk or optical disk and other media that can store programs.

[0132] The preferred embodiments of the present invention are described above with reference to the accompanying drawings, but the scope of the rights of the present invention is not limited thereto. Any modification, equivalent substitution and improvement made by a person skilled in the art without departing from the scope and essence of the present invention should be within the scope of the rights of the present invention.

Claims

1. A colorectal cancer grading method based on semantic token-aware attention, characterized in that: include: Acquire a histological image of a target area, segment the histological image into small image patches of a preset size and extract a feature map; The feature map is input into a semantic token-aware attention network to obtain an output feature map with semantic token-aware attention, wherein the output feature map includes a class semantic token feature map, a channel token feature map and a space semantic token map; wherein the semantic token-aware attention network includes at least a class token semantic-aware attention module, a channel token semantic-aware attention module and a space token semantic-aware attention module, wherein the class token semantic-aware attention module is used to acquire class tokens of different categories to learn the global information of the feature map feature map, and output the class semantic token feature map; the channel token semantic-aware attention module is used to model local pixel-level details to compensate for feature information, and output the channel token feature map; the space token semantic-aware attention module is used to capture dual attention space semantic information to fuse the microscopic region information of the histological image, and output the space semantic token map; The output feature map is input into a feed-forward network to obtain the predicted grade of colorectal cancer in the target area.

2. The method according to claim 1, characterized in that: The step of segmenting the histological image into small image patches of a preset size and extracting a feature map comprises: The histological image is segmented into small image patches of a preset size, and feature maps are extracted from the small image patches based on DenseNet161.

3. The method according to claim 1, characterized in that: The acquiring of class tokens of different categories to learn the global information of the feature map includes: calculating class tokens of different classifications for the feature map; Using self-attention to learn attention weights for all categories, processing the feature map to obtain an attention weight map; A class semantic token feature map of different classification information is obtained based on the attention weight map and the class token.

4. The method according to claim 3, characterized in that: The calculating of class tokens of different classifications for the feature map comprises: Class tokens for different categories are calculated for the feature map based on the following formula: Where X is the feature map, N is the length of the feature map, i is the index of the feature map, and CToken(X) represents the class token of the feature map X.

5. The method according to claim 3, characterized in that: The calculation formula of the attention weight map is: Among them, Att(Q,K,V) is the attention weight map, Q, K and V represent query, key and value respectively, d is the length of query Q, and Softmax() represents the Softmax activation function.

6. The method according to claim 3, characterized in that: The class semantic token feature map based on the attention weight map and the class token to obtain different classification information includes: Adding the attention weight map to the class token to obtain an integrated feature map Z; The integrated feature map Z is processed as follows to obtain a class semantic token: ResMLP(Z)=Z+MLP(Norm(Z)); Among them, MLP represents multi-layer perceptron, Norm represents normalization, ResMLP represents residual multi-layer perceptron, and ResMLP(Z) represents class semantic token map.

7. The method according to claim 1, characterized in that: The modeling of local pixel-level details to compensate for feature information includes: Feeding the feature map into a convolutional layer to obtain channel tokens; The channel tokens are flattened in the channel dimension and passed into the self-attention mechanism to generate channel token weights, the channel token weights are concatenated with the feature map, and then passed into the feed-forward network to obtain the channel token attention map; Multiplying the channel token attention map by the channel token and transmitting the result to the feedforward network to generate a channel token feature map; The calculation formula of the channel token feature map is: CHTSAA(X)=FFN{FFN{C[Att(CHToken(X),X)]}}; Among them, CHToken, Att and FFN represent channel token, self-attention and feed-forward network respectively, X is a feature map, CHToken(X) represents the channel token of feature map X, and CHTSAA(X) represents the channel token feature map.

8. The method according to claim 1, characterized in that: The method of capturing dual attention spatial semantic information to fuse the microscopic region information of the histological image includes: Flattening the feature map and feeding it to a first linear layer to generate a first query, a first key, and a first value; feeding the feature map to a second linear layer, generating a second query, and generating an attention weight for the feature map based on the second query and the first key and the first value; The feature map with attention weight added is input into a third linear layer to generate a second key and a second value; based on the first query, the second key and the second value generate a spatial semantic token map.

9. The method according to claim 8, characterized in that: Generating a spatial semantic token graph based on the first query, the second key, and the second value comprises: Transmitting a multiplication result of the first query, the second key and the second value to a forward transmission network to obtain a spatial semantic token graph; The calculation formula of the spatial semantic token map is: SPTSAA(X)=FFN(At(X),At T (X)); Among them, FFN and Att represent the feedforward network and attention mechanism respectively, and SPTSAA(X) represents the spatial semantic token map.

10. A colorectal cancer grading device based on semantic token-aware attention, characterized in that: The device comprises: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to run in the following modules of the device: A feature extraction module, used for acquiring a histological image of a target area, segmenting the histological image into small image patches of a preset size and extracting a feature map; An attention generation module is used to input the feature map into a semantic token-aware attention network to obtain an output feature map with semantic token-aware attention, wherein the output feature map includes a class semantic token feature map, a channel token feature map and a space semantic token map; wherein the semantic token-aware attention network includes at least a class token semantic-aware attention module, a channel token semantic-aware attention module and a space token semantic-aware attention module, wherein the class token semantic-aware attention module is used to obtain class tokens of different categories to learn the global information of the feature map, and output the class semantic token feature map; the channel token semantic-aware attention module is used to model local pixel-level details to compensate for feature information, and output the channel token feature map; the space token semantic-aware attention module is used to capture dual attention space semantic information to fuse the microscopic area information of the histological image, and output the space semantic token map; A graded prediction module is used to input the output feature map into a feedforward network to obtain a predicted grade of colorectal cancer in the target area.