High-resolution medical pathology WSI image analysis method based on scanning invariance Mama

Through the SMC-MIL network model, the scan-invariant packet features are generated using sequence contrast enhancement and Mamba aggregation module, combined with the consistency loss function, the problem of Mamba model being sensitive to scanning mode is solved, and the accuracy of medical pathological WSI image analysis is improved.

CN120070283AActive Publication Date: 2025-05-30CHONGQING UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510151871.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-12
Publication Date
2025-05-30
Estimated Expiration
2045-02-12

AI Technical Summary

Technical Problem

In the prior art, the Mamba model is very sensitive to scanning mode, resulting in its performance in medical pathological WSI image analysis and is not compatible with the MIL model with constant scanning.

Method used

A SMC-MIL network model is proposed. Through the sequence comparison enhancement module and the Mamba aggregation module, a scan-invariant packet characteristics are generated, combined with the consistency loss function, and network parameters are optimized to reduce the impact of scanning mode on Mamba performance.

Benefits of technology

It effectively reduces the impact of scanning mode on Mamba performance, improves the accuracy of high-resolution medical pathological WSI image analysis, making the Mamba model more suitable for MIL tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070283A_ABST
    Figure CN120070283A_ABST
Patent Text Reader

Abstract

The invention relates to a high-resolution medical pathology WSI image analysis method based on scanning invariance Mama, comprising the following steps: selecting a public database, pictures in the database needing to contain tissue information; constructing an SMC-MIL network model W, wherein the W comprises three parts of modules; and selecting a picture q from the database, inputting the picture q into the W, enabling the q to pass through the three modules in sequence, and finally obtaining a prediction packet label of the q. In actual detection, the accuracy of a picture analysis result can be improved according to the prediction packet label.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer image analysis, and particularly to a high-resolution medical pathology WSI image analysis method based on scan-invariant Mamba. Background Art

[0002] Histopathological image analysis is a key area of modern medicine, usually performed using whole slide images (WSIs). Due to the high resolution of WSIs (up to billions of pixels) and the lack of pixel-level annotations, traditional WSI analysis methods face unique challenges. In the most widely used MIL paradigm, instances are converted to offline features using a pre-trained model and then aggregated into bag-level representations for subsequent analysis; in this paradigm, WSI classification can be regarded as a long sequence modeling problem, where a large number of instances in a bag are treated as tokens, and the goal is to model the correlation between instances and the overall spatial context within the whole bag to capture discriminative information; some models based on the state space model (SSM) are designed to effectively model long sequences. Mamba is a recently proposed representative SSM-based model, which has better long sequence modeling performance compared to ViT but requires less computational resources; however, there is a fundamental difference between the Mamba model and the MIL task: the output of Mamba is sensitive to different scanning patterns (scan-sensitive), while in MIL, the model can produce invariant results regardless of the scanning pattern (scan-invariant).

[0003] To reconcile the differences between scan-sensitive Mamba and scan-invariant MIL, previous work mainly involved inputting sequences of image patches with different scanning patterns into Mamba and then aggregating multiple output sequences. For example, MamMIL combines a bi-SSM and a 2D-CAB module to integrate scanning pattern information from the original sequence, reverse sequence, and local sequence; MambaMIL reshapes the sequence into a rectangle and applies a sequence reordering operation to vertically scan the new sequence; however, the number of scanning patterns used to generate instance sequences is limited, restricting the model's ability to learn scan-pattern invariant features. Therefore, this does not fundamentally mitigate the impact of scanning patterns on the performance of Mamba. Summary of the Invention

[0004] Aiming at the above problems existing in the prior art, the technical problem to be solved by the present invention is: how to more fundamentally solve the impact of scanning patterns on the performance of Mamba, reconcile the differences between scan-sensitive Mamba and scan-invariant MIL, and thus improve the accuracy of high-resolution medical image pathology WSI image analysis.

[0005] To solve the above technical problems, the present invention adopts the following technical solutions: A high-resolution medical pathological WSI image analysis method based on scan-invariant Mamba, comprising the following steps:

[0006] S100: Select a public picture database Q. All pictures in Q contain tissue information, and each picture has a known label;

[0007] S200: Construct an SMC-MIL network model W, which includes a data offline preprocessing module, a sequence contrast enhancement module, and a Mamba aggregation module;

[0008] S300: Arbitrarily select a picture q from Q, regard q as a bag, and input q into the data offline preprocessing module to obtain a feature original scan sequence Z;

[0009] S400: Input Z into the sequence contrast enhancement module, and perform sequence enhancement branch operations on Z to obtain a new sequence Z easy and perform sequence masking branch operations on Z to obtain a new sequence Z hard :

[0010] Sequence enhancement branch operation: Pass Z through an attention mechanism to identify instances with high attention values, obtain a set of instance attention scores A, and then process A through the Mixup method and add it to Z to obtain a new sequence Z easy ;

[0011] Sequence masking branch operation: Pass Z through an attention masking mechanism to mask instances with low attention values to obtain a new masked sequence Z hard ;

[0012] S500: Pass Z easy and Z hard through the Mamba aggregation module respectively, and correspondingly obtain a bag feature b 1 and a bag feature b 2 , and the expressions are as follows:

[0013] b 1 = A 1 ·Ψ(Z easy ) = φ(Ψ(S A (Z)))·Ψ(S A (Z))

[0014] b 2 = A 2 ·Ψ(Z hard ) = φ(Ψ(S M (Z)))·Ψ(S M (Z))

[0015] where A1 and A 2 are the attention scores of Ψ(Z easy ) and Ψ(Z hard ) respectively; S A represents sequence enhancement, and S M represents sequence masking; Ψ(·) represents Mamba, and φ(·) represents the attention mechanism;

[0016] S600: Calculate the predicted packet labels corresponding to Z easy and Z hard respectively. The specific expressions are as follows: and The specific expressions are as follows:

[0017]

[0018]

[0019] Among them, represents the classifier;

[0020] S700: Calculate the network parameter optimization formula of the SMC-MIL model, which is expressed as follows:

[0021]

[0022] Among them, θ represents the network parameters of the SMC-MIL network model, γ represents the parameter for balancing the classification loss and the consistency loss, represents the overall loss function of W, represents the consistency loss of the packet features, represents the consistency loss of the packet logits / packet classification;

[0023] The classification loss of the sequence enhancement branch is expressed as follows:

[0024]

[0025]

[0026] Among them, Y represents the true value of the packet; the classification loss is the conventional cross-entropy loss, which belongs to the existing technology and aims to constrain the classification result of the model to make it closer to the true value; while the consistency loss is that the packet features and classification results input from different sequences to the same Mamba output can be consistent, so as to alleviate the difference between the sequence-related Mamba model and the sequence-independent multi-instance learning task.

[0027] S800: Set the maximum number of training times, execute steps S300 - S700 for all the pictures in Q, and according to Update the network parameter θ, stop training when the maximum number of training times is reached, and obtain the trained network model W'.

[0028] S900: Select an unknown picture containing tissue information and input it into W' to obtain the predicted packet label of the picture.

[0029] Preferably, the specific steps for obtaining the original feature scan sequence Z in S300 are as follows:

[0030] S310: Cut q into several tiles, discard the tiles that only contain the background area. Generally, when the information saturation of a tile or patch is less than 15, it is considered that the information contained in the tile or patch is the background area. Combine the remaining tiles into a set X, and each tile represents an instance x i , specifically expressed as follows:

[0031]

[0032] where i represents the i-th tile and I represents the total number of tiles;

[0033] S320: Use a feature extractor to extract features from all tiles in X, and then perform a linear projection on the extracted offline features to obtain the original feature scan sequence where z i represents the offline feature of the i-th instance; this feature extractor can be Resnet-50 or other feature extractors such as UNI, which belongs to the prior art.

[0034] Preferably, the operation of the sequence enhancement branch in S400 is as follows:

[0035] S410: Use an attention mechanism to calculate the attention score of each instance in Z, and define it as follows:

[0036] A = [a 1 , …, a i , …, a N = φ(Z)

[0037] where φ(·) represents the attention mechanism;

[0038] S411: Sort A in S410, and the expression is as follows:

[0039] I = [j 1 , j 2 , …, j N = Sort(A)

[0040] where j 1 represents the index of the instance with the highest attention score, and j N represents the index of the instance with the lowest score;

[0041] S412: Select the top α% instances with high attention scores in I as the amplification items of the Z sequence, and finally obtain the new sequence Z easy , and the expression is as follows:

[0042]

[0043]

[0044] where, is the Mixup on the selected instance, S A (·) represents the sequence enhancement module, represents the index of the amplification item; the attention mechanism here shares parameters with the subsequent aggregation module. Therefore, in the Figure 1 "easy" branch shown, the sequence length of Mamba input is

[0045] Preferably, the sequence masking branch operation in S400 is specifically as follows:

[0046] S420: Use the attention masking mechanism to generate an n-dimensional mask vector M from Z as shown below:

[0047] M = [m 1 , …, m r , …, m N ,

[0048] where, m r ∈{0, 1}. If m r = 1, the r-instance is unmasked; otherwise, it is masked;

[0049] S421: Set the mask rates of high attention values and low attention values to β h % and β l % respectively. Therefore, the sequence subscripts corresponding to high attention values are and the sequence subscripts corresponding to low attention values are Therefore,

[0050]

[0051] Finally, the new masked sequence Z hard is expressed as:

[0052] Z hard ←S M (Z): = M · Z,

[0053] where, S M (·) represents sequence masking; during the experiment, let β h = β l .

[0054] Compared with the prior art, the present invention has at least the following advantages:

[0055] 1. By improving the existing Mamba model, the present invention changes the "ordered" restriction requirement to an "unordered requirement" by increasing the highly concerned instances in the comparison sequence and masking the less concerned instances, thereby improving the adaptability of the improved model, more essentially solving the influence of the scanning pattern on the performance of Mamba, reconciling the differences between the scanning-sensitive Mamba and the scanning-invariant MIL, and thus improving the accuracy of pathological WSI image analysis in high-resolution medical images.

[0056] 2. The present invention proposes a new Mamba-based MIL framework named SMC-MIL for computational pathology; SMC-MIL teaches Mamba to learn scanning-invariant bag representations through instance sequence contrast learning, which is more suitable for the MIL-based WSI classification paradigm.

[0057] 3. A dual-branch instance sequence contrast learning framework is designed in the model of the present invention, aiming to achieve the consistency between two extremely different input instance sequences and eliminate the influence of sequence arrangement on the final bag decision; this framework provides a simple and effective solution for using SSM to process MIL tasks.

[0058] 4. A sequence contrast enhancement method is proposed in the model of the present invention, which deliberately shuffles, enhances, and masks instances based on the evaluation of the difficulty of instance classification; this method not only increases the differences in order and length between the two input comparison sequences, but also expands the difficulty differences in the recognition discriminant region; obviously, this method increases the difficulty of contrast learning, thereby forcing Mamba to learn the discriminant features of WSI without considering the arrangement of its patch sequences. Brief Description of the Drawings

[0059] Figure 1 It is a schematic overview diagram of the SMC-MIL model of the present invention.

[0060] Figure 2 It is a comparison between different variants of various Mambas.

[0061] Figure 3 It is a visual display of the patches made by the original Mamba (baseline) and SMC-MIL. Detailed Description of the Invention

[0062] The present invention will be further described in detail below.

[0063] The present invention proposes a high-resolution medical pathological WSI image analysis method based on scan-invariant Mamba. This method is a solution based on Mamba and scan-invariant MIL. First, based on the Mamba model and the MIL model, starting from two aspects of increasing instances of highly concerned sequences and masking instances of low-concerned virtual columns respectively, it can more essentially solve the performance impact of the scanning mode on Mamba, so as to reconcile the differences between scan-sensitive Mamba and scan-invariant MIL, thereby improving the accuracy of pathological WSI image analysis in high-resolution medical images.

[0064] See Figures 1 - 3 , a high-resolution medical pathological WSI image analysis method based on scan-invariant Mamba, comprising the following steps:

[0065] S100: Select a public image database Q, all images in Q contain tissue information, and each image has a known label;

[0066] S200: Construct an SMC-MIL network model W, which includes three parts: a data offline preprocessing module, a sequence contrast enhancement module, and a Mamba aggregation module;

[0067] S300: Select an arbitrary image q from Q, regard q as a bag, input q into the data offline preprocessing module, and obtain a feature original scan sequence Z;

[0068] The specific steps to obtain the feature original scan sequence Z in the S300 are as follows:

[0069] S310: Cut q into several tiles, discard the tiles that only contain background regions. Generally, when the information saturation of a tile or patch is less than 15, it is considered that the information contained in the tile or patch is the background region. Combine the remaining tiles into a set X, and each tile represents an instance x i , specifically expressed as follows:

[0070]

[0071] where i represents the i-th tile and I represents the total number of tiles;

[0072] S320: Use a feature extractor to extract features from all tiles in X, and then perform a linear projection on the extracted offline features to obtain a feature original scan sequence where z i represents the offline feature of the i-th instance; this feature extractor can be other feature extractors such as Resnet-50 or UNI, which belongs to the prior art.

[0073] S400: Input Z into the sequence contrast enhancement module, and perform sequence enhancement branch operations on Z respectively to obtain a new sequence Z easy and perform sequence masking branch operations to obtain a new sequence Z hard :

[0074] Sequence enhancement branch operation: Pass Z through the attention mechanism to identify instances with high attention values, obtain the instance attention score set A, and then process A through the Mixup method and add it to Z to obtain a new sequence Z easy ;

[0075] The operation of the sequence enhancement branch in S400 is as follows:

[0076] S410: Calculate the attention scores of each instance in Z using the attention mechanism and define as follows:

[0077] A = [a 1 , …, a i , …, a N = φ(Z)

[0078] where φ(·) represents the attention mechanism;

[0079] S411: Sort A in S410, and the expression is as follows:

[0080] I = [j 1 , j 2 , …, j N = Sort(A)

[0081] where j 1 represents the index of the instance with the highest attention score, and j N represents the index of the instance with the lowest score;

[0082] S412: Select the top α% of the instances with high attention scores in I as the amplification items of the Z sequence, and finally obtain a new sequence Z easy , and the expression is as follows:

[0083]

[0084]

[0085] where is the Mixup on the selected instances, and S A (·) represents the sequence enhancement module, represents the index of the amplification item; the attention mechanism here shares parameters with the subsequent aggregation module. Therefore, in the "easy" branch shown in Figure 1 , the sequence length input to the Mamba aggregation module is

[0086] Sequence mask branch operation: Pass Z through an attention masking mechanism to mask instances with low attention values, obtaining a new masked sequence Z hard ;

[0087] Preferably, the sequence mask branch operation in S400 is specifically as follows:

[0088] S420: Use the attention masking mechanism to generate an n-dimensional mask vector M from Z as follows:

[0089] M = [m 1 ,…,m r ,…,m N ,

[0090] where, m r ∈{0,1}, if m r = 1, then the r-instance is unmasked, otherwise it is masked;

[0091] S421: Set the mask rates for high and low attention values to β h % and β l %, respectively. Therefore, the sequence subscripts corresponding to high attention values are The sequence subscripts corresponding to low attention values are Therefore,

[0092]

[0093] Finally, the new masked sequence Z hard is expressed as:

[0094] Z hard ←S M (Z): = M·Z,

[0095] where, S M (·) represents sequence masking; during the experiment, let β h = β l .

[0096] S500: Pass Z easy and Z hard through the Mamba aggregation module respectively, and correspondingly obtain packet features b 1 and packet feature b 2 , and the expression is as follows:

[0097] b 1 = A 1 ·Ψ(Z easy ) = φ(Ψ(S A (Z)))·Ψ(S A (Z))

[0098] b 2 = A 2 ·Ψ(Z hard ) = φ(Ψ(S M (Z)))·Ψ(S M (Z))

[0099] Among them, A 1 and A 2 are the attention scores of Ψ(Z easy ) and Ψ(Z hard ) respectively; S A represents sequence enhancement, and S M represents sequence masking; Ψ(·) represents Mamba, and φ(·) represents the attention mechanism;

[0100] S600: Calculate the predicted packet labels corresponding to Z easy and Z hard respectively. The specific expressions are as follows: and Specifically, the expressions are as follows:

[0101]

[0102]

[0103] Among them, represents the classifier;

[0104] S700: Calculate the network parameter optimization formula of the SMC - MIL model, which is expressed as follows:

[0105]

[0106] Among them, θ represents the network parameters of the SMC - MIL network model, γ represents the parameter for balancing the classification loss and the consistency loss, represents the overall loss function of W, represents the consistency loss of the packet features, represents the consistency loss of the packet logits / packet classification;

[0107] represents the classification loss of the sequence enhancement branch, and the expression is as follows:

[0108]

[0109] Among them, Y represents the true value of the package; the classification loss is the conventional cross-entropy loss, which belongs to the existing technology and aims to constrain the classification result of the model to make it closer to the true value; while the consistency loss is that the packages of different sequences can maintain consistency when input into the package features and classification results output by the same Mamba, so as to alleviate the difference between the sequence-related Mamba model and the sequence-independent multi-instance learning task.

[0110] S800: Set the maximum number of training times, execute steps S300 - S700 for all pictures in Q, and according to Update the network parameter θ, and stop training when the training reaches the maximum number of times to obtain the trained network model W'.

[0111] S900: Select unknown pictures containing tissue information and input them into W' to obtain the predicted package labels of these pictures.

[0112] Experimental design and content

[0113] Experimental dataset

[0114] The present invention evaluates the SMC-MIL model through three WSI analysis tasks: cancer diagnosis, subtype classification, and survival prediction. For the cancer diagnosis task, the present invention uses the CAMELYON dataset, which is composed of the CAMELYON-16 and CAMELYON-17 datasets. For the typing task, the present invention uses the TCGA-NSCLC, TCGA-BRCA, and BRACS datasets. To evaluate the accuracy of survival prediction, the present invention uses TCGA-LUAD, TCGA-LUSC, and TCGA-BLCA to evaluate the performance of the survival prediction task.

[0115] Experimental settings

[0116] For diagnosis and typing, the present invention uses Accuracy, AUC, and F1-score to evaluate the performance of the model. For survival prediction, the present invention reports the C-index of all datasets. By default, the present invention adopts 5-fold cross-validation, which is consistent with the baseline method for performance comparison. Specifically, for the BRACS dataset, the present invention also follows the partitioning scheme from, and divides the dataset into a training set, a validation set, and a test set of 395:65:87 in the official partition. The present invention conducts three independent repeated experiments to minimize the random variation of the official division. In addition, to maintain consistency with other datasets in the experiment, the present invention also conducts 5-fold cross-validation experiments on the 7-class classification task of BRACS, which is marked as ★ in Table \ref{tab:main_comp2}.

[0117] Implementation details

[0118] Following the work of predecessors, the present invention uses ResNet50 pre-trained on ImageNet-1k and the latest base model UNI as an offline feature extractor. UNI is pre-trained on a large internal histological dataset of 1×10 6 WSIs of 1×10 8 patches.

[0119] Experimental results

[0120] Cancer diagnosis and subtyping

[0121] Table 1 Comparison of cancer diagnosis performance on the CAMELYON dataset and subtype classification performance on the TCGA-NSCLC dataset

[0122]

[0123] Table 1 presents the cancer diagnosis and subtyping performance of various MIL methods on the CAMELYON and NSCLC datasets. These results show that the method of the present invention achieves the best performance under all metrics of all benchmark tests. Specifically, on the CAMELYON dataset, the method of the present invention improves by 0.55%, 1.26%, and 0.28% respectively in terms of Accuracy, AUC, and F1-score compared to the second-best method. These figures are 0.57%, 0.4%, and 0.5% respectively in the NSCLC dataset. It is also worth noting that compared with SRMamba, which is also based on the Mamba architecture, the performance of the proposed SMC-MIL has been significantly improved. The AUC performance of SMC-MIL is 1.99% and 1.58% higher than that of SRMamba on the CAMELYON dataset and the NSCLC dataset respectively. This verifies the hypothesis of the present invention that eliminating the dependence of Mamba on the scanning pattern can better adapt to the MIL task.

[0124] Table 2 BRACS-7 ★ and subtyping results of TCGA-BRCA.

[0125]

[0126] Table 2 presents the comparison of subtype classification performance on the BRACS dataset (a 7-class classification task is selected in this experiment) and the BRCA dataset. The observations are consistent with those shown in Table 1. The model of the present invention achieves state-of-the-art performance on all metrics. Compared with the sub-optimal method, this method in BRACS-7 ★On the dataset, Accuracy, AUC, and F1-score increased by 1.81%, 1.51%, and 3.68% respectively. On the BRCA dataset, improvements of 1.63%, 0.67%, and 2.50% were achieved respectively. Compared with the srmanba model, BRACS-7 ★ The AUC of the dataset increased by 1.22%.

[0127] Table 3 Subtype classification results of BRACS data under the official division

[0128]

[0129]

[0130] Table 3 reports the subtype results of 3-class and 7-class classification using the official division on the BRACS dataset. The results show that the method of the present invention still has advantages compared with the baseline. For example, the method of the present invention improved srmanba in AUC by 4.62% and 4.65% respectively on these two tasks. This clearly shows that the significant performance improvement of the method of the present invention is not only due to the strong modeling ability of Mamba itself, and it again confirms that the scan invariance of the method of the present invention can break the bottleneck of Mamba in solving the MIL problem.

[0131] Table 4 Cancer diagnosis and typing performance with UNI as the offline feature extractor

[0132]

[0133] Since the baseline method using the UNI feature extractor for the diagnosis and typing tasks showed strong enough performance, making further comparison meaningless (see Table 4), the present invention only focuses on the impact of UNI on survival prediction in the main text.

[0134] Survival prediction

[0135] Table 5 Survival prediction results on three major datasets.

[0136]

[0137]

[0138] Table 5 presents the experimental results of three survival prediction datasets. The SMC-MIL model proposed in the present invention exhibits strong performance, with a c-index score of 61.96% on the BLCA dataset, 65.05% on LUAD, and 61.64% on LUSC. It significantly outperforms the comparative methods, improving by 0.98%, 1.00%, and 0.73% respectively over the second-best method on each dataset. Additionally, even when using high-quality features extracted by the base model, SMC-MIL yields substantial benefits. In particular, compared with SRMamba, its performance on the BLCA, LUAD, and LUSC datasets improves by 0.56%, 1.87%, and 4.17% respectively. These results emphasize the consistency and reliability of the strategy proposed in the present invention and its effectiveness in predicting survival outcomes.

[0139] Ablation experiment

[0140] Effectiveness analysis of different components

[0141] Table 6 shows the impact of different components on performance on three major datasets.

[0142]

[0143] Table 6 reports the impact of different modules in SMC-MIL on three datasets. OriginMamba means simply integrating Mamba into ABMIL. Aug. indicates that sequence contrast learning is applied. Samp. means using a sampling strategy to extend the "hard" branch to match the length of the "easy" branch. Con. indicates applying consistency constraints. The baseline of the present invention includes integrating the vanilla Mamba architecture into ABMIL. First, the contrast learning and consistency loss framework are integrated into the baseline model. This strategy improves the AUC by 1.50%, 1.18%, and 0.88% on the CAMELYON, NSCLC, and BRCA datasets respectively, indicating that the proposed consistency constraints effectively guide Mamba to learn more detailed and sequence-order-independent features from different scanning sequences. The experimental results show that compared with OriginMamba, the introduction of the sequence enhancement and masking modules respectively enhances Mamba's ability to distinguish significant regions within the sequence. After introducing these two modules into OriginMamba, the complete SMC-MIL achieves the best performance (AUC of 93.81% on CAMELYON, 96.65% on TCGA, and 93.98% on BRCA).

[0144] Scan invariance verification analysis

[0145] Table 7 Scan invariance evaluation of different Mambas

[0146]

[0147] To verify that the method of the present invention enables Mamba to construct more scan-invariant features, the present invention inputs three randomly selected scan sequences into the trained Mamba variant for inference, and calculates the average value and standard deviation of its AUC respectively. The experimental results are shown in Table 7. Compared with the baseline, the standard deviation of the method of the present invention is reduced by 2 times on the CAMELYON and NSCLC datasets, and the standard deviation is reduced by 3 times on the BRACS-7 dataset. It is worth noting that the method of the present invention not only achieves the highest average AUC in all datasets, but also shows the smallest standard deviation. This emphasizes the effectiveness of the method of the present invention in guiding the Mamba model to learn more invariant features of scan sequences, which is superior to other Mamba variant methods.

[0148] Comparison between different variants of Mamba

[0149] To further verify that the superior performance of the method of the present invention stems from the architecture of the present invention rather than Mamba itself, the present invention conducts a comparative experiment on three variants: Figure 2 Mamba Block, BiMamba, and SRMamba in. The results show that SMC-MIL achieves the best performance in all six tasks on all four datasets. The method of the present invention not only outperforms the original single-branch Mamba in terms of performance, but also shows many advantages in multi-sequence (or scan) aggregation methods (e.g., BiMamba and SRMamba). The experimental results show that the scan-invariant Mamba algorithm has stronger instance modeling ability than the multi-scan integrated Mamba algorithm and can extract more discriminative features.

[0150] Different branch comparison strategies

[0151] Table 8 Different branch comparison strategies

[0152]

[0153] Table 8 lists the performance of SMC-MIL under five different branch comparison strategies. The results show that the "easy vs hard" strategy produces the best performance. The main reason for such excellent performance is that the "easy" branch helps Mamba generate reliable bag representations and classifications, while the "hard" branch forces the model to not only learn scan-independent features, but also explore more discriminative realities from the masked sequences

[0154] For example, it is to achieve consistency between two branches. In addition, the differences between the two branches are also beneficial to performance improvement. The "hard vs hard" strategy performs the worst. The present invention attributes its failure to the lack of reliable guidance from the "easy" branch in bag representation and classification.

[0155] Visual analysis

[0156] To visually verify the interpretability of the method of the present invention, the present invention visualizes the instances with high attention scores generated by OriginMamba and the method of the present invention on the Camelyon-16 dataset, as Figure 3 shown. The blue lines outline the tumor area. The brighter the patches indicate the higher the attention score. The visualization results clearly show that compared with the original Mamba, SMC-MIL can more accurately select significant patches.

[0157] The present invention proposes a MIL framework based on Mamba and contrast learning to address the gap between scan-sensitive Mamba and scan-invariant MIL. By applying consistency constraints to the results of different scan branches, it guides Mamba to learn the scan-invariant features of bags from different sequences. Sequence contrast enhancement is respectively used to enhance and mask sequences, forcing Mamba to identify more discriminative patches. A large number of experiments show that this framework is superior to other recent methods, and the method of the present invention effectively alleviates the limitations of Mamba in MIL and other visual tasks, making it more suitable for these tasks and paving the way for the further application of Mamba in visual tasks.

[0158] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the purpose and scope of the technical solutions of the present invention, and they should all be covered by the scope of the claims of the present invention.

Claims

1. A high-resolution medical pathology WSI image analysis method based on scan invariance Mamba, characterized by: The steps include: S100: Select a public image database Q, where all images in Q contain organizational information and each image has a known label; S200: constructing the SMC-MIL network model W, which includes three parts: data offline preprocessing module, sequence contrast enhancement module and Mamba aggregation module; S300: select a picture q from Q, regard q as a package, input q into the data offline preprocessing module, and obtain the feature original scanning sequence Z; S400: Input Z into the sequence contrast enhancement module, and perform sequence enhancement branch operations on Z to obtain a new sequence Z easy And perform sequence mask branch operation to obtain a new sequence Z hard : Sequence enhancement branch operation: Z is passed through the attention mechanism to identify instances with high attention values, and the instance attention score set A is obtained. Then A is processed by the Mixup method and added to Z to obtain a new sequence Z easy ; Sequence mask branch operation: Z is passed through the attention masking mechanism to mask the instances with low attention values, and the masked new sequence Z is obtained. hard ; S500: Z easy and Z hard After passing through the Mamba aggregation module, the corresponding packet features b1 and b2 are obtained, and the expressions are as follows: b1=A1·Ψ(Z easy )=φ(Ψ(S A (Z)))·Ψ(S A (WITH)) b2=A2·Ψ(Z hard )=φ(Ψ(S M (X)))·Ψ(S M (Z)) Among them, A1 and A2 are Ψ(Z easy ) and Ψ(Z hard )’s attention score; S A represents sequence enhancement, S M represents sequence masking; Ψ(·) represents Mamba, and φ(·) represents the attention mechanism; S600: Calculate Z separately easy and Z hard Corresponding prediction package label and The specific expression is as follows: in, represents a classifier; S700: Calculate the network parameter optimization formula of the SMC-MIL model, as shown below: Among them, θ represents the network parameters of the SMC-MIL network model, γ represents the parameters for balancing classification loss and consistency loss, represents the overall loss function W, represents the consistency loss of the packet features, Represents the consistency loss of bag logits / bag classification; represents the classification loss of the sequence enhancement branch, and the expression is as follows: Among them, Y represents the true value of the packet; S800: Set the maximum number of training times, execute steps S300-S700 for all images in Q, and Update the network parameters θ, stop training when the maximum number of training times is reached, and obtain the trained network model W'; S900: Select an unknown image containing tissue information and input it into W' to obtain a predicted bag label of the image.

2. A high-resolution medical pathology WSI image analysis method based on scan invariance Mamba as claimed in claim 1, characterized in that: The specific steps of obtaining the characteristic original scanning sequence Z in S300 are as follows: S310: Cut q into several tiles, discard the tiles containing only the background area, and form the remaining tiles into a set X, where each tile represents an instance x i , specifically expressed as follows: Among them, i represents the i-th tile, and I represents the total number of tiles; S320: Use the feature extractor to extract features from all the blocks in X, and then perform linear projection on the extracted offline features to obtain the original feature scanning sequence Among them, z i represents the offline features of the i-th instance.

3. A high-resolution medical pathology WSI image analysis method based on scan invariance Mamba as claimed in claim 2, characterized in that: The operation of the sequence enhancement branch in S400 is as follows: S410: Use the attention mechanism to calculate the attention score of each instance in Z and define it as follows: A=[a1,…,a i ,…,a N ]=φ(Z) Among them, φ(·) represents the attention mechanism; S411: Sort A in S410. The expression is as follows: I=[j1,j2,…,j N ]=Sort(A) Among them, j1 represents the index of the instance with the highest attention score, j N represents the index of the instance with the lowest score; S412: Select the first α% instances with high attention scores in I as the expansion items of the Z sequence, and finally obtain the new sequence Z easy , the expression is as follows: where φ(·) is the Mixup on the selected instance, S A (·) represents the sequence enhancement module, Indicates the index of the augmentation item.

4. A high-resolution medical pathology WSI image analysis method based on scan invariance Mamba as claimed in claim 3, characterized in that: The sequence mask branch operation in S400 is specifically as follows: S420: Use the attention masking mechanism to generate an n-dimensional mask vector M from Z as follows: M=[m1,…,m r ,…,m N ], Among them, m r ∈{0,1}, if m r =1, the r-instance is unmasked, otherwise it is masked; S421: Set the masking rates of the high attention value and the low attention value to β respectively h % and β l %, so the sequence subscript corresponding to the high attention value is The sequence subscript corresponding to the low attention value is therefore, Finally, the new masked sequence Z hard It is expressed as: From hard ←S M (Z):=M·Z, Among them, S M (·) indicates a sequence mask.

Citation Information

Patent Citations

  • Mask-based digital pathological image classification method

    CN116486159A

  • Full-view digital slice image classification method based on spatial context perception

    CN118570534A

  • Pathological image classification method based on state space duality

    CN119048825A