Method for identifying ulcerative colitis inflammatory activity level based on dynamic graph multi-instance learning assistance

Through the method based on dynamic graph multi-instance learning, the problem of accuracy and inefficiency of inflammatory activity level assessment of ulcerative colitis is solved, and the efficient evaluation of inflammatory activity level in UC patients is achieved, reducing the work burden of pathologists.

CN120047411AActive Publication Date: 2025-05-27泰州学院
View PDF 14 Cites 0 Cited by

Patent Information

Application Number
CN202510122578.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-26
Publication Date
2025-05-27
Estimated Expiration
2045-01-26

AI Technical Summary

Technical Problem

The prior art is difficult to effectively assist in the evaluation of inflammatory activity levels in patients with ulcerative colitis (UC), resulting in a large work burden for pathologists and inaccurate assessment and inefficient evaluation.

Method used

Using a multi-instance learning method based on dynamic graphs, we collect and digitize full-slice images at different UC levels, perform data preprocessing and feature extraction, and build a multi-instance learning model of dynamic graphs, and combine the spatial relationship between image blocks to predict and classify inflammatory activity.

Benefits of technology

It effectively reduces the work burden of pathologists, improves the accuracy and efficiency of evaluating inflammatory activity levels in UC patients, enhances the prediction effect of the model, and can assist pathologists in diagnosis and evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047411A_ABST
    Figure CN120047411A_ABST
Patent Text Reader

Abstract

The invention provides a method for assisting in identifying ulcerative colitis inflammatory activity levels based on dynamic graph multi-instance learning. The method comprises the steps of data acquisition, data preprocessing, dynamic graph multi-instance learning model construction of inflammatory activity level diagnosis and model evaluation. The data acquisition mainly comprises the steps of collecting pathological images of a patient with ulcerative colitis, digitalizing the pathological images through a digital scanner, excluding a part of pathological images containing problems such as blurring, fading and abnormal staining, and finally determining a grading label of each pathological image by a deep gastrointestinal pathology expert. The data preprocessing mainly comprises the steps of processing a digital pathological image, converting the digital pathological image into image blocks, then extracting features from each image block by using a visual basic model, and finally constructing a dynamic graph multi-instance learning model for training. According to the method, exploration of the internal relation of the image blocks input to the WSI is promoted, and the prediction effect of the model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer vision, and particularly relates to a method for assisting in identifying the inflammatory activity level of ulcerative colitis based on dynamic graph multi-instance learning. Background Art

[0002] Ulcerative colitis (UC) is a chronic, recurrent, and non-specific inflammatory bowel disease (IBD). It mainly affects the mucosal layer of the colon and rectum, resulting in long-term inflammation and ulcers. Severe complications, such as massive bleeding and toxic megacolon, may increase the risk of life and require total colectomy in extreme cases. Therefore, evaluating the remission status of UC is very important for subsequent treatment and prognosis.

[0003] Prior studies have shown that observing a higher degree of mucosal healing during endoscopy is closely related to better clinical outcomes and prognosis in UC patients. Therefore, the clinical goals of UC treatment usually include symptom remission and achieving mucosal healing during endoscopy. Recent studies have also pointed out that persistent histological inflammation is closely related to the recurrence of UC, the risk of surgery, and the occurrence of colorectal cancer (CRC). It has been determined that the degree of histological inflammation plays a key role in patient treatment. In addition, during the histological mucosal healing process, the degree of active inflammation characterized by neutrophil infiltration is most closely related to endoscopic features and patient prognosis outcomes. Therefore, it is necessary to predict the inflammatory activity of UC.

[0004] Currently, artificial intelligence (AI) methods have been successfully applied in oncology, including pan-cancer tumor location detection, subtype classification of various diseases, prediction of tumor treatment response and prognosis, and comprehensive radiological analysis. Therefore, these AI methods provide a solid foundation for the research of IBD. Significant progress has been made in using AI technology to assist in the diagnosis and differential diagnosis of IBD. In the field of molecular genetics, genome-wide association studies have identified more than 240 gene loci associated with an increased risk of IBD, and these gene loci also help distinguish UC and CD. Genetic risk scores generated using disease-related gene information from large multi-center datasets of IBD patients show a strong correlation with IBD subtypes. In addition, analyzing protein features in colon tissue using a support vector machine learning model can identify CD and UC with an accuracy of 76.9%. Therefore, it provides technical feasibility in the aspect of diagnosing at the inflammatory level. Summary of the Invention

[0005] Objective of the Invention: Inflammatory bowel disease (IBD) is a global disease with an increasing incidence. However, there is a scarcity of research on computer-aided diagnosis of IBD based on pathological images. Therefore, in order to reduce the workload of pathologists and improve the accuracy and efficiency of grading the inflammatory activity of UC patients, and to assist pathologists in diagnosing and evaluating UC more precisely and consistently, the present invention proposes a method for assisting in identifying the inflammatory activity level of ulcerative colitis based on dynamic graph multi-instance learning.

[0006] Technical Solution: A method for assisting in identifying the inflammatory activity level of ulcerative colitis based on dynamic graph multi-instance learning, comprising the following steps:

[0007] Step 1, data collection: Collect whole-slide images at different ulcerative colitis (UC) levels and digitize them using a digital scanner;

[0008] Step 2, data preprocessing of whole-slide images: Control the image content and quality, divide the tissue regions in the whole-slide images into image patches to form the training data set and test data set of the model;

[0009] Step 3, quantify the similarity between the training data set and the external data set, and use the vision-based model UNI to perform screening feature extraction to obtain feature vectors;

[0010] Step 4, construct a prediction model for the inflammatory activity degree of ulcerative colitis (UC) at the image level, perform model training based on the feature vectors obtained in Step 3, use the method of multi-instance learning (MIL) of dynamic graphs to process the spatial relationship between different image patches, construct a directed graph between the image patches, so as to fuse the dynamic graph features between the image patches and perform image classification;

[0011] Step 5, use 5-fold cross-validation to evaluate the performance of the DGMIL model for two tasks, compare it with traditional MIL and AttMIL algorithms, and use Gradient-weighted Class Activation Mapping (Grad-CAM) for visualization; the two tasks refer to determining whether inflammation exists and further subdividing the level of inflammation.

[0012] Step 1 includes: collecting whole-slide images at different ulcerative colitis (UC) levels, grading each whole-slide image (WSI), with four grades: L0 (no activity: no intraepithelial neutrophils, ulcers or erosions); L1 (mild activity: cryptitis involving <= 25% of the crypts or crypt abscesses involving <= 10% of the crypts); L2 (moderate activity: cryptitis affecting > 25% of the crypts, crypt abscesses involving > 10% of the crypts, or focal, small erosions); L3 (severe activity: presence of ulcers or extensive erosions); and finally digitizing the whole-slide images using a digital scanner.

[0013] Step 2 includes: separately processing the tissue regions of each whole-slide image, removing the background of the WSI using median filtering and the Otsu algorithm and then performing foreground extraction on each WSI to obtain the tissue regions, usually using a magnification closest to 64x downsampling. This process involves extracting the foreground and eliminating larger cavities within the tissue regions. The extraction method uses median filtering, the Otsu algorithm, and other morphological closing techniques, with different threshold settings for tissue region extraction and contour delineation. Since the size of the whole-slide image (WSI) is very large, usually reaching the GB level in terms of storage, it is very difficult to directly process it. In the present invention, at the maximum magnification of 40x of the WSI, the tissue regions are cropped into image patches of size 512×512. These small patches are saved using a coordinate-based method to optimize computing resources and facilitate retrieval.

[0014] Step 3 includes: in order to quantify the similarity between the training dataset and the external dataset, the present invention uses the ψ (psi) metric analysis. This non-parametric and distribution-free technique can quantify the similarity between two datasets, thus better evaluating the generalization ability and robustness of the model. In order to extract representative features from each image patch. In the present invention, for each image patch, feature extraction is performed using the vision base model UNI through transfer learning. The vision base model UNI converts all input image patches into 1024-dimensional feature vectors, and then an independent linear layer is used to reduce the feature dimension of the image patches to 512 dimensions. The formula is as follows:

[0015] h i =W h f(X)

[0016] where X represents the input feature matrix, here representing the feature vector of the image patch (patch), f(X) represents extracting the feature vector of the original image patch into 1024 dimensions through the vision base model UNI, and W h is a weight matrix, here representing a linear transformation that reduces the 1024-dimensional feature vector extracted by the vision base model UNI to 512 dimensions; h iIt represents the feature vector obtained after the i-th image patch passes through the visual base model UNI and the linear layer.

[0017] Step 4 includes: for the labels of all image patches, using the label at the whole-slide image WSI level as the label of the image patch to construct an automatic grading model for ulcerative colitis UC inflammation activity at the slide level;

[0018] The automatic grading model for ulcerative colitis UC inflammation activity is a dynamic graph-based multiple instance learning model DGMIL (Dynamic Graph-based Multiple Instance Learning, DGMIL). The multiple instance learning model DGMIL includes a dynamic graph module and a multiple instance learning module MIL. The multiple instance learning model DGMIL combines the dynamic graph structure between image patches and the multiple instance learning MIL method, enhances the interactive processing of the spatial information of the whole-slide image WSI, and effectively mines the correlation between image patches;

[0019] The dynamic graph module first calculates the similarity score between image patches, and the formula is:

[0020]

[0021] where i≠j, N is the set of all image patches, h j represents the feature vector obtained after the j-th image patch passes through UNI and the linear layer, represents calculating the dot product similarity of the two feature vectors h i , h j After that, the obtained similarity is passed through the normalized exponential function softmax to obtain the similarity score w between the i-th and j-th image patches i,j ,

[0022] After obtaining the similarity score w i,j , each image patch selects the top-k image patches with the highest similarity score as the adjacent image patches of the i-th image patch, thereby enhancing the global feature representation. The formula is:

[0023]

[0024] where represents selecting the k w with the highest similarity score between the i-th image patch and all other image patches i,j , represents the set of the k image patches with the highest similarity score w i,j between the i-th image patch and the i-th image patch selected from all image patches, as the adjacent image patches of the i-th image patch;

[0025] The representation of the directed topological structure between image patches is as follows:

[0026] d i,j = w i,j h j +(1 - w i,j )h i

[0027] where d i,j represents the spatial embedding representation between the i-th image patch and the j-th image patch. Information flows from the k image patches with the largest similarity scores to the i-th image patch, and the features of the i-th image patch are updated by combining the features of these k image patches. In the graph structure, this means that the nodes corresponding to these k image patches are connected to the node corresponding to the i-th image patch, and the direction is from the nodes of the k image patches to the node of the i-th image patch;

[0028] For the i-th image patch, calculate the linear combination of the features of the set N(i) of the k image patches with the largest similarity scores w i,j selected from all image patches and the i-th image patch to characterize the first-order connection structure of the i-th image patch:

[0029]

[0030] where τ is a weight used to guide the information propagation of the top-k image patches to the i-th image patch, and h N(i) is the updated feature vector obtained by weighted combination of the feature vectors of the i-th image patch and its adjacent image patches;

[0031] The dot product and summation methods are used to enable the interaction information between nodes, which is expressed as:

[0032] h i = α(w 1 (h i + h N(i) )) + β(w 2 (h i ⊙ h N(i) ))

[0033] where α and β represent the LeakyReLU activation function, and w 1 and w 2 represent learnable transformation matrices;

[0034] Finally, the output of the dynamic graph structure is used as the input to the multi-instance learning module MIL, and the label category probability of each WSI image is obtained through Softmax and MaxPooling. The formula is:

[0035]

[0036] Among them, G is the updated feature vector obtained after information interaction of all image patches output by the dynamic graph, which is expressed as the class probability.

[0037] The dynamic graph structure effectively alleviates the loss of global and spatial information caused by dividing whole slide images (WSIs) into multiple image patches. Its core function enables each element in the input sequence to interact with other elements, facilitating the exploration of the internal relationships of the image patches input to the WSI. In addition, an important advancement of the dynamic graph structure is its ability to allow the decoder to access the entire encoded information and assign weights to the input data. This process captures the importance of each token, prioritizing them when generating output labels at each step. By covering the receptive field of the entire image, the dynamic graph utilizes self-captured global and spatial features to predict the class of the WSI.

[0038] In step 4, the formula for the softmax of the normalization exponential function is:

[0039]

[0040] where z i is the i-th element in the input vector, and exp(z i ) represents the exponential function of z i , and the denominator part is the sum of the exponential functions of all elements in the input vector.

[0041] In step 4, for image-level prediction, a multi-instance learning MIL method combined with the dynamic graph structure is adopted. The configuration of hyperparameters is as follows: the optimizer uses Adam, the loss function uses softmax cross-entropy, that is, the combination of standard cross-entropy loss and softmax; the initial learning rate is set to 2e-4, and the dropout rate is 0.25. During the training iteration process, the loss is calculated, and the weights of the multi-instance learning model DGMIL are updated based on the minimum loss achieved at the end of each iteration. The maximum number of training iterations is set to 200, and the minimum is 50. Subsequently, the decision on whether to continue training is based on the change in the loss value. If the loss value does not decrease after 20 consecutive training iterations, the training process is terminated.

[0042] In step 5, to evaluate the variability of the dataset and the robustness of the multi-instance learning model DGMIL, 5-fold cross-validation was used: the multi-instance learning model DGMIL and two classical multi-instance learning (MIL) algorithms, MIL and AttMIL, were both tested on internal and external test sets. In the dynamic graph structure of the network in step 4, ReLU was used as the activation function, and cross-entropy was used as the loss function. The positional relationship between patches was quantified through learnable implicit features. The whole-slide image WSI was obtained by non-overlapping patching to get image patches X = {x 1 , x 2 , …, x n}, where x n represents the nth image patch, n is the total number of image patches, and all image patches are used as nodes of the graph;

[0043] The multi-instance learning model DGMIL processes two different tasks (determining the presence of inflammation and further classifying the level of inflammation). It aggregates patch features through directed graph weights to predict the label of the entire whole-slide image WSI. The training of the multi-instance learning model DGMIL uses cross-entropy loss and combines softmax to calculate classification probabilities;

[0044] The multi-instance learning model DGMIL generates a heatmap to show the regions of interest that the multi-instance learning model DGMIL focuses on during the diagnosis process, in order to explain the decision-making basis of the model. This visualization method can help understand the degree of attention of the model in each region and compare the results with pathologists.

[0045] The DGMIL model of the present invention processes different tasks through multiple classification branches. Specifically, after feature extraction, it inputs the features into different task branches, and each branch is responsible for the classification of a specific task. This design enables the model to perform refined processing according to the requirements of the tasks on the basis of sharing underlying features. It can simultaneously complete two tasks: determining the presence of inflammation and further classifying the level of inflammation.

[0046] The present invention also provides an electronic device, including a processor and a memory. The memory stores program code, and when the program code is executed by the processor, the processor is caused to execute the steps of the method.

[0047] The present invention also provides a storage medium storing a computer program or instruction, and when the computer program or instruction runs on a computer, it executes the steps of the method.

[0048] Beneficial effects: The present invention uses image-level labels. Pathologists do not need to label the inflammation level at the pixel level, but only need to label at the image level, reducing the workload of pathologists involved in constructing the deep learning model. At the same time, the dynamic graph structure effectively reduces the loss of global and spatial information caused by splitting whole slide images (WSIs) into multiple image patches. Its core function enables each element in the input sequence to interact with other elements, facilitating the exploration of the internal relationships of the image patches input to the WSI and improving the prediction effect of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] The following further describes the present invention in detail with reference to the drawings and specific embodiments, and the above and / or other advantages of the present invention will become clearer.

[0050] Figure 1 It is a flowchart of the method of the present invention;

[0051] Figure 2 It is a bar chart including the patient selection process and the number of patients and patches at each inflammation activity level in each cohort;

[0052] Figure 3 It is a schematic diagram visualizing the differences between the datasets DTHUC and ZJH by the density and cumulative distribution function (CDF) graph of the present invention;

[0053] Figure 4 It is a schematic diagram evaluating the diagnostic performance of predicting whether UC patients have inflammation based on DGMIL on internal and external test sets through 5-fold cross-validation of the present invention;

[0054] Figure 5 It is a schematic diagram evaluating the diagnostic performance of predicting the histological activity grade of stage 4 in UC patients based on DGMIL on internal and external test sets through 5-fold cross-validation of the present invention;

[0055] Figure 6 It is a schematic diagram of the results of comparing the diagnostic accuracy of the present invention in two tasks and two pathologists;

[0056] Figure 7 It is a prediction heat map of the DGMIL model of the present invention; DETAILED DESCRIPTION OF THE EMBODIMENTS

[0057] This embodiment discloses a method for assisting in identifying the inflammatory activity level of ulcerative colitis based on dynamic graph multi-instance learning, including the following steps:

[0058] Step 1: The experimental dataset used in this invention is to collect whole-slide images (WSIs) with different ulcerative colitis (UC) levels from Nanjing Drum Tower Hospital Affiliated to Nanjing Medical University as the model training set and internal test set, and collect UC whole-slide images from Zhujiang Hospital of Southern Medical University as the external test set. For convenience, the two datasets are named DTHUC and ZJH respectively. To ensure the quality of WSIs, 33 WSIs with problems such as blurring, fading, and abnormal staining were excluded from the DTHUC dataset. In the ZJH dataset, 26 WSIs were excluded. Finally, the DTHUC dataset contains 570 WSIs of UC, among which there are 76 WSIs at L0 level, 135 at L1 level, 250 at L2 level, and 109 at L3 level. The training set, validation set, and test set are divided by 10-fold Monte Carlo cross-validation according to the ratio of 7:1:2. The ZJH dataset is used as the external test set, which contains 186 WSIs of UC, among which there are 12 WSIs at L0 level, 94 at L1 level, 58 at L2 level, and 22 at L3 level. After digitizing the tissue sections, the WSIs are preprocessed into image patches of 512×512 pixels at the maximum magnification. All WSIs are digitized by a Leica GT450 scanner at a pixel resolution of 263 nanometers / pixel.

[0059] Step 2: After extracting the foreground region of the WSI image, the foreground region is processed into an image patch of 512×512 pixels, as Figure 1 (A) shows. By analyzing the ψ (psi) index of the training dataset and the external dataset, the ψ value is obtained as 0.341, indicating that there is a moderate difference in the feature distribution between the training set and the external dataset. As Figure 3 shown, this invention uses density and cumulative distribution function graphs to visualize the differences between the DTHUC dataset and the ZJH dataset. Through this similarity quantification, the differences between the external dataset and the training dataset can be shown more clearly, so as to better evaluate the generalization ability and robustness of the model.

[0060] Step 3: The grading labels of each WSI are jointly determined by three senior gastrointestinal pathologists ( Figure 1 (B)). To compare the AI evaluation results with those of pathologists with different experiences, two other junior pathologists also participated in this study, one with 2 years of work experience and the other with 5 years of work experience. Pathologists do not need to label the inflammation level at the pixel level, but only need to label at the image level, reducing the workload of pathologists participating in the construction of the deep learning model.

[0061] Step 4: As a feature extractor, UNI converts all image patches into 1024-dimensional structured data through a vision-based model and fuses them to represent each WSI.

[0062] Step 5: In the model training stage, the vectorized features from the training set and the validation set are constructed into a dynamic graph representation, combined with the softmax loss for continuous optimization and iteration of the model. The DGMIL model uses feature vector encoding based on a vision-based foundation model, combines a dynamic graph structure to process the feature vectors, and uses softmax to output hierarchical scores. Finally, the test set is used to evaluate the algorithm performance to verify the effectiveness of the algorithm model.

[0063] To evaluate the variability of the dataset and the robustness of the algorithm model, 5-fold cross-validation is used. In the task of determining whether there is inflammation in UC patients (considering L0 as the non-inflammation class and L1, L2, and L3 as the inflammation classes), the average accuracy (ACC) of the algorithm model in the internal test set is 87.3%, and the average AUC is 0.863 (95% [CI] 0.829, 0.898). The average sensitivity is 0.913 (95% [CI] 0.866, 0.961), and the average specificity is 0.816 (95% [CI] 0.771, 0.861). In the external dataset, the average accuracy (ACC) of the model of the present invention is 88.7%, and the average AUC is 0.947 (95% [CI] 0.939, 0.955). The average sensitivity is 0.889 (95% [CI] 0.837, 0.940), and the average specificity is 0.858 (95% [CI] 0.777, 0.939). Figure 4 The table shows the results of 10-fold cross-validation of the internal and external test sets in the task of determining whether there is inflammation in UC patients, including ACC, Macro-AUC, Micro-AUC, sensitivity, and specificity for each fold. In the task of four-grade histological activity grading of UC patients, the average ACC of the algorithm model in the internal test set is 76.9%, the average Macro-AUC is 0.827 (95% [CI] 0.803, 0.850), the average Micro-AUC is 0.816 (95% [CI] 0.792, 0.840), the average sensitivity is 0.770 (95% [CI] 0.705, 0.835), and the average specificity is 0.856 (95% [CI] 0.834, 0.878). In the external dataset, the average accuracy (ACC) of the model of the present invention is 70.0%, the average Macro-AUC is 0.908 (95% [CI] 0.882, 0.935), and the average Micro-AUC is 0.898 (95% [CI] 0.869, 0.926). The average sensitivity is 0.678 (95% [CI] 0.626, 0.731), and the average specificity is 0.883 (95% [CI] 0.861, 0.906). Figure 5The table shows the results of 10-fold cross-validation of the internal and external test sets in the task of histological activity grading of four levels for UC patients.

[0064] Step 6: Comparison with the evaluation results of pathologists with different experiences. The main purpose of designing the algorithm model in this study is to assist pathologists in diagnosis, reduce repetitive workload, and allow them to have more time to focus on other complex medical conditions. During the model training stage, each WSI was graded by three senior gastrointestinal pathologists (each with more than 15 years of clinical experience) to ensure the high quality of the training data. To evaluate the potential application of the AI model in clinical practice, two junior pathologists (with 2-5 years of clinical experience) were selected in this invention for result comparison. The model mentioned in this invention being superior to the pathologists' diagnostic results refers to the comparison with the diagnostic results of these two junior pathologists. Pathologist 1 has 2 years of pathological diagnosis experience, and Pathologist 2 has 5 years of experience. In the internal test set and the external test set, the diagnostic accuracy rates of the two pathologists in the two tasks are shown in Figure 6 . Therefore, the algorithm model developed in this study can be used to guide pathologists with 2 years of experience in diagnosing the inflammatory activity level of UC, and can also assist pathologists with 5 years of experience in diagnosing the histological activity grading of UC.

[0065] Step 7: Comparative analysis of DGMIL and two other algorithms in calculating UC pathological activity. DGMIL was compared with two classical multi-instance learning (MIL) algorithms, MIL and AttMIL, in various experimental tasks. MIL and AttMIL each have their own advantages and disadvantages in different studies, and the test processes of the three algorithms are kept consistent.

[0066] Step 8: To provide a visual interpretation of DGMIL. The visualization results show that, as Figure 7 shown, the model's focus is concentrated on the glandular area, which coincides with the key points of pathologists' diagnosis. The histological activity grading criteria for pathological evaluation depend on the degree to which the crypt glands are damaged by active inflammation, which emphasizes that the model of this invention can identify the key pathological areas required for accurate diagnosis. The model's concentration of attention on these specific areas further indicates that the insights of DGMIL are consistent with the careful considerations made by pathologists during evaluation. In a typical case, the red highlighted area is concentrated in the lamina propria of the colon mucosa, which is the area of greatest interest to pathologists and the most common site of active enteritis. In contrast, active inflammation appears less in the muscularis mucosa and submucosa, so the AI pays less attention to these areas. The model of this invention always focuses on these key indicators, enhancing its diagnostic ability and being consistent with the expertise of pathologists, who also rely on similar visual cues for accurate and comprehensive evaluation. It helps to deepen the understanding of the algorithm's decision-making process.

[0067] The present invention provides a method for assisting in identifying the inflammatory activity level of ulcerative colitis based on dynamic graph multi-instance learning. There are many methods and approaches to specifically implement this technical solution. The above description is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art of this technology, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention. Each component not clearly defined in this embodiment can be implemented by using the prior art.

Claims

1. A method for identifying the inflammatory activity level of ulcerative colitis based on dynamic graph multi-instance learning, characterized in that: The following steps are involved: Step 1, data collection: whole-slide images of different UC levels were collected and digitized using a digital scanner; Step 2: Data preprocessing of the whole slice image: image content and quality control, dividing the tissue area in the whole slice image into image blocks to form the training data set and test data set of the model; Step 3: quantify the similarity between the training dataset and the external dataset, and use the visual basic model UNI to perform screening feature extraction to obtain a feature vector; Step 4: construct an image-level ulcerative colitis UC inflammatory activity prediction model, train the model based on the feature vector obtained in step 3, use the dynamic graph multiple instance learning (MIL) method to process the spatial relationship between different image blocks, construct a directed graph between image blocks, and fuse the dynamic graph features between image blocks to classify the image; Step 5, use 5-fold cross validation to evaluate the performance of the DGMIL model for two tasks, compare it with the traditional MIL and AttMIL algorithms, and visualize it using gradient weighted class activation mapping; the two tasks are to determine whether inflammation exists and further subdivide the level of inflammation.

2. The method according to claim 1, characterized in that: Step 1 includes: collecting whole-slice images of different UC levels, grading each whole-slice image WSI into four levels: L0 no activity; L1 mild activity; L2 moderate activity; L3 severe activity; and finally digitizing the whole-slice images using a digital scanner.

3. The method according to claim 2, characterized in that Step 2 includes: processing the tissue area of ​​each full-slice image separately, using median filtering and Otsu algorithm to remove the background of WSI and then extract the foreground of each WSI to obtain the tissue area, and then cropping the tissue area into an image block of 512×512 size.

4. The method according to claim 3, characterized in that Step 3 includes: for each image block, using the visual base model UNI to extract features through transfer learning. The visual base model UNI converts all input image blocks into 1024-dimensional feature vectors, and then uses an independent linear layer to reduce the feature dimension of the image block to 512 dimensions. The formula is as follows: h i =W h f(X) Where X represents the input feature matrix, which represents the feature vector of the image patch, f(X) represents the feature vector of the original image patch extracted into 1024 dimensions through the visual basic model UNI, and W h is a weight matrix representing a linear transformation that reduces the 1024-dimensional feature vector extracted by the visual basis model UNI to 512 dimensions; h i Represents the feature vector obtained after the i-th image block passes through the visual basis model UNI and the linear layer.

5. The method according to claim 4, characterized in that Step 4 includes: for labels of all image blocks, using the labels at the WSI level of the whole slice image as labels of the image blocks, for building an automatic grading model of inflammatory activity of ulcerative colitis (UC) at the slice level; The automatic classification model of ulcerative colitis UC inflammatory activity is a dynamic graph-based multiple instance learning model DGMIL, and the multiple instance learning model DGMIL includes a dynamic graph module and a multiple instance learning module MIL; The dynamic image module first calculates the similarity score between image blocks using the formula: Where i≠j, N is the set of all image blocks, h j represents the feature vector obtained after the jth image block passes through the UNI and linear layers, Represents the calculation h i ,h j The dot product similarity of the two feature vectors is then normalized by the exponential function softmax to obtain the similarity score w between the i-th and j-th image blocks. i,j , Get the similarity score w i,j After that, each image block selects the top-k image blocks with the highest similarity score as the adjacent image blocks of the i-th image block. The formula is: in Indicates that the k w with the highest similarity scores between the i-th image block and all other image blocks are selected i,j , Represents the similarity score w between the i-th image block and the image blocks selected from all the image blocks i,j The set of the largest k image blocks is used as the neighboring image blocks of the i-th image block; The representation of the directed topological structure between image blocks is: d i,j =w i,j h j +(1-w i,j )h i where d i,j Represents the spatial embedding representation between the i-th image block and the j-th image block. Information flows from the k image blocks with the largest similarity scores to the i-th image block. The features of the i-th image block are updated by combining the features of the k image blocks. In the graph structure, the nodes corresponding to the k image blocks are connected to the nodes corresponding to the i-th image block, and the direction is from the nodes of the k image blocks to the nodes of the i-th image block. For the i-th image block, calculate the similarity score w between the i-th image block and the i-th image block selected from all image blocks i,j The linear combination of the features of the largest set of k image patches N(i) is used to characterize the first-order connection structure of the i-th image patch: Where τ is a weight used to guide the information of the top-k image blocks to propagate to the i-th image block, and h N(i) is the updated feature vector obtained by weighted combination of the feature vectors of the i-th image block and the adjacent image blocks; The point product and summation method is used to exchange information between nodes, which can be expressed as: h i =α(w1(h i +h N(i) ))+β(w2(h i ⊙h N(i) )) Among them, α and β represent LeakyReLU activation functions, w1 and w2 represent learnable transformation matrices; Finally, the output of the dynamic graph structure is used as input to the multi-instance learning module MIL, and the label category probability of each WSI image is obtained through Softmax and MaxPooling. The formula is: Where G is the updated feature vector obtained after information interaction of all image blocks output by the dynamic graph. Expressed as class probability.

6. The method according to claim 5, characterized in that In step 4, the formula of the normalized exponential function softmax is: where z i is the i-th element in the input vector, exp(z i ) represents z i The exponential function of is the sum of the exponential functions of all elements in the input vector.

7. The method according to claim 6, characterized in that In step 4, for image-level prediction, a multi-instance learning MIL method combined with a dynamic graph structure is used, and the hyperparameters are configured as follows: the optimizer uses Adam, and the loss function uses softmax cross entropy, that is, a combination of standard cross entropy loss and softmax; the initial learning rate is set to 2e-4, and the drop rate is 0.25; during the training iteration, the loss is calculated, and the weights of the multi-instance learning model DGMIL are updated based on the minimum loss achieved at the end of each iteration.

8. The method according to claim 7, characterized in that In step 5, in order to evaluate the variability of the dataset and the robustness of the multi-instance learning model DGMIL, a 5-fold cross validation was used: the multi-instance learning model DGMIL and the two multi-instance learning algorithms MIL and AttMIL were tested on both internal and external test sets. In the dynamic graph structure of the network in step 4, ReLU was used as the activation function and cross entropy as the loss function. The positional relationship between patches was quantified by learnable implicit features. The full slice image WSI was obtained by non-overlapping block extraction to obtain a 512×512 image block X={x1,x2,…,x n }, where x n represents the nth image block, n is the total number of image blocks, and all image blocks are nodes of the graph; The multi-instance learning model DGMIL processes two tasks. It summarizes the image block features through directed graph weights to predict the label of the entire full slice image WSI. The multi-instance learning model DGMIL is trained using cross entropy loss and combined with softmax to calculate the classification probability; The multiple instance learning model DGMIL generates a heat map to show the regions of interest that the multiple instance learning model DGMIL focuses on during the diagnosis process to explain the decision basis of the model.

9. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores program codes, and when the program codes are executed by the processor, the processor executes the steps of the method according to any one of claims 1 to 8.

10. A storage medium, characterized in that: A computer program or instruction is stored, and when the computer program or instruction is run on a computer, the steps of the method according to any one of claims 1 to 8 are executed.

Citation Information

Patent Citations

  • Deep-network-based tissue segmentation method of panoramic digital colorectum pathology image

    CN107665492A

  • Ulcerative colitis severity assessment method and system based on deep learning

    CN110993099A

  • Multi-modal ultrasonic image RA activity deep learning method and device

    CN115439701A

  • Mammary gland pathological image molecular typing prediction method based on multi-attribute embedding model

    CN116310553A

  • Intestinal epithelium metaplasia region image segmentation system based on multi-instance learning

    CN116342627A