Instance segmentation method and device for early intestinal cancer vessel infringement phenomenon analysis and readable storage medium thereof

By using the Mask2Former instance segmentation network and a multimodal hybrid expert network enhanced by pathological large model in the analysis of early intestinal cancer, the problem of relying on subjective interpretation and single staining detection in the prior art is solved, and automated detection with high sensitivity and precision is achieved.

CN120235889AActive Publication Date: 2025-07-01SHENZHEN SHENGQIANG TECH

Patent Information

Application Number
CN202510712776.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-07-01
Estimated Expiration
2045-05-30

AI Technical Summary

Technical Problem

In the analysis of early bowel cancer vascular invasion phenomenon, the existing technology has problems such as relying on the subjective interpretation of pathologists, single HE staining detection, high misdiagnosis and missed diagnosis, large differences in observers, and insufficient detection ability of micro/irregular vascular invasion structures.

Method used

A multimodal hybrid expert network based on Mask2Former instance segmentation network and pathological large model enhancement is adopted to achieve end-to-end automated high-sensitivity detection and accurate identification of early intestinal cancer vascular invasion through multi-scale image preprocessing, full-field mask merging postprocessing, ROI expansion completion and multi-staining modal fusion classification.

Benefits of technology

It realizes automated analysis throughout the process, significantly reduces the rate of misdiagnosis and missed diagnosis, improves the high sensitivity and accuracy of detection, and meets the clinical needs for high-throughput and high-precision diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120235889A_ABST
    Figure CN120235889A_ABST
Patent Text Reader

Abstract

The invention provides an instance segmentation method and device for early intestinal cancer vessel infringement phenomenon analysis and a readable storage medium, and the method comprises the steps: constructing an end-to-end automatic process: firstly, carrying out the multi-magnification overlapping cutting of a digital pathological image, and carrying out the high-sensitivity instance segmentation through a Mask2form instance segmentation network; combining the overlapped masks by adopting an efficient union lookup algorithm, and recursively complementing incomplete instances in combination with an ROI (Region of Interest) expansion algorithm; hE, CD-31 and D2-40 dyeing full slice level alignment is realized through rigid-non-rigid-micro registration, and a three-mode instance level graph block is cut based on a registration result; and finally, fusing multi-modal features by using a hybrid expert network enhanced by a large pathological model, and dynamically screening expert paths to carry out true and false positive judgment. The invention provides an efficient technical scheme for early accurate diagnosis of intestinal cancer, and is suitable for pathological big data analysis and clinical auxiliary diagnosis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of medical diagnosis technology, and in particular to an instance segmentation method, device and readable storage medium for analyzing vascular invasion in early colorectal cancer. Background Art

[0002] Colorectal cancer (CRC) is a malignant tumor. Although the popularization of early screening technology has improved the detection rate, vascular invasion (tumor cells invading blood vessels or lymphatic vessels) remains a major challenge in clinical diagnosis. Vascular invasion is a key indicator of cancer metastasis and poor prognosis, and accurate detection is crucial for the selection of treatment plans. However, traditional pathological detection relies on pathologists' subjective interpretation of HE-stained sections, resulting in significant observer differences (for example, international research shows that the κ value is only 0.518). Moreover, blood vessels / lymphatic vessels are easily confused with similar structures (such as serum-filled lumens, carcinoma in situ) under HE staining, leading to misdiagnosis and missed diagnosis. Existing automated detection technologies are mostly based on single HE staining. Due to the blurred boundaries of blood vessels / lymphatic vessels and the tiny and irregular morphology of vascular invasion, problems such as low detection rate and high false positive rate exist. In addition, the lack of efficient multi-modal fusion (such as CD-31, D2-40 immunohistochemical staining) and a full-process automated analysis framework makes it difficult to meet the clinical requirements for high-throughput and high-precision diagnosis.

[0003] Therefore, there is an urgent need for an instance segmentation method, device and readable storage medium for analyzing vascular invasion in early colorectal cancer to solve the problems existing in the prior art. Summary of the Invention

[0004] Embodiments of the present invention provide an instance segmentation method, device and readable storage medium for analyzing vascular invasion in early colorectal cancer, aiming at the problems existing in the current technology, such as relying on pathologists' subjective interpretation or single-modal detection, having high misdiagnosis rate, missed diagnosis rate, large observer differences, and insufficient detection ability for tiny / irregular vascular invasion structures.

[0005] The core technology of the present invention is mainly based on the Mask2Former instance segmentation network and a multi-modal hybrid expert network enhanced by a pathological large model. Through multi-scale image preprocessing, full-field mask merging postprocessing, ROI expansion and complementation, and multi-staining modal fusion classification, end-to-end automated high-sensitivity detection and accurate identification of vascular invasion in early colorectal cancer are achieved.

[0006] In a first aspect, the present invention provides an instance segmentation method for analyzing vascular invasion in early colorectal cancer, and the method includes the following steps: S00. Generate multi-scale image patches adapted to the input of the deep learning model by overlapping cropping of digital pathology images at multiple magnifications; S10. Use the Mask2former instance segmentation network to perform instance segmentation of vascular invasion on multi-scale image patches. The Mask2former instance segmentation network uses ResNet-101 as the backbone network, extracts multi-scale feature maps of four levels and inputs them into the segmentation decoding network, and performs instance prediction through a preset number of query tokens and the multi-head attention mechanism; S20. Merge the segmentation masks within the full field of view based on the efficient union-find algorithm, group the masks with intersections into the same connected component and retain the maximum confidence score, and then perform recursive expansion inference on the instances with incomplete edges through the ROI expansion algorithm until there are no new incomplete instances; S30. Perform rigid registration, non-rigid registration, and micro-registration on the whole-slide images stained with HE, CD-31, and D2-40 through the registration algorithm to achieve instance-level alignment; S40. Taking the suspected vascular invasion area on the HE staining as a reference, crop the corresponding same area in the registered CD-31 staining and D2-40 staining images to form a three-modal instance-level aligned tile; S50. Input the three-modal instance-level aligned tiles into the mixture-of-experts network enhanced by the pathology large model, dynamically select the expert path through the gated activation network, and fuse multi-modal features for true / false positive discrimination.

[0007] Further, in step S00, the multi-magnifications include 2.5X, 5X, and 10X, the resolution of the image patch is 1024×1024, and the overlapping area is 512 pixels.

[0008] Further, in step S10, the Mask2former instance segmentation network uses a low confidence threshold during inference to retain more suspicious vascular invasion areas.

[0009] Further, in step S20, the union-find algorithm processes the overlapping masks through the path compression and union by rank optimization strategies, and the merging threshold Tm is 0; the ROI expansion algorithm identifies incomplete edges by judging that the ratio Re of the number of consecutive overlapping points of the instance segmentation mask contour to the total number of points of the detection box line is greater than the threshold Te.

[0010] Further, in step S30, taking the HE staining image as the reference image, first filter the background through foreground segmentation, and then perform rigid registration, non-rigid registration, and micro-registration in sequence to achieve whole-slide-level alignment of different staining modalities.

[0011] Further, in step S50, the pathology large model is the pre-trained VIT-Large model. Freeze the parameters of the first 23 Transformer modules, and only update the parameters of the last L - 23 layers to extract feature vectors for the three-modal tiles respectively, where L is the total number of layers of the VIT model.

[0012] Furthermore, the mixture-of-experts network contains 4 experts, with each staining modality activating 2 experts. The activation probabilities of the expert pathways are calculated through a weight-sharing gating activation network. Additive aggregation is used for single-staining expert features, and concatenation aggregation is used for multi-staining features.

[0013] In a second aspect, the present invention provides an instance segmentation device for analyzing vascular invasion in early-stage colorectal cancer, including: A preprocessing module that generates multi-scale image patches adapted for input to a deep learning model by overlapping cropping of digital pathology images at multiple magnifications; An instance segmentation module that performs instance segmentation of vascular invasion on the multi-scale image patches using the Mask2former instance segmentation network. The Mask2former instance segmentation network uses ResNet-101 as the backbone network, extracts multi-scale feature maps at four levels and inputs them into a segmentation decoding network, and performs instance prediction through a preset number of query tokens and a multi-head attention mechanism; A post-processing module that merges the segmentation masks within the full field of view based on an efficient union-find algorithm, classifies the masks with intersections into the same connected component and retains the maximum confidence score, and then recursively expands and infers the instances with incomplete edges through a ROI expansion algorithm until there are no new incomplete instances; A registration module that performs rigid registration, non-rigid registration, and micro-registration on the whole-slide images of HE staining, CD-31 staining, and D2-40 staining through a registration algorithm to achieve instance-level alignment; A region cropping module that takes the suspected vascular invasion region on the HE staining as a reference and correspondingly crops the same region in the registered CD-31 staining and D2-40 staining images to form tri-modal instance-level aligned tiles; A classification module that inputs the tri-modal instance-level aligned tiles into a mixture-of-experts network enhanced by a pathological large model, dynamically selects expert pathways through a gating activation network, and fuses multi-modal features for true / false positive discrimination.

[0014] In a third aspect, the present invention provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the above-mentioned instance segmentation method for analyzing vascular invasion in early-stage colorectal cancer.

[0015] In a fourth aspect, the present invention provides a readable storage medium, in which a computer program is stored. The computer program includes program codes for controlling a process to execute the process, and the process includes the above-mentioned instance segmentation method for analyzing vascular invasion in early-stage colorectal cancer.

[0016] The main contributions and innovations of the present invention are as follows: 1. Full-process automation and high efficiency: An end-to-end automated analysis framework is constructed, which requires no manual intervention from pathological image preprocessing, instance segmentation to multi-modal classification decision-making. It solves the problems of observer differences caused by the subjective interpretation of pathologists in traditional methods (such as the κ value is only 0.518 in international research) and low diagnostic efficiency, and meets the needs of high-throughput pathological image analysis.

[0017] 2. High-sensitivity instance segmentation and improved detection rate: The Mask2former instance segmentation network is combined with a multi-scale overlapping cropping strategy. Through low-confidence threshold inference and the Transformer multi-head attention mechanism, high-recall detection of tiny and irregular vascular invasion structures is achieved. Experimental data show that compared with existing methods (such as MaskRCNN+r101 Recall=89.1%), the Recall of the present invention is increased to 96.3%, significantly reducing missed diagnoses.

[0018] 3. Multi-modal joint diagnosis reduces false positives: The three staining modalities of HE, CD-31, and D2-40 are fused. Through whole-slide registration and a mixture-of-experts network enhanced by a pathological large model, multi-modal features are dynamically fused for true / false positive discrimination. Compared with single HE staining detection (FPR=0.542), multi-modal joint reduces the false positive rate to 38.8%, solving the misjudgment problem caused by the blurred boundary between blood vessels / lymphatic vessels in HE staining.

[0019] 4. Post-processing optimization improves segmentation accuracy: Based on an efficient union-find algorithm, overlapping masks are merged to reduce redundant detections; through the ROI expansion algorithm, recursively complete incomplete edge instances to ensure the complete detection of complex morphological vascular invasions, avoid contour missing caused by slice cropping, and improve the integrity and accuracy of the segmentation results.

[0020] 5. Multi-scale adaptability and robustness: Aiming at the large difference in the size of vascular invasions, through 2.5X-10X multi-magnification cropping and ResNet-101 multi-scale feature fusion, both small-scale fine structures and large-scale tissue backgrounds are taken into account, enhancing the robustness of the model to different morphological vascular invasions and breaking through the dependence of traditional methods on a single resolution or modality.

[0021] Details of one or more embodiments of the present invention are set forth in the following drawings and description to make other features, objects, and advantages of the present invention more concise and understandable. Description of the Drawings

[0022] The accompanying drawings described herein are used to provide a further understanding of the present invention and form a part of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings: Figure 1 is a flowchart of an instance segmentation method for analyzing vascular invasion in early-stage colorectal cancer according to an embodiment of the present invention; Figure 2 is a flowchart of a method for post-processing ROI expansion of an instance segmentation network according to an embodiment of the present invention; Figure 3 is a schematic diagram of the registration result of multi-stained slices for constructing a multi-modal hybrid expert classification network according to an embodiment of the present invention; Figure 4 is an architecture diagram of a multi-modal hybrid expert network enhanced by a pathological large model according to an embodiment of the present invention; Figure 5 is a diagram of the instance segmentation result of a suspected vascular invasion under HE staining by a Mask2former instance segmentation model according to an embodiment of the present invention; Figure 6 is a schematic diagram of a tile cropped from the corresponding region of an instance segmentation network under HE staining after post-processing on CD-31 and D2-40 staining according to an embodiment of the present invention; Figure 7 is a schematic diagram of the hardware structure of an electronic device according to an embodiment of the present invention. Detailed implementation manners

[0023] Here, the exemplary embodiments will be described in detail, and the examples are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all the implementation manners consistent with one or more embodiments of this specification. On the contrary, they are only examples of devices and methods consistent with some aspects of one or more embodiments of this specification as detailed in the appended claims.

[0024] It should be noted that: in other embodiments, the steps of the corresponding methods are not necessarily executed in the order shown and described in this specification. In some other embodiments, the steps included in the method may be more or less than those described in this specification. In addition, a single step described in this specification may be decomposed into multiple steps for description in other embodiments; and multiple steps described in this specification may also be combined into a single step for description in other embodiments.

[0025] Existing early colorectal cancer vascular invasion detection technologies lack a fully automated process, suffer from observer differences, and have problems such as long diagnostic times for pathologists. Therefore, there is an urgent need for highly sensitive automated vascular invasion detection technologies.

[0026] Based on this, the present invention is based on the Mask2former network and a multi-modal hybrid expert network enhanced by a pathological large model to solve the problems existing in the prior art.

[0027] Example 1 The present invention aims to propose an instance segmentation method for analyzing early colorectal cancer vascular invasion phenomena. Specifically, a workstation configured with an RTX3090 GPU or a more advanced model is used for model inference and analysis to ensure that the computing resources meet the requirements of large-scale instance segmentation and multi-modal joint classification. Refer to Figure 1 , the method includes: S00, Data preprocessing The digital pathology images are generated into multi-scale image patches adapted to the input of the deep learning model by overlapping cropping at multiple magnifications; specifically as follows: 1) Sample preparation: Collect high-resolution colorectal cancer digital pathology images at 80X as the input data for early colorectal cancer vascular invasion analysis.

[0028] 2) Image cropping: Under the multi-magnification conditions of 2.5X, 5X, and 10X, the entire image is segmented into 1024×1024 pixel image patches using the overlapping cropping technique, with an overlapping size of 512 pixels, and standardized processing is performed to adapt to the input requirements of the Mask2former segmentation model.

[0029] S10, Vascular invasion instance segmentation with a low confidence threshold: The Mask2former instance segmentation network is used to perform instance segmentation of vascular invasion on the multi-scale image patches. The Mask2former instance segmentation network uses ResNet-101 as the backbone network, extracts multi-scale feature maps at four levels and inputs them into the segmentation decoding network, and performs instance prediction through a preset number of query Tokens and the multi-head attention mechanism; specifically as follows: 1) Instance segmentation execution: The instance segmentation algorithm can obtain the number and contour of vascular invasion, providing instance-level information for subsequent multi-modal classification. The pre-trained ResNet-101 is used as the front backbone network of the vascular invasion instance segmentation network, and the overall instance segmentation network uses the Mask2former architecture. The pathological image patches containing vascular invasion are first encoded into low-resolution feature vectors by the feature extractor ResNet-101, and then input into the pixel decoder. The pixel decoder is a fully connected neural network that can decode the feature vectors into pixel-level prediction results, which can better retain image details compared with the traditional decoder that decodes to the original image size. The pixel-level decoder decodes the features to 4 different scales, and then the feature maps of the first 3 scales are fed into the Transformer decoder for decoding. The Transformer decoder is a decoding module stacked by masked attention and self-attention. The formula of the self-attention mechanism is: (1) X l represents the output of the l -th layer. Q l , K l , V l represent the query, key, and value matrices respectively, which are usually linear transformations of the input of the current layer. is the transpose of the key matrix. Softmax is a normalization function used to convert the attention scores into a probability distribution. represents the input of the l -th layer, which is the output of the previous layer. This formula describes the calculation process of the self-attention layer, where softmax( ) calculates the attention weights, and then these weights are used to weighted sum the value matrix V l . Finally, the result of this weighted sum is added to the input to achieve the residual connection. The residual connection helps to alleviate the vanishing gradient problem in deep networks and promotes the flow of information. In practical applications, this formula usually comes with some scaling factors, such as d k (where d k is the dimension of the key vector) to prevent the input of the softmax function from being too large and causing the gradient to vanish. The formula of the masked attention mechanism is: (2) (3) In masked cross-attention, the feature map serves as the Value vector and the Key vector, and the target query vector serves as the Query vector. Through the masking operation, the cross-attention can be restricted to the foreground area of the prediction mask rather than the entire feature map. After fusion through cross-attention, the self-attention mechanism is used to capture global dependencies. The Transformer decoder achieves multi-scale object detection by processing feature maps of different resolutions. In addition to multi-scale object detection at a fixed magnification, the present invention also enhances the multi-scale adaptability of the model over a larger range by inputting images of different magnifications into the network. The instance segmentation framework uses the cross-entropy loss function and the Dice loss function during training. During inference, multi-scale 1024x1024 tiles (with an overlapping area of 512) at 2.5X, 5X, and 10X are input for multi-scale inference.

[0030] S20. Post-processing of the instance segmentation result of vascular invasion: Based on the efficient union-find algorithm, the segmentation masks in the full field of view are merged. The masks with intersections are grouped into the same connected component and the maximum confidence score is retained. Then, through the ROI expansion algorithm, the instances with incomplete edges are recursively expanded and inferred until there are no new incomplete instances. Specifically as follows: 1) Post-processing and merging of the instance segmentation result based on the union-find: Under the conditions of multi-scale inference and overlapping inference in step S10, many overlapping regions will be generated in the full field of view of the pathological image, which will affect the subsequent judgment of the pathologist.

[0031] The present invention relates to a union-find algorithm for pathological image analysis. The algorithm aims to solve the problem that a large number of overlapping regions generated in the full field of view of the pathological image under the conditions of multi-scale inference and overlapping inference affect the subsequent judgment of the pathologist. Due to the huge size of the pathological image, a large number of mask regions exist in the inference result. The present invention proposes to use an efficient union-find algorithm to merge these mask regions, object detection boxes, and confidence scores.

[0032] In the specific implementation process, first, a unique identifier is assigned to each mask, object detection box, and the corresponding confidence score, and the union-find is initialized so that each element exists independently as its own root node at the beginning. Subsequently, by performing the find operation, the connected components to which each element belongs are determined. If it is found that two masks have an intersection ( IOU > T m), the masked regions are grouped into the same connected component through the merging operation, and the maximum value of the confidence scores in the two masks or object detection boxes is selected as the confidence of the merged mask or object detection box during the merging. Due to the huge scale difference of vascular invasion, large-object vascular invasion will obtain a high confidence during inference at low magnification, while small-object vascular invasion will obtain a high confidence during inference at high magnification. Therefore, selecting the maximum value in the confidence scores as the confidence of the merged suspected vascular invasion instance when merging overlapping vascular regions is consistent with the confidence distribution characteristics of multi-scale inference.

[0033] where T m is the threshold, which is an adjustable hyperparameter during post-processing and is set to 0 in this application case. However, for some specific scenarios, such as the phenomenon of vascular invasion in a dense distribution form, T m the value can be set to a value greater than 0 to avoid merging different vascular invasion instances. In addition, to further improve the algorithm efficiency, the present invention also adopts an optimization strategy of path compression and union by rank to reduce the complexity of the lookup and merging operations, ensuring that the algorithm can efficiently and accurately process the overlapping regions in the pathological images, thereby providing clearer and more accurate analysis results for pathologists.

[0034] 2) ROI expansion mask completion based on: Referring to Figure 2 , the present invention provides a method for ROI expansion mask completion to address the problem of incomplete-edge instances that may occur when cropping image patches from whole-slide pathological images for inference. Under the conditions of multi-scale inference and overlapping inference, many overlapping regions will be generated in the whole slide of the pathological image, and these overlapping regions may affect the subsequent judgment of pathologists. Therefore, the present invention first detects the incomplete edges, and then expands the regions outside the incomplete edges to obtain new image patches for inference until complete edges are obtained. The ROI expansion algorithm is based on an assumption: for the same vascular invasion instance, the inference under the complete field of view will obtain a higher confidence than that under the local field of view. Therefore, for the predicted confidence C e after ROI expansion, if it is greater than or equal to the original predicted confidence C o , then the new prediction result is used to replace the old prediction result. If C e is less than C o This process ensures the accuracy of the inference result, avoids the influence of incomplete-edge instances on pathological diagnosis, and thus improves the reliability and effectiveness of pathological image analysis. The determination condition for incomplete edges is the number of consecutive overlapping points of the straight line corresponding to the detection box in the instance segmentation mask contour No Ratio with the total number of straight-line points of the detection box N a of R e ( N o / N a )is greater than T e 。 where T e is the threshold value, an adjustable hyperparameter in the post-processing process, which is set to 0.2 in this application case.

[0035] S30. Rigid registration: Perform rigid registration, non-rigid registration, and micro-registration on the whole-slide images of HE staining, CD-31 staining, and D2-40 staining through the registration algorithm to achieve instance-level alignment; specifically as follows: In order to align the pathological sections of different staining modalities, the present invention uses an image registration algorithm at the whole-slide level to perform whole-slide level registration on HE, CD-31, and D2-40 sections, and the registration process is independent of the steps of S00, S10, and S20.

[0036] Figure 3 shows the alignment results of multi-stained sections after the registration algorithm. The present invention uses the HE section as the reference image, and aligns the CD-31 and D2-40 sections with the HE image respectively. First, use the foreground segmentation algorithm to filter the background blank area of the WSI image. This process generates a mask by calculating the difference between the color of each pixel and the background color. The specific steps are as follows: First, convert the image to the CAM16-UCS color space to obtain three channels: L (luminance), A, and B. For bright-field images, assuming the background is bright, the background color is defined as the average LAB value of the pixels whose luminance is higher than 99% of all pixels.

[0037] Next, calculate the Euclidean distance between the LAB color of each pixel and the background LAB color to generate a new image D, where larger values indicate greater differences between the pixel color and the background color.

[0038] Then, apply Otsu threshold segmentation to the image D. Pixels above this threshold are identified as the foreground, thereby generating a binary mask. Subsequently, standardize the image and detect the features in each image.

[0039] Among them, the standardization operation is mainly to make the images to be registered as similar as possible. For immunofluorescence (IF) images, the DAPI channel is the best choice for registration. However, if immunohistochemistry (IHC) images are used, standardized preprocessing is required to make them look more similar. The specific steps are as follows: Convert the RGB image to the CAM16-UCS color space, set the parameters and then convert back to RGB, and then convert to grayscale and invert it to make the background darker and the tissue brighter. This method is more accurate and robust than simple grayscale conversion or histogram equalization.

[0040] After preprocessing, all images (IHC and / or IF) will be further normalized to make the pixel intensity distributions more similar. By calculating the 5th percentile, average value, and 95th percentile of the pixel values, three interpolation fittings are performed, and the total variation (TV) denoising is used to slightly smooth the image, retain the edges, and reduce noise interference.

[0041] Subsequently, clustering and sorting methods are used to align the features in the reference image, and a preliminary rigid registration transformation is calculated. The rigid registration process uses the BRISK detector and the VGG descriptor to extract features, and most of the outliers are removed through brute-force matching and RANSAC.

[0042] To further improve the accuracy, the Tukey box plot method is used to filter out the mis-matched points. Subsequently, the registration algorithm constructs a similarity matrix based on the number of feature matches, and infers the optimal image order through hierarchical clustering, so that each image is surrounded by its two most similar images. After determining the image order, the images are sequentially aligned to the reference image (such as the middle image) to form a continuous registration sequence. To ensure the stability of the tissue structure matching, the algorithm introduces neighborhood matching filtering, and only retains the feature points shared within the local image window.

[0043] Finally, mutual information can be selectively combined to optimize the alignment accuracy. After rigid registration, non-rigid registration is continued at a higher resolution to correct some morphological changes and distortions of the tissue pathological structure. Finally, micro-registration is performed on the basis of non-rigid registration. Micro-registration also belongs to non-rigid registration and aims to correct the minor deformations and distortions at a higher resolution. After passing through the registration network, the aligned WSI of different staining modalities can achieve instance-level alignment of vascular invasion, laying the foundation for the subsequent three-modal true and false vascular classification network.

[0044] S40. Multi-staining cropping of suspected vascular invasion areas: Based on the suspected vascular invasion area in the HE staining, the same area is correspondingly cropped in the registered CD-31 staining and D2-40 staining images to form a tile with instance-level alignment of three modalities; specifically as follows: Figure 5The Mask2former instance segmentation model is shown for the instance segmentation results of suspected vascular invasion under HE staining. Based on the registration results at the whole-slide level in step S30, multi-stained paired regions of suspected vascular invasion are obtained by cropping according to the corresponding coordinates of suspected vascular invasion.

[0045] Since the registration process aligns the spatial positions of slices with different stainings, the coordinates of suspected vascular invasion under HE staining correspond to regions cropped under immunohistochemical staining with the same anatomical scope and spatial position. As Figure 6 Shown is a schematic diagram of the tiles cropped from the corresponding regions of the instance segmentation network under HE staining after post-processing on CD-31 and D2-40 stainings. It can be seen in the figure that the alignment of multi-stained vascular invasion regions at the tile level is in good condition, providing a basis for the construction of subsequent multi-modal classification models.

[0046] S50. Pathological large model enhanced multi-modal mixture of experts classification framework: Input the tiles with three-modal instance-level alignment into the mixture of experts network enhanced by the pathological large model, dynamically select the expert pathways through the gated activation network, and fuse multi-modal features for true / false positive discrimination; specifically as follows: Figure 4 Shows the architecture of the multi-modal mixture of experts network enhanced by the pathological large model. The general pathological large model has been self-supervised pre-trained on a large-scale pathological image dataset, so it has obtained the ability to represent general pathological features. Based on the capabilities of the pre-trained pathological basic model, more accurate identification of true / false positive vascular invasion can be achieved. Since the scales of vascular invasion itself vary greatly, the scales of the cropped tiles of suspected vascular invasion regions are not uniform. Therefore, in the present invention, the size of the suspected vascular invasion tiles is limited to the range of 224x224 to 1024x1024 through the method of adaptive hierarchical selection of the pathological image pyramid.

[0047] For the tiles of vascular invasion with different stainings, the present invention uses a pathological basic model with non-shared weights for pathological feature extraction. In order to enable the pathological basic model to have a better perception of the feature space of vascular invasion data, the parameters of the first N Transformer modules in the pathological basic model (VIT architecture) are frozen during the training stage, and only the last L-N Transformer modules are given the ability to update parameters, enabling it to achieve domain adaptation during the training process, where L is the number of layers of the Transformer in the VIT model. The suspected vascular invasion tiles of CD-31, D2-40, and HE stainings obtain corresponding feature vectors through the pathological basic model E CD-31 , E D2-40 andE HE , the encoding dimension of the feature vector is D , D dependent on the encoding dimension of the pathological large model. The feature vectors of different stains first pass through a gated activation network with shared weights to obtain the activation probabilities for different expert paths. The number of experts in the mixture-of-experts model is set to N e , and the number of activated experts in the single-stain modality is N s ( N s <=N e ). First, the features of the same stain from different experts are feature-fused, and then the mixture-of-experts features of different stains are feature-fused. The expert network uses a feed-forward neural network with an input dimension of D i , a hidden layer dimension of D h , and an output layer dimension of D o . The random dropout regularization technique is used during the training process to prevent overfitting.

[0048] 1) Model training parameter table: Table 1

[0049] 2) Training dataset construction table: Table 2

[0050] 3) Experimental verification: As shown in Table 3, the Mask2former vascular invasion preliminary screening network based on multi-scale inference has an extremely high vascular invasion detection rate.

[0051] Table 3

[0052] Note: ms stands for multi-scale inference (multi-scale) As shown in Table 4, multi-stain joint prediction can reduce the detection rate of false positives for vascular invasion.

[0053] Table 4

[0054] 4) Data saving and visualization analysis: After the true / false positive vasculature inference in step S30, by setting an appropriate confidence threshold, a large number of false positives (blood vessels, glands, etc.) obtained from the mask2former model inference in step S10 can be excluded. Subsequently, all the results of a single pathological section are stored in the pickle file format, which contains the confidence, object detection bounding box, and segmentation mask for each instance of vascular invasion. The present invention provides a format conversion tool that can convert the pickle results into a format compatible with the Qupath engineering software for visual analysis.

[0055] Through the above steps, the present invention realizes the fully automated analysis of vascular invasion. The complete process from suspicious vasculature extraction to multi-modal classification ensures sensitivity and accuracy, providing reliable data support and technical tools for pathological research and diagnosis.

[0056] The present invention is applicable to various pathological laboratories, hospitals, and research institutions, and has broad application prospects in the following fields: 1. Vascular invasion diagnosis: By performing high-sensitivity instance segmentation on the vascular invasion of early-stage colorectal cancer and jointly optimizing the false positive screening through multi-modalities, it provides a reliable basis for the early diagnosis and treatment plan formulation of colorectal cancer.

[0057] 2. Multi-organ generalization: The vascular invasion phenomena in different tissues and organs are similar. Therefore, the vascular invasion detection algorithm can be fine-tuned and migrated to multiple organs for application.

[0058] 3. Pathological big data processing: It has the ability to process a large amount of high-resolution pathological images, providing a solid technical foundation for artificial intelligence research and big data analysis in pathology.

[0059] 4. Basic medical research: It provides an efficient tool for colorectal pathology and ultrastructural biology research, promoting the in-depth study of the pathogenesis of basement membrane-related diseases.

[0060] 5. Education and training: It provides an automated auxiliary tool for pathology teaching and training, reducing the learning difficulty of complex microscopic operations and improving teaching efficiency.

[0061] The present invention takes efficient and accurate automated technology as the core, is widely applicable to various application scenarios in the medical and scientific research fields, and has important practical value and promotion prospects.

[0062] Embodiment 2 Based on the same concept, the present invention also proposes an instance segmentation device for analyzing the vascular invasion phenomenon of early-stage colorectal cancer, including: A preprocessing module that generates multi-scale image patches adapted to the input of the deep learning model by overlapping and cropping digital pathological images at multiple magnifications; The instance segmentation module uses the Mask2former instance segmentation network to perform instance segmentation of vascular invasion on multi-scale image patches. The Mask2former instance segmentation network uses ResNet-101 as the backbone network, extracts multi-scale feature maps of four levels and inputs them into the segmentation decoding network, and performs instance prediction through a preset number of query tokens and the multi-head attention mechanism; The post-processing module merges the segmentation masks within the full field of view based on the efficient union-find algorithm, classifies the masks with intersections into the same connected component and retains the maximum confidence score, and then recursively expands and infers the instances with incomplete edges through the ROI expansion algorithm until there are no new incomplete instances; The registration module performs rigid registration, non-rigid registration, and micro-registration on the whole-slide images of HE staining, CD-31 staining, and D2-40 staining through the registration algorithm to achieve instance-level alignment; The region cropping module takes the suspected vascular invasion region on the HE staining as a reference, and correspondingly crops the same region in the registered CD-31 staining and D2-40 staining images to form a tile with three-modal instance-level alignment; The classification module inputs the tiles with three-modal instance-level alignment into the mixture-of-experts network enhanced by the pathological large model, dynamically selects the expert path through the gated activation network, and fuses multi-modal features to distinguish true and false positives.

[0063] Embodiment III This embodiment also provides an electronic device, refer to Figure 7 , including a memory 404 and a processor 402. The memory 404 stores a computer program, and the processor 402 is configured to run the computer program to execute the steps in any one of the above method embodiments.

[0064] Specifically, the above-mentioned processor 402 may include a central processing unit (CPU), or a specific integrated circuit (Application Specific Integrated Circuit, abbreviated as ASIC), or may be configured as one or more integrated circuits implementing the embodiments of the present invention.

[0065] Among them, the memory 404 may include a mass storage 404 for data or instructions. By way of example and not limitation, the memory 404 may include a hard disk drive (HDD), a floppy disk drive, a solid state drive (SSD), a flash memory, an optical disc, a magneto-optical disc, a magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. Where appropriate, the memory 404 may include removable or non-removable (or fixed) media. Where appropriate, the memory 404 may be internal or external to the data processing device. In a particular embodiment, the memory 404 is non-volatile memory. In a particular embodiment, the memory 404 includes a read-only memory (ROM) and a random access memory (RAM). Where appropriate, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically alterable ROM (EAROM), or a flash memory, or a combination of two or more of these. Where appropriate, the RAM may be a static random access memory (SRAM) or a dynamic random access memory (DRAM), where the DRAM may be a fast page mode dynamic random access memory (FPMDRAM), an extended data output dynamic random access memory (EDODRAM), a synchronous dynamic random access memory (SDRAM), etc.

[0066] The memory 404 can be used to store or cache various data files required for processing and / or communication, as well as possible computer program instructions executed by the processor 402.

[0067] By reading and executing the computer program instructions stored in the memory 404, the processor 402 implements any one of the instance segmentation methods for the analysis of early colorectal cancer vascular invasion phenomena in the above embodiments.

[0068] Optionally, the above electronic device may further include a transmission device 406 and an input / output device 408. Among them, the transmission device 406 is connected to the above processor 402, and the input / output device 408 is connected to the above processor 402.

[0069] The transmission device 406 can be used to receive or send data via a network. Specific examples of the above network may include wired or wireless networks provided by the communication provider of the electronic device. In one example, the transmission device includes a Network Interface Controller (NIC), which can be connected to other network devices through a base station and thus communicate with the Internet. In one example, the transmission device 406 can be a Radio Frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0070] The input / output device 408 is used to input or output information.

[0071] Embodiment 4 This embodiment also provides a readable storage medium, in which a computer program is stored. The computer program includes program codes for controlling a process to execute the process, and the process includes the instance segmentation method for the analysis of early colorectal cancer vascular invasion phenomena according to Embodiment 1.

[0072] It should be noted that the specific examples in this embodiment can refer to the examples described in the above embodiments and optional implementation manners, and will not be elaborated here.

[0073] Generally, various embodiments can be implemented in hardware or special circuits, software, logic, or any combination thereof. Some aspects of the present invention can be implemented in hardware, while other aspects can be implemented by firmware or software executed by a controller, a microprocessor, or other computing devices, but the present invention is not limited thereto. Although the various aspects of the present invention can be shown and described as block diagrams, flowcharts, or using some other graphical representation, it should be understood that, as a non-limiting example, the blocks, devices, systems, technologies, or methods described herein can be implemented in hardware, software, firmware, special circuits or logic, general hardware or a controller, or other computing devices, or some combination thereof.

[0074] Embodiments of the present invention can be implemented by computer software, which can be executed by a data processor of a mobile device, such as in a processor entity, or implemented by hardware, or implemented by a combination of software and hardware. A computer software or program (also referred to as a program product), including software routines, applets, and / or macros, can be stored in any device-readable data storage medium, and they include program instructions for performing specific tasks. The computer program product can include one or more computer-executable components configured to execute the embodiments when the program runs. The one or more computer-executable components can be at least one software code or a part thereof. Additionally, at this point, it should be noted that any box in the logical flow, such as Figure 1 in, can represent a program step, or interconnected logic circuits, boxes, and functions, or a combination of program steps and logic circuits, boxes, and functions. The software can be stored on physical media such as memory chips or storage blocks implemented within a processor, magnetic media such as hard disks or floppy disks, and optical media such as, for example, DVDs and their data variants, CDs. The physical media is a non-transitory medium.

[0075] Those skilled in the art should understand that the technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as within the scope described in this specification.

[0076] The above embodiments only represent several implementation manners of the present invention. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the appended claims.

Claims

1. An instance segmentation method for analyzing vascular invasion in early-stage colorectal cancer, characterized in that, Including the following steps: S00. Generating multi-scale image patches adapted to the input of the deep learning model by overlapping cropping of the digital pathology image at multiple magnifications; S10. Using the Mask2former instance segmentation network to perform instance segmentation of vascular invasion on the multi-scale image patches. The Mask2former instance segmentation network uses ResNet-101 as the backbone network, extracts multi-scale feature maps of four levels and inputs them into the segmentation decoding network, and performs instance prediction through a preset number of query Tokens and the multi-head attention mechanism; S20. Merging the segmentation masks within the full field of view based on the efficient union-find algorithm, classifying the masks with intersections into the same connected component and retaining the maximum confidence score, and then recursively expanding and inferring the instances with incomplete edges through the ROI expansion algorithm until there are no new incomplete instances; S30. Performing rigid registration, non-rigid registration, and micro-registration on the whole-slide images stained with HE, CD-31, and D2-40 through the registration algorithm to achieve instance-level alignment; S40. Taking the suspected vascular invasion area on the HE staining as a reference, cropping the corresponding same area in the registered CD-31 staining and D2-40 staining images to form a three-modal instance-level aligned tile; S50. Inputting the three-modal instance-level aligned tiles into the hybrid expert network enhanced by the pathology large model, dynamically selecting the expert path through the gated activation network, and fusing multi-modal features to distinguish true and false positives.

2. The instance segmentation method for analyzing vascular invasion in early-stage colorectal cancer according to claim 1, wherein In step S00, the multiple magnifications include 2.5X, 5X, and 10X, the resolution of the image patch is 1024×1024, and the overlapping area is 512 pixels.

3. An instance segmentation method for analyzing vascular invasion in early-stage colorectal cancer as described in claim 1, characterized in that, In step S10, the Mask2former instance segmentation network uses a low confidence threshold during inference to retain more suspected vascular invasion areas.

4. An instance segmentation method for analyzing vascular invasion in early colorectal cancer according to claim 1, characterized in that, In step S20, the union-find algorithm processes the overlapping masks through the path compression and union by rank optimization strategies, and the merging threshold Tm is 0; the ROI expansion algorithm identifies incomplete edges by judging that the ratio Re of the number of continuous overlapping points of the instance segmentation mask contour to the total number of points of the detection frame line is greater than the threshold Te.

5. The instance segmentation method for analyzing the vascular invasion phenomenon of early-stage colorectal cancer according to claim 1, wherein In step S30, taking the HE staining image as the reference image, first filtering the background through foreground segmentation, and then performing rigid registration, non-rigid registration, and micro-registration in sequence to achieve whole-slide-level alignment of different staining modalities.

6. The instance segmentation method for analyzing the vascular invasion phenomenon of early-stage colorectal cancer according to claim 1, wherein, In step S50, the pathology large model is a pre-trained VIT-Large model. The parameters of the first 23 Transformer modules are frozen, and only the parameters of the last L-23 layers are updated to extract feature vectors for the three-modal tiles respectively, where L is the total number of layers of the VIT model.

7. An instance segmentation method for analyzing vascular invasion in early-stage colorectal cancer according to any one of claims 1-6, characterized in that, The hybrid expert network contains 4 experts, each staining modality activates 2 experts, calculates the activation probability of the expert path through the weight-sharing gated activation network, and uses additive aggregation for single-staining expert features and concatenation aggregation for multi-staining features.

8. An instance segmentation device for analyzing the phenomenon of vascular invasion in early-stage colorectal cancer, characterized in that, Including: A preprocessing module that generates multi-scale image patches adapted to the input of the deep learning model by overlapping cropping of the digital pathology image at multiple magnifications; The instance segmentation module uses the Mask2former instance segmentation network to perform instance segmentation of vascular invasion on the multi-scale image patches. The Mask2former instance segmentation network uses ResNet-101 as the backbone network, extracts multi-scale feature maps of four levels and inputs them into the segmentation decoding network, and performs instance prediction through a preset number of query Tokens and the multi-head attention mechanism; The post-processing module merges the segmentation masks within the full field of view based on the efficient union-find algorithm, classifies the masks with intersections into the same connected component and retains the maximum confidence score, and then recursively expands and infers the instances with incomplete edges through the ROI expansion algorithm until there are no new incomplete instances; The registration module performs rigid registration, non-rigid registration, and micro-registration on the whole-slide images stained with HE, CD-31, and D2-40 through the registration algorithm to achieve instance-level alignment; The region cropping module takes the suspected vascular invasion region on the HE stain as a reference, and correspondingly crops the same region in the registered CD-31 stain and D2-40 stain images to form a three-modal instance-level aligned tile; The classification module inputs the three-modal instance-level aligned tiles into the mixture-of-experts network enhanced by the pathological large model, dynamically selects the expert path through the gated activation network, and fuses multi-modal features to distinguish true and false positives.

9. An electronic device, comprising a memory and a processor, characterized in that, The memory stores a computer program, and the processor is set to run the computer program to execute the instance segmentation method for analyzing the phenomenon of vascular invasion in early-stage colorectal cancer according to any one of claims 1 to 7.

10. A readable storage medium, characterized in that, The readable storage medium stores a computer program, and the computer program includes program codes for controlling a process to execute the process, and the process includes the instance segmentation method for analyzing the phenomenon of vascular invasion in early-stage colorectal cancer according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Multi-dyed image accurate matching method and device based on same wax block and application of multi-dyed image accurate matching method and device

    CN118691655A

  • Pathological image segmentation method and device based on deep learning and readable storage medium thereof

    CN120031899A

Cited By

  • Automatic segmentation system for colorectal tumor CT image

    CN121073823A

  • An automatic segmentation system for colorectal tumor CT images

    CN121073823B

  • Tumor pathological image classification method and system based on artificial intelligence

    CN121170436A