A pathological image anomaly detection system based on a double-branch multi-scale inverse distillation
Patent Information
- Application Number
- CN202611216853.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-12
- Publication Date
- 2026-09-25
AI Technical Summary
然而,这类方法主要依赖局部图像块信息,难以充分利用全切片层面的组织结构和空间上下文信息,容易造成局部异常判断与全局异常状态不一致
[0099]1、本发明采用基于双分支多尺度逆向蒸馏的病理图像异常检测系统,在减少病灶级人工标注依赖的情况下实现病理图像异常检测,降低了模型训练对精细标注数据的需求。
Smart Images

Figure CN122822246A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the technical field of pathological image anomaly detection, and in particular to a pathological image anomaly detection system based on dual-branch multi-scale reverse distillation. Background Technology
[0002] With the development of digital pathology technology, pathological images are widely used in disease screening, assisted diagnosis, and pathological research. Pathological images are typically characterized by high resolution, large size, and complex tissue structures. Abnormal areas may only occupy a small portion of the entire image. Therefore, accurately identifying abnormal areas and determining whether the entire slide is abnormal from large-scale pathological images is a crucial issue in intelligent digital pathology analysis.
[0003] Existing methods for detecting abnormalities in pathological images largely rely on manual annotation at the lesion or region level. However, this type of annotation typically requires professional pathologists to complete it region by region, resulting in high annotation costs, long processing times, and potential subjective differences between different doctors. To reduce reliance on annotation, existing methods usually divide the whole pathological slide image into multiple image blocks and calculate the abnormality score for each block separately. Then, they use methods such as averaging, maximum value, or Top-k aggregation to obtain the abnormality judgment result for the entire slide. However, these methods mainly rely on local image block information and cannot fully utilize the tissue structure and spatial context information at the whole slide level, which can easily lead to inconsistencies between local abnormality judgments and global abnormality states.
[0004] Furthermore, existing anomaly detection methods based on reconstruction or feature distillation typically focus on single-scale feature modeling, making it difficult to simultaneously consider both fine-grained pathological morphological information at the image patch level and macroscopic tissue structure information at the whole-slice level. When pathological abnormalities manifest as cross-regional distribution, altered tissue structure, or weak local differences, relying solely on image patch features or simple score aggregation methods is insufficient to obtain stable and reliable detection results. Therefore, there is an urgent need for a method that can simultaneously model local image patch information and global information across the whole slice, and achieve pathological image anomaly detection with minimal manual annotation. Summary of the Invention
[0005] The purpose of this invention is to overcome the shortcomings and deficiencies of the prior art and propose a pathological image anomaly detection system based on dual-branch multi-scale reverse distillation. This system can reduce the dependence on manual annotation at the lesion level, improve the accuracy of pathological image anomaly detection, strengthen the collaborative modeling of local image block features and global pathological image features, and simultaneously realize anomaly judgment and anomaly region localization.
[0006] To achieve the above objectives, the technical solution provided by this invention is: a pathological image anomaly detection system based on dual-branch multi-scale reverse distillation, comprising:
[0007] The data acquisition and preprocessing module is used to acquire and preprocess whole pathological slide images to obtain multiple image blocks and the position information of each image block in the whole pathological slide image.
[0008] The feature extraction module includes a pathological visual language basic model with fixed parameters and a pathological whole-slice multimodal basic model with fixed parameters. The pathological visual language basic model is used to extract features from each image block to obtain corresponding image block-level features, and to form an image block-level feature set from all image block-level features. The pathological whole-slice multimodal basic model is used to extract whole-slice-level features based on the image block-level feature set and the position information of each image block in the pathological whole-slice image to obtain whole-slice-level features.
[0009] The feature fusion module performs multi-scale feature fusion between the image block-level features and the corresponding full-slice-level features to obtain joint multi-scale features.
[0010] The feature transformation module inputs the joint multi-scale features into the shared bottleneck layer and performs feature transformation to obtain the bottleneck layer output features.
[0011] The feature reconstruction module includes an asymmetric dual-branch decoder, which comprises an image patch decoder, a full slice decoder, and a cross-attention layer. The feature reconstruction module inputs the output features of the bottleneck layer into the asymmetric dual-branch decoder, forming image patch branch representations and full slice branch representations. The cross-attention layer facilitates feature interaction between the image patch branch representations and the full slice branch representations. The image patch decoder reconstructs image patch-level features to obtain image patch-level reconstructed features; the full slice decoder reconstructs full slice-level features to obtain full slice-level reconstructed features.
[0012] The training module trains the shared bottleneck layer and the asymmetric dual-branch decoder based on the image block-level reconstruction error between the image block-level features and the image block-level reconstruction features, and the full slice-level reconstruction error between the full slice-level features and the full slice-level reconstruction features.
[0013] The inference module, based on the trained shared bottleneck layer and the asymmetric dual-branch decoder, obtains the image block-level reconstruction features and the full-slice-level reconstruction features of the pathological whole-slice image to be tested, and then obtains the image block-level reconstruction error and the full-slice-level reconstruction error of the pathological whole-slice image to be tested. Based on the image block-level reconstruction error, a local anomaly heatmap is generated, and based on the full-slice-level reconstruction error, the anomaly detection result of the pathological whole-slice image is obtained.
[0014] Furthermore, the data acquisition and preprocessing module includes the following steps:
[0015] 1) Obtain whole pathological slide images and the pathological whole slide images Perform downsampling processing to obtain a downsampled image. ,in, Indicates the downsampling factor;
[0016] 2) For the downsampled image Perform organizational region segmentation to obtain the organizational region mask. ,in, Used to distinguish between the tissue area and the background area;
[0017] 3) Based on the tissue region mask Remove the whole pathological slide image The background area retains the effective organizational area;
[0018] 4) According to the preset image block size The effective tissue region is divided into image blocks to obtain an image block set. ,in, Indicates the first Image blocks, Indicates the total number of image patches. and These represent the height and width of the image block, respectively. ;
[0019] 5) Record the first Image blocks The whole pathological slide image Location information and obtain a set of location information. ,in, Indicates the first Location information of each image patch and Each represents an image block The whole pathological slide image The horizontal and vertical coordinates in the diagram.
[0020] Furthermore, the feature extraction module includes the following steps:
[0021] 1) The pathological visual language basic model uses a CONCH model with fixed parameters to analyze the image patches. Feature extraction is performed to obtain image block-level features. The CONCH model is used to extract cell morphology, local tissue structure, and pathological semantic information from image patches, and its expression is:
[0022] ;
[0023] In the formula, Indicates the first Image blocks, , Indicates the total number of image patches. This represents the model parameters of the CONCH model, and It remains constant during both the training and inference phases. Indicates the first Image blocks Corresponding image block-level features;
[0024] 2) Combine all image patch-level features into an image patch-level feature set. Its expression is:
[0025] ;
[0026] In the formula, This represents the pathological whole slide image. The image block-level feature set consists of all image block-level features;
[0027] 3) The multimodal basic model of the pathological whole slide uses the TITAN model with fixed parameters to analyze the image block-level feature set. and location information set Full-slice-level feature extraction was performed to obtain the pathological full-slice image. Corresponding full-slice level features The TITAN model is a pre-trained pathological whole-slice model used to model the macroscopic tissue structure and long-range spatial dependencies of pathological whole-slice images based on image block-level features and their location information. Its expression is:
[0028] ;
[0029] In the formula, This represents a set of location information composed of the location information of each image patch. Indicates the first Image blocks The whole pathological slide image Location information in This represents the model parameters of the TITAN model, and It remains constant during both the training and inference phases. The image represents the whole pathological slide. Corresponding full-slice level features;
[0030] 4) The image block-level feature set and the full-slice level features As output.
[0031] Furthermore, the feature fusion module includes the following steps:
[0032] 1) Obtain the image block-level feature set output by the feature extraction module. and full slice level features ;
[0033] 2) The image block-level feature set With the full slice level features Perform concatenated multi-scale feature fusion to obtain joint multi-scale features. Its expression is:
[0034] ;
[0035] In the formula, This indicates a feature concatenation operation. This represents the set of image block-level features composed of all image block-level features in the pathological whole-section image. This represents the whole-slice-level features corresponding to the pathological whole-slice image;
[0036] 3) Combine the joint multi-scale features As output, it is used as input for subsequent shared bottleneck layers.
[0037] Furthermore, the feature transformation module includes the following steps:
[0038] 1) Obtain the joint multi-scale features output by the feature fusion module. ;
[0039] 2) Combine the joint multi-scale features Input shared bottleneck layer Perform feature transformation to obtain the bottleneck layer output features. Its expression is:
[0040] ;
[0041] In the formula, Indicates a shared bottleneck layer. This indicates that the joint multi-scale features Through the shared bottleneck layer The output features of the bottleneck layer after transformation;
[0042] The shared bottleneck layer It includes a random deactivation layer, a first linear layer, a nonlinear activation function, a second linear layer, and an output random deactivation layer, used to process the joint multi-scale features. Feature perturbation and nonlinear mapping are performed to reduce the direct learning of identity mappings;
[0043] 3) Output features of the bottleneck layer As output, it is used as input to the subsequent asymmetric dual-branch decoder.
[0044] Furthermore, the feature reconstruction module includes the following steps:
[0045] 1) Obtain the bottleneck layer output features from the feature transformation module. ;
[0046] 2) Output features of the bottleneck layer Located in the image block-level feature set The feature portion at the corresponding location is denoted as the image patch branch representation. The bottleneck layer outputs features Mid-level features at the full slice level The feature portion at the corresponding position is denoted as the full slice branch representation. ;
[0047] 3) Output features of the bottleneck layer Input an asymmetric dual-branch decoder, the asymmetric dual-branch decoder including an image patch decoder Full slice decoder and cross-attention layers, where, This represents the image block decoder used to reconstruct image block-level features. This refers to the full-slice decoder used to reconstruct full-slice level features; the image patch decoder and the full slice decoder Each layer comprises eight sequentially stacked decoding layers. Each decoding layer includes a first normalization layer, a cross-attention layer, a second normalization layer, and a multilayer perceptron layer. The cross-attention layer is used to implement image patch branch representation. Compared with full slice branch representation Feature interactions between them;
[0048] 4) Define the query projection function in the cross-attention layer. Key projection function Sum projection function Its expression is:
[0049] , , ;
[0050] In the formula, Indicates input features, Represents the query mapping matrix. Represents the key mapping matrix, Represents a value mapping matrix;
[0051] 5) In the image block decoder In the image, the branches are represented by image patches. As a query, it is represented by a full slice branch. As keys and values, generate an image patch query matrix. Fully sliced bond matrix and the full slice value matrix Its expression is:
[0052] , , ;
[0053] Image patch enhancement representation is obtained through cross-attention calculation. Its expression is:
[0054] ;
[0055] In the formula, This represents the feature dimension of each attention head. This represents the image patch enhancement representation after fusing full-slice-level contextual information;
[0056] 6) In the full slice decoder In this context, the full slice branch is used to represent... As a query, it is represented by image patch branches. Use these as keys and values to generate a full-slice query matrix. Image block key matrix and image patch value matrix Its expression is:
[0057] , , ;
[0058] The full-slice augmented representation is obtained through cross-attention computation. Its expression is:
[0059] ;
[0060] In the formula, This represents the full-slice augmented representation after fusing block-level local information from the image;
[0061] 7) The image block decoder Based on the output characteristics of the bottleneck layer Using the full slice branch representation And combined with the image patch enhancement representation obtained by the above cross-attention calculation Feature reconstruction is performed to obtain image block-level reconstructed features. The full-slice decoder Based on the output characteristics of the bottleneck layer Using the image block branch representation And combined with the full-slice augmented representation obtained by the above cross-attention calculation Feature reconstruction is performed to obtain full-slice level reconstructed features. Its expression is:
[0062] , ;
[0063] 8) Reconstruct the image block-level features and the full-slice-level reconstruction features As output, it is used for subsequent reconstruction error calculation.
[0064] Furthermore, the training module includes the following steps:
[0065] 1) During the training phase, whole-slice images of pathological sections without abnormalities are used as training samples. The parameters of the image patch encoder and the whole-slice encoder are kept fixed, and the shared bottleneck layer is adjusted according to the reconstruction loss. The trainable parameters of the asymmetric dual-branch decoder are updated.
[0066] 2) Calculate the reconstruction error based on the cosine distance between the original feature vector and the corresponding reconstructed feature vector. For any original feature vector... and its corresponding reconstructed feature vector Cosine reconstruction error Represented as:
[0067] ;
[0068] In the formula, Represents the image block-level features or the full slice-level features. This represents the image block-level reconstruction features or the full-slice-level reconstruction features. Denotes the 2-norm;
[0069] 3) For any original feature sequence and its corresponding reconstructed feature sequence Calculate the reconstruction error set for each feature unit. Its expression is:
[0070] ;
[0071] in, This represents the original feature sequence from which the reconstructed distance is to be calculated. This represents the reconstructed feature sequence corresponding to the original feature sequence; Represents the first in the original feature sequence Each feature unit is either an image block-level feature or a full slice-level feature; Indicates the first Each feature unit corresponds to a reconstruction feature unit, wherein the reconstruction feature unit is the image block-level reconstruction feature or the full slice-level reconstruction feature; This indicates the number of feature units in the original feature sequence;
[0072] 4) To avoid feature units with low reconstruction error dominating the training process, a hard sample mining strategy is adopted from the reconstruction error set. Select the largest value from There are 1 reconstruction errors, and the set of indices corresponding to the selected reconstruction errors is denoted as . ,in, The value can be:
[0073] ;
[0074] In the formula, This indicates the proportion of low-error feature units that are ignored. This indicates a round-down operation;
[0075] 5) Based on the index set The corresponding reconstruction error is used to calculate the reconstruction distance for the hard sample mining strategy. Its expression is:
[0076] ;
[0077] In the formula, Represents the set of indexes The number of elements in the middle, Indicates the first The reconstruction error corresponding to each feature unit;
[0078] 6) Calculate the image patch branch reconstruction loss based on the reconstruction distance of the hard sample mining strategy. and full slice branch reconstruction loss Its expression is:
[0079] , ;
[0080] In the formula, This indicates the distance reconstructed using a difficult sample mining strategy. This indicates the calculation of reconstruction loss. Represents a set of image block-level features. Represents image block-level reconstruction features. Represents full-slice level features. Indicates full-slice level reconstruction features;
[0081] 7) Reconstruct loss based on the image patch branch and the full slice branch reconstruction loss Construct the total training loss Its expression is:
[0082] ;
[0083] Based on the total training loss For the shared bottleneck layer The trainable parameters of the asymmetric dual-branch decoder are updated via backpropagation.
[0084] Furthermore, the inference module includes the following steps:
[0085] 1) During the inference phase, based on the trained shared bottleneck layer and the asymmetric dual-branch decoder, image block-level reconstruction features and whole-slice-level reconstruction features of the pathological whole-slice image to be tested are obtained. Image block-level reconstruction error is obtained based on the difference between the image block-level features and the image block-level reconstruction features, and whole-slice-level reconstruction error is obtained based on the difference between the whole-slice-level features and the whole-slice-level reconstruction features; for the first... For each image patch, calculate its image patch-level anomaly response sequence based on the image patch-level reconstruction error. The image block-level anomaly response sequence The Middle Anomaly response of image patch-level features Represented as:
[0086] ;
[0087] In the formula, Indicates the first The corresponding image patch of the _th Image patch-level features, Indicates the first Image block-level reconstructed features corresponding to each image block-level feature Indicates the first The first image patch Anomaly responses of image block-level features; and the sequence of image block-level anomaly responses. Mapping to the image patch space yields an anomaly heatmap for a single image patch. Combine all image patch anomaly heatmaps into an image patch anomaly heatmap set. Its expression is:
[0088] ;
[0089] In the formula, Represents a set of abnormal heatmaps for image patches. Indicates the first Image patch anomaly heatmap corresponding to each image patch Indicates the total number of image blocks;
[0090] 2) Based on the image block-level anomaly response sequence Select the abnormal response with the highest preset ratio to form an index set. And calculate the first Anomaly score of each image patch Its expression is:
[0091] , ;
[0092] In the formula, Indicates the first Image block-level anomaly response sequence of image blocks The set of indices corresponding to the top 1% of anomaly responses in terms of median value;
[0093] 3) Based on the anomaly score of each image patch And the recorded position information of each image block in the pathological whole slide image. The anomaly score of each image patch The spatial location of the pathological whole-slice image is mapped back to the original image and then Gaussian smoothed to obtain a local anomaly heatmap. ;
[0094] 4) Based on full-slice level features and full-slice-level reconstruction features Calculate the full-slice-level reconstruction error and use it as the full-slice anomaly score. Its expression is:
[0095] ;
[0096] In the formula, This represents the whole-slice abnormality score of the pathological whole-slice image;
[0097] 5) The aforementioned local anomaly thermal map As the result of local abnormality localization in the whole pathological slide image, the abnormality score of the whole slide is used. This serves as the anomaly detection result for the pathological whole-section image.
[0098] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0099] 1. This invention employs a pathological image anomaly detection system based on dual-branch multi-scale reverse distillation, which achieves pathological image anomaly detection while reducing reliance on lesion-level manual annotation, thereby reducing the need for finely labeled data for model training.
[0100] 2. This invention obtains both image block-level features and full-slice-level features simultaneously through the feature extraction module, which can take into account both image block-level information and full-slice-level information, avoiding the problem of insufficient anomaly detection information caused by relying solely on single-scale features.
[0101] 3. This invention uses a feature fusion module and a feature transformation module to perform multi-scale feature fusion of image block-level features and full-slice-level features, and obtains bottleneck layer output features by sharing a bottleneck layer, which is beneficial to enhance the correlation between features of different scales.
[0102] 4. This invention achieves feature interaction between image block branch representation and full slice branch representation through an asymmetric dual-branch decoder, and obtains image block-level reconstruction features and full slice-level reconstruction features respectively, thereby improving the collaborative reconstruction capability of image block-level features and full slice-level features.
[0103] 5. This invention uses both image block-level reconstruction error and whole-slice-level reconstruction error for training through a training module, and generates a local anomaly heatmap based on the image block-level reconstruction error through an inference module, and obtains the anomaly detection result of the whole pathological slice image based on the whole-slice-level reconstruction error, thereby taking into account both the localization of abnormal regions and the anomaly judgment of the whole pathological slice image.
[0104] In summary, this invention can reduce the reliance on manual annotation at the lesion level, enhance the collaborative modeling capability between image block-level features and whole-slice-level features, improve the effectiveness of image block-level reconstruction features and whole-slice-level reconstruction features, enhance the accuracy and stability of anomaly detection in pathological whole-slice images, and simultaneously achieve the generation of local anomaly heatmaps and the detection of anomalies in pathological whole-slice images. Attached Figure Description
[0105] Figure 1 This is a framework diagram of the system of the present invention. Among them, CONCH is the basic model of pathological visual language, which is used to extract features from image blocks obtained by dividing a whole pathological slide image to obtain image block-level features; TITAN is the basic model of pathological whole slide multimodal model, which is used to extract whole slide-level features based on image block-level features and their spatial location information. Detailed Implementation
[0106] The present invention will be further described in detail below with reference to the embodiments and accompanying drawings, but the embodiments of the present invention are not limited thereto.
[0107] like Figure 1As shown, this embodiment discloses a pathological image anomaly detection system based on bi-branch multi-scale reverse distillation. Taking a whole-slice pathological image as input, it sequentially performs tissue region extraction, image block segmentation, image block-level feature extraction, whole-slice-level feature extraction, multi-scale feature fusion, shared bottleneck layer feature transformation, asymmetric bi-branch decoding and reconstruction, and anomaly detection result output, ultimately obtaining a set of image block anomaly heatmaps, local anomaly heatmaps, and whole-slice anomaly scores. Specifically, it includes:
[0108] The data acquisition and preprocessing module is used to acquire and preprocess whole pathological slide images to obtain multiple image blocks and the position information of each image block in the whole pathological slide image.
[0109] The feature extraction module includes a pathological visual language basic model with fixed parameters and a pathological whole-slice multimodal basic model with fixed parameters. The pathological visual language basic model is used to extract features from each image block to obtain corresponding image block-level features, and to form an image block-level feature set from all image block-level features. The pathological whole-slice multimodal basic model is used to extract whole-slice-level features based on the image block-level feature set and the position information of each image block in the pathological whole-slice image to obtain whole-slice-level features.
[0110] The feature fusion module performs multi-scale feature fusion between the image block-level features and the corresponding full-slice-level features to obtain joint multi-scale features.
[0111] The feature transformation module inputs the joint multi-scale features into the shared bottleneck layer and performs feature transformation to obtain the bottleneck layer output features.
[0112] The feature reconstruction module includes an asymmetric dual-branch decoder, which comprises an image patch decoder, a full slice decoder, and a cross-attention layer. The feature reconstruction module inputs the output features of the bottleneck layer into the asymmetric dual-branch decoder, forming image patch branch representations and full slice branch representations. The cross-attention layer facilitates feature interaction between the image patch branch representations and the full slice branch representations. The image patch decoder reconstructs image patch-level features to obtain image patch-level reconstructed features; the full slice decoder reconstructs full slice-level features to obtain full slice-level reconstructed features.
[0113] The training module trains the shared bottleneck layer and the asymmetric dual-branch decoder based on the image block-level reconstruction error between the image block-level features and the image block-level reconstruction features, and the full slice-level reconstruction error between the full slice-level features and the full slice-level reconstruction features.
[0114] The inference module, based on the trained shared bottleneck layer and the asymmetric dual-branch decoder, obtains the image block-level reconstruction features and the full-slice-level reconstruction features of the pathological whole-slice image to be tested, and then obtains the image block-level reconstruction error and the full-slice-level reconstruction error of the pathological whole-slice image to be tested. Based on the image block-level reconstruction error, a local anomaly heatmap is generated, and based on the full-slice-level reconstruction error, the anomaly detection result of the pathological whole-slice image is obtained.
[0115] Specifically, the data acquisition and preprocessing module performs the following operations:
[0116] First, obtain the whole-section image of the pathological tissue to be examined. and the pathological whole slide images Perform downsampling processing to obtain a downsampled image. ,in, This indicates the downsampling magnification. Since whole-slice pathological images typically contain a large amount of background, directly processing the entire image would increase computational load and introduce invalid information. Therefore, this embodiment first processes the downsampled image. Perform organizational region segmentation to obtain the organizational region mask. The organization region mask Used to distinguish between the organization area and the background area.
[0117] Subsequently, based on the organization's area mask Remove whole pathology slide images The background area is selected, and the effective tissue area is preserved. The effective tissue area is then processed according to a preset image block size. The image is divided into patches to obtain a set of image patches:
[0118] ;
[0119] in, Indicates the first Image blocks, Indicates the total number of image patches. and These represent the height and width of the image block, respectively. .
[0120] During the image block segmentation process, the first... Image blocks In pathological whole slide images Location information Its expression is:
[0121] ;
[0122] And obtain a set of location information:
[0123] ;
[0124] in, Indicates the first Image blocks In pathological whole slide images Location information in and Each represents an image block In pathological whole slide images The horizontal and vertical coordinates are used in the image. This location information is used to subsequently map the abnormal response of the image patch back to the spatial location of the whole pathological slice image.
[0125] Specifically, the feature extraction module performs the following operations:
[0126] For the image patch set 𝑃, this embodiment uses a CONCH model with fixed parameters to extract image patch-level features for each image patch. For the 𝑃... Image blocks Its image block-level features Represented as:
[0127] ;
[0128] Among them, the CONCH model is a basic model for pathological visual language, used to extract cell morphology, local tissue structure and pathological semantic information in image patches; This represents the model parameters of the CONCH model, and It remains constant during both the training and inference phases; Indicates the first Image blocks The corresponding image patch-level features. All image patch-level features are combined into an image patch-level feature set. Its expression is:
[0129] ;
[0130] in, Represents images of whole pathological sections The image patch-level feature set is composed of all image patch-level features. Then, a TITAN model with fixed parameters is used to process the image patch-level feature set. and location information set Full-slice-level feature extraction was performed to obtain pathological whole-slice images. Corresponding full-slice level features Its expression is:
[0131] ;
[0132] Among them, the TITAN model is the basic model for multimodal pathological whole slices, which is used to model the macroscopic tissue structure and long-distance spatial dependencies of pathological whole slice images based on image block-level features and their location information. This represents the model parameters of the TITAN model, and It remains constant during both the training and inference phases; Image of whole pathological slide The corresponding full-slice level features.
[0133] Through the above steps, this embodiment obtains image block-level features and full-slice-level features respectively, providing input for subsequent multi-scale feature fusion and reverse distillation reconstruction.
[0134] Specifically, the feature fusion module performs the following operations:
[0135] Obtain the image block-level feature set output by the feature extraction module. and full slice level features The two are then combined using a splicing multi-scale feature fusion method to obtain joint multi-scale features. Its expression is:
[0136] ;
[0137] In the formula, This indicates a feature concatenation operation. This represents the set of image block-level features composed of all image block-level features in the pathological whole-section image. This represents the full-slice-level features corresponding to the pathological whole-slice image. Through this stitched multi-scale feature fusion, image block-level local information and full-slice-level global information are combined into a unified feature representation, enabling subsequent models to simultaneously process local pathological morphological information and full-slice tissue structure information.
[0138] Specifically, the feature transformation module performs the following operations:
[0139] The joint multi-scale features output by the feature fusion module Input shared bottleneck layer The bottleneck layer output features are obtained. Its expression is:
[0140] ;
[0141] in, Indicates a shared bottleneck layer. Indicates joint multi-scale features Through the shared bottleneck layer The output features of the bottleneck layer obtained after transformation.
[0142] In this embodiment, the shared bottleneck layer It includes a random deactivation layer, a first linear layer, a non-linear activation function, a second linear layer, and an output random deactivation layer. A shared bottleneck layer is also included. Used for joint multi-scale features By performing feature perturbation and nonlinear mapping, the possibility of the model directly learning the identity mapping is reduced, so that the subsequent reconstruction error can more effectively reflect the difference between normal and abnormal features.
[0143] Specifically, the feature reconstruction module performs the following operations:
[0144] Obtain the bottleneck layer output features from the feature transformation module. Due to the combined multi-scale features From image block-level feature set and full slice level features Therefore, the bottleneck layer outputs features obtained by splicing. Located in the image block-level feature set The feature portion at the corresponding location is denoted as the image patch branch representation. Located in the full slice level features The feature portion at the corresponding position is denoted as the full slice branch representation. .
[0145] Output features of the bottleneck layer Input an asymmetric dual-branch decoder. The asymmetric dual-branch decoder includes an image patch decoder. Full slice decoder And cross-attention layers. Among them, This represents the image block decoder used to reconstruct image block-level features. This represents the full-slice decoder used to reconstruct full-slice level features.
[0146] In this embodiment, the image block decoder and full slice decoder Each layer comprises eight sequentially stacked decoding layers. Each decoding layer includes a first normalization layer, a cross-attention layer, a second normalization layer, and a multilayer perceptron layer. The cross-attention layer is used to implement image patch branch representation. Compared with full slice branch representation The characteristic interactions between them.
[0147] In the cross-attention layer, define the query projection function. Key projection function Sum projection function Its expression is:
[0148] , , ;
[0149] in, Indicates input features, Represents the query mapping matrix. Represents the key mapping matrix, Represents a value mapping matrix.
[0150] In image block decoder In the image, the branches are represented by image patches. As a query, it is represented by a full slice branch. As keys and values, generate an image patch query matrix. Fully sliced bond matrix and the full slice value matrix Its expression is:
[0151] , , ;
[0152] Image patch enhancement representation is obtained through cross-attention calculation. Its expression is:
[0153] ;
[0154] in, This represents the feature dimension of each attention head. This represents the enhanced image patch representation after fusing full-slice-level contextual information. Through this process, the image patch decoder... When reconstructing image block-level features, it is possible to incorporate full-slice-level contextual information.
[0155] In full slice decoder In this context, the full slice branch is used to represent... As a query, it is represented by image patch branches. Use these as keys and values to generate a full-slice query matrix. Image block key matrix and image patch value matrix Its expression is:
[0156] , , ;
[0157] The full-slice augmented representation is obtained through cross-attention computation. Its expression is:
[0158] ;
[0159] In the formula, This represents the enhanced full-slice representation after fusing block-level local information from the image. Through this process, the full-slice decoder... The image block decoder can incorporate local image block information when reconstructing full-slice-level features; Based on the output characteristics of the bottleneck layer Using the full slice branch representation And combined with the image patch enhancement representation obtained by the above cross-attention calculation Feature reconstruction is performed to obtain image block-level reconstructed features. The full-slice decoder Based on the output characteristics of the bottleneck layer Using the image block branch representation And combined with the full-slice augmented representation obtained by the above cross-attention calculation Feature reconstruction is performed to obtain full-slice level reconstructed features. Its expression is:
[0160] , ;
[0161] Specifically, the training module performs the following operations: During the training phase, it uses whole-section images of pathological tissue without abnormalities as training samples, keeps the parameters of the CONCH and TITAN models fixed, and adjusts the shared bottleneck layer according to the reconstruction loss. The trainable parameters of the asymmetric dual-branch decoder are updated.
[0162] For any original feature vector and its corresponding reconstructed feature vector The reconstruction error is calculated using cosine distance. Represented as:
[0163] ;
[0164] in, Represents the image block-level features or the full slice-level features. This represents the image block-level reconstruction features or the full-slice-level reconstruction features. This represents the L2 norm.
[0165] For any original feature sequence and its corresponding reconstructed feature sequence Calculate the reconstruction error set for each feature unit. Its expression is:
[0166] ;
[0167] In the formula, This represents the original feature sequence from which the reconstructed distance is to be calculated. This represents the reconstructed feature sequence corresponding to the original feature sequence; Represents the first in the original feature sequence Each feature unit is either an image block-level feature or a full slice-level feature; Indicates the first Each feature unit corresponds to a reconstruction feature unit, wherein the reconstruction feature unit is the image block-level reconstruction feature or the full slice-level reconstruction feature; This indicates the number of feature units in the original feature sequence.
[0168] To avoid feature units with low reconstruction error dominating the training process, a hard sample mining strategy is adopted from the reconstruction error set. Select the largest value from There are 1 reconstruction errors, and the set of indices corresponding to the selected reconstruction errors is denoted as . ,in, The value can be:
[0169] ;
[0170] In the formula, This indicates the proportion of low-error feature units that are ignored. This indicates a round-down operation;
[0171] According to the index set The corresponding reconstruction error is used to calculate the reconstruction distance for the hard sample mining strategy. Its expression is:
[0172] ;
[0173] In the formula, Represents the set of indexes The number of elements in the middle, Indicates the first The reconstruction error corresponding to each feature unit;
[0174] The image patch branch reconstruction loss is calculated based on the reconstruction distance according to the hard sample mining strategy. and full slice branch reconstruction loss Its expression is:
[0175] , ;
[0176] In the formula, This indicates the distance reconstructed using a difficult sample mining strategy. This indicates the calculation of reconstruction loss. Represents a set of image block-level features. Represents image block-level reconstruction features. Represents full-slice level features. Indicates full-slice level reconstruction features;
[0177] Further construct the total training loss Its expression is:
[0178] ;
[0179] Based on the total training loss For the shared bottleneck layer The trainable parameters of the asymmetric dual-branch decoder are updated via backpropagation.
[0180] Specifically, the inference module performs the following operations: During the inference phase, based on the trained shared bottleneck layer and the asymmetric dual-branch decoder, it obtains the image block-level reconstruction features and the whole-slice-level reconstruction features of the pathological whole-slice image to be tested, and obtains the image block-level reconstruction error based on the difference between the image block-level features and the image block-level reconstruction features, and obtains the whole-slice-level reconstruction error based on the difference between the whole-slice-level features and the whole-slice-level reconstruction features; for the first... For each image patch, calculate its image patch-level anomaly response sequence based on the image patch-level reconstruction error. The image block-level anomaly response sequence The Middle Anomaly response of image patch-level features Represented as:
[0181] ;
[0182] In the formula, Indicates the first The corresponding image patch of the _th Image patch-level features, Indicates the first Image block-level reconstructed features corresponding to each image block-level feature Indicates the first The first image patch Anomaly responses of image block-level features; and the sequence of image block-level anomaly responses. Mapping to the image patch space yields an anomaly heatmap for a single image patch. Combine all image patch anomaly heatmaps into an image patch anomaly heatmap set. Its expression is:
[0183] ;
[0184] In the formula, Represents a set of abnormal heatmaps for image patches. Indicates the first Image patch anomaly heatmap corresponding to each image patch Indicates the total number of image blocks;
[0185] Based on the image block-level anomaly response sequence Select the abnormal response with the highest preset ratio to form an index set. And calculate the first Anomaly score of each image patch Its expression is:
[0186] , ;
[0187] In the formula, Indicates the first Image block-level anomaly response sequence of image blocks The set of indices corresponding to the top 1% of anomaly responses in terms of median value;
[0188] Based on the anomaly score of each image patch And the recorded position information of each image block in the pathological whole slide image. The anomaly score of each image patch The spatial location of the pathological whole-slice image is mapped back to the original image and then Gaussian smoothed to obtain a local anomaly heatmap. ;
[0189] Based on full slice level features and full-slice-level reconstruction features Calculate the full-slice-level reconstruction error and use it as the full-slice anomaly score. Its expression is:
[0190] ;
[0191] In the formula, This represents the whole-slice abnormality score of the pathological whole-slice image;
[0192] The local abnormal heat map As the result of local abnormality localization in the whole pathological slide image, the abnormality score of the whole slide is used. This serves as the anomaly detection result for the pathological whole-section image.
[0193] Therefore, this embodiment can achieve collaborative modeling between local features at the image block level and global features at the whole slice level while reducing the reliance on manual annotation at the lesion level, and simultaneously output a set of image block abnormality heatmaps, local abnormality heatmaps and whole slice abnormality scores, thereby completing the localization of abnormal areas in pathological images and the judgment of abnormalities in whole slices.
[0194] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any changes, modifications, substitutions, combinations, or simplifications made without departing from the spirit and principle of the present invention shall be considered equivalent substitutions and shall be included within the protection scope of the present invention.
Claims
1. A pathological image anomaly detection system based on dual-branch multi-scale reverse distillation, characterized in that, include: The data acquisition and preprocessing module is used to acquire and preprocess whole pathological slide images to obtain multiple image blocks and the position information of each image block in the whole pathological slide image. The feature extraction module includes a pathological visual language basic model with fixed parameters and a pathological whole-slice multimodal basic model with fixed parameters. The pathological visual language basic model is used to extract features from each image block to obtain corresponding image block-level features, and to form an image block-level feature set from all image block-level features. The pathological whole-slice multimodal basic model is used to extract whole-slice-level features based on the image block-level feature set and the position information of each image block in the pathological whole-slice image to obtain whole-slice-level features. The feature fusion module performs multi-scale feature fusion between the image block-level features and the corresponding full-slice-level features to obtain joint multi-scale features. The feature transformation module inputs the joint multi-scale features into the shared bottleneck layer and performs feature transformation to obtain the bottleneck layer output features. The feature reconstruction module includes an asymmetric dual-branch decoder, which comprises an image patch decoder, a full-slice decoder, and a cross-attention layer. The feature reconstruction module is used to input the output features of the bottleneck layer into the asymmetric dual-branch decoder, form image patch branch representation and full slice branch representation in the asymmetric dual-branch decoder, and realize feature interaction between the image patch branch representation and the full slice branch representation through the cross attention layer; The image block decoder is used to reconstruct image block-level features to obtain image block-level reconstructed features; The full-slice decoder is used to reconstruct full-slice-level features to obtain full-slice-level reconstructed features; The training module trains the shared bottleneck layer and the asymmetric dual-branch decoder based on the image block-level reconstruction error between the image block-level features and the image block-level reconstruction features, and the full slice-level reconstruction error between the full slice-level features and the full slice-level reconstruction features. The inference module, based on the trained shared bottleneck layer and the asymmetric dual-branch decoder, obtains the image block-level reconstruction features and the full-slice-level reconstruction features of the pathological whole-slice image to be tested, and then obtains the image block-level reconstruction error and the full-slice-level reconstruction error of the pathological whole-slice image to be tested. Based on the image block-level reconstruction error, a local anomaly heatmap is generated, and based on the full-slice-level reconstruction error, the anomaly detection result of the pathological whole-slice image is obtained.
2. The pathological image anomaly detection system based on dual-branch multi-scale reverse distillation according to claim 1, characterized in that, The data acquisition and preprocessing module includes the following steps: 1) Obtain whole pathological slide images and the pathological whole slide images Perform downsampling processing to obtain a downsampled image. ,in, Indicates the downsampling factor; 2) For the downsampled image Perform organizational region segmentation to obtain the organizational region mask. ,in, Used to distinguish between the tissue area and the background area; 3) Based on the tissue region mask Remove the whole pathological slide image The background area retains the effective organizational area; 4) According to the preset image block size The effective tissue region is divided into image blocks to obtain an image block set. ,in, Indicates the first Image blocks, Indicates the total number of image patches. and These represent the height and width of the image block, respectively. ; 5) Record the first Image blocks The pathological whole slide image Location information and obtain a set of location information. ,in, Indicates the first Location information of each image patch and Each represents a different image block. The whole pathological slide image The horizontal and vertical coordinates in the diagram.
3. The pathological image anomaly detection system based on dual-branch multi-scale reverse distillation according to claim 2, characterized in that, The feature extraction module includes the following steps: 1) The pathological visual language basic model uses a CONCH model with fixed parameters to analyze the image patches. Feature extraction is performed to obtain image block-level features. The CONCH model is used to extract cell morphology, local tissue structure, and pathological semantic information from image patches, and its expression is: ; In the formula, Indicates the first Image blocks, , Indicates the total number of image patches. This represents the model parameters of the CONCH model, and It remains constant during both the training and inference phases. Indicates the first Image blocks Corresponding image block-level features; 2) Combine all image patch-level features into an image patch-level feature set. Its expression is: ; In the formula, This represents the pathological whole slide image. The image block-level feature set consists of all image block-level features; 3) The multimodal basic model of the pathological whole slide uses the TITAN model with fixed parameters to analyze the image block-level feature set. and location information set Full-slice-level feature extraction was performed to obtain the pathological full-slice image. Corresponding full-slice level features The TITAN model is a pre-trained pathological whole-slice model used to model the macroscopic tissue structure and long-range spatial dependencies of pathological whole-slice images based on image block-level features and their location information. Its expression is: ; In the formula, This represents a set of location information composed of the location information of each image patch. Indicates the first Image blocks The whole pathological slide image Location information in This represents the model parameters of the TITAN model, and It remains constant during both the training and inference phases. The image represents the whole pathological slide. Corresponding full-slice level features; 4) The image block-level feature set and the full-slice level features As output.
4. The pathological image anomaly detection system based on dual-branch multi-scale reverse distillation according to claim 3, characterized in that, The feature fusion module includes the following steps: 1) Obtain the image block-level feature set output by the feature extraction module. and full slice level features ; 2) The image block-level feature set With the full slice level features Perform concatenated multi-scale feature fusion to obtain joint multi-scale features. Its expression is: ; In the formula, This indicates a feature concatenation operation. This represents the set of image block-level features composed of all image block-level features in the pathological whole-section image. This represents the whole-slice-level features corresponding to the pathological whole-slice image; 3) Combine the joint multi-scale features As output, it is used as input for subsequent shared bottleneck layers.
5. The pathological image anomaly detection system based on dual-branch multi-scale reverse distillation according to claim 4, characterized in that, The feature transformation module includes the following steps: 1) Obtain the joint multi-scale features output by the feature fusion module. ; 2) Combine the joint multi-scale features Input shared bottleneck layer Perform feature transformation to obtain the bottleneck layer output features. Its expression is: ; In the formula, Indicates a shared bottleneck layer. This indicates that the joint multi-scale features Through the shared bottleneck layer The output features of the bottleneck layer after transformation; The shared bottleneck layer It includes a random deactivation layer, a first linear layer, a nonlinear activation function, a second linear layer, and an output random deactivation layer, used to process the joint multi-scale features. Feature perturbation and nonlinear mapping are performed to reduce the direct learning of identity mappings; 3) Output features of the bottleneck layer As output, it is used as input to the subsequent asymmetric dual-branch decoder.
6. The pathological image anomaly detection system based on dual-branch multi-scale reverse distillation according to claim 5, characterized in that, The feature reconstruction module includes the following steps: 1) Obtain the bottleneck layer output features from the feature transformation module. ; 2) Output features of the bottleneck layer Located in the image block-level feature set The feature portion at the corresponding location is denoted as the image patch branch representation. The bottleneck layer outputs features Mid-level features at the full slice level The feature portion at the corresponding position is denoted as the full slice branch representation. ; 3) Output features of the bottleneck layer Input an asymmetric dual-branch decoder, the asymmetric dual-branch decoder including an image patch decoder Full slice decoder and cross-attention layers, where, This represents the image block decoder used to reconstruct image block-level features. This refers to the full-slice decoder used to reconstruct full-slice level features; the image patch decoder and the full slice decoder Each layer comprises eight sequentially stacked decoding layers. Each decoding layer includes a first normalization layer, a cross-attention layer, a second normalization layer, and a multilayer perceptron layer. The cross-attention layer is used to implement image patch branch representation. Compared with full slice branch representation Feature interactions between them; 4) Define the query projection function in the cross-attention layer. Key projection function Sum projection function Its expression is: , , ; In the formula, Indicates input features, Represents the query mapping matrix. Represents the key mapping matrix, Represents a value mapping matrix; 5) In the image block decoder In the image, the branches are represented by image patches. As a query, it is represented by a full slice branch. As keys and values, generate an image patch query matrix. Fully sliced bond matrix and the full slice value matrix Its expression is: , , ; Image patch enhancement representation is obtained through cross-attention calculation. Its expression is: ; In the formula, This represents the feature dimension of each attention head. This represents the image patch enhancement representation after fusing full-slice-level contextual information; 6) In the full slice decoder In this context, the full slice branch is used to represent... As a query, it is represented by image patch branches. Use these as keys and values to generate a full-slice query matrix. Image block key matrix and image patch value matrix Its expression is: , , ; The full-slice augmented representation is obtained through cross-attention computation. Its expression is: ; In the formula, This represents the full-slice augmented representation after fusing block-level local information from the image; 7) The image block decoder Based on the output characteristics of the bottleneck layer Using the full slice branch representation And combined with the image patch enhancement representation obtained by the above cross-attention calculation Feature reconstruction is performed to obtain image block-level reconstructed features. The full-slice decoder Based on the output characteristics of the bottleneck layer Using the image block branch representation And combined with the full-slice augmented representation obtained by the above cross-attention calculation Feature reconstruction is performed to obtain full-slice level reconstructed features. Its expression is: , ; 8) Reconstruct the image block-level features and the full-slice-level reconstruction features As output, it is used for subsequent reconstruction error calculation.
7. A pathological image anomaly detection system based on dual-branch multi-scale reverse distillation according to claim 6, characterized in that, The training module includes the following steps: 1) During the training phase, whole-slice images of pathological sections without abnormalities are used as training samples. The parameters of the image patch encoder and the whole-slice encoder are kept fixed, and the shared bottleneck layer is adjusted according to the reconstruction loss. The trainable parameters of the asymmetric dual-branch decoder are updated. 2) Calculate the reconstruction error based on the cosine distance between the original feature vector and the corresponding reconstructed feature vector. For any original feature vector... and its corresponding reconstructed feature vector Cosine reconstruction error Represented as: ; In the formula, Represents the image block-level features or the full slice-level features. This represents the image block-level reconstruction features or the full-slice-level reconstruction features. Denotes the 2-norm; 3) For any original feature sequence and its corresponding reconstructed feature sequence Calculate the reconstruction error set for each feature unit. Its expression is: ; in, This represents the original feature sequence from which the reconstructed distance is to be calculated. This represents the reconstructed feature sequence corresponding to the original feature sequence; Represents the first in the original feature sequence Each feature unit is either an image block-level feature or a full slice-level feature; Indicates the first Each feature unit corresponds to a reconstruction feature unit, wherein the reconstruction feature unit is the image block-level reconstruction feature or the full slice-level reconstruction feature; This indicates the number of feature units in the original feature sequence; 4) To avoid feature units with low reconstruction error dominating the training process, a hard sample mining strategy is adopted from the reconstruction error set. Select the largest value from There are 1 reconstruction errors, and the set of indices corresponding to the selected reconstruction errors is denoted as . ,in, The value can be: ; In the formula, This indicates the proportion of low-error feature units that are ignored. This indicates a round-down operation; 5) Based on the index set The corresponding reconstruction error is used to calculate the reconstruction distance for the hard sample mining strategy. Its expression is: ; In the formula, Represents the set of indexes The number of elements in the middle, Indicates the first The reconstruction error corresponding to each feature unit; 6) Calculate the image patch branch reconstruction loss based on the reconstruction distance of the hard sample mining strategy. and full slice branch reconstruction loss Its expression is: , ; In the formula, This indicates the distance reconstructed using a hard sample mining strategy. This indicates the calculation of reconstruction loss. Represents a set of image block-level features. Represents image block-level reconstruction features. Represents full-slice level features. Indicates full-slice level reconstruction features; 7) Reconstruct loss based on the image patch branch and the full slice branch reconstruction loss Construct the total training loss Its expression is: ; Based on the total training loss For the shared bottleneck layer The trainable parameters of the asymmetric dual-branch decoder are updated via backpropagation.
8. The pathological image anomaly detection system based on dual-branch multi-scale reverse distillation according to claim 7, characterized in that, The reasoning module includes the following steps: 1) During the inference phase, based on the trained shared bottleneck layer and the asymmetric dual-branch decoder, image block-level reconstruction features and whole-slice-level reconstruction features of the pathological whole-slice image to be tested are obtained. Image block-level reconstruction error is obtained based on the difference between the image block-level features and the image block-level reconstruction features, and whole-slice-level reconstruction error is obtained based on the difference between the whole-slice-level features and the whole-slice-level reconstruction features; for the first... For each image patch, calculate its image patch-level anomaly response sequence based on the image patch-level reconstruction error. The image block-level anomaly response sequence The Middle Anomaly response of image patch-level features Represented as: ; In the formula, Indicates the first The corresponding image patch of the _th Image patch-level features, Indicates the first Image block-level reconstructed features corresponding to each image block-level feature Indicates the first The first image patch Anomaly responses of image block-level features; and the sequence of image block-level anomaly responses. Mapping to the image patch space yields an anomaly heatmap for a single image patch. Combine all image patch anomaly heatmaps into an image patch anomaly heatmap set. Its expression is: ; In the formula, Represents a set of abnormal heatmaps for image patches. Indicates the first Image patch anomaly heatmap corresponding to each image patch Indicates the total number of image blocks; 2) Based on the image block-level anomaly response sequence Select the abnormal response with the highest preset ratio to form an index set. And calculate the first Anomaly score of each image patch Its expression is: , ; In the formula, Indicates the first Image block-level anomaly response sequence of image blocks The set of indices corresponding to the top 1% of anomaly responses in terms of median value; 3) Based on the anomaly score of each image patch And the recorded position information of each image block in the pathological whole slide image. The anomaly score of each image patch The spatial location of the pathological whole-slice image is mapped back to the original image and then Gaussian smoothed to obtain a local anomaly heatmap. ; 4) Based on full-slice level features and full-slice-level reconstruction features Calculate the full-slice-level reconstruction error and use it as the full-slice anomaly score. Its expression is: ; In the formula, This represents the whole-slice abnormality score of the pathological whole-slice image; 5) The aforementioned local anomaly thermal map As the result of local abnormality localization in the whole pathological slide image, the abnormality score of the whole slide is used. This serves as the anomaly detection result for the pathological whole-section image.