Three-dimensional microscopic image recognition model training method based on cross-dimension knowledge distillation

By employing a cross-dimensional knowledge distillation method, a two-dimensional microscopic image recognition model is used as a teacher model. Combined with three-dimensional microscopic image sample data, hybrid supervised training is conducted, which solves the problem of insufficient recognition accuracy of the three-dimensional microscopic image recognition model and realizes the effective utilization of three-dimensional information and performance improvement.

CN120932069AActive Publication Date: 2025-11-11WUHAN SMARTVIEW BIOTECHNOLOGY CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511046723.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-29
Publication Date
2025-11-11
Estimated Expiration
2045-07-29

AI Technical Summary

Technical Problem

Existing 3D microscopic image recognition models suffer from insufficient labeled data and limited validation data, making it difficult to overcome performance bottlenecks, resulting in insufficient recognition accuracy and an inability to effectively utilize 3D information.

Method used

A cross-dimensional knowledge distillation method is adopted, using a pre-trained two-dimensional microscopic image recognition model as a teacher model, and combining it with three-dimensional microscopic image sample data for hybrid supervised training. Knowledge transfer is performed through soft label and hard label loss functions to train the three-dimensional microscopic image recognition model.

Benefits of technology

It significantly improves the reasoning ability of the 3D microscopic image recognition model, breaks through the performance bottleneck, realizes the effective utilization of 3D information, reduces the computational cost, and improves the training stability and accuracy of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120932069A_ABST
    Figure CN120932069A_ABST
Patent Text Reader

Abstract

The invention discloses a three-dimensional microscopic image recognition model training method and system based on cross-dimension knowledge distillation. According to the method, a to-be-trained three-dimensional microscopic image recognition model is used as a student model, a pre-trained two-dimensional microscopic image recognition model is used as a teacher model, and three-dimensional microscopic image sample data for strong supervision training in three-dimensional microscopic image sample data for supervision training is manually labeled to obtain a hard tag; a teacher model is adopted to identify the three-dimensional microscopic image sample data for supervised training to obtain a soft label, the three-dimensional microscopic image sample data for supervised training is utilized to perform mixed supervised knowledge distillation training, and a student model with training convergence is obtained to be used as a three-dimensional microscopic image identification model. According to the method, cross-dimension knowledge migration from a two-dimensional pre-trained image recognition large model to a three-dimensional student model is realized, and three-dimensional information is effectively utilized to improve the reasoning ability of the three-dimensional student model, so that the performance bottleneck of a three-dimensional microscopic image recognition model is broken through.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of three-dimensional image technology, and more specifically, relates to a training method for a three-dimensional microscopic image recognition model based on cross-dimensional knowledge distillation. Background Technology

[0002] The evolution of microscopic imaging technology from two-dimensional to three-dimensional represents a technological revolution spanning centuries in scientific history. The earliest microscopic imaging can be traced back to the single-lens optical microscope invented by the 17th-century Dutch scientist Antonie van Leeuwenhoek, which first revealed the two-dimensional morphology of cells and microorganisms. In the late 20th century, the advent of the confocal microscope marked the beginning of three-dimensional imaging. It utilized pinhole filtering technology to achieve optical sectioning, reconstructing the three-dimensional structure through layer-by-layer scanning; however, due to speed limitations, it was mainly used for stationary samples. It was only after the 21st century that two-photon microscopy, super-resolution fluorescence microscopy, and light-sheet microscopy were developed and applied.

[0003] Microscopic imaging technology has revolutionized applications in life sciences and medicine by breaking through the limits of resolution and dynamic observation. For example, in assisted diagnosis, microscopic images of samples are used as input, and image recognition models are employed to identify pathological samples, thus aiding doctors in making diagnoses. Theoretically, three-dimensional images contain more information, making it easier for microscopic image recognition models to accurately identify samples. However, due to its longer development history and wider application, manually annotated two-dimensional microscopic imaging data far surpasses three-dimensional microscopic imaging data in both scale and quality. Limited by the scale and quality of manually annotated data, the recognition accuracy of three-dimensional microscopic image recognition models needs improvement, thus restricting their application scope.

[0004] Traditional two-dimensional microscopic image recognition models suffer from significant dimensionality limitations. While two-dimensional microscopic images are relatively easy to acquire and have low annotation costs, their inherent planar characteristics fail to accurately reflect true three-dimensional spatial distribution features, such as tumor morphology and distribution, leading to systematic biases in model applications. Image recognition models based on three-dimensional microscopic images and trained end-to-end using 3D pathological data face severe challenges. Due to the difficulty in annotating pathological data and the limited sample size, models trained from scratch often struggle to overcome performance bottlenecks, exhibiting significant deficiencies in feature extraction and generalization capabilities. Furthermore, the limited scale of validation data makes it difficult to comprehensively evaluate the model's generalization ability.

[0005] Diba et al. proposed a "Spatiotemporal Channel Correlation (STC)" module, which models the spatiotemporal correlation between channels of a 3D CNN through spatial and temporal branches to enhance feature representation. They also introduced a knowledge transfer method from a pre-trained 2D CNN to a 3D CNN, achieving cross-dimensional knowledge transfer. However, this method requires the inclusion of temporal dimension information and is not applicable to knowledge transfer from a static 2D image recognition model to a 3D image recognition model. Summary of the Invention

[0006] To address the aforementioned deficiencies or improvement needs of existing technologies, this invention provides a training method and system for a three-dimensional microscopic image recognition model based on cross-dimensional knowledge distillation. The aim is to use a two-dimensional microscopic image recognition model trained with a large amount of data as the teacher model, and employ three-dimensional microscopic image sample data, combined with soft and hard labels, for hybrid supervised knowledge distillation training. This solves the technical problem that the three-dimensional microscopic recognition model cannot overcome the limitations of the reasoning ability of the two-dimensional microscopic image recognition model used as the teacher model, and that insufficient three-dimensional labeled data prevents the three-dimensional microscopic recognition model from effectively utilizing the information of three-dimensional images, thus limiting its reasoning ability.

[0007] To achieve the above objectives, according to one aspect of the present invention, a method for training a three-dimensional microscopic image recognition model based on cross-dimensional knowledge distillation is provided, comprising the following steps:

[0008] The three-dimensional microscopic image recognition model to be trained is used as the student model, and the pre-trained two-dimensional microscopic image recognition model is used as the teacher model. Hard labels are obtained by manually annotating the strongly supervised three-dimensional microscopic image sample data in the supervised training data. The teacher model is used to recognize the supervised three-dimensional microscopic image sample data to obtain soft labels. Hybrid supervised knowledge distillation training is performed using the supervised three-dimensional microscopic image sample data to obtain the converged student model, which is used as the three-dimensional microscopic image recognition model.

[0009] Preferably, the training data for the three-dimensional microscopic image recognition model training method based on cross-dimensional knowledge distillation is obtained as follows:

[0010] (1) Data hard label preparation: The three-dimensional microscopic image sample data for strong supervision training is manually labeled to obtain the category of the three-dimensional microscopic image sample data, which is used as the hard label of the three-dimensional microscopic image sample data for strong supervision training; the three-dimensional microscopic image sample for strong supervision training is the manually labeled sample.

[0011] (2) Data soft label preparation: The supervised training three-dimensional microscopic image sample data, including the three-dimensional microscopic image sample data for strongly supervised training, is compressed into multiple two-dimensional microscopic image samples through two-dimensional microscopic imaging simulation; the feature vector of the multiple two-dimensional microscopic image samples is obtained by feature extraction using the teacher model; the feature vector of the multiple two-dimensional image samples is then pooled by mean pooling and fused into the feature vector of the three-dimensional microscopic image, which serves as the data soft label of the three-dimensional microscopic image sample data.

[0012] The supervised training 3D microscopic image sample data consists of strongly supervised training 3D microscopic image sample data and weakly supervised training 3D microscopic image sample data. The strongly supervised training 3D microscopic image sample data has both hard and soft labels, while the weakly supervised training 3D microscopic image sample data only has soft labels.

[0013] Preferably, in the training method for the three-dimensional microscopic image recognition model based on cross-dimensional knowledge distillation, the amount of three-dimensional microscopic image sample data used for strong supervision training is not less than 1 / 10 of the amount of three-dimensional microscopic image sample data used for weak supervision training. In a preferred embodiment, the ratio of the amount of three-dimensional microscopic image sample data used for strong supervision training to the amount of three-dimensional microscopic image sample data used for weak supervision training is between 1:2 and 1:10.

[0014] When the supervised training three-dimensional microscopic image sample data is a pathological sample, the amount of three-dimensional microscopic image sample data for strong supervision training is greater than or equal to 20,000, and the amount of three-dimensional microscopic image sample data for weak supervision training is greater than or equal to 40,000.

[0015] Preferably, the training method for the three-dimensional microscopic image recognition model based on cross-dimensional knowledge distillation trains the student model with the goal of minimizing the loss function; the loss function is used to characterize the comprehensive difference between the three-dimensional microscopic image recognition model as the student model and the two-dimensional microscopic image recognition model as the teacher model, as well as the difference between the student model and the manually labeled results.

[0016] Preferably, in the training method for the three-dimensional microscopic image recognition model based on cross-dimensional knowledge distillation, the loss function is expressed as a weighted sum of hard label loss and soft label loss; the loss function LOSS is denoted as:

[0017] LOSS=λLOSS h +LOSS s

[0018] Among them, LOSS h λ is the hard label loss, i.e., the binary cross-entropy loss; λ is the weight coefficient of the hard label loss, taking values ​​between [0,1]; LOSS s This is the soft label loss.

[0019] Preferably, the training method for the three-dimensional microscopic image recognition model based on cross-dimensional knowledge distillation has a low loss function. s The soft-label loss is calculated as follows:

[0020] LOSS s =αD KL +(1-α)MSE

[0021] Where α is the weighting coefficient, with a value range of (0,1), and DKL The divergence soft label loss and the MSE soft label loss are calculated using the following methods:

[0022]

[0023] During training, the KL divergence between the feature vector P output by the teacher model and the feature vector Q output by the student model for supervised training 3D microscopic image sample data is calculated; where x represents all dimensions of each image feature vector, z is a certain dimension of the feature vector, P(x) and Q(x) are the values ​​corresponding to dimension x in the feature vector, respectively; n is the number of samples.

[0024]

[0025] Where P is the feature vector output by the teacher model for supervised training using 3D microscopic image sample data, Q is the feature vector output by the student model for supervised training using 3D microscopic image sample data, ||·||2 is the Euclidean norm, and n is the number of samples.

[0026] Preferably, the training method for the three-dimensional microscopic image recognition model based on cross-dimensional knowledge distillation adopts progressive training for knowledge distillation training. That is, during the knowledge distillation training process, the weights of the prediction difference of the soft-label data samples and the prediction difference of the student model for the hard-label data samples are adjusted according to the training level: the higher the training level of the student model, the smaller the weight of the prediction difference of the soft-label data samples in the loss function.

[0027] Preferably, in the training method for the three-dimensional microscopic image recognition model based on cross-dimensional knowledge distillation, the student model and the teacher model are based on the same architecture.

[0028] Preferably, the training method for the three-dimensional microscopic image recognition model based on cross-dimensional knowledge distillation includes a student model specifically comprising an image preprocessing module for mapping the input image to the embedding space to obtain a dense vector of high-dimensional semantic features; the image preprocessing module includes a one-dimensional position encoder and / or an adaptor.

[0029] The introduced learnable one-dimensional positional encoding has the same dimension as the embedding vector, and its output positional encoding is directly superimposed on the extracted feature vector.

[0030] The adaptive takes a dense vector of an encoded input image as input and outputs a dense vector of the same dimension, comprising a lower projection layer, a deep 3D convolutional layer, and an upper projection layer connected in series, and has trainable parameters.

[0031] According to another aspect of the present invention, a training system for a three-dimensional microscopic image recognition model based on cross-dimensional knowledge distillation is provided, which is an electronic device or a non-transitory computer-readable storage medium.

[0032] The electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the steps of the training method for a three-dimensional microscopic image recognition model based on cross-dimensional knowledge distillation provided by the present invention.

[0033] The non-transitory computer-readable storage medium stores a computer program thereon, which, when executed by a processor, implements the steps of the training method for a three-dimensional microscopic image recognition model based on cross-dimensional knowledge distillation provided by the present invention.

[0034] In summary, compared with the prior art, the above-described technical solutions conceived by this invention can achieve the following beneficial effects:

[0035] The present invention provides a training method and system for a three-dimensional microscopic image recognition model based on cross-dimensional knowledge distillation. It uses a two-dimensional microscopic image recognition model trained on a large amount of data as the teacher model. For three-dimensional microscopic image sample data, soft and hard labels are created respectively, and strong and weak mixed supervised knowledge distillation training is performed. This achieves cross-dimensional knowledge transfer from a two-dimensional pre-trained image recognition model to a three-dimensional student model and effectively utilizes three-dimensional information to improve the reasoning ability of the three-dimensional student model. This overcomes the performance bottleneck of three-dimensional microscopic image recognition models, especially pathological models, where the scarcity of manually labeled training data makes it difficult to obtain a three-dimensional microscopic image recognition model with superior reasoning ability through direct training.

[0036] In a preferred embodiment, this invention employs a multidimensional loss function and a dynamic training strategy, innovatively combining the multidimensional loss function of soft-label loss and hard-label loss, and adopting a dynamic loss switching mechanism to adjust the relative weights of soft-label loss and hard-label loss as training progresses, thereby avoiding the problem of knowledge forgetting, significantly improving the model's convergence speed and training stability, and enabling the model to maintain excellent inference performance even under limited supervised training conditions using three-dimensional microscopic image sample data.

[0037] The preferred approach is to add an adaptor to the student model to achieve dimensional adaptation from a 2D model to a 3D model. This adaptor has lightweight parameters, and only a few parameters need to be fine-tuned to achieve dimensional transformation, which significantly reduces the computational cost and avoids the risk of overfitting caused by full-parameter training. This allows the model to efficiently learn the feature representation of 3D images even under limited supervised training using 3D microscopic image sample data. Attached Figure Description

[0038] Figure 1This is a schematic diagram of the training method for a three-dimensional microscopic image recognition model based on cross-dimensional knowledge distillation provided by the present invention;

[0039] Figure 2 This is a schematic diagram of the student model provided in an embodiment of the present invention;

[0040] Figure 3 These are the test results of the student model recognition accuracy provided in Embodiment 1 of the present invention;

[0041] Figure 4 This is the test result of the student model recognition accuracy provided in Embodiment 2 of the present invention. Detailed Implementation

[0042] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention. Furthermore, the technical features involved in the various embodiments of this invention described below can be combined with each other as long as they do not conflict with each other.

[0043] The present invention provides a method for training a three-dimensional microscopic image recognition model based on cross-dimensional knowledge distillation, comprising the following steps:

[0044] The 3D microscopic image recognition model to be trained is used as the student model, and the pre-trained 2D microscopic image recognition model is used as the teacher model. Hard labels are obtained by manually annotating the strongly supervised 3D microscopic image sample data in the supervised training data. Soft labels are obtained by using the teacher model to recognize the supervised 3D microscopic image sample data. Hybrid supervised knowledge distillation training is then performed using the supervised 3D microscopic image sample data to obtain a converged student model, which is then used as the 3D microscopic image recognition model. Specifically:

[0045] The training data was obtained using the following method:

[0046] (1) Data hard label preparation: The three-dimensional microscopic image sample data for strong supervision training is manually labeled to obtain the category of the three-dimensional microscopic image sample data, which is used as the hard label of the three-dimensional microscopic image sample data for strong supervision training; the three-dimensional microscopic image sample for strong supervision training is the manually labeled sample.

[0047] (2) Data soft label preparation: The supervised training three-dimensional microscopic image sample data, including the three-dimensional microscopic image sample data for strongly supervised training, is compressed into multiple two-dimensional microscopic image samples through two-dimensional microscopic imaging simulation; the feature vector of the multiple two-dimensional microscopic image samples is obtained by feature extraction using the teacher model; the feature vector of the multiple two-dimensional image samples is averaged and pooled to form the feature vector of the three-dimensional microscopic image, which is used as the data soft label of the three-dimensional microscopic image sample data.

[0048] The supervised training 3D microscopic image sample data consists of strongly supervised training 3D microscopic image sample data and weakly supervised training 3D microscopic image sample data. The strongly supervised training 3D microscopic image sample data has both hard and soft labels, while the weakly supervised training 3D microscopic image sample data only has soft labels.

[0049] Strongly supervised training requires manually labeled 3D microscopic image samples, which limits their quantity. Weakly supervised training uses a larger number of 3D microscopic image samples than supervised training. To ensure better utilization of 3D information, the amount of 3D microscopic image samples used in strongly supervised training should be no less than 1 / 10 of that used in weakly supervised training. Ideally, the ratio of strongly supervised to weakly supervised training 3D microscopic image samples should be between 1:2 and 1:10 to balance the learning effect of large amounts of data and the training effect of 3D information. For pathological samples, to improve training accuracy, the amount of 3D microscopic image samples used in strongly supervised training should be greater than or equal to 20,000, and the amount used in weakly supervised training should be greater than or equal to 40,000.

[0050] The student model is trained with the goal of minimizing the loss function. The loss function is used to characterize the comprehensive difference between the 3D microscopic image recognition model (as the student model) and the 2D microscopic image recognition model (as the teacher model), as well as the difference between the student model and the manually labeled results. Specifically, the loss function is expressed as a weighted sum of hard label loss and soft label loss. By using the comprehensive difference as the loss function, the 3D student model can not only distill the teacher model to achieve rapid training, but also learn 3D information, achieving cross-dimensional distillation. This breaks through the limitation of the teacher model's training effect on the student model, and achieves the training effect of the student model surpassing that of the teacher model.

[0051] To accommodate the differences in data properties between hard and soft labels, the hard label loss is the binary cross-entropy loss of the student model on the strongly supervised training 3D microscopic image sample data, while the soft label loss is the weighted sum of the KL divergence and MSE loss of the student model on the supervised training 3D microscopic image samples. Hard labels are black-and-white, with limited information content, only forcing the student model to fit the true distribution. Cross-entropy, on the other hand, forces the student model to predict a probability close to 1 for class k and close to 0 for other classes, meeting the "hard classification" requirement of hard labels. Soft labels, however, carry the teacher model's "tacit knowledge" (such as class correlation), guiding the student model to learn more generalized feature representations. KL divergence emphasizes fitting the teacher distribution P from the student distribution Q, rather than asymmetrically penalizing the difference between Q and P. Furthermore, KL divergence is sensitive to the "shape" of the distribution, effectively capturing the relative confidence between categories in the teacher model, making it suitable for conveying the implicit knowledge of the teacher model. MSE, on the other hand, symmetrically penalizes differences across all categories, which is suitable for scenarios requiring a uniform fit of all class probabilities, but may overemphasize low-confidence categories, diluting the transmission of high-confidence knowledge. Therefore, KL divergence combined with MSE is used to form a soft-label loss, achieving a complementary effect. This soft-label loss, a composite of KL divergence and MSE loss, avoids overlearning of the teacher model's predictions and absorbs the implicit knowledge of the teacher model. Combined with hard-label loss, it achieves cross-dimensional distillation from a 2D image recognition model to a 3D image recognition model.

[0052] Specifically, the loss function LOSS is denoted as:

[0053] LOSS=λLOSS h +LOSS s

[0054] Among them, LOSS h λ is the hard label loss, i.e., the binary cross-entropy loss; λ is the weight coefficient of the hard label loss, taking values ​​between [0,1]; LOSS s The soft-label loss is calculated as follows:

[0055] LOSS s =αD KL +(1-α)MSE

[0056] Where α is the weighting coefficient, with a value range of (0,1), and D KL The divergence soft label loss and the MSE soft label loss are calculated using the following methods:

[0057]

[0058] During training, the KL divergence between the feature vector P output by the teacher model and the feature vector Q output by the student model for supervised training using 3D microscopic image sample data is calculated; where, This represents all dimensions of each image feature vector, where x is a certain dimension of the feature vector, P(x) and Q(x) are the values ​​corresponding to dimension x in the feature vector, respectively; n is the number of samples.

[0059]

[0060] Where P is the feature vector output by the teacher model for supervised training using 3D microscopic image sample data, Q is the feature vector output by the student model for supervised training using 3D microscopic image sample data, ||·||2 is the Euclidean norm, and n is the number of samples.

[0061] The preferred approach employs progressive training for knowledge distillation. This involves adjusting the weights of the prediction differences for soft-labeled samples and the prediction differences for hard-labeled samples in the student model according to the training level: the higher the training level of the student model, the smaller the weight of the prediction differences for soft-labeled samples in the loss function. In the initial stage of model training, only soft-label loss is used to constrain the model, i.e., the weight coefficient λ of the hard-label loss is 0. As the number of training rounds increases, the hard-label loss constraint is activated, and its weight coefficient λ is continuously increased in subsequent training. To ensure that the model does not forget the knowledge learned from the teacher model due to overlearning hard labels, the weight of the hard labels typically does not exceed the weight of the soft-label loss.

[0062] Dynamically adjusting the weights of hard-label loss and soft-label loss during knowledge distillation training can accelerate model convergence and prevent student models from hitting performance bottlenecks due to insufficient ability or over-reliance on the teacher. The main adjustment strategy is "emphasize soft labels in the early stage and hard labels in the later stage," that is, "learn the thought process in the early stage and learn the judgment in the later stage." In the early stage of knowledge distillation, the student model is not capable enough and needs to quickly absorb the experience of the teacher model (such as feature representation and category association) through soft labels. At this time, soft labels are the main source of knowledge. In the later stage of knowledge distillation, the student model has acquired certain capabilities and needs to combine hard labels to correct potential errors in the teacher model and form decision-making capabilities independent of the teacher. At this time, hard labels are the key calibration signal. The student model is based on the same architecture as the teacher model and generally includes convolutional layers. Unlike the teacher model, which uses two-dimensional convolutional kernels, the student model's convolutional layers are correspondingly extended to three-dimensional convolutional kernels. The student model specifically includes an image preprocessing module, and generally also includes feature extraction modules such as Transformer encoders, as well as a classifier.

[0063] The image preprocessing module is used to map the input image to the embedding space, obtain a dense vector of high-dimensional semantic features and output it. The output dense vector is used for feature extraction and classification.

[0064] In a preferred embodiment, the image preprocessing module includes a one-dimensional position encoder (3D Position Embedding). The one-dimensional position encoder introduces a learnable one-dimensional position code with the same dimension as the embedding vector. The position code output by the encoder is directly superimposed on the extracted feature vector to compensate for the insensitivity of the Transformer encoder to the input order, preserve the differences brought about by the three-dimensional image information, and facilitate cross-dimensional distillation.

[0065] In a preferred embodiment, the image preprocessing module includes an adaptor; the input of the adaptor is a dense vector encoding the input image, and its output is a dense vector of the same dimension, including a lower projection layer, a deep 3D convolutional layer, and an upper projection layer connected in sequence, and has trainable parameters; the adaptor enhances the information fusion of 3D images through the deep 3D convolutional layer, so that the student model suitable for 3D images and the teacher model suitable for 2D images adopt the same architecture, ensuring the distillation training effect while realizing cross-dimensional information preservation and training from 2D images to 3D images, and achieving cross-dimensional distillation by combining soft label and hard label joint training.

[0066] The Transformer encoder is used to extract feature sequences containing global information, including:

[0067] Multiple Transformer sub-layers, each of which includes a multi-head self-attention layer for self-attention encoding, and a feedforward network concatenated with the multi-head self-attention layer for non-linear transformation of features;

[0068] In the preferred embodiment, a multi-head self-attention layer and a feedforward network are followed by residual connections to link the inputs and outputs of each sub-layer, and a layer normalization layer to mitigate gradient vanishing.

[0069] The classifier is used to identify the category of the input 3D image based on a sequence of features.

[0070] The following is an example:

[0071] Example 1: Training of a 3D Microscopic Image Recognition Model for Predicting Lymph Node Metastasis in Breast Cancer

[0072] Accurate assessment of tumor metastasis characteristics is a key challenge in the diagnosis and treatment of various solid tumors. Lymph node metastasis, as one of the main routes of spread for breast cancer and other malignant tumors, requires early screening and diagnosis. This is not only crucial for determining cancer staging and developing treatment plans, but also a key indicator for assessing patient prognosis. Traditional two-dimensional pathological sections cannot fully reflect the three-dimensional spatial characteristics of tumors, potentially leading to missed diagnoses of small lymph node metastases and delaying optimal treatment. High-resolution light section microscopy, which acquires three-dimensional pathological data, overcomes this limitation. Its non-destructive imaging characteristics preserve the complete spatial structure of the tumor, significantly improving sample utilization and enabling more accurate identification of early small metastatic lesions that are difficult to detect using traditional methods. This facilitates the early detection of small lesions. It demonstrates unique diagnostic advantages in detecting key sites such as sentinel lymph nodes in breast cancer. Combined with innovative cross-dimensional knowledge distillation technology, this method fully utilizes prior knowledge from existing two-dimensional data and improves diagnostic accuracy through three-dimensional feature analysis, providing a reliable basis for the early diagnosis of lymph node metastasis. Furthermore, by accurately assessing the degree of tumor invasion and spatial distribution characteristics, it provides important references for prognostic analysis and individualized treatment of various solid cancers, demonstrating broad clinical application prospects.

[0073] The method for training a three-dimensional microscopic image recognition model based on cross-dimensional knowledge distillation provided in this embodiment includes the following steps:

[0074] Teacher Model Selection: The teacher model selected is the Virchow 2.0 large-scale model pre-trained on massive two-dimensional pathological data. This model is based on the Vision Transformer architecture and has 632 million parameters. It was pre-trained on 1.5 million whole-slice histopathological images of approximately 100,000 patients, has powerful feature extraction capabilities, and can achieve state-of-the-art performance in a variety of downstream computational pathology tasks.

[0075] The student model's structure: The student model uses the same VIT architecture as the teacher model, employing a three-dimensional embedding layer; in this embodiment, the student model is as follows... Figure 2 As shown, it specifically includes an image preprocessing module, a Transformer encoder, and a classifier:

[0076] The image preprocessing module is used to map the input image to the embedding space, obtain a dense vector of high-dimensional semantic features, and input it into the Transformer encoder; the image preprocessing module includes an embedding layer, a one-dimensional position encoder, and an adaptor;

[0077] The embedding layer maps the flattened vectors of the input image patches obtained from 3D image segmentation to the embedding space via linear projection. The embedding layer has a learnable weight matrix E∈R. 768×D(D is the embedding dimension, default 768) This implementation yields a sequence of dimension N×D, where N is the number of input image blocks obtained by 3D segmentation of the input image.

[0078] In this embodiment, the input image (e.g., 56×224×224×3, where 3 represents the RGB channels) is uniformly divided into fixed-size image patches. In this embodiment, each patch is 7×14×14×3, generating a total of N = 56 / 7×224 / 14×224 / 14 = 2048 patches. Each patch is flattened to obtain a one-dimensional vector of length 4116 (7×14×14×3), which is then mapped to the embedding space using linear projection. This process utilizes a learnable weight matrix E∈R. 4116×D (D is the embedding dimension, default 1280) This implementation yields a sequence of dimension N×D.

[0079] A one-dimensional position encoder (3D Position Embedding) introduces a learnable one-dimensional position code with the same dimension as the embedding vector. The output position code is directly superimposed on the image block sequence projected by the Transformer encoder.

[0080] The adaptor takes a dense vector of the encoded input image as input and outputs a dense vector of the same dimension, consisting of a series of concatenated lower projection layers, deep 3D convolutional layers, and upper projection layers. It has trainable parameters and enhances the information fusion of 3D images through deep 3D convolutional layers. This allows the student model suitable for 3D images and the teacher model suitable for 2D images to use the same architecture, ensuring the distillation training effect while achieving cross-dimensional information preservation and training from 2D to 3D images. Combined with joint training using soft and hard labels, cross-dimensional distillation is achieved.

[0081] An adaptor is a lightweight module whose core goal is to enable a pre-trained model to efficiently adapt to different tasks or data modalities by inserting a small number of trainable parameters, without modifying most of the parameters of the original model. Our model architecture uses a large 2D model to distill a 3D student model. The Adapter module is added to our student model to increase its adaptability to data of different dimensions. In summary, the Adapter module and the student model architecture, which is as similar as possible to the teacher model, together determine that our architecture is more advantageous in cross-dimensional distillation.

[0082] The Transformer encoder is used to extract feature sequences containing global information, including:

[0083] Multiple Transformer sub-layers, each including a multi-head self-attention layer for self-attention encoding, and a feedforward network concatenated with the multi-head self-attention layer for non-linear transformation of features; the multi-head self-attention layer and the feedforward network are followed by residual connections for linking the input and output of each sub-layer, and a layer normalization layer for mitigating gradient vanishing.

[0084] The Multi-Head Self-Attention (MSA) layer splits the input into h heads (e.g., h=12), calculates attention weights for each head, and concatenates the results, integrating multi-view features through linear projection. This mechanism allows the model to capture long-range dependencies between image patches.

[0085] Feed-Forward Network (FFN): Consists of two fully connected layers, with the GELU activation function used in between to perform a non-linear transformation on the features at each location.

[0086] Residual connections and layer normalization layers: The output and input of each sub-layer (MSA, FFN) are added through residual connections, and a layer normalization layer (LayerNorm) is applied to alleviate the gradient vanishing problem.

[0087] The embedded token sequence is input to a Transformer encoder consisting of L layers (12 layers in this embodiment). The final output of the encoder is a feature sequence containing global information, where the mean pooling result of the token sequence serves as the overall representation of the image.

[0088] The classifier is used to identify the category of the input 3D image based on the feature sequence. In this embodiment, an MLP classification head is used.

[0089] After all tokens are processed by all encoding layers, they are sent to the classification head (MLP Head) through a layer normalization layer. In this embodiment, the classification head consists of two fully connected layers, with Dropout and activation functions in between. The final output dimension is consistent with the number of target categories (such as the diseased / non-diseased classification of pathological images), thus completing the classification task.

[0090] Training data is obtained using the following method:

[0091] This embodiment employs a high-resolution light-sheet microscopy system to acquire three-dimensional pathological data of lymphoma tissue. This imaging technique enables subcellular resolution three-dimensional imaging, fully preserving the spatial structural information of the tumor tissue. To address the characteristics of the original data, an anisotropy correction algorithm is used to eliminate resolution differences along different axes. To adapt to the input requirements of deep learning models, the large volume of data is segmented into standardized three-dimensional block data, serving as supervised training samples for three-dimensional microscopic images, totaling approximately 60,000 image blocks. Image specifications: 56×224×224 pixels.

[0092] (1) Data hard label preparation: supervised training three-dimensional microscopic image sample data, of which 20,000 image blocks are used as strongly supervised training three-dimensional image samples, which are manually labeled. The hard label value of the supervised training three-dimensional microscopic image sample for breast cancer lymph node metastasis is 1, and the hard label value of other supervised training three-dimensional microscopic image samples is 0.

[0093] (2) Data soft label preparation: Input all supervised training 3D microscopic image sample data into the teacher model to obtain soft labels. The specific steps are as follows:

[0094] The supervised training 3D microscopic image sample data with a specification of 56×224×224 pixels is sliced ​​pixel by pixel along the Z-axis (the scanning direction of the light sheet) to obtain 56 224×224 pixel 2D microscopic images. These images are then input into the teacher model for feature extraction, resulting in feature vectors from the 56 2D microscopic images. These feature vectors are then fused using mean pooling to obtain the feature vectors of the supervised training 3D microscopic image sample data, which are the soft labels of the supervised training 3D microscopic image sample data.

[0095] The student model is trained with the goal of minimizing the loss function. The loss function used in this embodiment is:

[0096] The loss function LOSS is denoted as:

[0097] LOSS=λLOSS h +LOSS s

[0098] Among them, LOSS h λ is the hard label loss, i.e., the binary cross-entropy loss; λ is the weight coefficient of the hard label loss, taking values ​​between [0,1]; LOSS s The soft-label loss is calculated as follows:

[0099] LOSS s =αD KL +(1-α)MSE

[0100] Where α is the weighting coefficient, with a value range of (0,1), and D KLThe divergence soft label loss and the MSE soft label loss are calculated using the following methods:

[0101]

[0102] During training, the KL divergence between the feature vector P output by the teacher model and the feature vector Q output by the student model for supervised training using 3D microscopic image sample data is calculated; where, This represents all dimensions of each image feature vector, where x is a certain dimension of the feature vector, P(x) and Q(x) are the values ​​corresponding to dimension x in the feature vector, respectively; n is the number of samples.

[0103]

[0104] Where P is the feature vector output by the teacher model for supervised training using 3D microscopic image sample data, Q is the feature vector output by the student model for supervised training using 3D microscopic image sample data, ||·||2 is the Euclidean norm, and n is the number of samples.

[0105] Kullback-Leibler divergence (KL divergence) is a metric that measures the difference between two probability distributions and is widely used in information theory, statistics, and machine learning. It quantifies the amount of information lost when approximating one probability distribution P with another, such as a probability distribution Q. In machine learning, KL divergence is often used as a loss function to optimize model distributions (such as generative models and variational autoencoders) to approximate the true data distribution.

[0106] MSE is a commonly used metric to measure the difference between predicted and true values, and is widely used in regression tasks. It calculates the average of the squared differences between predicted and true values, emphasizing larger errors (square effect) and being sensitive to outliers.

[0107] In this embodiment, α is set to 0.5. Progressive training is used for knowledge distillation training: in the initial stage of 0 to 50 rounds, the weight coefficient λ of the hard label loss is 0; in the 51st to 99th rounds, the weight coefficient λ of the hard label loss is 0.02; and in the final stage after 100 rounds, the weight coefficient λ of the hard label loss is 0.2.

[0108] Example 2: Training of a 3D Microscopic Image Recognition Model for Structural Grading Prediction of Three-Dimensional Prostate Biopsy Samples

[0109] Accurately assessing the histological features of prostate cancer is a significant challenge in clinical diagnosis and treatment. Unlike breast cancer, which focuses on minute lesions, prostate cancer diagnosis relies more heavily on the accurate identification of glandular structural heterogeneity and spatial arrangement patterns. Traditional two-dimensional pathological sections struggle to fully represent the three-dimensional structural features of prostate tissue, potentially leading to the loss of crucial diagnostic information. High-resolution light section microscopy overcomes this limitation by acquiring three-dimensional pathological data. Its non-destructive imaging characteristics fully preserve the spatial topological relationships of glandular structures, significantly improving the utilization rate of tissue samples and enabling the precise identification of subtle structural abnormalities that are difficult to detect using traditional methods. This demonstrates a unique advantage in the critical diagnostic process of prostate cancer.

[0110] The three-dimensional microscopic image recognition model training method based on cross-dimensional knowledge distillation provided in this embodiment includes the following steps:

[0111] Teacher model selection: The teacher model is a large model pre-trained on a large amount of two-dimensional pathological data, such as the Prov-GigaPath model, or the Virchow model, the same as in Example 1. This example uses the Virchow model, which has powerful feature extraction capabilities.

[0112] The structure of the student model: The student model is based on the 3D vision Transformer architecture, and it has the same image preprocessing module as the student model in Example 1.

[0113] Training data is obtained using the following method:

[0114] This study used a high-resolution light-sheet microscopy system to acquire three-dimensional pathological data of prostate tissue. This imaging technique enabled three-dimensional overall imaging of the glandular structure, fully preserving key histological features.

[0115] (1) Data hard label preparation: supervised training three-dimensional microscopic image sample data, of which about 10,000 image blocks are used as three-dimensional image samples for strong supervised training and are manually labeled. Based on the gold standard of Gleason classification of prostate puncture samples, we divide the tissue structure of the samples into four types: low risk, grade three, grade four and grade five, thereby obtaining hard labels for the images.

[0116] (2) Preparation of data soft tags: Same as in Example 1.

[0117] The student model is trained with the goal of minimizing the loss function. The loss function and training process used in this embodiment are similar to those in Embodiment 1:

[0118] In this embodiment, α is set to 0.5. Progressive training is used for knowledge distillation training: in the initial stage of 0 to 50 rounds, the weight coefficient λ of the hard label loss is 0; in the 51st to 99th rounds, the weight coefficient λ of the hard label loss is 0.02; and in the final stage after 100 rounds, the weight coefficient λ of the hard label loss is 0.2.

[0119] like Figure 3 , 4 As shown, in Examples 1 and 2, the knowledge distillation model outperformed both the standalone two-dimensional teacher model and the three-dimensional student model obtained directly through strongly supervised training on the test set. Furthermore, the adaptor improved the performance of the knowledge distillation model. Data from Examples 1 and 2 demonstrate that using high-resolution 3D pathological data to replace traditional 2D slides fully preserves the spatial structural information of tumor tissue, more realistically reflecting the three-dimensional distribution characteristics of lesions. This innovation overcomes the information limitations of traditional 2D pathological analysis, providing more accurate spatial morphological evidence for cancer prediction and significantly improving diagnostic reliability.

[0120] Those skilled in the art will readily understand that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A training method for a three-dimensional microscopic image recognition model based on cross-dimensional knowledge distillation, characterized in that, Includes the following steps: The three-dimensional microscopic image recognition model to be trained is used as the student model, and the pre-trained two-dimensional microscopic image recognition model is used as the teacher model. Hard labels are obtained by manually annotating the strongly supervised three-dimensional microscopic image sample data in the supervised training data. The teacher model is used to recognize the supervised three-dimensional microscopic image sample data to obtain soft labels. Hybrid supervised knowledge distillation training is performed using the supervised three-dimensional microscopic image sample data to obtain the converged student model, which is used as the three-dimensional microscopic image recognition model.

2. The training method for a three-dimensional microscopic image recognition model based on cross-dimensional knowledge distillation as described in claim 1, characterized in that, The training data was obtained using the following method: (1) Data hard label preparation: The three-dimensional microscopic image sample data for strong supervision training is manually labeled to obtain the category of the three-dimensional microscopic image sample data, which is used as the hard label of the three-dimensional microscopic image sample data for strong supervision training; the three-dimensional microscopic image sample for strong supervision training is the manually labeled sample. (2) Data soft label preparation: The supervised training three-dimensional microscopic image sample data, including the three-dimensional microscopic image sample data for strongly supervised training, is compressed into multiple two-dimensional microscopic image samples through two-dimensional microscopic imaging simulation; the feature vector of the multiple two-dimensional microscopic image samples is obtained by feature extraction using the teacher model; the feature vector of the multiple two-dimensional image samples is then pooled by mean pooling and fused into the feature vector of the three-dimensional microscopic image, which serves as the data soft label of the three-dimensional microscopic image sample data. The supervised training 3D microscopic image sample data consists of strongly supervised training 3D microscopic image sample data and weakly supervised training 3D microscopic image sample data. The strongly supervised training 3D microscopic image sample data has both hard and soft labels, while the weakly supervised training 3D microscopic image sample data only has soft labels.

3. The training method for a three-dimensional microscopic image recognition model based on cross-dimensional knowledge distillation as described in claim 2, characterized in that, The amount of 3D microscopic image sample data used for strongly supervised training shall not be less than 1 / 10 of the amount of 3D microscopic image sample data used for weakly supervised training. In the preferred embodiment, the ratio of the amount of 3D microscopic image sample data used for strongly supervised training to the amount of 3D microscopic image sample data used for weakly supervised training shall be between 1:2 and 1:

10. When the supervised training three-dimensional microscopic image sample data is a pathological sample, the amount of three-dimensional microscopic image sample data for strong supervision training is greater than or equal to 20,000, and the amount of three-dimensional microscopic image sample data for weak supervision training is greater than or equal to 40,000.

4. The training method for a three-dimensional microscopic image recognition model based on cross-dimensional knowledge distillation as described in claim 1, characterized in that, The student model is trained with the goal of minimizing the loss function; the loss function is used to characterize the comprehensive difference between the three-dimensional microscopic image recognition model as the student model and the two-dimensional microscopic image recognition model as the teacher model, as well as the difference between the student model and the manually labeled results.

5. The training method for a three-dimensional microscopic image recognition model based on cross-dimensional knowledge distillation as described in claim 4, characterized in that, The loss function is expressed as a weighted sum of hard-label loss and soft-label loss; the loss function LOSS is denoted as: LOSS=λLOSS h +LOSS s Among them, LOSS h λ is the hard label loss, i.e., the binary cross-entropy loss; λ is the weight coefficient of the hard label loss, taking values ​​between [0,1]; LOSS s This is the soft label loss.

6. The training method for a three-dimensional microscopic image recognition model based on cross-dimensional knowledge distillation as described in claim 5, characterized in that, LOSS s The soft-label loss is calculated as follows: LOSS s =αD KL +(1-a)MSE Where α is the weighting coefficient, with a value range of (0,1), and D KL The divergence soft label loss and the MSE soft label loss are calculated using the following methods: During training, the KL divergence between the feature vector P output by the teacher model and the feature vector Q output by the student model for supervised training using 3D microscopic image sample data is calculated; where, This represents all dimensions of each image feature vector, where x is a certain dimension of the feature vector, P(x) and Q(x) are the values ​​corresponding to dimension x in the feature vector, respectively; n is the number of samples. Where P is the feature vector output by the teacher model for supervised training using 3D microscopic image sample data, Q is the feature vector output by the student model for supervised training using 3D microscopic image sample data, ||·||2 is the Euclidean norm, and n is the number of samples.

7. The training method for a three-dimensional microscopic image recognition model based on cross-dimensional knowledge distillation as described in claim 5, characterized in that, Knowledge distillation training is conducted using progressive training, which means that during the knowledge distillation training process, the weights of the prediction difference of the soft-label data samples and the prediction difference of the student model for the hard-label data samples are adjusted according to the training level: the higher the training level of the student model, the smaller the weight of the prediction difference of the soft-label data samples in the loss function.

8. The training method for a three-dimensional microscopic image recognition model based on cross-dimensional knowledge distillation as described in claim 1, characterized in that, The student model and the teacher model are based on the same architecture.

9. The training method for a three-dimensional microscopic image recognition model based on cross-dimensional knowledge distillation as described in claim 8, characterized in that, The student model specifically includes an image preprocessing module, which maps the input image to the embedding space to obtain a dense vector of high-dimensional semantic features; the image preprocessing module includes a one-dimensional position encoder and / or an adaptor; The introduced learnable one-dimensional positional encoding has the same dimension as the embedding vector, and its output positional encoding is directly superimposed on the extracted feature vector. The adaptive takes a dense vector of an encoded input image as input and outputs a dense vector of the same dimension, comprising a lower projection layer, a deep 3D convolutional layer, and an upper projection layer connected in series, and has trainable parameters.

10. A training system for a three-dimensional microscopic image recognition model based on cross-dimensional knowledge distillation, characterized in that, For electronic devices or non-transitory computer-readable storage media; The electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the steps of the training method for a three-dimensional microscopic image recognition model based on cross-dimensional knowledge distillation as described in any one of claims 1 to 9. The non-transitory computer-readable storage medium stores a computer program thereon, which, when executed by a processor, implements the steps of the training method for a three-dimensional microscopic image recognition model based on cross-dimensional knowledge distillation as described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Model training method and device, image processing method and device and electronic equipment

    CN115063875A

  • Online vector map construction method and device based on volume rendering knowledge distillation

    CN118644603A

  • Image processing method and apparatus, and electronic device and storage medium

    WO2024036847A1