Pathological section image auxiliary analysis method and device based on pre-training

The pathology slide core encoder trained through multimodal fusion solves the problem of data limitations in pathology slide diagnosis, improves the efficiency and accuracy of rare disease diagnosis, and is suitable for resource-constrained clinical scenarios.

CN121095147APending Publication Date: 2025-12-09JINAN LEI ZANCHEN MEDICAL TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511129905.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-13
Publication Date
2025-12-09

AI Technical Summary

Technical Problem

Existing technologies are limited by data in pathological slide diagnosis, especially due to insufficient clinical data for rare diseases, resulting in poor model training performance and difficulty in effectively applying them to resource-constrained clinical scenarios.

Method used

A multimodal fusion training approach is adopted, which extracts general feature representations of pathological slides through visual self-supervised learning and visual language alignment. Combined with pathology reports and histological morphology descriptions, the core encoder of pathological slides is trained to achieve auxiliary analysis of pathological slide images.

Benefits of technology

It improves the efficiency and accuracy of auxiliary analysis of pathological slide images, supports clinical scenarios such as diagnosis of rare diseases and prediction of cancer prognosis, and solves the problem of data limitations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121095147A_ABST
    Figure CN121095147A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an auxiliary analysis method and device for a pathological section image based on pre-training, and the method comprises the following steps: training an image encoder based on a visual self-supervision strategy in a first pre-training stage, and capturing the tissue structure features of the pathological section image; in a second pre-training stage, based on alignment training of image and pathological histomorphology description, through image-text comparison loss and image-text generation loss, joint training is performed on the image encoder and the text encoder, a pre-trained pathological section core encoder is obtained, and histomorphology semantic fusion is realized; and in a third pre-training stage, carrying out cross-modal alignment training on the image encoder and the pathological report text, optimizing the semantic expression capability of the image encoder, and realizing semantic fusion of the pathological report. According to the embodiment of the invention, the efficiency and accuracy of auxiliary analysis of the pathological section image can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of computer technology, and specifically relates to an auxiliary analysis method and device based on pre-trained pathological slide images. Background Technology

[0002] With recent advancements in foundational models, the field of computational pathology has undergone a transformation, gradually shifting from the traditional paradigm of supervised training based on CNNs (Convolutional Neural Networks) or ViTs (Vision Transformers) to a new paradigm that uses a foundational model to encode regions of interest in histopathology into general and transferable feature representations through self-supervised learning. This paradigm of training a foundational model through self-supervised learning has gradually become a mainstream AI-assisted pathology diagnostic method.

[0003] However, this paradigm remains limited by the limited clinical data from disease-specific cohorts, particularly for rare clinical diseases, in addressing the complex clinical challenges of patient and pathological slide diagnosis.

[0004] Application content

[0005] The purpose of this application is to provide an auxiliary analysis method and apparatus based on pre-trained pathological slide images to overcome the shortcomings of existing technologies due to data limitations.

[0006] To solve the above-mentioned technical problems, this application is implemented as follows:

[0007] Firstly, an auxiliary analysis method based on pre-trained pathological slide images is provided, comprising the following steps:

[0008] The pathological slide image is divided into multiple image patches, and the image patches are vector-embedded using an image encoder to obtain a general image representation of the pathological slide image; the pathological tissue morphology description corresponding to the pathological slide image is input into a text encoder to obtain a text vector representation.

[0009] In the first pre-training stage, the image encoder is trained based on a visual self-supervised strategy to capture the tissue structure features of the pathological slide images;

[0010] In the second pre-training stage, based on the alignment training between the image and the pathological tissue morphology description, the image encoder and the text encoder are jointly trained through image-text contrast loss and image-text generation loss to obtain the pre-trained pathological slice core encoder, thereby realizing tissue morphology semantic fusion.

[0011] In the third pre-training stage, the image encoder and the pathology report text are trained across modalities to optimize the semantic expression capability of the image encoder and achieve semantic fusion of the pathology report.

[0012] The pathological slide core encoder, which has completed three stages of pre-training, is deployed on a computing power integrated machine, and the image-assisted analysis model is deployed on a portable device. The image-assisted analysis model receives the image feature expression output by the pathological slide core encoder and performs auxiliary analysis tasks on the pathological slide image based on the image feature expression to generate auxiliary analysis results.

[0013] Secondly, an auxiliary analysis device based on pre-trained pathological slide images is provided, comprising:

[0014] The encoding module is used to divide the pathological slide image into multiple image patches, use an image encoder to perform vector embedding on the image patches to obtain a general image representation of the pathological slide image; and input the pathological tissue morphology description corresponding to the pathological slide image into a text encoder to obtain a text vector representation.

[0015] The first training module is used to train the image encoder based on a visual self-supervised strategy during the first pre-training stage to capture the tissue structure features of the pathological slide image.

[0016] The second training module is used in the second pre-training stage to perform alignment training based on the image and pathological tissue morphology description. Through image-text contrast loss and image-text generation loss, the image encoder and the text encoder are jointly trained to obtain the pre-trained pathological slice core encoder, thereby realizing tissue morphology semantic fusion.

[0017] The third training module is used in the third pre-training stage to perform cross-modal alignment training between the image encoder and the pathology report text, optimize the semantic expression capability of the image encoder, and realize semantic fusion of the pathology report.

[0018] The analysis module is used to deploy the pathological slide core encoder, which has completed three-stage pre-training, on a computing all-in-one machine, and to deploy the image-assisted analysis model on a portable device. The image-assisted analysis model receives the image feature expression output by the pathological slide core encoder, and performs auxiliary analysis tasks on the pathological slide image based on the image feature expression to generate auxiliary analysis results.

[0019] This application embodiment addresses the limitations of existing technologies due to data constraints by encoding pathological slide images into general image representations and pathological tissue morphological descriptions into text vector representations, and trains the core encoder of pathological slides. It also supports multimodal fusion training methods, thereby improving the efficiency and accuracy of auxiliary analysis of pathological slide images. Attached Figure Description

[0020] Figure 1 This is a flowchart of an auxiliary analysis method based on pre-trained pathological slide images provided in an embodiment of this application;

[0021] Figure 2 This is a schematic diagram of the structure of the single-modal image encoder, single-modal text encoder, and multimodal text decoder provided in the embodiments of this application;

[0022] Figure 3 This is a schematic diagram of the deployment of the pathological slide core encoder provided in the embodiments of this application;

[0023] Figure 4 This is a schematic diagram of the structure of an auxiliary analysis device based on pre-trained pathological slide images provided in an embodiment of this application. Detailed Implementation

[0024] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0025] Artificial intelligence foundational models are gradually changing the paradigm of computational pathology by accelerating the development of AI tools for diagnosis, prognosis, and biomarker prediction from digitized pathological tissue slides. These foundational models are trained using self-supervised learning (SSL) on millions of image patches (or regions of interest) of pathological tissue slides, capturing embedded expressions of the tissue morphology of the digitized slides, such as tissue and cellular structures. These expressions then serve as the basis for clinical diagnosis, feeding into subsequent models to make pathological diagnoses or biomarker predictions. However, current methods still have many limitations, primarily because the slide-level expressions in current foundational models are formed from pudding-level expressions, and these models need to be trained from scratch for different rare diseases. Furthermore, due to the limited number of case samples for rare diseases, training from scratch is often ineffective.

[0026] To address the above issues, this application proposes a novel multimodal tumor pathology diagnostic paradigm. The goal is to develop a universal feature representation for tumor pathology slides, pre-training the model by extracting specific pathological knowledge from a large number of slides. The multimodal tumor pathology diagnostic model in this application is a multimodal pathology foundational model. Through visual self-supervised learning and visual-language alignment, this model can extract universal digital pathology slide representations and generate pathology reports suitable for resource-constrained clinical scenarios, such as rare disease retrieval and cancer prognosis prediction.

[0027] Specifically, in this application embodiment, a pathological slide of any size is segmented into a group of patch slides and embedded to form a universal representation of the entire pathological slide. The pre-training mode supports not only visual pre-training methods—such as masked image reconstruction or intra-slice contrastive learning—but also multimodal fusion training methods—involving multimodal image-text alignment methods related to pathology reports, large-scale transcriptomics, or immunohistochemistry. This application embodiment addresses the shortcomings of previous common pathological diagnostic models: previous models were primarily based on visual models, thus ignoring other supervisory signals from pathology reports or clinical prognoses. This application embodiment also addresses the cross-model retrieval capability of pathological models, a fundamental capability of pathological models.

[0028] The auxiliary analysis method based on pre-trained pathological slide images provided in this application will be described in detail below with reference to the accompanying drawings, through specific embodiments and application scenarios.

[0029] like Figure 1 The diagram shown is a flowchart of an auxiliary analysis method based on pre-trained pathological slide images provided in an embodiment of this application. The method includes the following steps:

[0030] Step 101: Divide the pathological slide image into multiple image patches, use an image encoder to perform vector embedding on the image patches to obtain a general image representation of the pathological slide image; input the pathological tissue morphology description corresponding to the pathological slide image into a text encoder to obtain a text vector representation.

[0031] Specifically, the pathological slide image can be divided into multiple image patches of 512×512 pixels each.

[0032] Step 102: In the first pre-training stage, the image encoder is trained based on a visual self-supervised strategy to capture the tissue structure features of the pathological slide image.

[0033] Step 103: In the second pre-training stage, based on the alignment training between the image and the pathological tissue morphology description, the image encoder and the text encoder are jointly trained through image-text contrast loss and image-text generation loss to obtain the pre-trained pathological slide core encoder, thereby realizing tissue morphology semantic fusion.

[0034] The image encoder employs a visual Transformer architecture, comprising six Transformer layers, each with twelve attention heads. Each attention head has a dimension of 64, and the output dimension is 768. During training, the image encoder utilizes a joint training strategy based on image mask reconstruction and knowledge distillation, employing cross-entropy loss with category labels for different cropped regions of the pathological image. The text encoder uses a BERT architecture, with an output vector dimension of 768. The input is a pathological tissue morphological description corresponding to the pathological slide image. The text vector representation and the image patch are subsequently aligned using a cross-modal alignment loss during training.

[0035] In this embodiment, the image-text contrast loss and the image-text generation loss are jointly optimized by a multimodal text decoder, which adopts the standard Transformer decoder structure.

[0036] Step 104: In the third pre-training stage, the image encoder and the pathology report text are trained across modalities to optimize the semantic expression capability of the image encoder and achieve semantic fusion of the pathology report.

[0037] In this embodiment, the pathological slide core encoder can perform position encoding through a two-dimensional linear bias attention mechanism using Eulerian distance.

[0038] Step 105: Deploy the pathological slide core encoder, which has completed three-stage pre-training, on a computing power integrated machine, and deploy the image-assisted analysis model on a portable device. The image-assisted analysis model receives the image feature expression output by the pathological slide core encoder and performs auxiliary analysis tasks on the pathological slide image based on the image feature expression to generate auxiliary analysis results.

[0039] This application embodiment addresses the limitations of existing technologies due to data constraints by encoding pathological slide images into general image representations and pathological tissue morphological descriptions into text vector representations, and trains the core encoder of pathological slides. It also supports multimodal fusion training methods, thereby improving the efficiency and accuracy of auxiliary analysis of pathological slide images.

[0040] In this embodiment of the application, a specific implementation process for auxiliary analysis of pathological slide images includes the following steps:

[0041] 1. Data preparation: The training data in this embodiment includes a large number of pathological slides and corresponding pathological reports, as well as a large number of pathological ROIs (regions of interest) and corresponding pathological tissue morphological descriptions of the ROI regions.

[0042] 2. Train an image vector embedding encoder to encode slices and patches into vector representations. The image vector embedding encoder is trained using image-text alignment and multimodal text generation. The specific steps are as follows:

[0043] (1) Specifically, two Transformer-based encoders are defined: a unimodal text encoder T(θ), selected as the BERT model; and a unimodal image encoder I(θ), selected as the ViT model. Here, θ represents the training parameters. Given a pair of image-text pairs X... I and X T The corresponding encoder distribution is: T θ (X T ) and I θ (X I The output vectors of the two encoders are ensured to have the same dimension, defined as 768. The input patch size of the image encoder is 512*512.

[0044] (2) Define the image-text set: Where b is the size of the set.

[0045] (3) Definition Image-text contrast loss:

[0046] in It is defined as the vector cosine similarity between the k-th image and the j-th text.

[0047] (4) Definition Image-text contrast loss:

[0048] in Defined as the vector cosine similarity between the k-th text and the j-th image.

[0049] (5) Loss functions of single-modal image encoder and single-modal text encoder:

[0050]

[0051] (6) The training data image is the ROI image. The ROI region is cut out from the pathological slice of any size according to 8192*8192 pixels. Then the 8192*8192 pixel region is cut into 16*16 non-overlapping 512*512 pixel puddings. The text is selected as the pathological tissue morphology description corresponding to the ROI image.

[0052] (7) Define a multimodal text decoder, using a standard Transformer decoder model. The output vector of the Vit model after encoding the image and the output vector of the Bert model after encoding the text are calculated through a cross-modal attention mechanism using the standard Transformer decoder model. The loss function is defined as:

[0053] (8) The overall loss function of the three encoders is L = λ C ·L C +λ G ·L G .

[0054] (9) Cross-modal attention is defined as: Q I K is obtained by multiplying the image encoding vector by the projection matrix. T The text encoding for Zhihu is obtained by multiplying projection matrices. d represents the dimensions of Q and K.

[0055] (10) The parameters of the single-modal image encoder are fixed by training through the above algorithm, thus forming a single-modal image encoder.

[0056] 3. The entire pre-training phase is divided into three sub-phases to ensure that the entire model captures the morphological semantics of pathological tissues at the ROI level and includes the semantic information of visual and pathological reports at the entire pathological slice level.

[0057] 4. Define a pathological slide core encoder. This encoder has a ViT structure, including 6 Transformer layers, 12 attention heads with a dimension of 64, an output layer with 768 dimensions, and a hidden layer with 3072 dimensions.

[0058] 5. The first pre-training stage is the single-modal visual pre-training stage of the pathology slide core encoder. This model is pre-trained on ROI slides. Pathology slides of arbitrary size are cropped into ROI regions of 8192*8192 pixels. Then, the 8192*8192 pixel region is divided into 16*16 non-overlapping 512*512 pixel patches. Each patch is processed by the patch encoder to generate a 768-dimensional vector representation. The entire ROI region forms a 16*16 vector representation matrix. From the 16*16 vector representation matrix, two 14*14 global cropping regions and ten 6*6 local cropping regions are randomly sampled.

[0059] (1) Feature enhancement is performed on 2 global clipping regions and 10 local clipping regions. The enhancement algorithm is horizontal flip and vertical inversion.

[0060] (2) The model trained in step 2 above is used as an image encoder to perform vector embedding representation of the ROI.

[0061] (3) Based on the defined pathological slide core encoder, a training framework was designed using knowledge distillation technology. The pathological slide core encoder is defined as a student model and a teacher model, with cross-entropy loss between the extracted feature labels. Both feature labels of the teacher and student models are derived from the class labels of the view transform model (ViT) (obtained from different croppings of the same image). The student class label is processed through a multilayer perceptron model that outputs a score vector to obtain an original student score vector, and then the softmax function is applied to calculate the student probability P. s Similarly, teacher category labeling computes a teacher's score vector using a multilayer perceptron model with one output score vector, and then applies a softmax function to calculate the teacher probability P. t The cross-entropy loss for training in the ROI region is as follows: L ROI =-∑p t logp s .

[0062] (4) Image Patch Hierarchical Network Based on Masked Image Modeling and Knowledge Distillation: Some input patches given to students are randomly masked. Then, the student category head is applied to the student mask token. Similarly, the teacher category head is applied to the teacher mask token, which corresponds to the tokens masked in the students. Then, the above softmax is applied to obtain the cross-entropy loss of the patch region:

[0063] L patch =-∑p ti logp si .

[0064] (5) Both of the above cross-entropy loss functions use a learnable multilayer vector perceptron for projection, output tokens, and calculate the cross-entropy loss. The final model's cross-entropy is as follows:

[0065] L loss =(1-α)L ROI +αL patch α is a hyperparameter, which was chosen to be 0.7 in the experiment.

[0066] (6) The pre-training position encoding in this stage adopts a two-dimensional linear bias attention mechanism based on the Eulerian distance of the feature region, which effectively improves the model training efficiency without sacrificing accuracy. The position encoding algorithm is as follows:

[0067] Where i x i y j xj y This represents the x-coordinates and y-coordinates of the i-th and j-th patches.

[0068] (7) Through the above training, a core encoder for pathological slides is formed.

[0069] 6. The second pre-training stage enhances the semantic fusion capability of the pathological slide core encoder for pathological tissue morphology by aligning the ROI representation with the pathological tissue morphological description in both image and text. Specifically, after embedding the entire 8192*8192 ROI region to form the ROI representation, the method described in step 2 above is used to perform cross-modal fusion of the ROI image and the pathological tissue description text. The specific method is as follows:

[0070] (1) Two attention pooling components are added to the pathology slide core encoder. The first attention pooling component is used to interact with the single-modal text encoder in step 2 based on L. C Training is performed using a loss function of the form [formula missing].

[0071] (2) The second attention pooling component and the multimodal text decoder in step 2 are based on L G The loss function is used for training.

[0072] (3) After training, the pathological slice core encoder integrates pathological tissue morphological semantics into the encoding of images.

[0073] 7. The third pre-training stage enhances the semantic fusion capability of the pathology slide core encoder for pathology reports by aligning the whole slide representation with the pathology report text. Specifically, after embedding the entire pathology slide to form the whole slide representation, the cross-modal fusion of the whole pathology slide image and the pathology report descriptive text is performed using a method similar to step 6.

[0074] 8. Fine-tuning stage: Based on the core encoder of the pathological slides trained above, a small model such as a linear classifier, KNN classifier or MLP is added, and a specific pathological task is selected for fine-tuning.

[0075] 9. The core encoder for pathological slides is deployed inside the integrated computing machine, while the small model fine-tuned in step 8 is deployed to a portable terminal device with lower computing power, such as... Figure 3 As shown.

[0076] 10. Small models on portable devices communicate with large models with greater computing power through a custom protocol to ultimately complete pathological-specific tasks.

[0077] like Figure 4 The diagram shown is a schematic representation of an auxiliary analysis device based on pre-trained pathological slide images provided in an embodiment of this application, comprising:

[0078] The encoding module 410 is used to divide the pathological slide image into multiple image patches, use an image encoder to perform vector embedding on the image patches to obtain a general image representation of the pathological slide image; and input the pathological tissue morphology description corresponding to the pathological slide image into a text encoder to obtain a text vector representation.

[0079] Specifically, the encoding module 410 is used to divide the pathological slide image into multiple image patches of size 512×512 pixels.

[0080] The first training module 420 is used to train the image encoder based on a visual self-supervised strategy during the first pre-training stage to capture the tissue structure features of the pathological slide image.

[0081] The second training module 430 is used in the second pre-training stage to perform alignment training based on the image and pathological tissue morphology description. Through image-text contrast loss and image-text generation loss, the image encoder and the text encoder are jointly trained to obtain the pre-trained pathological slice core encoder, thereby realizing tissue morphology semantic fusion.

[0082] The image encoder employs a visual Transformer architecture, comprising six Transformer layers, each with twelve attention heads. Each attention head has a dimension of 64, and the output dimension is 768. During training, the image encoder utilizes a joint training strategy based on image mask reconstruction and knowledge distillation, employing cross-entropy loss with category labels for different cropped regions of the pathological image. The text encoder uses a BERT architecture, with an output vector dimension of 768. The input is a pathological tissue morphological description corresponding to the pathological slide image. The text vector representation and the image patch are subsequently aligned using a cross-modal alignment loss during training.

[0083] In this embodiment, the image-text contrast loss and the image-text generation loss are jointly optimized by a multimodal text decoder, which adopts the standard Transformer decoder structure.

[0084] The third training module 440 is used to perform cross-modal alignment training between the image encoder and the pathology report text in the third pre-training stage, optimize the semantic expression capability of the image encoder, and realize semantic fusion of the pathology report.

[0085] In this embodiment, the pathological slide core encoder performs position encoding using a two-dimensional linear bias attention mechanism based on Eulerian distance.

[0086] The analysis module 450 is used to deploy the pathological slide core encoder, which has completed three-stage pre-training, on a computing power all-in-one machine, and to deploy the image-assisted analysis model on a portable device. The image-assisted analysis model receives the image feature expression output by the pathological slide core encoder, and performs auxiliary analysis tasks on the pathological slide image based on the image feature expression to generate auxiliary analysis results.

[0087] This application embodiment addresses the limitations of existing technologies due to data constraints by encoding pathological slide images into general image representations and pathological tissue morphological descriptions into text vector representations, and trains the core encoder of pathological slides. It also supports multimodal fusion training methods, thereby improving the efficiency and accuracy of auxiliary analysis of pathological slide images.

[0088] This application also provides a computer-readable storage medium storing a computer program. When executed by a processor, this computer program implements the various processes of the above-described auxiliary analysis method for pathological slide images and achieves the same technical effects. To avoid repetition, it will not be described again here. The computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0089] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0090] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0091] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

Claims

1. An auxiliary analysis method based on pre-trained pathological slide images, characterized in that, Includes the following steps: The pathological slide image is divided into multiple image patches, and the image patches are vector-embedded using an image encoder to obtain a general image representation of the pathological slide image; the pathological tissue morphology description corresponding to the pathological slide image is input into a text encoder to obtain a text vector representation. In the first pre-training stage, the image encoder is trained based on a visual self-supervised strategy to capture the tissue structure features of the pathological slide images; In the second pre-training stage, based on the alignment training between the image and the pathological tissue morphology description, the image encoder and the text encoder are jointly trained through image-text contrast loss and image-text generation loss to obtain the pre-trained pathological slice core encoder, thereby realizing tissue morphology semantic fusion. In the third pre-training stage, the image encoder and the pathology report text are trained across modalities to optimize the semantic expression capability of the image encoder and achieve semantic fusion of the pathology report. The pathological slide core encoder, which has completed three stages of pre-training, is deployed on a computing power integrated machine, and the image-assisted analysis model is deployed on a portable device. The image-assisted analysis model receives the image feature expression output by the pathological slide core encoder and performs auxiliary analysis tasks on the pathological slide image based on the image feature expression to generate auxiliary analysis results.

2. The method according to claim 1, characterized in that, The process of dividing the pathological slide image into multiple image patches specifically includes: The pathological slide image is divided into multiple image patches of 512×512 pixels each; The image encoder adopts a visual Transformer structure, which includes 6 Transformer layers, each layer includes 12 attention heads, the attention head dimension is 64, and the output dimension is 768. The image encoder is trained based on a joint training strategy of image mask reconstruction and knowledge distillation, and uses cross-entropy loss with category labels for different cropped regions of pathological images.

3. The method according to claim 1, characterized in that, The text encoder adopts a BERT structure, with an output vector dimension of 768. The input is a pathological tissue morphological description corresponding to the pathological slide image. The text vector representation and the image patch are aligned using cross-modal alignment loss during subsequent training.

4. The method according to claim 1, characterized in that, The image-text contrast loss and the image-text generation loss are jointly optimized by a multimodal text decoder, which adopts the standard Transformer decoder structure.

5. The method according to claim 1, characterized in that, The pathological slide core encoder performs position encoding using a two-dimensional linear bias attention mechanism based on Eulerian distance.

6. An auxiliary analysis device based on pre-trained pathological slide images, characterized in that, include: The encoding module is used to divide the pathological slide image into multiple image patches, use an image encoder to perform vector embedding on the image patches to obtain a general image representation of the pathological slide image; and input the pathological tissue morphology description corresponding to the pathological slide image into a text encoder to obtain a text vector representation. The first training module is used to train the image encoder based on a visual self-supervised strategy during the first pre-training stage to capture the tissue structure features of the pathological slide image. The second training module is used in the second pre-training stage to perform alignment training based on the image and pathological tissue morphology description. Through image-text contrast loss and image-text generation loss, the image encoder and the text encoder are jointly trained to obtain the pre-trained pathological slice core encoder, thereby realizing tissue morphology semantic fusion. The third training module is used in the third pre-training stage to perform cross-modal alignment training between the image encoder and the pathology report text, optimize the semantic expression capability of the image encoder, and realize semantic fusion of the pathology report. The analysis module is used to deploy the pathological slide core encoder, which has completed three-stage pre-training, on a computing all-in-one machine, and to deploy the image-assisted analysis model on a portable device. The image-assisted analysis model receives the image feature expression output by the pathological slide core encoder, and performs auxiliary analysis tasks on the pathological slide image based on the image feature expression to generate auxiliary analysis results.

7. The apparatus according to claim 6, characterized in that, The encoding module is specifically used to divide the pathological slide image into multiple image patches of 512×512 pixels in size. The image encoder adopts a visual Transformer structure, which includes 6 Transformer layers, each layer including 12 attention heads, with an attention head dimension of 64 and an output dimension of 768. The image encoder is trained based on a joint training strategy of image mask reconstruction and knowledge distillation, and uses cross-entropy loss with category labels for different cropped regions of the pathological image.

8. The apparatus according to claim 6, characterized in that, The text encoder adopts a BERT structure, with an output vector dimension of 768. The input is a pathological tissue morphological description corresponding to the pathological slide image. The text vector representation and the image patch are aligned using cross-modal alignment loss during subsequent training.

9. The apparatus according to claim 6, characterized in that, The image-text contrast loss and the image-text generation loss are jointly optimized by a multimodal text decoder, which adopts the standard Transformer decoder structure.

10. The apparatus according to claim 6, characterized in that, The pathological slide core encoder performs position encoding using a two-dimensional linear bias attention mechanism based on Eulerian distance.