Biomarker prediction method and device, equipment and storage medium

Through a biomarker prediction method combining self-supervised learning model and weakly supervised learning method, the problems of difficult classification of tumor and normal tissue boundary, high labeling cost and poor adaptability in the prior art are solved, and high-accurate biomarker prediction is achieved.

CN120013931AActive Publication Date: 2025-05-16SUZHOU KEBANG GENE TECH CO LTD

Patent Information

Application Number
CN202510481096.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-05-16
Estimated Expiration
2045-04-17

AI Technical Summary

Technical Problem

The existing artificial intelligence prediction methods are difficult to accurately classify at the boundaries of tumors and normal tissues, resulting in reduced diagnostic accuracy, high labeling costs and poor adaptability, which limits its wide application in clinical practice.

Method used

The self-supervised learning model is used to extract image film features, and the tumor segmentation model is trained in combination with weakly supervised learning methods. A new biomarker data set is generated through screening and optimization of image blocks, thereby predicting the biomarker status.

Benefits of technology

It improves the accuracy of biomarker prediction, reduces the dependence on artificially labeled data, reduces the burden on clinicians, and improves the prediction accuracy of tumor-related biomarkers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120013931A_ABST
    Figure CN120013931A_ABST
Patent Text Reader

Abstract

The invention discloses a biomarker prediction method and device, equipment and a storage medium, and the method comprises the steps: collecting a pathological image of a patient and corresponding labeling information, carrying out the preprocessing, and building a biomarker data set and a tumor segmentation data set; training a self-supervised learning model by using the biomarker data set; training a tumor segmentation model by using the image patch features of the tumor segmentation data set; predicting the tumor probability of each image block and the corresponding image slice in the biomarker data set by adopting the trained tumor segmentation model; based on a prediction result, screening and optimizing image blocks in the biomarker data set, and generating a new biomarker data set; training a biomarker prediction model by using the image patch features of the new biomarker data set; and applying the trained model to a to-be-analyzed pathological image. According to the method, the state of the biomarker can be effectively predicted, the prediction precision is high, and the burden of clinical pathologists can be effectively relieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of pathological image processing, and specifically relates to a biomarker prediction method, device, equipment and storage medium. Background Art

[0002] Timely and accurate diagnosis of malignant tumors is crucial for patients' treatment selection and prognosis assessment. Biomarkers play a vital role in the diagnosis and treatment of cancer. They can not only improve the accuracy of diagnosis, but also play an important role in risk assessment, treatment decision-making, and efficacy monitoring. The detection of biomarkers is usually carried out by polymerase chain reaction (PCR), sequencing or immunohistochemistry analysis, which is the current mainstream method. However, for many patients in low-income and middle-income countries, the detection of genetic biomarkers has the problems of high cost and complex infrastructure. At the same time, due to the complexity of the detection methods of genetic biomarkers, the detection cycle is long, resulting in delayed treatment plans.

[0003] For example, the prostate cancer biomarker AR-V7 (androgen receptor splice variant 7) is an important focus in the current field of prostate cancer treatment. AR-V7 is a splice variant of the androgen receptor (AR), which plays a key role in treatment resistance and disease progression in prostate cancer. qRT-PCR is currently the most commonly used method for AR-V7 mRNA detection, but the technical route is complex and requires professional laboratory conditions and technicians, which limits its widespread clinical application.

[0004] The application of artificial intelligence (AI) in predicting biomarkers is developing rapidly, mainly including data mining, machine learning, deep learning, etc. Deep learning technology can process and analyze a large amount of bioinformatics and clinical data, mine the association between biomarkers and diseases, and discover new biomarkers and potential therapeutic targets. However, the existing artificial intelligence prediction methods have the following problems.

[0005] First, it is difficult to identify tumors and normal tissues, and the technology is difficult. Classification errors are prone to occur at the boundaries between tumors and normal tissues, reducing the accuracy of diagnosis. Conventional classification methods fail to accurately reflect the actual distribution of tumor cells and cannot provide the specific proportion of tumor cells. The lack of key information affects treatment decisions.

[0006] Second, high annotation cost. Accurate pixel-level segmentation requires a large amount of manual annotation by experts, which is costly and time-consuming.

[0007] Third, poor adaptability. The high resolution and staining inconsistency of pathological images make accurate pixel-level segmentation extremely challenging. Individual differences in cell morphology and staining quality make it difficult for the algorithm to be universally applicable to different samples. Summary of the invention

[0008] In order to solve the above technical problems, the present invention proposes a biomarker prediction method, device, equipment and storage medium.

[0009] In order to achieve the above object, the technical solution of the present invention is as follows: In a first aspect, the present invention discloses a biomarker prediction method, comprising: Step S1: Collect the patient's pathological images and their corresponding annotation information; Step S2: Divide each collected pathological image into several image blocks and perform preprocessing; Step S3: establishing a biomarker dataset and a tumor segmentation dataset based on the preprocessed image blocks; The biomarker dataset includes: image patches annotated with biomarker diagnostic results; The tumor segmentation dataset includes: image patches annotated with tumors and normal tissues; Step S4: using the biomarker dataset to train a self-supervised learning model, where the self-supervised learning model is used to divide each image block into a number of image slices and extract image slice features of each image slice; Step S5: using the trained self-supervised learning model to divide each image block in the tumor segmentation data set into a number of image slices, and extracting image slice features of each image slice; Step S6: using the image slice features of the tumor segmentation data set obtained in step S5 to train a tumor segmentation model, the tumor segmentation model is used to predict the tumor probability of each image block and the corresponding image slice; Step S7: using the trained tumor segmentation model to predict the tumor probability of each image block and the corresponding image slice in the biomarker data set; Step S8: Based on the prediction result of step S7, the image blocks in the biomarker data set are screened and optimized to generate a new biomarker data set; Step S9: using the trained self-supervised learning model to divide each image block in the new biomarker dataset into a number of image slices, and extracting image slice features of each image slice; Step S10: training a biomarker prediction model using the image features of the new biomarker data set obtained in step S9, where the biomarker prediction model is used to predict the biomarker status of the patient; Step S11: Apply the trained model to the pathological image to be analyzed to predict the biomarker status.

[0010] Based on the above technical solution, the following improvements can be made: As a preferred solution, step S2 includes: Step S2.1: Divide each collected pathological image into a number of image blocks according to a fixed size; Step S2.2: Adjust the resolution of all image blocks to be unified; Step S2.3: Screen and exclude invalid image blocks.

[0011] As a preferred solution, step S11 includes: Step S11.1: Collect the patient's pathological images to be analyzed; Step S11.2: Divide each collected pathological image into a number of image blocks and perform preprocessing; Step S11.3: establishing a data set to be analyzed based on the preprocessed image blocks; Step S11.4: using the trained self-supervised learning model to divide each image block in the data set to be analyzed into a number of image slices, and extracting image slice features of each image slice; Step S11.5: Based on the image slice features of the data set to be analyzed, the trained tumor segmentation model is used to predict the tumor probability of each image block and the corresponding image slice; Step S11.6: Based on the prediction result of step S11.5, the image blocks in the data set to be analyzed are screened and optimized to generate a new data set to be analyzed; Step S11.7: using the trained self-supervised learning model to divide each image block in the new data set to be analyzed into a number of image slices, and extracting the image slice features of each image slice; Step S11.8: Based on the image features of the new data set to be analyzed, the biomarker status of the patient is predicted using the trained biomarker prediction model.

[0012] As a preferred solution, step S8 and step S11.6 respectively screen and optimize the image blocks in the corresponding data set through the following steps, specifically including: Step A: According to the tumor probability predicted for each image block, each image block is judged to be a tumor image block or a normal tissue image block; According to the tumor probability predicted for each image slice, each image slice is judged to be a tumor image slice or a normal tissue image slice; Step B: Screen out all image blocks that are judged to be normal tissues in the corresponding data set; Step C: Evaluate the tumor content t of each remaining image block in the corresponding dataset, t = m / n; Where: m is the number of image slices judged as tumors in the image block; n is the number of all image slices in the image block; Step D: According to the tumor content of each image block, the image blocks whose tumor content exceeds the content threshold are screened out from the remaining image blocks in the corresponding data set; Step E: For the screened image blocks, use masking technology to remove all image slices judged as normal tissues to obtain new image blocks and form a new corresponding data set.

[0013] In a second aspect, the present invention further discloses a biomarker prediction device, comprising: A collection module, used to collect the patient's pathological images and their corresponding annotation information; A preprocessing module, used for dividing each collected pathological image into a number of image blocks and performing preprocessing; A data set building module, used to build a biomarker data set and a tumor segmentation data set based on the preprocessed image blocks; The biomarker dataset includes: image patches annotated with biomarker diagnostic results; The tumor segmentation dataset includes: image patches annotated with tumors and normal tissues; A self-supervised learning model training module, used for training a self-supervised learning model using a biomarker dataset, wherein the self-supervised learning model is used for dividing each image block into a number of image slices and extracting image slice features of each image slice; A first feature extraction module is used to divide each image block in the tumor segmentation data set into a plurality of image slices using a trained self-supervised learning model, and extract image slice features of each image slice; A tumor segmentation model training module is used to train a tumor segmentation model using the image slice features of the tumor segmentation data set obtained by the first feature extraction module, and the tumor segmentation model is used to predict the tumor probability of each image block and the corresponding image slice; A segmentation prediction module is used to predict the tumor probability of each image block and the corresponding image slice in the biomarker data set using the trained tumor segmentation model; A screening and optimization module, used for screening and optimizing image blocks in the biomarker data set based on the prediction results of the segmentation prediction module to generate a new biomarker data set; A second feature extraction module is used to divide each image block in the new biomarker data set into a number of image slices using a trained self-supervised learning model, and extract image slice features of each image slice; A biomarker prediction model training module, used to train a biomarker prediction model using the image features of the new biomarker data set obtained by the second feature extraction module, wherein the biomarker prediction model is used to predict the biomarker status of the patient; The application module is used to apply the trained model to the pathological image to be analyzed to predict the biomarker status.

[0014] As a preferred solution, the preprocessing module includes: An image block division unit, used for dividing each collected pathological image into a number of image blocks according to a fixed size; A resolution adjustment unit, used for adjusting and unifying the resolution of all image blocks; The invalid image block screening unit is used to screen and exclude invalid image blocks.

[0015] As a preferred solution, the application module includes: An application collection unit is used to collect the patient's pathological images to be analyzed; A preprocessing unit is applied to divide each collected pathological image into a number of image blocks and perform preprocessing; An application data set establishment unit is used to establish a data set to be analyzed based on the preprocessed image blocks; Applying a first feature extraction unit, for dividing each image block in the to-be-analyzed data set into a plurality of image slices using the trained self-supervised learning model, and extracting image slice features of each image slice; Applying a segmentation prediction unit, for predicting the tumor probability of each image block and the corresponding image slice using the trained tumor segmentation model based on the image slice features of the data set to be analyzed; An application screening and optimization unit is used to screen and optimize the image blocks in the data set to be analyzed based on the prediction result of the application segmentation prediction unit to generate a new data set to be analyzed; Applying a second feature extraction unit, for dividing each image block in a new data set to be analyzed into a plurality of image slices using a trained self-supervised learning model, and extracting image slice features of each image slice; The prediction unit is applied to predict the biomarker status of the patient using the trained biomarker prediction model based on the image slice features of the new data set to be analyzed.

[0016] As a preferred solution, the screening optimization module and the application screening optimization unit respectively include: A judgment unit, used to judge whether each image block is a tumor image block or a normal tissue image block according to the tumor probability predicted by each image block; According to the tumor probability predicted for each image slice, each image slice is judged to be a tumor image slice or a normal tissue image slice; A screening unit, used to screen out all image blocks judged to be normal tissues in the corresponding data set; The tumor content evaluation unit is used to evaluate the tumor content t of each remaining image block in the corresponding data set. t = m / n; Where: m is the number of image slices judged as tumors in the image block; n is the number of all image slices in the image block; A screening unit, used for screening, according to the tumor content of each image block, image blocks having a tumor content exceeding a content threshold from the remaining image blocks of the corresponding data set; The forming unit is used to remove all image slices judged as normal tissues from the screened image blocks by using a masking technique to obtain new image blocks and form a new corresponding data set.

[0017] In a third aspect, the present invention further discloses a computing device, comprising: one or more processors; Memory; and one or more programs, wherein the one or more programs are stored in a memory and configured to be executed by one or more processors, and the one or more programs include instructions for any of the above-mentioned biomarker prediction methods.

[0018] In a fourth aspect, the present invention further discloses a storage medium, characterized in that the storage medium stores one or more computer-readable programs, and the one or more programs include instructions, and the instructions are suitable for being loaded by a memory and executing any of the above-mentioned biomarker prediction methods.

[0019] The present invention discloses a biomarker prediction method, device, equipment and storage medium, which have the following beneficial effects: First, the present invention uses a self-supervised learning model to extract image features, reducing the reliance on manually labeled data.

[0020] Second, the present invention uses a weakly supervised learning method to train a tumor segmentation model based on image slice features, and can accurately segment tumor tissue and normal tissue.

[0021] Third, the present invention screens and optimizes image blocks according to the prediction results of the tumor segmentation model, so that the prediction of biomarkers is more accurate.

[0022] In summary, the present invention can effectively predict the status of biomarkers with high prediction accuracy, which can effectively reduce the burden of clinical pathologists and improve the prediction accuracy of tumor-related biomarkers, and has significant clinical significance. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments are briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without creative work.

[0024] Figure 1A flow chart of a biomarker prediction method provided in an embodiment of the present invention.

[0025] Figure 2 A schematic diagram of the process of pathological image segmentation and preprocessing provided by an embodiment of the present invention.

[0026] Figure 3 A schematic diagram of the process of extracting image block features provided by an embodiment of the present invention.

[0027] Figure 4 A schematic diagram of the flow of a tumor segmentation model provided in an embodiment of the present invention.

[0028] FIG5 (a) is an original image block provided by an embodiment of the present invention; FIG5( b ) is an image block of a predicted tumor region provided by an embodiment of the present invention.

[0029] Figure 6 The AUROC curve provided by the embodiment of the present invention.

[0030] Figure 7 A schematic diagram of the process flow of the application phase provided by an embodiment of the present invention.

[0031] Figure 8 A block diagram of a biomarker prediction device provided in an embodiment of the present invention.

[0032] Fig. 9 A block diagram of a computing device provided for an embodiment of the present invention.

[0033] Among them: 201-collection module, 202-preprocessing module, 203-dataset establishment module, 204-self-supervised learning model training module, 205-first feature extraction module, 206-tumor segmentation model training module, 207-segmentation prediction module, 208-screening optimization module, 209-second feature extraction module, 210-biomarker prediction model training module, 211-application module, 301-processor, 302-memory. DETAILED DESCRIPTION

[0034] The preferred embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0035] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0036] Using ordinal numbers “first,” “second,” “third,” etc. to describe common objects merely indicates that different instances of similar objects are involved and is not intended to imply that the objects so described must have a given order in time, space, order, or in any other manner.

[0037] In addition, the expression of “comprising” an element is an “open” expression, which merely means that corresponding components or steps exist, and should not be interpreted as excluding additional components or steps.

[0038] In order to achieve the purpose of the present invention, in some embodiments of the biomarker prediction method, the prostate cancer biomarker AR-V7 is taken as an example. Figure 1 As shown, biomarker prediction methods include: Step S101: Collecting the patient's pathological images and their corresponding annotation information; Step S102: Divide each collected pathological image into a number of image blocks and perform preprocessing; Step S103: establishing a biomarker data set and a tumor segmentation data set based on the preprocessed image blocks; The biomarker dataset includes: image patches annotated with biomarker diagnostic results; The tumor segmentation dataset includes: image patches annotated with tumors and normal tissues; Step S104: using the biomarker data set to train a self-supervised learning model, the self-supervised learning model is used to divide each image block into a number of image slices, and extract image slice features of each image slice; Step S105: using the trained self-supervised learning model to divide each image block in the tumor segmentation data set into a number of image slices, and extracting image slice features of each image slice; Step S106: using the image slice features of the tumor segmentation data set obtained in step S105 to train a tumor segmentation model, the tumor segmentation model is used to predict the tumor probability of each image block and the corresponding image slice; Step S107: using the trained tumor segmentation model to predict the tumor probability of each image block and the corresponding image slice in the biomarker data set; Step S108: Based on the prediction result of step S107, the image blocks in the biomarker data set are screened and optimized to generate a new biomarker data set; Step S109: using the trained self-supervised learning model to divide each image block in the new biomarker data set into a number of image slices, and extracting image slice features of each image slice; Step S110: training a biomarker prediction model using the image features of the new biomarker data set obtained in step S109, where the biomarker prediction model is used to predict the biomarker status of the patient; Step S111: Apply the trained model to the pathological image to be analyzed to predict the biomarker status.

[0039] Each of the above steps is explained in detail below.

[0040] Step S101 collects H&E stained pathological images and biomarker diagnosis results of prostate cancer patients.

[0041] Specifically, in this embodiment, the detection status of AR-V7 in pathological images of 400 prostate cancers in the TCGA dataset was collected to determine the status of each sample: AR-V7 positive or AR-V7 negative, including: 100 positive cases and 300 negative cases.

[0042] Step S102 divides each collected pathological image into a number of image blocks and performs preprocessing, specifically including: Step S102.1: Divide each collected pathological image into a number of image blocks according to a fixed size; Step S102.2: adjusting and unifying the resolution of all image blocks; Step S102.3: Screen and exclude invalid image blocks.

[0043] like Figure 2 As shown, specifically, in this embodiment, the pathological image is divided according to a fixed physical size of 256um×256um to obtain a number of image blocks (tlie). The resolution of the image block is adjusted to 512x512 pixels, that is, the image block is normalized to 0.5um / pixel. The image block with a tissue content ratio lower than a preset threshold (such as: 30%) is identified as an invalid image block, and the image block higher than the preset threshold is identified as a valid image block, indicating that the amount of information is rich, and the invalid image blocks are screened and excluded to reduce the amount of calculation.

[0044] Step S103 establishes a biomarker dataset and a tumor segmentation dataset.

[0045] Specifically, in this embodiment, The biomarker dataset is used for model training and final evaluation of biomarker prediction. The biomarker dataset is divided into a training set and a test set according to a 7:3 ratio, with 280 cases in the training set, including 70 positive cases and 210 negative cases; and 120 cases in the test set, including 30 positive cases and 90 negative cases.

[0046] The tumor segmentation dataset is used to train the tumor segmentation model. 500 tumor image patches and 500 normal tissue image patches are selected from the image patches in the training set of the biomarker dataset.

[0047] The tumor segmentation dataset is divided into 7:3, where the training set has 700 image blocks, including 350 tumor image blocks and 350 normal tissue image blocks, and the test set has 300 image blocks, including 150 tumor images and 150 normal tissue images.

[0048] Step S104 uses the biomarker data set to train a self-supervised learning model. The Dinov2 model is used as the backbone network of the self-supervised learning model. The self-supervised learning model can efficiently extract multi-level pathological features from pathological images.

[0049] Specifically, in this embodiment, Figure 3 As shown, the feature extraction of the self-supervised learning model includes: First, extract the image blocks , , h represents the image block height, w represents the image block width, and the scaled image block is obtained by the scaling function.

[0050]

[0051] in, Represents a scaling function, and in the embodiment, the scaled image block The size is 224×224.

[0052] Then, the scaled image blocks are The input is sent to the pre-trained model to extract features. The pre-trained model divides the image block into 16×16×3 non-overlapping image patches. Each image patch is flattened into a vector and mapped to generate the initial feature representation:

[0053] in, Indicates 16×16 non-overlapping image slices; is the linear mapping matrix,

[0054] is the bias term, , d is the feature dimension; Represents a vector dot product operation.

[0055] The pre-trained model considers the relative position between each image slice and the feature vector of each image slice. Perform positional encoding:

[0056] in, is the positional encoding, ; N is the number of image patches.

[0057] Finally, the feature vector after position encoding is Input into the multi-layer Transformer through the multi-layer attention mechanism, and obtain the feature vector of each image piece in the penultimate layer .

[0058]

[0059] Among them, MAS stands for Multi-head Self-Attention. Indicates the number of layers of Transformer.

[0060] Self-supervised learning can extract valuable features from unlabeled data. In the face of a few-sample learning environment where samples are scarce, self-supervised learning generates supervisory signals by constructing proxy tasks, effectively reducing the reliance on manually labeled data. In addition, the feature vector of the image slice not only contains the location and detail information of the image, but also covers detailed information inside the cell, which is helpful for the subsequent training of the tumor segmentation model.

[0061] Step S105 uses the trained self-supervised learning model to divide each image block in the tumor segmentation data set into a number of image slices, and extracts the image slice features of each image slice.

[0062] like Figure 4 Specifically, in this embodiment, each image block obtains 196 image slice feature vectors. , as a feature bag.

[0063] Step S106 uses a weakly supervised learning method to train a tumor segmentation model, and segments the tumor and normal tissue using the feature vectors of the image slices.

[0064] The training set is divided according to the tumor segmentation dataset, and the evaluation is performed on the test set, and the optimal tumor segmentation model is saved.

[0065] The weakly supervised learning method is to use the feature vector in the feature bag In order to influence the prediction results, the feature package is first encoded by a multi-layer perceptron (MLP), then weighted by the attention mechanism, and finally classified by the decoder.

[0066] The intermediate layer of the tumor segmentation model outputs the predicted tumor probability for each image slice , the last layer outputs the predicted tumor probability for each image block .

[0067] The number of pixels in pathological images is usually in the hundreds of thousands or millions, and it is difficult to process pathological images using pixel-based segmentation models. Based on the image slice features, a weakly supervised learning method is used to train the tumor segmentation model. This method does not require manual labeling, but only requires defining the category of the image block. This not only solves the problem of high labeling costs and time consumption, but also obtains pixel-level segmentation results through decoding of the tumor segmentation model, accurately distinguishing normal tissue from tumor tissue.

[0068] Furthermore, FIG5(a) is an original image block, and FIG5(b) is an image block showing a predicted tumor area. The red part represents the predicted tumor area, and the blue part represents the predicted normal tissue area. The present invention can achieve pixel-level segmentation.

[0069] Step S107 uses the trained tumor segmentation model to predict the tumor probability of each image block and the corresponding image slice in the biomarker data set.

[0070] Step S108 is to screen and optimize the image blocks in the biomarker data set based on the prediction result of step S107 to generate a new biomarker data set.

[0071] Step S108 includes: Step S108.1: Predicting tumor probability based on each image block , determining each image block to be a tumor image block or a normal tissue image block; The predicted tumor probability for each image slice , determining each image slice as a tumor image slice or a normal tissue image slice; For example: and The images are judged with probability thresholds (such as 0.5) respectively. When the probability threshold is exceeded, it is judged as a tumor image block or a tumor image slice. Otherwise, it is a normal tissue image block or a normal tissue image slice. Step S108.2: Screening out all image blocks in the biomarker data set that are judged to be normal tissues; Step S108.3: Evaluate the tumor content t of each remaining image block in the biomarker dataset, t = m / n; Where: m is the number of image slices judged as tumors in the image block; n is the number of all image slices in the image block; Step S108.4: based on the tumor content t of each image block, select the image blocks whose tumor content exceeds the content threshold from the remaining image blocks in the biomarker data set; For example, the tumor content t is compared with the content threshold (such as 0.85), and when the content threshold is exceeded, the corresponding image block is screened out; Step S108.5: For the screened image blocks, use masking technology to remove all image slices judged to be normal tissues to obtain new image blocks and form a new biomarker data set.

[0072] For the image blocks selected in step S108.4, a masking technique is applied to remove all image slices predicted to be normal tissues. This step is performed by setting a mask value for each image slice. Implementation, where Depends on the prediction result of the image patch: Mask value corresponding to the tumor image slice , =1; Mask value for normal tissue image slice , =0.

[0073] These mask values ​​are then used to modify the representation of the image patch to ensure that only information about the tumor tissue is included: Where: T represents the pixel value corresponding to the image block, M represents the The matrix formed, Represents a new image patch containing only the pixel values ​​of tumor tissue.

[0074] Using screening-optimized datasets can lead to more accurate predictions of biomarkers.

[0075] Step S109 uses the trained self-supervised learning model to divide each image block in the new biomarker data set into a number of image slices, and extracts the image slice features of each image slice.

[0076] Step S110 uses the image features of the new biomarker data set obtained in step S109 to train a biomarker prediction model, and the biomarker prediction model is used to predict the prostate cancer biomarker status of the patient.

[0077] Specifically, in this embodiment, the biomarker prediction model consists of a feature encoder (Encoder), an attention mechanism (Attention), and an output layer (Head).

[0078] The feature encoder converts the input feature vector into a low-dimensional vector.

[0079] Assume the input feature is ,and , where B is the batch size, N is the number of instances in each feature bag, and F represents the feature dimension. The output of the feature after passing through the encoder is:

[0080] Where: d represents the feature dimension after encoding. The specific feature encoder structure is as follows:

[0081] in: represents the weight matrix, ; represents the bias term, .

[0082] The attention mechanism is used to calculate the importance weight of each instance so as to weight the instance features in the subsequent steps. It calculates the importance score by mapping the features to a smaller dimension.

[0083] Assume the input feature is , the calculation of attention score can be expressed as:

[0084] The attention mechanism involves three parts: linear transformation, activation function, and output layer.

[0085] Among them, the linear transformation maps the input features to a low dimension. Let the input features be , output features after linear transformation:

[0086] In the above formula , Represents the hidden layer feature dimension and inputs the linear transformation into the activation function:

[0087] The output of the activation function is passed into another linear change layer to obtain the final attention score of the attention mechanism, as follows: .

[0088]

[0089] For each instance in a feature bag, the attention mechanism performs a mask based on the number (length) of instances.

[0090] Masked attention scores Normalize through the Softmax function to ensure that the sum of the weights is 1:

[0091] The weights obtained through the attention mechanism are weighted and summed up for the encoded features to generate a global weighted feature representation for each feature package:

[0092] In the above formula, is the global weighted feature representation of each feature bag, ; is the ith encoded feature, is the attention score corresponding to the i-th encoded feature.

[0093] The output layer is used to classify the weighted global features and obtain the classification results of each feature package. The output layer contains a linear layer and is trained using cross entropy loss. The output calculation formula for classification is:

[0094] In the above formula is the weight matrix of the output layer, is the bias term, is the number of output categories.

[0095] During training, cross entropy loss is used to measure the difference between the model's predictions and the true labels:

[0096] In the above formula, is the true label, is the predicted probability of the model.

[0097] The prediction results of the biomarker prediction model are verified, and the model is optimized based on the verification results to improve the accuracy and reliability of the prediction.

[0098] like Figure 6 As shown, the AUROC (Area Under the Receiver Operating Characteristic Curve) of the biomarker AR-v7 is the area under the ROC curve, which is used to evaluate the performance of the binary classification model, and it reached 0.772.

[0099] like Figure 7 As shown, step S111 is a clinical application step, which specifically includes: Step S111.1: Collect the patient's pathological images to be analyzed; Step S111.2: Divide each collected pathological image into a number of image blocks and perform preprocessing; Step S111.3: establishing a data set to be analyzed based on the preprocessed image blocks; Step S111.4: using the trained self-supervised learning model to divide each image block in the data set to be analyzed into a number of image slices, and extracting image slice features of each image slice; Step S111.5: Based on the image slice features of the data set to be analyzed, the trained tumor segmentation model is used to predict the tumor probability of each image block and the corresponding image slice; Step S111.6: Based on the prediction result of step S111.5, the image blocks in the data set to be analyzed are screened and optimized to generate a new data set to be analyzed; Step S111.7: using the trained self-supervised learning model to divide each image block in the new data set to be analyzed into a number of image slices, and extracting image slice features of each image slice; Step S111.8: Based on the image features of the new data set to be analyzed, the trained biomarker prediction model is used to predict the patient's biomarker status, which is positive or negative.

[0100] Doctors can make reasonable clinical decisions based on the predicted results.

[0101] The above step S111.6 can filter and optimize the image blocks in the data set to be analyzed by the following steps, specifically including: Step S111.6.1: judging each image block as a tumor image block or a normal tissue image block according to the tumor probability predicted for each image block; According to the tumor probability predicted for each image slice, each image slice is judged to be a tumor image slice or a normal tissue image slice; Step S111.6.2: Screen out all image blocks that are judged to be normal tissues in the data set to be analyzed; Step S111.6.3: Evaluate the tumor content t of each remaining image block in the data set to be analyzed, t = m / n; Where: m is the number of image slices judged as tumors in the image block; n is the number of all image slices in the image block; Step S111.6.4: based on the tumor content of each image block, filter out the image blocks whose tumor content exceeds the content threshold from the remaining image blocks in the data set to be analyzed; Step S111.6.5: For the screened image blocks, use masking technology to remove all image slices judged to be normal tissues, obtain new image blocks, and form a new data set to be analyzed.

[0102] The details are similar to the above step S108 and will not be repeated here.

[0103] In other embodiments, Figure 8 As shown, the present invention also discloses a biomarker prediction device, comprising: A collection module 201 is used to collect the patient's pathological images and their corresponding annotation information; A preprocessing module 202, used to divide each collected pathological image into a number of image blocks and perform preprocessing; A data set establishment module 203, used to establish a biomarker data set and a tumor segmentation data set based on the preprocessed image blocks; The biomarker dataset includes: image patches annotated with biomarker diagnostic results; The tumor segmentation dataset includes: image patches annotated with tumors and normal tissues; A self-supervised learning model training module 204, used to train a self-supervised learning model using a biomarker dataset, wherein the self-supervised learning model is used to divide each image block into a plurality of image slices and extract image slice features of each image slice; A first feature extraction module 205 is used to divide each image block in the tumor segmentation data set into a plurality of image slices using a trained self-supervised learning model, and extract image slice features of each image slice; A tumor segmentation model training module 206 is used to train a tumor segmentation model using the image slice features of the tumor segmentation data set obtained by the first feature extraction module. The tumor segmentation model is used to predict the tumor probability of each image block and the corresponding image slice; The segmentation prediction module 207 is used to predict the tumor probability of each image block and the corresponding image slice in the biomarker data set using the trained tumor segmentation model; A screening and optimization module 208, configured to screen and optimize image blocks in the biomarker data set based on the prediction results of the segmentation prediction module to generate a new biomarker data set; A second feature extraction module 209, for dividing each image block in the new biomarker data set into a plurality of image slices using a trained self-supervised learning model, and extracting image slice features of each image slice; A biomarker prediction model training module 210, used to train a biomarker prediction model using the image features of the new biomarker data set obtained by the second feature extraction module, the biomarker prediction model being used to predict the biomarker status of the patient; The application module 211 is used to apply the trained model to the pathological image to be analyzed to predict the status of the biomarker.

[0104] Furthermore, the preprocessing module includes: An image block division unit, used for dividing each collected pathological image into a number of image blocks according to a fixed size; A resolution adjustment unit, used for adjusting and unifying the resolution of all image blocks; The invalid image block screening unit is used to screen and exclude invalid image blocks.

[0105] Furthermore, the application module includes: An application collection unit is used to collect the patient's pathological images to be analyzed; A preprocessing unit is applied to divide each collected pathological image into a number of image blocks and perform preprocessing; An application data set establishment unit is used to establish a data set to be analyzed based on the preprocessed image blocks; Applying a first feature extraction unit, for dividing each image block in the to-be-analyzed data set into a plurality of image slices using the trained self-supervised learning model, and extracting image slice features of each image slice; Applying a segmentation prediction unit, for predicting the tumor probability of each image block and the corresponding image slice using the trained tumor segmentation model based on the image slice features of the data set to be analyzed; An application screening and optimization unit is used to screen and optimize the image blocks in the data set to be analyzed based on the prediction result of the application segmentation prediction unit to generate a new data set to be analyzed; Applying a second feature extraction unit, for dividing each image block in a new data set to be analyzed into a plurality of image slices using a trained self-supervised learning model, and extracting image slice features of each image slice; The prediction unit is applied to predict the biomarker status of the patient using the trained biomarker prediction model based on the image slice features of the new data set to be analyzed.

[0106] Furthermore, the screening optimization module and the application screening optimization unit respectively include: A judgment unit, used to judge whether each image block is a tumor image block or a normal tissue image block according to the tumor probability predicted by each image block; According to the tumor probability predicted for each image slice, each image slice is judged to be a tumor image slice or a normal tissue image slice; A screening unit, used to screen out all image blocks judged to be normal tissues in the corresponding data set; The tumor content evaluation unit is used to evaluate the tumor content t of each remaining image block in the corresponding data set. t = m / n; Where: m is the number of image slices judged as tumors in the image block; n is the number of all image slices in the image block; A screening unit, used for screening, according to the tumor content of each image block, image blocks having a tumor content exceeding a content threshold from the remaining image blocks of the corresponding data set; The forming unit is used to remove all image slices judged as normal tissues from the screened image blocks by using a masking technique to obtain new image blocks and form a new corresponding data set.

[0107] Furthermore, it should be noted that: the biomarker prediction device provided in the above embodiments only uses the division of the above functional modules as an example when performing biomarker prediction. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the biomarker prediction device is divided into different functional modules to complete all or part of the functions described above.

[0108] In addition, the biomarker prediction device and the biomarker prediction method provided in the above embodiments belong to the same concept, and their specific implementation processes are detailed in the method embodiments, which will not be repeated here.

[0109] In addition, in some other embodiments, Fig. 9 As shown, the present invention also discloses a computing device, including: One or more processors 301; Memory 302; and one or more programs, wherein the one or more programs are stored in the memory 302 and configured to be executed by the one or more processors 301, and the one or more programs include instructions of the biomarker prediction method disclosed in the above embodiments.

[0110] The processor 301 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 301 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor 301 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state, also known as a CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 301 may be integrated with a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 301 may also include an AI (Artificial Intelligence) processor, which is used to process computing operations related to machine learning.

[0111] The memory 302 may include one or more computer-readable storage media, which may be non-transitory. The memory 302 may also include a high-speed random access memory, and a non-volatile memory, such as one or more disk storage devices, flash memory storage devices. In some embodiments, the non-transitory computer-readable storage medium in the memory 302 is used to store at least one instruction, which is used to be executed by the processor 301 to implement the biomarker prediction method provided by the method embodiment of the present invention.

[0112] In addition, the computing device may also optionally include: a peripheral device interface and at least one peripheral device. The processor 301, the memory 302 and the peripheral device interface may be connected via a bus or a signal line. Each peripheral device may be connected to the peripheral device interface via a bus, a signal line or a circuit board. Schematically, the peripheral devices include but are not limited to: a radio frequency circuit, a touch display screen, an audio circuit, and a power supply.

[0113] Of course, the computing device may also include fewer or more components, which is not limited in this embodiment.

[0114] In addition, in some other embodiments, the present invention further discloses a storage medium, which stores one or more computer-readable programs, and the one or more programs include instructions, and the instructions are suitable for being loaded by a memory and executing the biomarker prediction method disclosed in the above embodiment.

[0115] The present invention discloses a biomarker prediction method, device, equipment and storage medium, which have the following beneficial effects: First, the present invention uses a self-supervised learning model to extract image features, reducing the reliance on manually labeled data.

[0116] Second, the present invention uses a weakly supervised learning method to train a tumor segmentation model based on image slice features, and can accurately segment tumor tissue and normal tissue.

[0117] Third, the present invention screens and optimizes image blocks according to the prediction results of the tumor segmentation model, so that the prediction of biomarkers is more accurate.

[0118] In summary, the present invention can effectively predict the status of biomarkers with high prediction accuracy, which can effectively reduce the burden of clinical pathologists and improve the prediction accuracy of tumor-related biomarkers, and has significant clinical significance.

[0119] The above shows and describes the basic principles and main features of the present invention and the advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The above embodiments and descriptions are only for illustrating the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention may have various changes and improvements, which shall fall within the scope of the present invention to be protected. The scope of protection of the present invention is defined by the attached claims and their equivalents.

Claims

1. A biomarker prediction method, characterized in that: include: Step S1: Collect the patient's pathological images and their corresponding annotation information; Step S2: Divide each collected pathological image into several image blocks and perform preprocessing; Step S3: establishing a biomarker dataset and a tumor segmentation dataset based on the preprocessed image blocks; The biomarker data set includes: image blocks annotated with biomarker diagnosis results; The tumor segmentation data set includes: image blocks annotated with tumors and normal tissues; Step S4: using the biomarker dataset to train a self-supervised learning model, wherein the self-supervised learning model is used to divide each image block into a number of image slices and extract image slice features of each image slice; Step S5: using the trained self-supervised learning model to divide each image block in the tumor segmentation data set into a number of image slices, and extracting image slice features of each image slice; Step S6: using the image slice features of the tumor segmentation data set obtained in step S5 to train a tumor segmentation model, wherein the tumor segmentation model is used to predict the tumor probability of each image block and the corresponding image slice; Step S7: using the trained tumor segmentation model to predict the tumor probability of each image block and the corresponding image slice in the biomarker data set; Step S8: Based on the prediction result of step S7, the image blocks in the biomarker data set are screened and optimized to generate a new biomarker data set; Step S9: using the trained self-supervised learning model to divide each image block in the new biomarker dataset into a number of image slices, and extracting image slice features of each image slice; Step S10: training a biomarker prediction model using the image features of the new biomarker data set obtained in step S9, wherein the biomarker prediction model is used to predict the biomarker status of the patient; Step S11: Apply the trained model to the pathological image to be analyzed to predict the biomarker status.

2. The biomarker prediction method according to claim 1, characterized in that: The step S2 comprises: Step S2.1: Divide each collected pathological image into a number of image blocks according to a fixed size; Step S2.2: Adjust the resolution of all image blocks to be unified; Step S2.3: Screen and exclude invalid image blocks.

3. The biomarker prediction method according to claim 1 or 2, characterized in that: The step S11 comprises: Step S11.1: Collect the patient's pathological images to be analyzed; Step S11.2: Divide each collected pathological image into a number of image blocks and perform preprocessing; Step S11.3: establishing a data set to be analyzed based on the preprocessed image blocks; Step S11.4: using the trained self-supervised learning model to divide each image block in the data set to be analyzed into a number of image slices, and extracting image slice features of each image slice; Step S11.5: Based on the image slice features of the data set to be analyzed, the trained tumor segmentation model is used to predict the tumor probability of each image block and the corresponding image slice; Step S11.6: Based on the prediction result of step S11.5, the image blocks in the data set to be analyzed are screened and optimized to generate a new data set to be analyzed; Step S11.7: using the trained self-supervised learning model to divide each image block in the new data set to be analyzed into a number of image slices, and extracting the image slice features of each image slice; Step S11.8: Based on the image features of the new data set to be analyzed, the biomarker status of the patient is predicted using the trained biomarker prediction model.

4. The biomarker prediction method according to claim 3, characterized in that: The steps S8 and S11.6 respectively screen and optimize the image blocks in the corresponding data sets through the following steps, specifically including: Step A: According to the tumor probability predicted for each image block, each image block is judged to be a tumor image block or a normal tissue image block; According to the tumor probability predicted for each image slice, each image slice is judged to be a tumor image slice or a normal tissue image slice; Step B: Screen out all image blocks that are judged to be normal tissues in the corresponding data set; Step C: Evaluate the tumor content t of each remaining image block in the corresponding dataset, t=m / n; Where: m is the number of image slices judged as tumors in the image block; n is the number of all image slices in the image block; Step D: According to the tumor content of each image block, the image blocks whose tumor content exceeds the content threshold are screened out from the remaining image blocks in the corresponding data set; Step E: For the screened image blocks, use masking technology to remove all image slices judged as normal tissues to obtain new image blocks and form a new corresponding data set.

5. A biomarker prediction device, characterized in that: include: A collection module, used to collect the patient's pathological images and their corresponding annotation information; A preprocessing module, used for dividing each collected pathological image into a number of image blocks and performing preprocessing; A data set building module, used to build a biomarker data set and a tumor segmentation data set based on the preprocessed image blocks; The biomarker data set includes: image blocks annotated with biomarker diagnosis results; The tumor segmentation data set includes: image blocks annotated with tumors and normal tissues; A self-supervised learning model training module, used to train a self-supervised learning model using a biomarker dataset, wherein the self-supervised learning model is used to divide each image block into a plurality of image slices and extract image slice features of each image slice; A first feature extraction module is used to divide each image block in the tumor segmentation data set into a plurality of image slices using a trained self-supervised learning model, and extract image slice features of each image slice; A tumor segmentation model training module, used to train a tumor segmentation model using the image slice features of the tumor segmentation data set obtained by the first feature extraction module, wherein the tumor segmentation model is used to predict the tumor probability of each image block and the corresponding image slice; A segmentation prediction module is used to predict the tumor probability of each image block and the corresponding image slice in the biomarker data set using the trained tumor segmentation model; A screening and optimization module, used for screening and optimizing image blocks in the biomarker data set based on the prediction results of the segmentation prediction module to generate a new biomarker data set; A second feature extraction module is used to divide each image block in the new biomarker data set into a number of image slices using a trained self-supervised learning model, and extract image slice features of each image slice; A biomarker prediction model training module, used to train a biomarker prediction model using the image features of the new biomarker data set obtained by the second feature extraction module, wherein the biomarker prediction model is used to predict the biomarker status of the patient; The application module is used to apply the trained model to the pathological image to be analyzed to predict the biomarker status.

6. The biomarker prediction device according to claim 5, characterized in that: The preprocessing module comprises: An image block division unit, used for dividing each collected pathological image into a number of image blocks according to a fixed size; A resolution adjustment unit, used for adjusting and unifying the resolution of all image blocks; The invalid image block screening unit is used to screen and exclude invalid image blocks.

7. The biomarker prediction device according to claim 5 or 6, characterized in that: The application module comprises: An application collection unit is used to collect the patient's pathological images to be analyzed; A preprocessing unit is applied to divide each collected pathological image into a number of image blocks and perform preprocessing; An application data set establishment unit is used to establish a data set to be analyzed based on the preprocessed image blocks; Applying a first feature extraction unit, for dividing each image block in the to-be-analyzed data set into a plurality of image slices using the trained self-supervised learning model, and extracting image slice features of each image slice; Applying a segmentation prediction unit, for predicting the tumor probability of each image block and the corresponding image slice using the trained tumor segmentation model based on the image slice features of the data set to be analyzed; An application screening and optimization unit is used to screen and optimize the image blocks in the data set to be analyzed based on the prediction result of the application segmentation prediction unit to generate a new data set to be analyzed; Applying a second feature extraction unit, for dividing each image block in a new data set to be analyzed into a plurality of image slices using a trained self-supervised learning model, and extracting image slice features of each image slice; The prediction unit is applied to predict the biomarker status of the patient using the trained biomarker prediction model based on the image slice features of the new data set to be analyzed.

8. The biomarker prediction device according to claim 7, characterized in that: The screening optimization module and the application screening optimization unit respectively include: A judgment unit, used to judge whether each image block is a tumor image block or a normal tissue image block according to the tumor probability predicted by each image block; According to the tumor probability predicted for each image slice, each image slice is judged to be a tumor image slice or a normal tissue image slice; A screening unit, used to screen out all image blocks judged to be normal tissues in the corresponding data set; The tumor content evaluation unit is used to evaluate the tumor content t of each remaining image block in the corresponding data set. t=m / n; Where: m is the number of image slices judged as tumors in the image block; n is the number of all image slices in the image block; A screening unit, used for screening, according to the tumor content of each image block, image blocks having a tumor content exceeding a content threshold from the remaining image blocks of the corresponding data set; The forming unit is used to remove all image slices judged as normal tissues from the screened image blocks by using a masking technique to obtain new image blocks and form a new corresponding data set.

9. A computing device, characterized in that include: one or more processors; Memory; and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by one or more processors, and the one or more programs include instructions for the biomarker prediction method described in any one of claims 1 to 4 above.

10. A storage medium, characterized in that The storage medium stores one or more computer-readable programs, and the one or more programs include instructions, and the instructions are suitable for being loaded by the memory and executing the biomarker prediction method described in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Systems and methods for biomarker detection in digitized pathological samples

    CN116868229A

  • Uterus smooth muscle tumor recognition system and method based on ensemble voting mechanism

    CN117689985A

  • Biomarker identification method for weak supervision multi-instance learning of lung CT (Computed Tomography) image

    CN118570604A

  • Pathological image analysis method and equipment based on self-supervised learning and multi-task learning

    CN119131018A

  • Tumor prognosis prediction model training method, and tumor prognosis prediction method and device

    CN119251141A

Cited By

  • Microdissection method, device and equipment for tumor tissue and storage medium

    CN122066704A