A method for predicting gene mutations using multi-modal slices based on deep learning

By combining multimodal pathological data from digestive tract HE sections and immunohistochemical sections, a gene mutation prediction model was constructed, which solved the problems of low detection accuracy and high cost in existing technologies, and achieved highly specific and accurate gene mutation detection.

CN117037909BActive Publication Date: 2025-12-12THE FIRST AFFILIATED HOSPITAL OF ZHENGZHOU UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310993781.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-08
Publication Date
2025-12-12
Estimated Expiration
2043-08-08

AI Technical Summary

Technical Problem

While existing gene mutation detection technologies are accurate, they are expensive and have low accuracy when using single-modality pathological slide data to predict gene mutations.

Method used

A deep learning-based multimodal slice prediction method was adopted, combining digestive tract HE slices and immunohistochemical slices. Through image patch segmentation, registration and multi-instance learning model, a gene mutation prediction model was constructed, and prediction was performed using multimodal pathological data.

Benefits of technology

It improves the accuracy of gene mutation detection, reduces detection costs, ensures high model specificity, and reduces the need for expensive testing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117037909B_ABST
    Figure CN117037909B_ABST
Patent Text Reader

Abstract

The application relates to a method for predicting gene mutation based on deep learning using multi-modal slices, image block division is performed on a digestive tract HE slice, and based on the image block division result, each image block of the digestive tract HE slice is classified into a cancerous region; the digestive tract HE slice and an immunohistochemical slice are registered; based on the digestive tract HE slice and the registered immunohistochemical slice, a gene mutation prediction training sample set is constructed, a multi-instance learning model is trained, the training standard is whether a gene is mutated, and a gene mutation prediction model is obtained; gene mutation prediction is performed according to the obtained gene mutation prediction model, a model for predicting gene mutation in combination with multi-modal pathological data of a patient is used, high specificity of the model is ensured, compared with a pathological gene prediction model using single modal data, more modal data is introduced, and the prediction accuracy is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to a method for predicting gene mutation based on deep learning using multi-modal slices. BACKGROUND

[0002] Gene mutation detection refers to analyzing DNA mutation by establishing a series of electrophoresis, analyzing DNA conformation or melting characteristics, or using DNA denaturation and renaturation characteristics. This detection method has important significance for early screening, diagnosis and prognosis of tumor patients such as lung cancer, breast cancer and colorectal cancer. Although the existing gene mutation detection technology is accurate, it is expensive.

[0003] Deep learning has developed rapidly in recent years and has been widely used in computer vision, speech recognition, medical diagnosis and many other practical fields. Using the powerful feature extraction capability of deep learning algorithm to realize the mining of massive digestive tract pathological slices, and to explore the potential rules in the data, has certain practical significance and scientific research value for gene mutation prediction research. There are more and more pathological prediction gene models based on deep learning, but most of them are for using single modal pathological slice data. Using single modal pathological slice data to predict gene mutation often has the problem of low detection accuracy. SUMMARY

[0004] In order to solve the existing problems, the present application provides a method for predicting gene mutation based on deep learning using multi-modal slices.

[0005] The present application adopts the following technical scheme:

[0006] A method for predicting gene mutation based on deep learning using multi-modal slices, comprising:

[0007] Obtaining a digestive tract HE slice and a corresponding immunohistochemical slice;

[0008] Dividing the digestive tract HE slice into image blocks, and based on the image block division result, classifying each image block of the digestive tract HE slice into a cancerous area;

[0009] Registering the digestive tract HE slice and the immunohistochemical slice;

[0010] Based on the digestive tract HE slice and the registered immunohistochemical slice, constructing a gene mutation prediction training sample set, and inputting the gene mutation prediction training sample set into a preset multi-instance learning model for training, the training standard being whether the gene is mutated, to obtain a gene mutation prediction model;

[0011] According to the trained gene mutation prediction model, gene mutation prediction is performed.

[0012] In one embodiment, the immunohistochemical slice includes an immunohistochemical ER, PR, HER2, KI67 slice.

[0013] In one embodiment, the image block division of the digestive tract HE slice includes:

[0014] The contours of the cancerous regions in the digestive tract HE slice are manually annotated, and the image block division of the digestive tract HE slice is performed based on the manual annotation information to obtain positive image blocks and negative image blocks, wherein the positive image blocks represent that the area of the cancerous region is greater than 50% of the area of the image block, and the negative image blocks represent that there is no cancerous region in the image block.

[0015] In one embodiment, the cancerous region classification of each image block of the digestive tract HE slice based on the image block division result includes:

[0016] According to the positive image blocks and the negative image blocks, an image block classification training sample set is constructed;

[0017] The image block classification training sample set is input into an image block classification model for training, and the positive and negative classification of the image block is performed according to the trained image block classification model.

[0018] In one embodiment, the learning rate of the image block classification model is set to 1e-3, the Adam optimization function is used, and the cross-entropy loss function is used as the loss function.

[0019] In one embodiment, the registration of the digestive tract HE slice and the immunohistochemical slice includes:

[0020] Based on a preset pathological image registration algorithm, the digestive tract HE slice and the immunohistochemical slice are registered rigidly and non-rigidly.

[0021] In one embodiment, the learning rate of the multi-instance learning model is set to 1e-3, the Adam optimization function is used, and the cross-entropy loss function is used as the loss function.

[0022] The beneficial effects of the present application include: the model for predicting gene mutations combined with multi-modal pathological data of patients in the present application ensures high specificity of the model, compared with the pathological gene prediction model using single modal data, more modal data is introduced, and the prediction accuracy is effectively improved. BRIEF DESCRIPTION OF DRAWINGS

[0023] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed in the embodiments will be briefly introduced as follows:

[0024] Figure 1is a whole process schematic diagram of a method for predicting gene mutation based on deep learning using multi-modal slices provided by an embodiment of the present application.

[0025] Figure 2 is an execution schematic diagram of a method for predicting gene mutation based on deep learning using multi-modal slices provided by an embodiment of the present application. DETAILED DESCRIPTION

[0026] In the following description, for the purpose of explanation and not limitation, specific details are set forth, such as particular system configurations, techniques, etc., in order to provide a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application can be practiced in other embodiments that depart from these specific details. In other instances, detailed descriptions of well-known systems, devices, circuits, and methods are omitted so as not to obscure the description of the present application with unnecessary detail.

[0027] It should be understood that the term "comprises" as used in the specification and the appended claims indicates the presence of the described features, integers, steps, operations, elements, and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0028] It should also be understood that the term "and / or" as used in the specification and the appended claims, means any one or more of the associated listed items, as well as all possible combinations of the items, and includes these combinations.

[0029] As used in the specification and the appended claims, the term "if" can be interpreted as meaning "when" or "upon" or "in response to a determination" or "in response to detecting" depending on the context. Similarly, the phrase "if determined" or "if detected [the described condition or event]" can be interpreted as meaning "upon determining" or "in response to determining" or "upon detecting [the described condition or event]" or "in response to detecting [the described condition or event]" depending on the context.

[0030] In addition, in the description of the specification and the appended claims, the terms "first", "second", "third", etc. are only used for differentiation of description, and cannot be understood as indicating or implying relative importance.

[0031] Reference within the specification of this application to "one embodiment" or "some embodiments" means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the application. The appearances of the phrases "in one embodiment" or "in some embodiments" in various places within specified sections of this specification are not necessarily all referring to the same embodiment, however, unless otherwise specifically stated

[0032] To illustrate the technical solutions described in the present application, the following will be described by specific embodiments.

[0033] The present embodiment provides a method for predicting gene mutations using multi-modal slices based on deep learning, as shown in Figure 1 . Figure 2 is an execution schematic diagram of the method for predicting gene mutations using multi-modal slices based on deep learning provided by the present embodiment.

[0034] As shown in Figure 1 , the method for predicting gene mutations using multi-modal slices based on deep learning provided by the present embodiment includes the following steps:

[0035] Step S1: Obtain a digestive tract HE slice and a corresponding immunohistochemical slice:

[0036] Obtain a digestive tract HE slice and a corresponding immunohistochemical slice of a patient. In the present embodiment, the immunohistochemical slice includes immunohistochemical ER, PR, HER2, and KI67 slices.

[0037] Step S2: Divide the digestive tract HE slice into image blocks, and based on the image block division result, classify each image block of the digestive tract HE slice into a cancerous area:

[0038] In the present embodiment, a deep learning classification algorithm (such as Resnet, Densenet, etc.) is used to construct a cancerous area classification algorithm, so that the classification model has the ability to predict whether an image block contains a cancerous area, specifically as follows:

[0039] The contours of the cancerous regions in the digestive tract HE section are manually labeled by artificial labeling. In this embodiment, the contours of the cancerous regions in the digestive tract HE section are manually labeled by professional doctors to ensure that the unlabeled regions in the section do not contain cancer. Then, based on the manual labeling information of the doctors, the digestive tract HE section is divided into image blocks (i.e., patch images) to obtain positive image blocks and negative image blocks. Among them, the positive image block represents that the area of the cancerous region is greater than 50% of the area of the image block, and the criterion formula of the positive image block is as follows:

[0040]

[0041] Wherein, sum(mask) represents the area of the cancerous region, h mask and w mask respectively represent the length and width of the image block.

[0042] The negative image block represents that there is no cancerous region in the image block, and the criterion formula of the negative image block is as follows:

[0043]

[0044] According to the obtained positive image blocks and negative image blocks, an image block classification training sample set is constructed, which can be represented by P={p1, p2,...p n}.

[0045] The image block classification training sample set is input into the image block classification model for training. The image block classification model uses a convolutional neural network classification model, including but not limited to Resnet, Densenet, Mobilenet, Mobilevit, etc. The learning rate of the image block classification model is set to 1e-3, the Adam optimization function is used, and the cross-entropy loss function is used as the loss function. Specifically as follows:

[0046] cross entropy loss=-[gt*logpred+(1-gt)*log(1-pred)]

[0047] The image block classification model is trained until the model converges. According to the trained image block classification model, the positive and negative classification of the image block can be effectively distinguished.

[0048] Step S3: Registering the digestive tract HE section and the immunohistochemical section:

[0049] Based on the preset pathological image registration algorithm, the gastrointestinal HE section and the immunohistochemical section are rigidly and non-rigidly registered. In this embodiment, the registration methods including but not limited to VALIS and the like are used to perform rigid registration such as translation and rotation on HE, ER, PR, HER2 and KI67 and non-rigid registration such as local area vector field transformation. After the registration is completed, the effective areas in the sections are one-to-one corresponding.

[0050] Step S4: Based on the gastrointestinal HE section and the registered immunohistochemical section, a gene mutation prediction training sample set is constructed, and the gene mutation prediction training sample set is input into a preset multi-instance learning model for training, the training standard is whether the gene is mutated, and a gene mutation prediction model is obtained:

[0051] Based on the gastrointestinal HE section and the registered immunohistochemical section, a gene mutation prediction training sample set is constructed. In this embodiment, the gastrointestinal HE section of the patient is inferred to obtain the coordinates of the image block predicted to be cancerous and the probability value of the image block predicted to be cancerous. The image block probability value is sorted, the coordinates of the image block with the Top K probability value are selected, the images in the HE section and the registered ER, PR, KI67 and HER2 sections are obtained through the coordinates, and a data set H={h1, h2,…h n} is constructed.

[0052] The gene mutation prediction training sample set is input into a preset multi-instance learning model for training, and the model includes but is not limited to Resnet combined with Lstm and Resnet combined with Rnn and the like. The model input is K*h*w*(3*5) dimensions, and 3*5 is the HE and IHC images spliced by registration. The data is input into the model, and the algorithm training gold standard is whether the gene is mutated 0 / 1. During the model training, the learning rate is set to 1e-3, the Adam optimization function is used, the cross-entropy loss function is used as the loss function, and the model is converged until the model is converged. After the training is completed, the gene mutation prediction model is obtained.

[0053] Step S5: Gene mutation prediction is performed according to the gene mutation prediction model obtained by training:

[0054] The image block data of the patient to be predicted is input into the gene mutation prediction model obtained by training, and the model directly outputs the probability value of gene mutation.

[0055] The method for predicting gene mutation based on deep learning using multi-modal slices provided by the embodiment uses patient digestive tract slice data and corresponding immunohistochemical ER, PR, HER2 and KI67 slices, and combines multi-instance learning deep learning technology to predict gene mutations such as BRAF and NRAS. The method has the following advantages: (1) The model for predicting gene mutation by combining patient multi-modal pathological data is proposed, threshold processing is performed on the model output, high specificity of the model is ensured, that is, the result of the model predicting no gene mutation is as correct as possible, and some patients do not need to undergo expensive gene mutation detection; (2) The model for predicting gene mutation by combining patient multi-modal pathological data is proposed, and compared with the pathological gene prediction model using single modal data, more modal data is introduced, and the prediction accuracy is effectively improved.

[0056] Through the application of the method for predicting gene mutation based on deep learning using multi-modal slices provided by the embodiment, the following functions can be realized: through analysis and reasoning of HE, ER, PR, KI67 and HER2 slices of the patient, whether the patient has a specific gene mutation can be obtained, the use of multi-modal data in gene prediction is solved, and an algorithm execution pipeline is implemented.

[0057] The above-described embodiments are only used to illustrate the technical solutions of the present application, but not limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that they can modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for some technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application, and should be included in the protection scope of the present application.

Claims

1. A method for predicting genetic mutations using multi-modal slices based on deep learning, characterized in that, The method comprises the following steps: obtaining a digestive tract HE slice and a corresponding immunohistochemical slice; dividing the digestive tract HE slice into image blocks, and classifying each image block of the digestive tract HE slice based on the image block division result; registering the digestive tract HE slice and the immunohistochemical slice; based on the digestive tract HE slice and the registered immunohistochemical slice, constructing a gene mutation prediction training sample set, and inputting the gene mutation prediction training sample set into a preset multi-instance learning model for training, the training standard being whether the gene is mutated, to obtain a gene mutation prediction model; According to the gene mutation prediction model obtained by training, the gene mutation is predicted.

2. The method for predicting a genetic mutation using multi-modal slices based on deep learning according to claim 1, wherein, The immunohistochemical slice includes immunohistochemical ER, PR, HER2 and KI67 slices. 3.The method of predicting a genetic mutation using multi-modal slices based on deep learning according to claim 1, wherein, The image block division of the digestive tract HE slice comprises: manually labeling the outline of the cancerous area in the digestive tract HE slice, and dividing the digestive tract HE slice into image blocks based on the manual labeling information to obtain positive image blocks and negative image blocks, wherein the positive image blocks represent that the cancerous area is greater than 50% of the area of the image block, and the negative image blocks represent that there is no cancerous area in the image block.

4. The method for predicting a genetic mutation using multi-modal slices based on deep learning according to claim 3, wherein, The cancerous area classification of each image block of the digestive tract HE slice based on the image block division result comprises: According to the positive image blocks and negative image blocks, an image block classification training sample set is constructed; inputting the image block classification training sample set into an image block classification model for training, and classifying the image blocks according to the trained image block classification model.

5. The method for predicting a genetic mutation using multi-modal slices based on deep learning according to claim 4, wherein, The learning rate of the image block classification model is set to 1e-3, the Adam optimization function is used, and the cross-entropy loss function is used as the loss function.

6. The method for predicting a genetic mutation using multi-modal slices based on deep learning according to claim 1, wherein, The registration of the digestive tract HE slice and the immunohistochemical slice comprises: based on a preset pathological image registration algorithm, rigid and non-rigid registration is performed on the digestive tract HE slice and the immunohistochemical slice.

7. The method for predicting a genetic mutation using multi-modal slices based on deep learning according to claim 1, wherein, The learning rate of the multi-instance learning model is set to 1e-3, the Adam optimization function is used, and the cross-entropy loss function is used as the loss function.

Citation Information

Patent Citations

  • Gastric cancer pathological image classification method and system based on HER2 gene detection

    CN115690056A

  • Model training device and model application device for histopathologic staining section image

    CN115708127A