Risk prediction method for endometrial carcinoma

By establishing a pathological multimodal database and integrating multimodal data using the support vector machine algorithm, the accuracy and applicability issues of endometrioid carcinoma risk prediction in existing technologies have been resolved, achieving more accurate prediction results.

CN120977584AInactive Publication Date: 2025-11-18JILIN UNIVERSITY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511240100.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-02
Publication Date
2025-11-18
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing risk prediction models for endometrioid carcinoma suffer from limited single-modal prediction accuracy, lack of multimodal prediction methods, lack of dynamic optimization mechanisms, and failure to effectively integrate image and text data, resulting in prediction results that are out of touch with real-world diagnostic and treatment scenarios.

Method used

A pathological multimodal database was built, and multimodal datasets, including clinical information, pathological images, and clinical text features, were accessed. A multimodal prediction model was established using the support vector machine algorithm, and the model was optimized based on doctor feedback.

Benefits of technology

It improves the accuracy and applicability of endometrioid carcinoma risk prediction, enabling more accurate screening of suitable candidates for fertility care treatment and reducing the risk of misjudgment of treatment response and recurrence.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120977584A_ABST
    Figure CN120977584A_ABST
Patent Text Reader

Abstract

The invention discloses a risk prediction method of endometrial carcinoma, and relates to the field of risk prediction of endometrial carcinoma, and the method comprises the steps: accessing a multi-modal data set, and building a pathological multi-modal database; analyzing the pathological multi-modal database to obtain multi-modal characteristic data; establishing a prediction model, and comparing the performance of the prediction model to obtain a comparison result; according to the risk prediction model for the endometrial carcinoma, a prediction model is selected, a pathological data set is predicted, prediction data is obtained, and the model and a pathological multi-modal database are optimized in combination with a doctor feedback result. The technical problems of lack of a multi-modal prediction model, lack of a dynamic optimization mechanism of the risk prediction model of the endometrial carcinoma and lack of image and text data support of the risk prediction model of the endometrial carcinoma in the prior art are solved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of risk prediction of endometrioid carcinoma, in particular to a risk prediction method of endometrioid carcinoma. BACKGROUND

[0002] As a common gynecological malignancy, the incidence of endometrioid carcinoma in young nulliparous women has increased significantly in recent years, and about 70% of patients have not completed childbirth when diagnosed. Traditional conservation therapy relies on progestin therapy, but the individual differences in therapeutic effect are large, and there is an urgent need for precise prediction models to screen treatment response populations and assess the risk of disease progression to balance fertility preservation and tumor control goals.

[0003] Current risk prediction is mainly based on single modal data, and pathologists subjectively assess lesion risk through morphological features, or use molecular typing and immunohistochemical markers to assist in stratification. Existing algorithms such as single-modal support vector machine models only use pathological images or molecular data to make independent predictions. Although a 2023 study attempted to use weakly supervised deep learning to predict progestin response, it has not integrated multi-dimensional clinical information. Pathological morphology is currently recognized and most commonly used for efficacy evaluation of conservation therapy, but morphological changes after progestin therapy are complex, leading to differences in interpretation by different pathologists. In recent years, there have been related biomarker research reports on the evaluation of endometrioid carcinoma conservative treatment. Most biomarkers can be evaluated by routine immunohistochemical staining, but some results are not certain, and their application value still needs more research to support. This study applies artificial intelligence deep learning methods to establish a conservation therapy prediction and evaluation model based on pathological morphology and histological phenotype, to screen the appropriate population for conservation therapy, and to evaluate the prediction effect of the model.

[0004] However, the existing technology has significant defects, the single-modal prediction accuracy is limited, the molecular typing and clinical text are not integrated, the cross-modal correlation cannot be captured, the dynamic optimization mechanism is lacking, the model cannot be updated iteratively combined with doctor feedback after training, leading to a decline in clinical applicability, and the image and text data are lacking, traditional methods are difficult to uniformly process structured clinical data, unstructured pathological images and text, leading to a disconnect between the prediction results and the real diagnosis and treatment scene. These defects directly cause screening bias for conservation patients, misjudgment of treatment response, and missed detection of recurrence risk. SUMMARY

[0005] The present application provides a risk prediction model for endometrioid carcinoma, which is used to solve the technical problems of limited single-modal prediction accuracy, lack of multi-modal prediction method, lack of dynamic optimization mechanism in the risk prediction model for endometrioid carcinoma, and lack of image and text data support in the risk prediction model for endometrioid carcinoma in the prior art.

[0006] In view of the above problems, the application provides a risk prediction method for endometrioid carcinoma, which comprises the following steps: accessing a multi-modal data set, building a pathological multi-modal database; analyzing the pathological multi-modal database to obtain multi-modal feature data; establishing a prediction model, comparing the performance of the prediction model to obtain a comparison result; selecting the prediction model, predicting the pathological data set to obtain prediction data, and optimizing the model and the pathological multi-modal database in combination with the feedback result of a doctor.

[0007] The one or more technical solutions provided in the application have at least the following technical effects or advantages: The embodiments of the application have the following advantages: accessing a multi-modal data set, building a pathological multi-modal database; analyzing the pathological multi-modal database to obtain multi-modal feature data; establishing a prediction model, comparing the performance of the prediction model to obtain a comparison result; selecting the prediction model, predicting the pathological data set to obtain prediction data, and optimizing the model and the pathological multi-modal database in combination with the feedback result of a doctor, which can solve the technical problems of limited single-modal prediction accuracy, lack of multi-modal prediction model, lack of dynamic optimization mechanism for the risk prediction model of endometrioid carcinoma, and lack of image and text data support for the risk prediction model of endometrioid carcinoma.

[0008] The above description is only a summary of the technical solutions of the application. In order to more clearly understand the technical means of the application, and in order to make the above and other purposes, characteristics and advantages of the application more apparent and easy to understand, the following specific embodiments of the application are described. BRIEF DESCRIPTION OF DRAWINGS

[0009] In order to more clearly illustrate the technical solutions in the embodiments of the application, the following will briefly introduce the drawings needed in the embodiment description. Obviously, the drawings in the following description are only some embodiments of the application, and other drawings can be obtained by those skilled in the art without creative labor.

[0010] Figure 1 The control flowchart provided in the application is provided. DETAILED DESCRIPTION

[0011] The application provides a risk prediction method for endometrioid carcinoma, which is used to solve the technical problems of limited single-modal prediction accuracy, lack of multi-modal prediction model, lack of dynamic optimization mechanism for the risk prediction model of endometrioid carcinoma, and lack of image and text data support for the risk prediction model of endometrioid carcinoma in the prior art.

[0012] With the introduction of the basic principles of the present application, the technical solutions in the present application will be described clearly and completely below with reference to the drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. It should be understood that the present application is not limited by the example embodiments described herein. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of protection of the present application. In addition, it should be noted that, for the convenience of description, only parts related to the present application are shown in the drawings, not all.

[0013] Embodiment one As shown in the drawings, the present application provides a risk prediction method for endometrial-like carcinoma, the method comprising: Figure 1 S100: access to a multi-modal data set, build a pathological multi-modal database; In the embodiments of the present application, the multi-modal data set of early endometrial carcinoma and complex atypical hyperplasia in the Department of Pathology of the Second Hospital of Jilin University is accessed, which contains clinical information, follow-up information, immunohistochemical information, molecular typing information, pathological image information, clinical text feature information and treatment effect information; pathological information is collected to obtain pathological data; clinical information is collected to obtain clinical data; follow-up information is collected to obtain follow-up data; immunohistochemical information is collected to obtain immunohistochemical data; molecular typing information is collected to obtain molecular typing data; treatment effect information is collected to obtain treatment effect data; pathological image information is collected to obtain pathological image data; clinical text feature information is collected to obtain clinical text feature data; the clinical data, follow-up data, immunohistochemical data, molecular typing data, pathological image data, clinical text feature data and treatment effect data are taken as a pathological data set; the pathological data set is integrated and sorted, and the pathological data set is preprocessed, including data cleaning, format unification and missing value filling, to ensure the integrity and consistency of the data, to obtain a pathological multi-modal database. The step S100 in the method provided by the embodiments of the present application comprises:

[0014] S101: access to a multi-modal data set; S102: collect pathological information to obtain pathological data; S103: collect clinical information to obtain clinical data; S104: collect follow-up information to obtain follow-up data; S105: collect immunohistochemical information to obtain immunohistochemical data; S106: collect molecular typing information to obtain molecular typing data; S107: collect treatment effect information to obtain treatment effect data; ​S108: Collect pathological image information to obtain pathological image data; S109: Collect clinical text feature information to obtain clinical text feature data; S110: Take the clinical data, follow-up data, immunohistochemical data, molecular typing data, pathological image data, clinical text feature data, and treatment effect data as a pathology data set; S111: Integrate and sort the pathology data set to obtain a pathology multi-modal database.

[0015] The pathological information includes but is not limited to lesion range, endometrial gland histological structure characteristics, endometrial gland cytological features, and endometrial epithelial metaplasia changes. The clinical information includes but is not limited to age, body mass index, comorbidity, tumor family history, marriage, pregnancy times, and birth times. The follow-up information includes but is not limited to tumor remission, progression, or recurrence of patients, whether there is a birth plan, and pregnancy outcome. The immunohistochemical information refers to immunohistochemical indexes of the effect of progesterone treatment, such as estrogen receptor, progesterone receptor, phosphatase and tensin homolog, paired box gene 2, and progesterone receptor B. The molecular typing information includes but is not limited to POLE mutant type, i.e., DNA polymerase epsilon hypermutation type, mismatch repair deficiency type, non-specific molecular spectrum type, and p53 gene mutation type. The treatment effect information refers to the evaluation of pathological specimens after conservation treatment, which is divided into complete remission, partial remission, non-response, or disease persistence, disease progression, and recurrence. The pathological data refers to the data of the unified format after the analysis of the pathological information.

[0016] Specifically, access to the early endometrial cancer and complex atypical hyperplasia of Jilin University Second Hospital pathology department multi-modal data set, which contains clinical information, follow-up information, immunohistochemical information, molecular typing information, pathological image information, clinical text feature information and treatment effect information; collect pathological information to obtain pathological data; collect clinical information to obtain clinical data; collect follow-up information to obtain follow-up data; collect immunohistochemical information to obtain immunohistochemical data; collect molecular typing information to obtain molecular typing data; collect treatment effect information to obtain treatment effect data; collect pathological image information to obtain pathological image data; collect clinical text feature information to obtain clinical text feature data; classify and label the collected data, and extract relevant quantitative features according to different methods according to the data type. For example, clinical text feature information is converted into structured data for analysis by natural language processing technology. The clinical data, follow-up data, immunohistochemical data, molecular typing data, pathological image data, clinical text feature data and treatment effect data are used as a pathological data set; the pathological data set is integrated and sorted, the priority is set according to the importance and relevance of the data, and the overall architecture of the database is optimized. The pathological data set is preprocessed, the data is cleaned, the outliers are detected and corrected, and the missing values are filled by statistical algorithm.

[0017] S200: analyzing the pathological multi-modal database to obtain multi-modal feature data; In the embodiment of the application, the pathological multi-modal database is analyzed by single factor analysis and multi-factor Logistic regression to obtain nursing data and dangerous pathological data; CNN is used to analyze pathological image data to extract key features in the pathological image; natural language processing technology is used to deeply mine clinical text feature data to extract key information related to the disease to obtain clinical text feature data; the nursing data, dangerous pathological data, pathological image feature data and clinical text feature data are integrated to obtain multi-modal feature data.

[0018] The method provided in the embodiment of the application comprises the following steps: S201: analyzing the pathological multi-modal database to obtain nursing data and dangerous pathological data; S202: analyzing pathological image data to obtain pathological image features; S203: analyzing clinical text feature data to obtain clinical text feature data; S204: integrating the nursing data, dangerous pathological data, pathological image feature data and clinical text feature data to obtain multi-modal feature data.

[0019] Among them, the conservation data refers to various types of information related to the patient's fertility, including but not limited to age, reproductive history, hormone level, endometrial thickness, used to evaluate the standard of patients who can be treated with conservation treatment; the dangerous pathological data refers to the pathological indicators closely related to the occurrence and development of endometrial carcinoma, including but not limited to cell atypia, tumor differentiation degree, invasion depth and lymph node metastasis; the pathological image feature refers to the key visual information extracted from the pathological section, including but not limited to cell morphology, tissue structure, staining intensity and distribution pattern of lesion area; the clinical text feature data refers to the clinical text feature data including but not limited to patient's medical history record, symptom description, laboratory examination result and doctor's diagnosis opinion.

[0020] Specifically, independent variables with differences in diagnosis of endometrial carcinoma (EC) / precancerous lesions (EAH) and conservation treatment response are screened out by single factor analysis, and whether each independent variable such as age, reproductive history, hormone level exists statistically significant difference between EC / EAH patients and non-patients, or between good response group and poor response group of conservation treatment is investigated one by one. The independent variables showing significant association with EC / EAH diagnosis results or conservation treatment response in single factor analysis are obtained by using multi-factor Logistic regression analysis, and the key risk factors that can independently affect EC / EAH diagnosis results and conservation treatment response, i.e. dangerous pathological data, and the key conservation factors that can be treated with conservation treatment, i.e. conservation data, are obtained. The key features in the pathological image are extracted by using CNN analysis of pathological image data, and the pathological image features are obtained. The key information related to the disease is extracted by using natural language processing technology to deeply mine the clinical text feature data, and the clinical text feature data is obtained. The conservation data, dangerous pathological data, pathological image feature data and clinical text feature data are integrated and analyzed by cross-modal information through deep learning method, multi-modal data fusion is realized, and multi-modal feature data is obtained.

[0021] S300: Establish a prediction model, compare the performance of the prediction model, and obtain a comparison result; In the embodiments of the application, the TCGA database is accessed as an external independent verification set; a single-modal feature prediction model based on a support vector machine is established, multi-modal feature data is input into a support vector machine classifier to obtain a single-modal prediction model; a pathology data set is input into the single-modal feature prediction model to obtain a single-modal prediction result; the single-modal prediction result is compared with the TCGA database to obtain single-modal model prediction performance; a multi-modal feature prediction model based on a support vector machine is established, multi-modal feature data is input into a support vector machine classifier to obtain a multi-modal prediction model; a pathology data set is input into the multi-modal feature prediction model to obtain a multi-modal prediction result; the multi-modal prediction result is compared with the TCGA database to obtain multi-modal model prediction performance; and the single-modal prediction performance and the multi-modal prediction performance are compared to obtain a comparison result.

[0022] The step S300 in the method provided by the embodiments of the application comprises: S301: accessing a TCGA database; S302: establishing a single-modal feature prediction model based on a support vector machine to obtain a single-modal prediction model; S303: inputting a pathology data set into the single-modal feature prediction model to obtain a single-modal prediction result; S304: comparing the single-modal prediction result with the TCGA database to obtain single-modal model prediction performance; S305: establishing a multi-modal feature prediction model based on a support vector machine to obtain a multi-modal prediction model; S306: inputting a pathology data set into the multi-modal feature prediction model to obtain a multi-modal prediction result; S307: comparing the multi-modal prediction result with the TCGA database to obtain multi-modal model prediction performance; S308: comparing the single-modal prediction performance with the multi-modal prediction performance to obtain a comparison result.

[0023] The single-modal prediction model refers to a model for prediction by using a single feature; the multi-modal prediction model refers to a model for prediction by using multiple features; the TCGA database refers to a Cancer Genome Atlas, which is a large public project jointly launched by the National Cancer Institute and the National Human Genome Research Institute of the United States in 2006, aiming to systematically analyze the molecular mechanism of cancer through multi-omics technology, and to provide data support for cancer prevention, diagnosis and treatment; and the support vector machine is a supervised machine learning algorithm, mainly used for classification problems, and also used for regression problems.

[0024] Specifically, the TCGA database is accessed as an external independent validation set to determine the accuracy of the prediction results of the feature prediction model; a single-modal feature prediction model is constructed using a support vector machine algorithm, multi-modal feature data is input into the support vector machine classifier for training and optimization, and through cross-validation and parameter tuning, a single-modal prediction model is obtained; the pathological data set is input into the single-modal feature prediction model, the trained single-modal prediction model is used to predict the data in the pathological data set, the predicted data is taken as the single-modal prediction result, and the single-modal prediction result is obtained; the single-modal prediction result is compared with the TCGA database, and the single-modal model prediction performance of the single-modal prediction model is obtained, including the accuracy, recall rate and F1 score; a multi-modal feature prediction model based on a support vector machine is established, multi-modal feature data is input into the support vector machine classifier for training and optimization, and through cross-validation and parameter tuning, a multi-modal prediction model is obtained; the pathological data set is input into the multi-modal feature prediction model, and a multi-modal prediction result is obtained; the multi-modal prediction result is compared with the TCGA database, and the multi-modal model prediction performance is obtained; the single-modal prediction performance and the multi-modal prediction performance are compared, the differences between the two models in terms of accuracy, recall rate and F1 score are analyzed, the improvement effect of multi-modal feature fusion on the prediction performance is evaluated, and a comparison result is obtained.

[0025] In S303, the pathological data set is input into the single-modal feature prediction model to obtain a single-modal prediction result, including: a plurality of single-modal prediction models based on single feature data are built, a pathological image single-modal prediction model is established, a convolutional neural network architecture is used, and pathological image feature data is input. The model includes a feature extractor composed of a convolutional layer and a pooling layer, and a classifier composed of a full connection layer, and outputs the prediction result of the pathological image; a clinical text single-modal prediction model is established, a pre-trained language model based on a Transformer architecture is used, and clinical text feature data is input. The pre-trained model is fine-tuned to adapt to the endometrial-like carcinoma risk prediction task, a specific classification layer is added, and the prediction result of the clinical text is output; a structured data single-modal prediction model is established, a gradient boosting tree model such as XGBoost or a support vector machine SVM is used, and nursing data and risk pathological data including structured features such as clinical, follow-up, immunohistochemistry, molecular typing, treatment effect are input. The model is trained based on the input features, and the output is a prediction result based on structured data.

[0026] Exemplarily, a clinical text single-modal prediction model is built and trained, the architecture is: a pre-trained language model based on the Transformer architecture, such as BERT or BioBERT; input: the clinical text feature data obtained in step S200, i.e. the structured text representation after natural language processing, such as word vector sequence, sentence vector or input format of specific task; processing: adding a task-specific classification layer on top of the pre-trained model; training: fine-tuning the model using the clinical text data and its corresponding labels in the pathology dataset. The fine-tuning process updates the classification layer parameters and selectively updates part or all of the pre-trained model parameters as needed. The optimizer, loss function and training strategy need to be set; output: after training is completed, the model receives new clinical text feature data and outputs the corresponding single-modal text prediction result.

[0027] S305: Establish a multi-modal feature prediction model based on support vector machine, obtain a multi-modal prediction model, including: input multi-modal feature data; pre-process the feature data, use Z-score standardization to make the feature mean 0 and the standard deviation 1, or Min-Max normalization to scale the feature to the interval [0, 1] or [-1, 1], improve the stability and performance of SVM training; perform feature fusion, use a feature-level fusion strategy. Concatenate the pre-processed feature vectors of each modality along the feature dimension to form a longer joint feature vector. This joint feature vector captures complementary information from pathological images, clinical text and structured data, and improves prediction ability through cross-modal association; build model architecture and model training, the core classifier uses support vector machine, select kernel function, considering that the fused feature space may be highly complex and nonlinearly separable, the radial basis function kernel is preferred. RBF can effectively map data to high-dimensional space to find a nonlinear decision boundary. Its formula is: where, is the kernel function parameter, which controls the influence range of a single sample, The larger the sample influence range is reduced, the more complex the decision boundary is, The smaller the sample influence range is expanded, the smoother the decision boundary is; are the sample feature vectors of two samples that need to calculate similarity, if each sample has d features, The kernel function measures similarity by calculating their distance; is the square of the Euclidean distance, its calculation formula is: The greater the value, the greater the sample difference, the lower the similarity, and the smaller the value, the closer the sample, the higher the similarity. Model training is performed, the pathological data set and its corresponding labels such as treatment effect and recurrence risk are taken as training samples, the joint feature vector is input, and the label is output. The training target is to find an optimal classification hyperplane for SVM to maximize the interval between different categories such as response / non-response and low risk / high risk samples in a high-dimensional space composed of joint feature vectors. The optimization of the hyperparameter adopts grid search combined with k-fold cross-validation to optimize the key hyperparameters of SVM: the penalty coefficient C, which controls the punishment degree of misclassification samples. A smaller C value allows more misclassification to make the decision boundary smoother, and a larger C value reduces training misclassification to make the decision boundary more complex, but may overfit. The search range is usually on a logarithmic scale such as [0.001, 0.01, 0.1, 1, 10, 100]. In the grid search, the defined (C, γ) parameter combinations are traversed. For each set of parameters, k-fold cross-validation is performed: the training set is randomly divided into k parts, and the model is trained using k-1 parts in turn, and the remaining 1 part is used to verify the model performance such as calculating the F1 score. Record the average performance indicator of each set of parameters in k validations. Finally, select the (C, γ) parameter combination with the optimal average performance on the cross-validation set; model output, the trained multi-modal SVM model with the optimal hyperparameters receives the joint feature vector of a new patient and outputs the corresponding multi-modal prediction result. This result can be a classification label, a lesion probability or a risk score.

[0028] S400: Select a prediction model, predict the pathological data set to obtain prediction data, and optimize the model and the pathological multi-modal database in combination with the doctor's feedback result.

[0029] In the embodiments of the present application, according to the comparison result, the performance of the single-modal prediction model and the multi-modal prediction model is judged, the prediction model is selected according to the use environment and the prediction performance, the pathological data set is input into the selected prediction model to obtain prediction data, the prediction data is compared with the pathological multi-modal database, the possibility of nursing and the possibility of lesion are obtained by comparing the patient parameters obtained by prediction with the parameters in the database, the possibility of nursing and the possibility of lesion are uploaded to the system, and the feedback result is obtained by waiting for the doctor's feedback. The prediction model and the pathological multi-modal database are optimized according to the feedback result and the prediction data.

[0030] The step S400 in the method provided in the embodiments of the present application comprises: S401: According to the comparison result, the prediction model is matched, and the prediction model is selected; S402: The pathological data set is input into the selected prediction model to obtain prediction data; S403: The prediction data is compared with the pathological multi-modal database to obtain the possibility of nursing and the possibility of lesion; S404: upload the conservation possibility and the lesion possibility to the system, wait for the feedback of the doctor, and get the feedback result; S405: optimize the prediction model and the pathology multi-modal database according to the feedback result and the prediction data.

[0031] The conservation possibility refers to the possibility of the patient retaining the reproductive function during the treatment process, and the calculation of the conservation possibility is based on the comparison of the prediction data and the pathology multi-modal database; the lesion possibility refers to the risk of the patient's disease progression or deterioration, and the calculation of the lesion possibility is based on the comparison of the prediction data and the pathology multi-modal database.

[0032] Specifically, according to the comparison result, the performance of the single-modal prediction model and the multi-modal prediction model is judged, the prediction model suitable for the current use environment is selected according to the use environment and the prediction performance, the accuracy, stability and calculation efficiency of the model are comprehensively considered, the pathology data set is input into the selected prediction model, the prediction data containing the basic pathological characteristics of the patient, the clinical parameters and molecular parameters of the patient are obtained through the analysis of the prediction model, the prediction data is compared with the pathology multi-modal database, the same parameters are found by comparing the predicted patient parameters with the parameters in the database, and the conservation possibility and the lesion possibility corresponding to the parameters are output; the conservation possibility and the lesion possibility are uploaded to the system, the doctor checks the conservation possibility and the lesion possibility of the patient and the related parameters in the background system, waits for the feedback of the doctor, and the doctor gives the feedback result according to the conservation possibility and the lesion possibility and the related parameters, wherein the feedback result refers to that the doctor evaluates the accuracy, rationality and clinical applicability of the prediction result according to his own experience and professional knowledge, and puts forward the corresponding adjustment suggestion or confirmation opinion. According to the user feedback result collected in the actual application process and the statistical analysis of the historical prediction data, the algorithm parameters of the existing prediction model are systematically optimized and finely adjusted, including but not limited to recalibrating the learning rate of the model, adjusting the weight coefficient of the regularization term, optimizing the threshold parameter of feature selection and other key links, improving the prediction accuracy and generalization ability of the model in the real scene, so that it can more accurately capture the potential rules and feature associations in the data, thereby providing more reliable data support for decision-making; the new pathology data and the corresponding feedback result are supplemented to the pathology multi-modal database, and the data in the pathology multi-modal database is updated and supplemented.

[0033] The steps of the methods or algorithms described in this application can be directly embedded in hardware, a software unit executed by a processor, or a combination of both. The software unit can be stored in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, hard disk, removable disk, CD-ROM, or any other storage medium of any form in the art. Exemplarily, the storage medium can be connected to the processor so that the processor can read information from the storage medium and write information to the storage medium. Optionally, the storage medium can also be integrated into the processor. The processor and storage medium can be disposed in an ASIC, which can be disposed in a terminal. Optionally, the processor and storage medium can also be disposed in different components within the terminal. These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable apparatus for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0034] Although this application has been described in conjunction with specific features and embodiments, it is obvious that various modifications and combinations can be made thereto without departing from the spirit and scope of this application. Accordingly, this specification and drawings are merely illustrative examples of this application and are considered to cover any and all modifications, variations, combinations, or equivalents within the scope of this application. Clearly, those skilled in the art can make various alterations and modifications to this application without departing from its scope. Thus, if such modifications and modifications fall within the scope of this application and its equivalents, this application intends to include such modifications and modifications.

Claims

1. A method for predicting the risk of endometrioid carcinoma, characterized in that, The method includes: Access multimodal datasets and build a pathological multimodal database; Analyze the pathological multimodal database to obtain multimodal feature data; Establish prediction models, compare the performance of prediction models, and obtain comparison results; Select a prediction model, make predictions on the pathology dataset, obtain the predicted data, and optimize the model and the pathology multimodal database by combining the doctor's feedback results.

2. The method according to claim 1, characterized in that, The process of accessing multimodal datasets and building a pathological multimodal database includes: Access to multimodal datasets; Collect pathological information and obtain pathological data; Collect clinical information to obtain clinical data; Collect follow-up information to obtain follow-up data; Immunohistochemical information was collected to obtain immunohistochemical data; Collect molecular typing information to obtain molecular typing data; Collect treatment effect information to obtain treatment effect data; Collect pathological image information to obtain pathological image data; Collect clinical text feature information to obtain clinical text feature data; Clinical data, follow-up data, immunohistochemical data, molecular subtyping data, pathological image data, clinical text feature data, and treatment effect data are used as a pathological dataset. The pathology datasets are integrated and sorted to obtain a pathology multimodal database.

3. The method according to claim 2, characterized in that, The analysis of the pathological multimodal database yields multimodal feature data, including: Analyze the pathological multimodal database to obtain conservation data and hazardous pathological data; Analyze pathological image data to obtain pathological image features; Analyze clinical text feature data to obtain clinical text feature data; By integrating conservation data, hazardous pathology data, pathological image feature data, and clinical text feature data, multimodal feature data is obtained.

4. The method according to claim 3, characterized in that, The process of establishing a prediction model, comparing the performance of the prediction models, and obtaining comparison results includes: Access to the TCGA database; A single-modal feature prediction model based on support vector machine is established to obtain the single-modal prediction model; The pathological dataset is input into the single-modal feature prediction model to obtain the single-modal prediction results; The single-mode prediction results are compared with the TCGA database to obtain the prediction performance of the single-mode model; A multimodal feature prediction model based on support vector machines is established to obtain the multimodal prediction model; The pathological dataset is input into the multimodal feature prediction model to obtain multimodal prediction results; The multimodal prediction results are compared with the TCGA database to obtain the prediction performance of the multimodal model; The performance of single-mode prediction and multi-mode prediction is compared to obtain the comparison results.

5. The method according to claim 4, characterized in that, The selected prediction model predicts the pathology dataset to obtain predicted data. The model is then optimized based on physician feedback and a multimodal pathology database, including: Based on the comparison results, match the prediction model and select the prediction model; Input the pathology dataset into the selected prediction model to obtain the prediction data; The predicted data are compared with a pathological multimodal database to obtain the probability of preservation and the probability of disease. Upload the potential for care and the potential for disease to the system, wait for the doctor's feedback, and receive the feedback results; The prediction model and pathological multimodal database were optimized based on feedback results and prediction data.

Citation Information

Cited By

  • Multi-modal medical image analysis system and method

    CN121601273A