Prostate cancer endocrine therapy drug resistance prediction system based on multi-model artificial intelligence

By integrating multi-omics data through multi-model artificial intelligence, a prostate cancer endocrine therapy resistance prediction system was constructed. This system solves the accuracy and generalization problems of traditional models in predicting prostate cancer drug resistance, achieving higher prediction accuracy and clinical applicability, while also meeting the requirements of data privacy protection and multi-center data sharing.

CN120527031BActive Publication Date: 2026-05-01HUNAN PROVINCIAL PEOPLES HOSPITAL
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HUNAN PROVINCIAL PEOPLES HOSPITAL
Filing Date
2025-05-27
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing technologies have limitations in the accuracy and generalization of predicting drug resistance in endocrine therapy for prostate cancer. The interaction of multiple factors leads to poor predictive performance of traditional single models, and there is insufficient cross-center validation and model interpretability.

Method used

A prostate cancer endocrine therapy resistance prediction system based on multi-model artificial intelligence was constructed. It integrates multi-omics data and adopts data acquisition and preprocessing, single-omics prediction, multi-modal feature integration and model training modules. Combined with knowledge enhancement and federated learning, an accurate multi-omics prediction model was constructed.

Benefits of technology

It improves the accuracy and clinical applicability of drug resistance prediction, overcomes the limitations of a single data source, identifies complex biomarker interactions, meets data privacy protection requirements, and supports multi-center data sharing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120527031B_ABST
    Figure CN120527031B_ABST
Patent Text Reader

Abstract

The application discloses a prostate cancer endocrine therapy drug resistance prediction system based on a multi-model artificial intelligence, which comprises a data acquisition and preprocessing module, a single-omics prediction module, a multi-modal feature integration module, a model training module and a prediction result output module. The multi-omics data of the patient is acquired and preprocessed, the data representation of the omics data with different labels is generated by using prior knowledge, the best single-omics prediction model is matched for the omics data with different labels based on the data representation enhanced by the knowledge, the best single-omics prediction model corresponding to the medical data with different labels is constructed into a multi-omics prediction model of the prostate cancer endocrine therapy drug resistance by using Stacking, the drug resistance risk of the patient within a preset time is predicted and an early warning is output, and the treatment scheme of the patient is adjusted according to the drug resistance risk. The application integrates the multi-omics data, constructs a precise multi-omics prediction model of the prostate cancer endocrine therapy drug resistance, and improves the accuracy and clinical applicability of the drug resistance prediction.
Need to check novelty before this filing date? Find Prior Art

Description

A Multi-Model Artificial Intelligence-Based Prediction System for Endocrine Therapy Resistance in Prostate Cancer Technical Field

[0001] This invention relates to the field of drug resistance prediction technology, and more specifically, to a prostate cancer endocrine therapy drug resistance prediction system based on multi-model artificial intelligence. Background Technology

[0002] Prostate cancer is one of the most common malignant tumors in men, and endocrine therapy (such as androgen deprivation therapy, ADT) is the main treatment for advanced prostate cancer. However, most patients eventually develop castration-resistant prostate cancer (CRPC), meaning they develop resistance to endocrine therapy. The mechanisms of resistance are complex, involving the interaction of multiple factors such as genomic variations, epigenetic regulation, and alterations in the tumor microenvironment, which limits the accuracy and generalization of predictions using traditional single models.

[0003] In recent years, the application of multi-model artificial intelligence (AI) technology in the medical field has provided new insights for predicting drug resistance in prostate cancer. By integrating multimodal data such as genomics, transcriptomics, proteomics, clinicopathological features, and imaging, multi-model AI can more comprehensively capture drug resistance-related biomarkers and molecular pathways. For example, deep learning-based convolutional neural networks (CNNs) can analyze pathological image features, graph neural networks (GNNs) can model gene interaction networks, and ensemble learning methods (such as random forests and XGBoost) can fuse heterogeneous data to improve prediction robustness. Furthermore, the introduction of transfer learning and federated learning helps address issues such as insufficient medical data sample size and privacy protection. Current challenges include the high-dimensional heterogeneity of multimodal data, insufficient model interpretability, and the need for cross-center validation for clinical translation. Therefore, how to use multi-model AI to improve the accuracy of predicting endocrine therapy resistance in prostate cancer is an urgent problem to be solved. Summary of the Invention

[0004] To address the aforementioned technical challenges, this invention proposes a multi-model artificial intelligence-based system for predicting endocrine therapy resistance in prostate cancer. This system integrates multi-omics data (genomics, transcriptomics, proteomics, clinicopathological features, and imaging data) to construct a precise multi-omics prediction model for endocrine therapy resistance in prostate cancer, thereby improving the accuracy and clinical applicability of resistance prediction.

[0005] This invention provides a prostate cancer endocrine therapy resistance prediction system based on multi-model artificial intelligence, including a data acquisition and preprocessing module, a single-omics prediction module, a multi-modal feature integration module, a model training module, and a prediction result output module;

[0006] The data acquisition and preprocessing module is responsible for collecting multi-omics data from patients and preprocessing the collected multi-omics data.

[0007] The single-omics prediction module uses prior knowledge to generate data representations of omics data with different labels, matches the best single-omics prediction model for omics data with different labels based on the knowledge-enhanced data representations, and trains the best single-omics prediction model corresponding to medical data with different labels.

[0008] The multimodal feature integration module uses Stacking to construct a multi-omics prediction model for endocrine therapy resistance in prostate cancer by using the best single-omics prediction model corresponding to different labeled medical data, and predicts the risk of drug resistance in patients within a preset time.

[0009] The model training module acquires training data, uses the balanced training data to train the best single-omics prediction model, and introduces federated learning to update the multi-omics prediction model for endocrine therapy resistance in prostate cancer.

[0010] The prediction result output module outputs an early warning of the patient's drug resistance risk within a preset time, and adjusts the patient's treatment plan based on the drug resistance risk.

[0011] In this solution, the data acquisition and preprocessing module specifically comprises:

[0012] Medical record data is obtained based on the patient's identity information. The obtained medical record data is cleaned and standardized. A medical data extraction model is built based on the BERT model and combined with the Bi-LSTM network. Relevant medical data is obtained through preset data category labels to train the medical data extraction model for span detection and category classification.

[0013] The preprocessed medical record data is imported into the trained medical data extraction model. The BERT model is used to segment words and generate text vectors. The text vectors are then encoded using a Bi-LSTM network to detect entity words in the medical record data. Continuous entity words are used as candidate spans.

[0014] An attention layer is introduced into the Bi-LSTM network to calculate the similarity between the text vector sequence and key value in the candidate span to obtain the attention weight. The attention weight is used to weight the vector, and the weighted result is used as the feature vector. Softmax is used for classification to extract data with preset data category labels from the medical record data.

[0015] Based on preset data category labels, corresponding medical data subsets are constructed to generate multi-omics data for patients, and different data preprocessing is performed on different medical data subsets.

[0016] In this scheme, the single-omics prediction module uses prior knowledge to generate data representations of omics data with different labels, specifically as follows:

[0017] We acquire multi-omics data of patients, construct a time-series encoder to encode medical data subsets with different data category labels to obtain the time dependencies of different data, and generate hidden layer representations of different data in the multi-omics data.

[0018] Based on the hidden layer representation of different data corresponding to the patient, similarity calculation is used to obtain prostate cancer endocrine therapy resistance record instances with similarity meeting preset requirements. From the prostate cancer endocrine therapy resistance record instances, patient groups similar to the patient are retrieved to construct context learning examples.

[0019] Spatially align the hidden layer representations and context learning examples of different patient data, and import them into a large medical model to generate personalized medical knowledge for the endocrine therapy resistance prediction task by labeling each data category.

[0020] The hidden representations of different data are encoded with personalized medical knowledge into embedding vectors, and the resulting data representations of omics data with different labels are enhanced by knowledge.

[0021] In this scheme, the optimal single-omics prediction model is matched to omics data with different labels based on the knowledge-enhanced data representation, and the optimal single-omics prediction model corresponding to medical data with different labels is trained, specifically as follows:

[0022] Based on the knowledge-enhanced data representation of patient multi-omics data, the predictive model categories that are associated with different data category labels are obtained in the task of predicting endocrine therapy resistance in prostate cancer.

[0023] A relationship graph is constructed using different data category labels, prediction model category labels, and associations. The relationship graph is pre-trained using Metapath random walk to generate metapaths containing different data category labels as head nodes.

[0024] A graph attention network is introduced to represent the meta-path with different data category labels as head nodes. The prediction model category nodes in the meta-path are used as the neighbor nodes of the head node. The node feature vector is obtained through graph convolution, and the attention weight of the neighbor nodes to the head node in different meta-paths is calculated.

[0025] The neighboring nodes in each meta-path are sorted by the attention weights, and the best single-omics prediction model in each meta-path is determined based on the sorting results. The best single-omics prediction model is trained by constructing training samples from subsets of medical data with different data category labels.

[0026] In this solution, within the multimodal feature integration module, Stacking is used to construct a multi-omics prediction model for prostate cancer endocrine therapy resistance by combining the best single-omics prediction models corresponding to different labeled medical data. Specifically:

[0027] Training samples with different data category labels are obtained to construct a base training dataset. Each best single-omics prediction model is trained. A 5-fold cross-validation strategy is adopted. The predicted probabilities output by each best single-omics prediction model are concatenated into a complete meta-feature matrix. The meta-feature matrix is ​​used as a new training dataset to train a meta-model based on an SVM network.

[0028] During the training of the meta-selector, the hyperparameters of the meta-model are adjusted by grid search. Each best single-omics prediction model is retrained using the entire training dataset. The contribution of each best single-omics prediction model is calculated using the non-negative least squares algorithm to generate the weight information of the meta-model. The prediction results of the meta-model are combined with the weight information to generate the final endocrine therapy resistance prediction result for prostate cancer.

[0029] Model validation is performed using an independent test dataset. When the final endocrine therapy resistance prediction result for prostate cancer meets the preset test criteria, a multi-omics prediction model for endocrine therapy resistance in prostate cancer is output.

[0030] Based on the patient's medical information, the medication data of the patient within a preset time period is determined. The data representation of the patient's multi-omics data after knowledge enhancement is used as the input of the model to predict the probability of drug resistance of the patient to the medication data within the preset time period and output the patient's drug resistance risk.

[0031] In this scheme, the model training module acquires training data and uses the balanced training data to train the optimal single-omics prediction model, specifically as follows:

[0032] Feature extraction is performed on prostate cancer endocrine therapy resistance record instances corresponding to different data category labels. Based on the real labels in the instances, resistance training data and non-resistance training data are constructed. The majority class training sample set and minority class training sample set are defined by the number of samples in the resistance training data and non-resistance training data.

[0033] In the minority class training samples, a training sample is selected as the root sample. The K training samples closest to the root sample are obtained to generate a neighbor sample set. From the neighbor sample set, one neighbor sample is randomly selected as an auxiliary sample for sample synthesis.

[0034] A new training sample is synthesized by random difference between the root sample and the auxiliary sample. The new training sample is added to a new training sample set. The minority class training sample is expanded using the new training sample set. The optimal single-omics prediction model is trained using the balanced training data.

[0035] In this approach, the federated learning concept is introduced into the model training module to update the multi-omics prediction model for endocrine therapy resistance in prostate cancer. Specifically:

[0036] The multi-omics prediction model of endocrine therapy resistance in prostate cancer was used as the initial global model. Training data from hospital nodes in different regions were extracted and local data standardization was performed to initialize federated parameters. The model was then trained and updated using a federated learning algorithm.

[0037] In federated training, the central server sends the parameters of the initial global model to hospital nodes in different regions to build local models. Each hospital node updates its local model using local data. The hospital nodes upload the trained local models, which are then weighted and aggregated on the central server. Finally, the aggregated and updated global model is retrained using the balanced training data.

[0038] The updated global model is redistributed, local model parameters are updated, and each hospital node performs drug resistance prediction. The drug resistance prediction results and confidence scores are used to guide federated training. The global model is validated on the central server. After successful validation, the multi-omics prediction model for endocrine therapy resistance in prostate cancer is updated.

[0039] A second aspect of this invention provides a method for predicting endocrine therapy resistance in prostate cancer based on multi-model artificial intelligence, comprising the following steps:

[0040] Collect multi-omics data from patients and preprocess the collected multi-omics data;

[0041] Data representations of omics data with different labels are generated using prior knowledge. Based on the knowledge-enhanced data representations, the best single-omics prediction model is matched for omics data with different labels, and the best single-omics prediction model corresponding to medical data with different labels is trained.

[0042] Stacking was used to construct a multi-omics prediction model for endocrine therapy resistance in prostate cancer by using the best single-omics prediction model corresponding to different labeled medical data. Federated learning was introduced to train the model and predict the risk of drug resistance in patients within a preset time.

[0043] The system outputs an early warning of the patient's drug resistance risk within a preset time period, and adjusts the patient's treatment plan based on the drug resistance risk.

[0044] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0045] This invention utilizes multi-model artificial intelligence technology to integrate multi-omics data (genomics, transcriptomics, proteomics, clinicopathological features, and imaging data) to construct a precise predictive model for endocrine therapy resistance in prostate cancer. The system adopts a modular design, including data preprocessing, feature extraction, multi-model fusion, predictive analysis, and visualization output, to improve the accuracy and clinical applicability of resistance prediction.

[0046] This invention integrates multi-dimensional data such as genomic mutations, transcriptome expression, and clinical indicators, overcoming the limitations of a single data source. The combination of base models such as XGBoost and neural networks with meta-models can effectively identify complex biomarker interactions, improve robustness, and allow future access to multi-center data (each hospital trains its base model locally, sharing only meta-features), which complies with data privacy requirements. Attached Figure Description

[0047] To more clearly illustrate the technical solutions in the embodiments or examples of the present invention, the drawings used in the embodiments or examples will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained according to these drawings without creative effort.

[0048] Figure 1 shows a block diagram of a prostate cancer endocrine therapy resistance prediction system based on multi-model artificial intelligence.

[0049] Figure 2 shows a flowchart for matching the best single-omics prediction model to omics data with different labels;

[0050] Figure 3 shows a flowchart of constructing a multi-omics prediction model for endocrine therapy resistance in prostate cancer;

[0051] Figure 4 shows a flowchart of a method for predicting endocrine therapy resistance in prostate cancer based on multi-model artificial intelligence; Detailed Implementation

[0052] To better understand the above-mentioned objectives, features, and advantages of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that, unless otherwise specified, the embodiments and features described in these embodiments can be combined with each other.

[0053] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and therefore the scope of protection of the invention is not limited to the specific embodiments disclosed below.

[0054] Figure 1 shows a block diagram of a prostate cancer endocrine therapy resistance prediction system based on multi-model artificial intelligence.

[0055] As shown in Figure 1, the first embodiment of the present invention provides a prostate cancer endocrine therapy resistance prediction system based on multi-model artificial intelligence, including: a data acquisition and preprocessing module 101, a single-omics prediction module 102, a multimodal feature integration module 103, a model training module 104, and a prediction result output module 105.

[0056] The data acquisition and preprocessing module is responsible for collecting multi-omics data from patients and preprocessing the collected multi-omics data.

[0057] The single-omics prediction module uses prior knowledge to generate data representations of omics data with different labels, matches the best single-omics prediction model for omics data with different labels based on the knowledge-enhanced data representations, and trains the best single-omics prediction model corresponding to medical data with different labels.

[0058] The multimodal feature integration module uses Stacking to construct a multi-omics prediction model for endocrine therapy resistance in prostate cancer by using the best single-omics prediction model corresponding to different labeled medical data, and predicts the risk of drug resistance in patients within a preset time.

[0059] The model training module acquires training data, uses the balanced training data to train the best single-omics prediction model, and introduces federated learning to update the multi-omics prediction model for endocrine therapy resistance in prostate cancer.

[0060] The prediction result output module outputs an early warning of the patient's drug resistance risk within a preset time, and adjusts the patient's treatment plan based on the drug resistance risk.

[0061] It should be noted that medical record data is obtained based on the patient's identity information. This data undergoes cleaning and standardization. A medical data extraction model is constructed based on the BERT model combined with a Bi-LSTM network. Relevant medical data is obtained through preset data category labels to train the model for span detection and category classification. These preset data category labels include: clinical data (patient age, PSA level, Gleason score, treatment plan, drug resistance (CRPC progression time)), genomic data (AR mutation, PTEN deletion, TP53 mutation, and other drug resistance-related variants), transcriptome data (differentially expressed genes (e.g., NKX3-1, ERG), non-coding RNA (lncRNA, miRNA)), proteome data (AR-V7, PSMA, PTEN protein expression levels), and imaging data (tumor volume, ADC value, metabolic characteristics). The medical data with the preset data category labels serves as the source domain dataset for the extraction model, used for training in the span detection and category classification stages. A contrastive learning method is employed to improve the accuracy and stability of the extraction model.

[0062] In the target domain, preprocessed medical record data is imported into a trained medical data extraction model. A BERT model is used for word segmentation to generate text vectors, which are then encoded using a Bi-LSTM network to detect entity words in the medical record data, using consecutive entity words as candidate spans. The BERT model extracts word vectors from the preprocessed medical record data, completely preserving the semantic information of the case data and improving the ability to extract bidirectional contextual features. The Bi-LSTM network, composed of a forward LSTM and a backward LSTM, is trained in two directions, possessing the ability to remember past and future information, thus overcoming the difficulty of capturing contextual information. An attention mechanism is added to the Bi-LSTM network to distinguish the importance of different features, ignoring unimportant features and focusing attention on important features, thereby improving classification accuracy. An attention layer is introduced into the Bi-LSTM network to calculate the similarity between text vector sequences and key values ​​in the candidate span, thereby obtaining attention weights. The attention weights are then used to weight the vectors, and the weighted result is used as a feature vector. Softmax is used for classification to extract data with preset data category labels from the medical record data. Based on the preset data category labels, corresponding medical data subsets are constructed to generate multi-omics data for patients. Different data preprocessing is performed on different medical data subsets. For example, variant detection (GATK, STAR) is performed on genomic and transcriptomic data to obtain variant annotation (ANNOVAR), and key genes are screened through differential expression analysis (DESeq2 / edgeR) (|logFC|>1, FDR<0.05). Image data is preprocessed with image registration, ROI segmentation, feature extraction, and other methods.

[0063] It should be noted that in the single-omics prediction module, multi-omics data of patients are acquired, and a time encoder is constructed to encode medical data subsets with different data category labels to obtain the time dependencies of different data and generate hidden layer representations of different data in multi-omics data. Preferably, HiTANet is used as the encoding model for multi-omics data. Based on the hidden representations of different data corresponding to patients, similarity calculations are used to obtain prostate cancer endocrine therapy resistance record instances with similarity meeting preset requirements. From these instances, patient groups similar to the patient are retrieved to construct context learning examples. A large model is guided to learn potential patterns through relevant information about similar patients. Predictive models using medical data with different data category labels are read, and the hidden representations of different data corresponding to patients and context learning examples are spatially aligned. These are then imported into large medical models (such as BioGPT, Med-PaLM, ClinicalBERT, etc.) to generate personalized medical knowledge for each data category label in the endocrine therapy resistance prediction task. By utilizing the rich medical knowledge within the large model, the distribution differences between the two types of data are aligned, and the global information of both types of data is combined to generate valuable personalized medical knowledge for the patient's prostate cancer endocrine therapy resistance prediction task. The hidden representations of different data and personalized medical knowledge are encoded into embedding vectors, and the output is a knowledge-enhanced data representation of omics data with different labels. For example, inputs could include: patient A's EMR text, AR mutation detection report, and MRI images. The output is enhanced by knowledge from a large medical model: Text embedding: prompts “rapid rise in PSA + bone metastasis” → high risk of drug resistance; Gene embedding: AR mutation is mapped to “androgen receptor signaling pathway activation”; Image embedding: PI-RADS score of 5 → strong tumor invasiveness.

[0064] Figure 2 shows a flowchart for matching the best single-omics prediction model to omics data with different labels.

[0065] According to an embodiment of the present invention, based on the knowledge-enhanced data representation, the optimal single-omics prediction model is matched for omics data with different labels, and the optimal single-omics prediction model corresponding to medical data with different labels is trained, specifically as follows:

[0066] S202, based on the knowledge-enhanced data representation of patient multi-omics data, obtain the predictive model categories that are associated with different data category labels in the task of predicting endocrine therapy resistance in prostate cancer;

[0067] S204, construct a relationship graph using different data category labels, prediction model category labels, and associations, and pre-train the relationship graph using Metapath random walk to generate metapaths containing different data category labels as head nodes;

[0068] S206. A graph attention network is introduced to represent the meta-path with different data category labels as head nodes. The prediction model category nodes in the meta-path are used as the neighbor nodes of the head node. The node feature vector is obtained through graph convolution, and the attention weight of the neighbor nodes to the head node in different meta-paths is calculated.

[0069] S208, the neighbor nodes in each meta-path are sorted by the attention weights, the best single-omics prediction model in each meta-path is determined according to the sorting results, and the best single-omics prediction model is trained and matched by a subset of medical data with different data category labels.

[0070] It should be noted that Metapath2vec++ pre-training is used to represent the relationship graph in high dimension, define path parameters, and generate walking sequences to construct corresponding meta-paths. A graph attention network is used to represent these meta-paths, and graph convolution is used to learn the feature information of neighboring entities. The weights of neighboring entities in each meta-path are obtained through an attention mechanism, representing the importance of neighboring entities to the head entity. The neighbor node with the largest weight in each meta-path is selected as the optimal single-omics prediction model. For example, for genomic / transcriptome data, a gene interaction network (STRING database) is constructed, and GraphSAGE is used to learn key pathways for prediction; for clinical data, XGBoost or LightGBM is used for prediction, and SHAP value analysis is used to screen important features and predict drug resistance risk scores; for imaging data, 3D CNN + Transformer is used for drug resistance prediction, and 3D ResNet is used to extract spatial features, while ViT (Vision Transformer) captures long-range dependencies.

[0071] Figure 3 shows a flowchart of constructing a multi-omics prediction model for endocrine therapy resistance in prostate cancer.

[0072] According to an embodiment of the present invention, in the multimodal feature integration module, Stacking is used to construct a multi-omics prediction model for endocrine therapy resistance in prostate cancer using the optimal single-omics prediction model corresponding to different labeled medical data. Specifically:

[0073] S302, obtain training samples with different data category labels to construct a base training dataset, train each best single-omics prediction model, adopt a 5-fold cross-validation strategy, concatenate the prediction probabilities output by each best single-omics prediction model into a complete meta-feature matrix, and use the meta-feature matrix as a new training dataset to train a meta-model based on the SVM network.

[0074] S304, during the training of the meta-selector, the hyperparameters of the meta-model are adjusted by grid search, each best single-omics prediction model is retrained using the entire training dataset, the contribution of each best single-omics prediction model is calculated using the non-negative least squares algorithm, the weight information of the meta-model is generated, and the prediction results of the meta-model are combined with the weight information to generate the final endocrine therapy resistance prediction result for prostate cancer.

[0075] S306. Use an independent test dataset to validate the model. When the final endocrine therapy resistance prediction result of prostate cancer meets the preset test criteria, output the multi-omics prediction model of endocrine therapy resistance of prostate cancer.

[0076] S308 determines the patient's medication data within a preset time period based on the patient's condition information, uses the patient's multi-omics data after knowledge enhancement as the data representation as the model input, predicts the probability of the patient's drug resistance to the medication data within the preset time period, and outputs the patient's drug resistance risk.

[0077] It should be noted that Stacking consists of two layers: a base model and a meta-model. The base model trains the optimal single-omics prediction model for different omics data. The meta-model trains the final classifier based on the output of the base model. Five-fold cross-validation is used for each base model, and the cross-validation prediction probabilities of the base models are concatenated into a complete meta-feature matrix. This meta-feature matrix is ​​used as a new training set, retaining the original true labels. An SVM network is selected as the meta-model, and the meta-model hyperparameters (such as learning rate and tree depth) are optimized through grid search. All optimal single-omics predictions are retrained using all training data. Stacking is used to construct a multi-omics prediction model for endocrine therapy resistance in prostate cancer from the optimal single-omics prediction models corresponding to different labeled medical data. This model integrates multi-dimensional information from genomic mutations, gene expression, and clinical indicators. When the quality of a certain omics data is poor, other omics can provide compensation, and the contribution weights of each omics are output to help doctors understand the prediction basis. It also supports the flexible addition of other modalities such as radiomics and proteomics data.

[0078] It should be noted that in the model training module, features are extracted from prostate cancer endocrine therapy resistance record instances corresponding to different data category labels. Based on the true labels in the instances, resistance training data and non-resistance training data are constructed. The data is divided into a training set and an independent test set in a 7:3 ratio. To ensure a consistent ratio of resistance / non-resistance samples in the two groups, the majority class training sample set and the minority class training sample set are defined by the number of samples in the resistance training data and the non-resistance training data. A training sample is selected as the root sample from the minority class training samples. The K nearest training samples to the root sample are used to generate a neighbor sample set. One neighbor sample is randomly selected from the neighbor sample set as an auxiliary sample for sample synthesis. A new training sample is synthesized by random difference between the root sample and the auxiliary sample. The new training sample is added to the new training sample set. The new training sample set is used to expand the minority class training samples. The optimal single-omics prediction model is trained using the balanced training data.

[0079] In the model training module, federated learning is introduced to update the multi-omics prediction model for endocrine therapy resistance in prostate cancer. Using this model as the initial global model, training data from hospital nodes in different regions is extracted and standardized locally. A data isolation zone is established, and all computations are performed in a secure local environment. The federated parameters are initialized, including the aggregation algorithm (FedAvg), communication rounds (T), and minimum number of participating nodes (K≥3). Training and updates are performed using a federated learning algorithm. During federated training, the central server sends the initial global model parameters to hospital nodes in different regions to build local models. Each hospital node updates its local model using local data. The trained local models are then uploaded to the central server for weighted aggregation. The aggregated and updated global model is retrained using the balanced training data to alleviate data sparsity to some extent. Data augmentation also protects patients' real data from leakage. The updated global model is redistributed, and the local model parameters are updated. After each hospital node performs drug resistance prediction, the results and confidence scores guide federated training. The global model is validated on the central server. Once validation is successful, the multi-omics prediction model for prostate cancer endocrine therapy resistance is updated. In the dynamic update mechanism of the multi-omics prediction model, model distillation (DistilFL) is used for rapid adaptation, allowing new hospitals to join. Federated training is automatically triggered monthly for model iteration.

[0080] It should be noted that the prediction results output module automatically generates a PDF report containing risk scores, decision-making basis, and literature support (example: This patient's lack of PTEN + high Gleason score leads to an 87% probability of drug resistance to a certain drug (95% CI: 82-91%)). The system predicts the patient's drug resistance risk within 1-2 years, adjusts the treatment plan in advance, and achieves visual early warning.

[0081] Figure 4 shows a flowchart of a method for predicting endocrine therapy resistance in prostate cancer based on multi-model artificial intelligence.

[0082] The second embodiment of the present invention provides a method for predicting endocrine therapy resistance in prostate cancer based on multi-model artificial intelligence, comprising the following steps:

[0083] S402, collect multi-omics data from patients and preprocess the collected multi-omics data;

[0084] S404 uses prior knowledge to generate data representations of omics data with different labels, matches the best single-omics prediction model for omics data with different labels based on the knowledge-enhanced data representations, and trains the best single-omics prediction model for medical data with different labels.

[0085] S406 uses Stacking to construct a multi-omics prediction model for endocrine therapy resistance in prostate cancer by using the best single-omics prediction model corresponding to different labeled medical data. It introduces federated learning to train the model and predict the risk of drug resistance in patients within a preset time.

[0086] S408 outputs an early warning of the patient's drug resistance risk within a preset time period, and adjusts the patient's treatment plan according to the drug resistance risk.

[0087] A third embodiment of the present invention provides a computer-readable storage medium, which includes a program for predicting prostate cancer endocrine therapy resistance based on multi-model artificial intelligence. When the program for predicting prostate cancer endocrine therapy resistance based on multi-model artificial intelligence is executed by a processor, it implements the steps of the method for predicting prostate cancer endocrine therapy resistance based on multi-model artificial intelligence.

[0088] In the embodiments provided in this application, it should be understood that the disclosed methods can be implemented in other ways. The embodiments described above are merely illustrative. For example, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods, such as: multiple modules or components can be combined, or integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the various components shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or modules, and can be electrical, mechanical, or other forms.

[0089] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0090] Alternatively, if the integrated modules of this invention are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this invention, or the parts that contribute to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, ROM, RAM, magnetic disks, or optical disks.

[0091] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention.

Claims

1. A prostate cancer endocrine therapy resistance prediction system based on multi-model artificial intelligence, characterized in that, The system includes a data acquisition and preprocessing module, a single-omics prediction module, a multimodal feature integration module, a model training module, and a prediction result output module. The data acquisition and preprocessing module collects multi-omics data from patients and preprocesses it. The single-omics prediction module uses prior knowledge to generate data representations of omics data with different labels, matches the optimal single-omics prediction model to the omics data with different labels based on the knowledge-enhanced data representations, and trains the optimal single-omics prediction model corresponding to the medical data with different labels. The multimodal feature integration module uses Stacking to construct a multi-omics prediction model for prostate cancer endocrine therapy resistance using the optimal single-omics prediction models corresponding to the medical data with different labels, predicting patient outcomes. The model training module acquires training data, uses the balanced training data to train the optimal single-omics prediction model, and introduces federated learning to update the multi-omics prediction model for endocrine therapy resistance in prostate cancer; the prediction result output module outputs a warning about the patient's drug resistance risk within the preset time, and adjusts the patient's treatment plan based on the drug resistance risk; in the single-omics prediction module, prior knowledge is used to generate data representations of omics data with different labels, specifically: acquiring the patient's multi-omics data, constructing a time-series encoder to encode medical data subsets with different data category labels to obtain the time dependencies of different data, and generating hidden layer representations of different data in the multi-omics data; based on the patient's corresponding The hidden representations of different data are used to obtain prostate cancer endocrine therapy resistance record instances with similarity that meet preset requirements using similarity calculation. From these instances, patient groups similar to the patient are retrieved to construct context learning examples. The hidden representations of different data corresponding to the patient and the context learning examples are spatially aligned and imported into a large medical model to generate personalized medical knowledge for each data category label in the endocrine therapy resistance prediction task. The hidden representations of different data and personalized medical knowledge are encoded into embedding vectors, outputting the knowledge-enhanced data representations of omics data with different labels. Based on the knowledge-enhanced data representations, the optimal single-omics prediction model is matched for omics data with different labels, and the medical data with different labels are analyzed. The training process involves: obtaining prediction model categories associated with different data category labels in the prostate cancer endocrine therapy resistance prediction task based on the knowledge-enhanced data representation of the patient's multi-omics data; constructing a relationship graph using different data category labels, prediction model category labels, and associations; pre-training the relationship graph using Metapath random walk to generate metapaths containing different data category labels as head nodes; introducing a graph attention network to represent the metapaths with different data category labels as head nodes, using the prediction model category nodes in the metapaths as neighbor nodes of the head nodes, obtaining node feature vectors through graph convolution, and calculating the attention weights of neighbor nodes to head nodes in different metapaths;The neighbor nodes in each meta-path are sorted using the attention weights. Based on the sorting results, the optimal single-omics prediction model for each meta-path is determined. Training samples are then constructed using subsets of medical data with different data category labels to train and match the optimal single-omics prediction model.

2. The prostate cancer endocrine therapy resistance prediction system based on multi-model artificial intelligence according to claim 1, characterized in that, The data acquisition and preprocessing module specifically includes: acquiring medical record data based on the patient's identity information; cleaning and standardizing the acquired medical record data; constructing a medical data extraction model based on the BERT model combined with a Bi-LSTM network; acquiring relevant medical data through preset data category labels to train the medical data extraction model for span detection and category classification; importing the preprocessed medical record data into the trained medical data extraction model; using the BERT model for word segmentation to generate text vectors; encoding the text vectors using a Bi-LSTM network; detecting entity words in the medical record data; and using consecutive entity words as candidate spans. An attention layer is introduced into the Bi-LSTM network to calculate the similarity between the text vector sequence and key value in the candidate span to obtain the attention weight. The attention weight is used to weight the vector, and the weighted result is used as the feature vector. Softmax is used for classification to extract data with preset data category labels from the medical record data. Based on preset data category labels, corresponding medical data subsets are constructed to generate multi-omics data for patients, and different data preprocessing is performed on different medical data subsets.

3. The prostate cancer endocrine therapy resistance prediction system based on multi-model artificial intelligence according to claim 1, characterized in that, In the multimodal feature integration module, Stacking is used to construct a multimodal prediction model for endocrine therapy resistance in prostate cancer using the best single-omics prediction models corresponding to different labeled medical data. Specifically, training samples with different data category labels are obtained to construct a base training dataset. Each best single-omics prediction model is trained using a 5-fold cross-validation strategy. The predicted probabilities output by each best single-omics prediction model are concatenated into a complete meta-feature matrix. The meta-feature matrix is ​​used as a new training dataset to train the meta-model based on an SVM network. During the training of the meta-selector, the hyperparameters of the meta-model are adjusted through grid search. Each best single-omics prediction model is retrained using the entire training dataset. The contribution of each best single-omics prediction model is calculated using a non-negative least squares algorithm to generate the weight information of the meta-model. The prediction results of the meta-model are combined with the weight information to generate the final prediction result of endocrine therapy resistance in prostate cancer. Model validation is performed using an independent test dataset. When the final endocrine therapy resistance prediction result for prostate cancer meets the preset test criteria, a multi-omics prediction model for prostate cancer endocrine therapy resistance is output. Based on the patient's condition information, the patient's medication data within a preset time period is determined. The patient's multi-omics data, after knowledge enhancement, is used as the data representation as the model input to predict the patient's drug resistance probability based on the medication data within the preset time period, and the patient's drug resistance risk is output.

4. The prostate cancer endocrine therapy resistance prediction system based on multi-model artificial intelligence according to claim 1, characterized in that, In the model training module, training data is acquired, and the optimal single-omics prediction model is trained using the balanced training data. Specifically, features are extracted from prostate cancer endocrine therapy resistance record instances corresponding to different data category labels. Based on the true labels in the instances, resistance training data and non-resistance training data are constructed. The majority class training sample set and minority class training sample set are defined by the number of samples in the resistance training data and non-resistance training data. A training sample is selected as the root sample from the minority class training samples. The K training samples closest to the root sample are obtained to generate a neighbor sample set. One neighbor sample is randomly selected from the neighbor sample set as an auxiliary sample for sample synthesis. A new training sample is synthesized by random difference between the root sample and the auxiliary sample. The new training sample is added to the new training sample set. The minority class training samples are expanded using the new training sample set. The optimal single-omics prediction model is trained using the balanced training data.

5. The prostate cancer endocrine therapy resistance prediction system based on multi-model artificial intelligence according to claim 4, characterized in that, In the model training module, the concept of federated learning is introduced to update the multi-omics prediction model for prostate cancer endocrine therapy resistance. Specifically, the multi-omics prediction model for prostate cancer endocrine therapy resistance is used as the initial global model. Training data from hospital nodes in different regions is extracted and standardized locally to initialize federated parameters. The model is then trained and updated using a federated learning algorithm. During federated training, the central server sends the parameters of the initial global model to hospital nodes in different regions to build local models. Each hospital node updates its local model using local data. The trained local models are then uploaded by the hospital nodes and weighted and aggregated on the central server. The aggregated and updated global model is then retrained using the balanced training data. The updated global model is then redistributed, and the local model parameters are updated. After each hospital node performs resistance prediction, it uses the resistance prediction results and confidence scores to guide federated training. The global model is validated on the central server. Once validation is successful, the multi-omics prediction model for prostate cancer endocrine therapy resistance is updated.

6. A method for predicting endocrine therapy resistance in prostate cancer based on multi-model artificial intelligence, characterized in that, The system for predicting prostate cancer endocrine therapy resistance based on multi-model artificial intelligence as described in any one of claims 1-5 includes the following steps: collecting multi-omics data from patients and preprocessing the collected multi-omics data; generating data representations of omics data with different labels using prior knowledge, matching the best single-omics prediction model for omics data with different labels based on the knowledge-enhanced data representations, and training the best single-omics prediction model corresponding to the medical data with different labels; constructing a multi-omics prediction model for prostate cancer endocrine therapy resistance using Stacking to construct the best single-omics prediction model corresponding to the medical data with different labels, introducing federated learning to train the model, and predicting the patient's drug resistance risk within a preset time period; outputting an early warning of the patient's drug resistance risk within the preset time period, and adjusting the patient's treatment plan according to the drug resistance risk.

Citation Information

Patent Citations

  • Method and device for predicting tumor prognosis through imaging omics machine learning survival model

    CN117173167A

  • Medical service platform hospital guide analysis method and system based on intelligent response

    CN118538430A