Method and device for predicting lifetime of tumor patient, computing equipment and storage medium

By training and fine-tuning a large language model, and combining it with multidimensional medical data and similar datasets, the problems of data integration and interpretability in the prediction of survival of cancer patients were solved, and accurate and transparent survival prediction was achieved.

CN121565495APending Publication Date: 2026-02-24CHONGQING YIHONG TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511769131.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-27
Publication Date
2026-02-24

AI Technical Summary

Technical Problem

Existing models for predicting the survival of cancer patients are unable to efficiently integrate multi-dimensional heterogeneous data, lack interpretability, and have insufficient adaptability and generalization ability, especially in cases of rare tumor types or when data is scarce, resulting in inaccurate predictions.

Method used

A large language model is used to train and fine-tune multidimensional medical data. The model is optimized using a hybrid loss function. Explanatory text is generated to predict survival by constructing reference datasets and similar datasets.

Benefits of technology

It enables accurate prediction of survival for cancer patients, enhances the transparency and reliability of the model, and improves the prediction accuracy and generalization ability under different tumor types and individual differences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121565495A_ABST
    Figure CN121565495A_ABST
Patent Text Reader

Abstract

The invention provides a method and device for predicting the lifetime of a tumor patient, computing equipment and a storage medium. The method comprises the following steps: acquiring multi-dimensional medical data related to the condition of a tumor patient, wherein the multi-dimensional medical data at least comprises data of at least two dimensions; determining a similar dataset of the multi-dimensional medical data from a reference dataset, the reference dataset including reference multi-dimensional medical data and corresponding lifetime of each of a plurality of reference patients having reached the end of life at different time points; and obtaining a prediction result of the lifetime of the tumor patient based on the multi-dimensional medical data and the similar data set by using a fine-tuned large language model, the prediction result indicating the lifetime of the tumor patient and an explanatory text of information on which the prediction result is based.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and more specifically, to a method, apparatus, computing device, and computer-readable storage medium for predicting the survival of cancer patients. Background Technology

[0002] In clinical oncology practice, accurately predicting the survival of cancer patients is a complex and urgent problem, directly limited by a variety of factors, including:

[0003] The diversity of data and the challenges of integration: The diversity of tumor data is reflected in its inclusion of a wide range of dimensions, including biomarker data, imaging data, clinical symptom data, and gene sequencing data. Furthermore, if survival prediction for cancer patients is required, other relevant information, such as electronic health record data, often needs to be integrated. This data is not only voluminous but also highly heterogeneous. Traditional predictive models often struggle to efficiently integrate this diverse data, especially when the data sources are varied and the formats differ, making the efficiency of information integration and utilization a major bottleneck.

[0004] Lack of interpretability in models: Existing predictive models, especially those based on complex algorithms such as deep learning networks, while improving predictive performance to some extent, often act as "black boxes," making it difficult to clearly demonstrate the specific reasons and logic behind the predictions to doctors and patients. This not only reduces the clinical acceptance of the models but also hinders the development of personalized treatment plans and the establishment of trust to some extent.

[0005] Insufficient adaptability and generalization: There are many types of tumors, and the progression of disease varies greatly from patient to patient. Existing predictive models often exhibit low adaptability and generalization when dealing with rare tumor types or when data is scarce, making it difficult to provide accurate and reliable survival predictions.

[0006] Therefore, there is a need for a scheme to predict the survival of cancer patients that can integrate multi-dimensional cancer-related medical data, explain the reasons and logic on which the prediction is based, and has good adaptability and generalization ability. Summary of the Invention

[0007] According to one aspect of this application, a method for predicting the survival of cancer patients is provided. The method may include: acquiring multidimensional medical data related to the cancer patient's condition, the multidimensional medical data including at least two dimensions of data; determining a similar dataset of the multidimensional medical data from a reference dataset, wherein the reference dataset includes reference multidimensional medical data and corresponding survival times of each of a plurality of reference patients who have reached the end of their lives at different time points; and obtaining a prediction result of the survival of the cancer patient based on the multidimensional medical data and the similar dataset using a finely tuned large language model, wherein the prediction result indicates the survival time of the cancer patient and explanatory text of the information on which the prediction result is based.

[0008] According to embodiments of this application, the large language model can be trained in the following manner: A training dataset associated with the reference dataset is obtained, wherein each training data includes input sample data and output sample data. The input sample data includes multidimensional medical data of a corresponding patient who has reached the end of life at a corresponding time point, and the output sample data includes the survival period label of the corresponding patient and reference explanatory text. For each training data, a similar training dataset similar to the multidimensional medical data in the training data is determined, wherein each similar multidimensional medical data and its corresponding survival period in the similar training dataset serve as a prompt associated with the training data during the training process. The input sample data of each training data is then input into the large language model, and the large language model is fine-tuned based on the difference between the predicted output of the large language model and the corresponding output sample data.

[0009] According to embodiments of this application, the input sample data is an input sequence with a preset structure, the input sequence including one or more of system instructions, reference context, corresponding patients, and interpretive triggers. Furthermore, fine-tuning the large language model includes using a hybrid loss function, wherein the hybrid loss function is constructed based on a causal language modeling loss associated with the fluency and logicality of the interpretive text, a numerical regression loss associated with differences in survival prediction, and a ranking consistency loss associated with the true survival ranking relationship between paired samples.

[0010] According to embodiments of this application, determining a similar dataset of the multidimensional medical data from a reference dataset may include: mapping the multidimensional medical data into a multimodal feature space, the multimodal feature space including hard constraint features, numerical clinical features, and unstructured text features corresponding to different dimensions; based on the hard constraint features, using an inverted index or Boolean logic to filter a candidate patient pool from the reference dataset, wherein the reference patients in the candidate patient pool match the tumor patients on the hard constraint features; for each reference patient in the candidate patient pool, determining a similar dataset of the multidimensional medical data based on the unstructured text features and numerical clinical features corresponding to the reference patient and the numerical clinical features and unstructured text features corresponding to the tumor patients. Optionally, each numerical clinical feature corresponding to the different dimensions is assigned a corresponding weight. Determining the similar dataset of the multidimensional medical data includes: calculating the weighted mixed distance with the tumor patient based on the semantic similarity between the unstructured text features corresponding to the reference patient and the unstructured text features corresponding to the tumor patient, and the weighted difference between the numerical clinical features corresponding to the reference patient and the numerical clinical features corresponding to the tumor patient; and retaining the reference multidimensional data corresponding to the reference patient whose weighted mixed distance is less than a preset threshold as the similar dataset based on a dynamic threshold strategy.

[0011] According to embodiments of this application, determining a similar dataset from a reference dataset for the multidimensional medical data may include: determining multidimensional medical data in the reference dataset that are identical or similar to the multidimensional medical data in core attributes, thus obtaining a filtered dataset; generating a corresponding reference feature vector for each piece of reference multidimensional medical data in the filtered dataset based on one or more selected dimensions; and using a similarity algorithm based on the feature vector of the multidimensional medical data and each reference feature vector, determining a preset number of reference multidimensional medical data that are most similar to the multidimensional medical data of the tumor patient from the filtered dataset, thus obtaining the similar dataset for the multidimensional medical data.

[0012] According to an embodiment of this application, the reference dataset is constructed in the following manner: acquiring multi-dimensional medical data for different patients at different time points from different data sources; selecting multiple reference patients who have reached the end of life from the different patients, and determining the time point of the end of life for each reference patient; and for each of the multiple reference patients, generating corresponding multi-dimensional medical data at different time points from the multi-dimensional medical data belonging to the reference patient at different time points, and determining the corresponding survival period of the reference patient at the different time points based on the time point of the end of life of the reference patient.

[0013] According to an embodiment of this application, each piece of medical data acquired includes timestamp information, and the multidimensional medical data includes medical data of preset dimensions. When the medical data of multiple dimensions acquired for a certain reference patient at a first time point lacks the first dimension of medical data relative to the preset dimensions, the first dimension of medical data at the first time point is determined based on the first dimension of medical data for the certain reference patient at other time points.

[0014] According to an embodiment of this application, a prediction result for the survival of a cancer patient is obtained using a fine-tuned large language model based on the multidimensional medical data and the similar dataset. This includes: integrating the multidimensional medical data and the similar data to generate integrated data with a predetermined input format; encoding the integrated data using an encoder associated with the fine-tuned large language model to obtain the model input data; and obtaining the prediction result for the survival of the cancer patient based on the model input data using the fine-tuned large language model.

[0015] According to an embodiment of this application, the encoder separately encodes the multidimensional medical data and the similar dataset in the integrated data, such that the encoding information of the similar dataset serves as a prompting information for the large language model during prediction.

[0016] According to embodiments of this application, obtaining multi-dimensional medical data for different patients at different time points from different data sources includes: obtaining text data for different patients at different time points from different data sources; and extracting key information from the obtained text data using natural language processing technology, and generating structured multi-dimensional medical data based on the extracted key information.

[0017] According to embodiments of this application, the different data sources include at least two of the following: electronic health records from different hospitals, laboratory test reports, imaging data, gene sequencing data, and clinical trial databases.

[0018] According to another aspect of this application, an apparatus for predicting the survival of cancer patients is also provided. The apparatus may include: an acquisition module for acquiring multidimensional medical data related to the cancer patient's condition, the multidimensional medical data including at least two dimensions of data; a determination module for determining a similar dataset of the multidimensional medical data from a reference dataset, wherein the reference dataset includes reference multidimensional medical data and corresponding survival times of each of a plurality of reference patients who have reached the end of their lives at different time points; and a prediction module for obtaining a prediction result of the survival of the cancer patient based on the multidimensional medical data and the similar dataset using a finely tuned large language model, wherein the prediction result indicates the survival time of the cancer patient and explanatory text of the information on which the prediction result is based.

[0019] According to another aspect of this application, a computing device is also provided, which may include: one or more processors; and one or more memories storing a computer program or instruction set thereon, which, when executed by the one or more processors, causes the one or more processors to perform the method as described above.

[0020] According to another aspect of this application, a computer-readable storage medium is also provided, on which a computer program or instruction set is stored, which, when executed by the one or more processors, causes the one or more processors to perform the method as described above.

[0021] The proposed method for predicting the survival of cancer patients utilizes a customized large language model to analyze processed and integrated multidimensional medical data. This model predicts the survival time of cancer patients after a specified time point, accurate to the month, and generates explanatory text that clearly explains the basis of the prediction, enhancing its transparency and reliability. Furthermore, by constructing a reference dataset, the model comprehensively analyzes multidimensional information from patients with different cancer types from various sources, continuously training and fine-tuning the large language model. This allows the model to deeply understand the language of oncology and the complexity of disease progression, thus providing more accurate survival predictions. Additionally, when predicting survival, similar datasets matched with similar patients are used as prompts, enabling more precise survival predictions and explanations, thereby enhancing the model's generalization ability and predictive accuracy. Attached Figure Description

[0022] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments of this application or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this application. For those skilled in the art, other drawings can be obtained based on these drawings of the embodiments of this application.

[0023] Figure 1 A schematic diagram of an application system for predicting the survival of cancer patients according to an embodiment of this application is shown.

[0024] Figure 2 A flowchart illustrating a method for predicting the survival of cancer patients according to an embodiment of this application is shown.

[0025] Figure 3 An example process for predicting the survival of cancer patients according to an embodiment of this application is shown.

[0026] Figure 4 A structural block diagram of a device 400 for predicting the survival of cancer patients according to an embodiment of this application is shown.

[0027] Figure 5 A structural block diagram of a computing device according to an embodiment of this application is shown. Detailed Implementation

[0028] The technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this disclosure. All other embodiments obtained by those skilled in the art based on the embodiments of this disclosure without creative effort are within the scope of protection of this disclosure.

[0029] Large language models or large models refer to deep learning models with a massive number of parameters, typically containing hundreds of millions, tens of billions, or even trillions of parameters. Large models are also known as foundational models (FM). They are pre-trained on large-scale unlabeled corpora, producing pre-trained models with hundreds of millions of parameters. These models can adapt to a wide range of downstream tasks and have good generalization ability. Examples include large-scale language models and multi-modal pre-training models. In practical applications, large models only require a small number of samples to fine-tune the pre-trained model before being applied to different tasks. Large models can be widely used in Natural Language Processing (NLP), computer vision, and other fields. Specifically, they can be applied to computer vision tasks such as visual question answering, image captioning, and image generation, as well as NLP tasks such as text-based sentiment classification, text summarization, and machine translation.

[0030] Depending on the specific application domain (e.g., oncology), the large language model can be specifically trained and fine-tuned to perform specific tasks within that domain. As mentioned earlier, a solution for predicting the survival of cancer patients is needed. Therefore, existing open-source large language models (e.g., Qwen1.5-MoE-A2.7B) can be used as the base model, and then fine-tuned using training data specific to the oncology domain to obtain a customized large language model for predicting the survival of cancer patients. The specific fine-tuning process will be described later.

[0031] It should be noted that the patient user information and data involved in this application (including but not limited to data used for analysis, stored data, and displayed data from various data sources) are all information and data authorized by the patient user or fully authorized by all parties. Furthermore, the collection, use, and processing of the relevant data must comply with relevant laws, regulations, and standards, and corresponding operation entry points are provided for the patient user to choose to authorize or refuse.

[0032] Figure 1 A schematic diagram of an application system for predicting the survival of cancer patients according to an embodiment of this application is shown.

[0033] like Figure 1As shown, the application system 10 may include a large language model 11, multiple different data sources 12 (12-1 / 12-2, ... 12-N), a data processing and integration device 13, and an edge device 14. Optionally, the application system 10 may also include a training device 15 for further training (e.g., fine-tuning) the large language model 11.

[0034] For current cancer patients, doctors can input relevant information about the patient's condition using the edge device 14, i.e., text-based medical data (e.g., this may include the patient's user information and clinical data, such as age, gender, lifestyle habits, cancer classification, stage (clinical and pathological), treatment effects, and complications, etc.). For example, "A 65-year-old male lung cancer patient, with no smoking history, is in an advanced stage and has received two cycles of chemotherapy. The most recent CT scan (time point: May 2023) shows that although the tumor has slightly shrunk, it is accompanied by mild pulmonary edema." This relevant information can be provided to the data processing and integration device 13, which processes and integrates it to generate input data with an input format suitable for the large language model 11. Explicit prediction task instructions can be added via the edge device 14, allowing the large language model to output a prediction of the current cancer patient's survival based on the input data and return the prediction result to the edge device 14. This prediction result may include the specific number of months corresponding to the survival period and explanatory text regarding the basis for the prediction.

[0035] The large language model 11 can be implemented based on existing open-source large language models, and can be fine-tuned according to the specific application of this application. For example, the training device 15 can input the input sample data (including input sample data and output sample data) from the training dataset into the large language model 11, obtain the prediction result from the large language model 11, and adjust the relevant parameters in the large language model 11 based on the difference between the prediction result and the output sample data. Since the large language model needs to predict the survival period of cancer patients based on their current condition information, each piece of training data used for training also needs to include the corresponding cancer patient's condition information and corresponding survival period. Specific details will be described in detail later. The training device 15 can be composed of hardware and corresponding software, such as a program or instruction set running on a processor.

[0036] The training data in the training dataset, or the reference data as described later, can be constructed based on data obtained from different data sources 12. For example, different data sources 12 could be electronic health records (EHRs) from different hospitals, laboratory test reports, imaging data, gene sequencing data, and clinical trial databases. Data from these different data sources can each provide information in different dimensions, and there may be information overlap between them. Data in different dimensions can be processed by the data processing and integration device 13 to obtain the training dataset or reference dataset.

[0037] The data processing and integration device 13 can be comprised of hardware and corresponding software, such as a program or instruction set running on a processor.

[0038] It should be noted that, Figure 1 Although the training device 14, the data processing and integration device 13 and the large language model 11 are shown separately, they can be physically located in the same device, for example, in a server.

[0039] The server can be a computing device deployed in the cloud or locally, such as a cloud cluster. The server stores a large language model 11, which can respond to the instructions of the prediction task, perform text processing based on text input data in a predetermined input format, and generate text processing results about the survival prediction results.

[0040] The edge device 14 can be an electronic device that runs downstream tasks, such as acquiring patient information and prediction task instructions input by the doctor for the current tumor patient, transmitting or receiving information and data from the server, or facilitating interaction with the user, etc. Specifically, the edge device 14 can be a hardware device with network communication, computing, and information display functions, including but not limited to smartphones, tablets, desktop computers, local servers, cloud servers, etc.

[0041] The following combination Figures 2 to 5 The specific details of the scheme for predicting the survival of cancer patients according to embodiments of this application are further described.

[0042] Figure 2 A flowchart illustrating a method for predicting the survival of cancer patients according to an embodiment of this application is shown. Optionally, the method may be... Figure 1 The server shown is used to execute this.

[0043] like Figure 2 As shown, in step S210, multidimensional medical data related to the condition of the cancer patient is obtained, and the multidimensional medical data includes data in at least two dimensions.

[0044] For example, the dimensions of multidimensional medical data can include various dimensions of basic user information and clinical parameter information of current cancer patients. Basic user information may include one or more dimensions such as age, gender, and lifestyle habits (e.g., smoking status), while clinical parameter information may include one or more dimensions such as tumor type, stage classification (clinical and pathological stages), treatment effects and complications, examinations or tests (e.g., imaging, biomarker measurements), and gene sequencing data. Typically, multidimensional medical data should include as many clinical parameter-related information as possible to more fully reflect the cancer patient's condition.

[0045] For example, doctors can input patient information through a terminal device, such as "A 65-year-old male lung cancer patient, no smoking history, in the advanced stage, has received two cycles of chemotherapy. The most recent CT scan (time point: May 2023) shows that although the tumor has shrunk slightly, it is accompanied by mild pulmonary edema." This patient information has multiple dimensions, such as the patient's age (65 years old), gender (male), tumor type (lung cancer), lifestyle habits (no smoking history), stage (advanced), CT scan (examination or test items), treatment effects, and complications (although the tumor has shrunk slightly, it is accompanied by mild pulmonary edema).

[0046] After obtaining the patient information input by the doctor, key information can be extracted using natural language processing techniques to facilitate computer processing. For example, a descriptive sentence can be formed: "Advanced lung cancer, 65-year-old male, no smoking history, CT scan after two cycles of chemotherapy showed slight tumor shrinkage, but accompanied by mild pulmonary edema." This descriptive sentence is equivalent to the multidimensional medical data mentioned in this application, and is shorter than the original text input by the doctor. Key information can correspond to multiple dimensions of the multidimensional medical data. In this application, any data ultimately input into a large language model or that needs to be processed by a large model is a feature representation (e.g., embedding vector) obtained after encoding or other similar methods. Methods of encoding and processing the input text are widely used in this field, so they will not be described in detail here to avoid obscuring the focus of this application.

[0047] In step S220, a similar dataset of the multidimensional medical data is determined from the reference dataset, wherein the reference dataset includes reference multidimensional medical data of each of a plurality of reference patients who have reached the end of their lives at different time points and their corresponding survival period.

[0048] For example, a reference dataset can be a collection of multiple multidimensional medical data corresponding to multiple different patients, and it can be stored in a database. For instance, medical data with different dimensions of information for different patients can be obtained from different data sources, and the obtained data can be processed and integrated to obtain the reference dataset.

[0049] Optionally, regarding the specific construction process of the reference dataset, multi-dimensional medical data for different patients at different time points can be obtained from different data sources. Multiple reference patients who have reached the end of their lives are then selected from these different patients, and the time point of the end of life for each reference patient is determined. In this application, "time point" can be in days or weeks. For each of the multiple reference patients, the multi-dimensional medical data belonging to that reference patient at different time points is used to generate corresponding multi-dimensional medical data for different time points, and the corresponding survival time of the reference patient at each different time point is determined based on the time point of the reference patient's end of life. That is, for each reference patient, multiple multi-dimensional medical data points at multiple different time points may be obtained, and these multi-dimensional medical data can be stored in a time point sequence. For example, multiple reference patients who have reached the end of life can be selected based on keyword recognition and / or manual methods, using strict medical definitions and standards, and the corresponding survival time calculated from each time point can be determined accordingly. Furthermore, the multiple reference patients who have reached the end of life can be labeled, and this label can be "reached the end of life" and the corresponding survival time (specific number of months of survival). Of course, after obtaining medical data of different patients at different time points, it is also possible to generate corresponding multidimensional medical data for each patient at different time points, and then select the corresponding multidimensional medical data and corresponding survival period of multiple reference patients who have reached the end of life at different time points.

[0050] Optionally, assuming that each piece of multidimensional medical data needs to include medical data of a preset dimension, when the medical data of multiple dimensions obtained for a certain reference patient at a certain time point (referred to as the first time point) lacks the first dimension of medical data relative to the preset dimension, then based on the first dimension of medical data for the reference patient at other time points, the first dimension of medical data at the first time point is determined (for example, interpolation, forward or backward filling strategies can be used) to ensure that the reference patient has the preset dimension of medical data at each time point.

[0051] Optionally, when acquiring medical data from different data sources, all potential data sources can be identified, including electronic health records (EHRs) from different hospitals, laboratory test reports, imaging data, gene sequencing data, and clinical trial databases. Data access schemes can be designed and implemented to ensure that privacy protection regulations such as the Health Insurance Portability and Accountability Act (HIPAA) are followed to securely collect data.

[0052] Furthermore, when constructing the final multidimensional medical data based on the acquired medical data of each patient across multiple dimensions, one can first obtain text data for different patients at different time points from different data sources. Natural language processing (NLP) techniques can then be used to extract key information from this text data. This extracted key information can then be used to generate structured multidimensional medical data. For example, data preprocessing can be performed on the multidimensional medical data for different patients at different time points obtained from multiple different data sources. This includes cleaning the raw data for each patient, removing irrelevant or erroneous information such as blank records, duplicate entries, and outliers, and converting unstructured data into structured data (corresponding to the corresponding dimensions). For instance, NLP techniques can be used to extract key dimension information from each piece of text data, and each key dimension includes timestamp information. After processing, as mentioned earlier, multiple dimensions of medical data can be determined for each patient at multiple time points, ensuring accurate correspondence between data for each dimension of the same patient at different time points, and addressing the issue of missing dimensions in the time-point series (e.g., using interpolation, forward or backward imputation strategies). This yields multidimensional medical data for each patient at each time point (with preset dimensions). Furthermore, data from different sources and modalities can be standardized, such as converting laboratory values ​​to uniform units to ensure consistency across data sources. Optional data normalization can also be implemented, such as encoding categorical variables and normalizing or standardizing continuous variables, to facilitate model processing. Additionally, as mentioned earlier, since the reference dataset requires multidimensional medical data from multiple selected patients who have reached the end of their lives, strict medical definitions and standards can be used to identify those who have reached the end of their lives. During screening, patient records must be carefully reviewed to confirm the accurate date and cause of the end-of-life event, ensuring the data quality used for survival analysis (this step can be performed manually or by a machine learning model). In addition, relevant personnel can conduct regular data quality checks, including audits of data integrity, consistency, and timeliness, to ensure the reliability of the reference dataset.

[0053] In other words, by constructing a comprehensive data collection network spanning different medical institutions and the field of oncology, multi-dimensional medical data of each patient is systematically aggregated. After rigorous quality control, this data is carefully organized into a continuous time-series format, ensuring the temporal continuity and integrity of the data, thus forming a high-quality database for the diagnosis and treatment of cancer patients.

[0054] For example, a sample of multidimensional medical data from one reference dataset could be: male, 69 years old, non-smoker, advanced lung cancer, condition stabilized after three cycles of chemotherapy, with a survival of 24 months from the end of the last chemotherapy session (May 2021). It should be noted that the reference dataset may also include other dimensions of disease-related information.

[0055] Optionally, in order to accurately locate similar patients with high clinical reference value to the current cancer patients from the massive data of the reference dataset, and thereby determine the similar dataset, an "Expert-Guided Hierarchical Hybrid Retrieval" algorithm is proposed in some exemplary embodiments of this application. This algorithm differs from traditional single-vector retrieval methods by combining explicit medical rules with implicit semantic features, and the steps are as follows.

[0056] 1. Construct a multimodal feature space and expert weight matrix.

[0057] First, the multidimensional medical data of cancer patients is mapped into a multimodal feature space, such as three sub-feature spaces:

[0058] A. Hard constraint characteristics ( This includes, for example, core attributes that medical experts define as “must match,” such as the primary tumor site, pathological type (adenocarcinoma, squamous cell carcinoma, etc.), and / or key gene mutations (EGFR, ALK, etc.).

[0059] B. Numerical clinical characteristics ( This includes patient-related characteristics expressed in explicit numerical values, such as age, body mass index (BMI), tumor diameter, and / or laboratory indicators (LDH, CEA, etc.), after normalization.

[0060] C. Features of unstructured text ( This includes textual descriptions related to the patient's condition, such as imaging descriptions and discharge summaries, and uses medical BERT (such as ClinicalBERT) to extract high-dimensional semantic vectors.

[0061] Furthermore, the system can assign weights to features corresponding to different dimensions. For example, by constructing a feature weight matrix set by experts or optimized through meta-learning, weights can be assigned to features corresponding to each dimension (e.g., primary tumor site, pathological type (adenocarcinoma, squamous cell carcinoma, etc.), key gene mutations, age, body mass index (BMI), tumor diameter, and / or laboratory indicators, etc.).

[0062] For example, in the context of advanced lung cancer, the weight of gene mutations. Weight far greater than age Optionally, as will be described later, weighting can be applied only to numerical clinical features across different dimensions.

[0063] 2. Coarse Filtering Stage

[0064] Based on the hard constraint features, a pool of candidate patients is selected from the reference dataset using an inverted index or Boolean logic, wherein the reference patients in the pool of candidate patients match the tumor patients on the hard constraint features.

[0065] For example, using inverted indexes or Boolean logic, a coarse-screening process is performed on the reference dataset based on the aforementioned hard constraint features to obtain coarsely selected candidate patients. For instance, only samples that are completely identical or highly matched to the target patients in terms of hard constraint features (e.g., tumor location, pathological type, and / or key driver genes) are retained to form a candidate pool. This step ensures that subsequent similar patients are biologically comparable.

[0066] 3. Fine-grained Re-ranking and dynamic K-value truncation

[0067] For each reference patient in the candidate patient pool, a similar dataset of the multidimensional medical data is determined based on the unstructured text features and numerical clinical features corresponding to the reference patient, and the numerical clinical features and unstructured text features corresponding to the tumor patient. For example, a weighted mixed distance with the tumor patient is calculated based on the semantic similarity between the unstructured text features corresponding to the reference patient and the unstructured text features corresponding to the tumor patient, and the weighted difference between the numerical clinical features corresponding to the reference patient and the numerical clinical features corresponding to the tumor patient; and based on a dynamic threshold strategy, reference multidimensional data corresponding to reference patients whose weighted mixed distance is less than a preset threshold are retained as the similar dataset.

[0068] Specifically, as an example, for each candidate patient Calculation with target patients Weighted mixing distance:

[0069]

[0070] in:

[0071] First item Cosine distance for text semantic similarity based on unstructured text features;

[0072] Second item Weighted Euclidean distance to measure differences in numerical features

[0073] α, β: balance coefficients.

[0074] Subsequently, this algorithm employs a dynamic thresholding strategy, for example, retaining only the relevant datasets of reference patients that satisfy D < δ as the final similarity dataset, where δ is a preset threshold. This ensures that the reference cases input into the large model have sufficiently high clinical relevance, avoiding noisy samples interfering with training and inference.

[0075] Alternatively, in other embodiments, the process can be performed simply based on feature comparison, i.e., a single-vector retrieval method. Key data from one or more selected dimensions (e.g., age, gender, lifestyle habits, tumor type and / or treatment efficacy and complications) can be extracted from the multidimensional medical data of the current cancer patient, forming, for example, a descriptive sentence such as "Lung cancer, stage T3N2M1, 68-year-old female, human epidermal growth factor receptor 2 (HER2) negative, poor response to chemotherapy." This text can then be processed using a pre-trained BERT model. The BERT model can transform the text into a high-dimensional vector form through contextual understanding and lexical relationship analysis. In practice, the embedding vector obtained from the initial encoding of the text corresponding to the '[CLS]' tag output by the BERT model is directly adopted as the patient's comprehensive feature vector.

[0076] Then, when selecting similar datasets, a comprehensive feature vector can be created for each reference data point. For example, a corresponding comprehensive feature vector can be created for each reference data point in the reference dataset based on one or more selected dimensions (e.g., age, gender, lifestyle habits, tumor type and / or treatment effectiveness and complications). Then, a similarity algorithm (e.g., cosine similarity algorithm) can be used to measure the similarity between the comprehensive feature vector of each reference data point and the comprehensive feature vector of the current cancer patient. Finally, based on the calculated similarity scores, the top K comprehensive feature vectors most similar to the comprehensive feature vector of the current cancer patient (corresponding to K similar patients) are selected. The complete multidimensional medical data corresponding to these K comprehensive feature vectors can be determined (the value of K can be set according to actual needs, usually a small value such as 5 to ensure high relevance), forming a closely matched dataset as a similar dataset for the multidimensional medical data of the current cancer patient.

[0077] Furthermore, within this context, since each patient may have corresponding multidimensional medical data at different points in time, a single set of similar data can be identified for each similar patient. For example, the multidimensional medical data with the highest similarity to the current cancer patient for each similar patient can be used as the similar data for that patient. Each set of similar data also includes the patient's survival time from the corresponding point in time.

[0078] Optionally, to reduce computational load, a pre-screening mechanism can be incorporated to quickly identify potential similar patient groups from a large patient database. This can be partially similar to the expert-guided hierarchical hybrid retrieval described earlier. For example, multidimensional medical data of other patients (referred to as similar patients) with medical data that are identical or similar to the current tumor patient in core attributes (corresponding to dimensions) can be filtered from the reference dataset to obtain a filtered dataset. Then, a predetermined number of multidimensional medical data similar to the current tumor patient's multidimensional medical data can be determined from the filtered dataset using the similarity algorithm described above, as a similar dataset. For example, filtering conditions can be set according to core attributes (corresponding to dimensions) to ensure that candidate similar patients are highly consistent with the target patient in core attributes. These core attributes may include, but are not limited to, the same tumor type, similar stages (clinical and pathological stages) (e.g., differing by one stage), and the same age and / or gender within a predetermined range. Thus, corresponding comprehensive feature vectors can be created for the multidimensional medical data of these candidate similar patients based on one or more selected dimensions, and the similarity calculation method described above can be used to determine the similar dataset of the current tumor patient's multidimensional medical data.

[0079] In step S230, a prediction result of the survival of the cancer patient is obtained using a finely tuned large language model based on the multidimensional medical data of the cancer patient and the similar dataset, wherein the prediction result indicates the survival of the cancer patient and the explanatory text of the information on which the prediction result is based.

[0080] Optionally, multidimensional medical data of the current cancer patient and similar datasets can be combined as input to a large language model. For example, the multidimensional medical data and the similar datasets are integrated to generate integrated data with a predetermined input format; the integrated data is encoded using an encoder associated with the fine-tuned large language model (included in the large language model or as a pre-encoder, such as a linear projection layer, etc.) to obtain the model input data; and the fine-tuned large language model is used to obtain a prediction result of the survival time of the cancer patient based on the model input data.

[0081] Optionally, the encoder can encode the multidimensional medical data and the similar datasets in the integrated data separately, so that the encoded information of the similar datasets serves as prompt information for the large language model during prediction.

[0082] In this context, similar datasets can be viewed as prompts on the multidimensional medical data of current cancer patients. By comparing historical cases with similar conditions to the current cancer patient, the model's generalization ability and predictive accuracy can be enhanced. In large language models, prompts are typically short text snippets used to guide the model in generating specific outputs. For a given large language model, well-designed prompts can yield excellent results.

[0083] In this way, the finely tuned large language model can predict the current patient's survival time from the current point in time by taking into account the multidimensional medical data of the current cancer patient and the similar multidimensional medical data and survival time of similar patients in similar datasets (the direct result may be the survival time from the date of diagnosis, but it can be further converted to be calculated from the current point in time). It can also include explanatory text, such as explaining the logic behind the prediction, such as which key characteristics of similar patients have the greatest impact on the prediction results, and how the patient's personal condition affects their prognosis, thus providing doctors with a scientific and easy-to-understand basis for decision-making.

[0084] The key to fine-tuning large language models lies in the construction of the training dataset. As mentioned earlier, the training dataset can be associated with the reference dataset, for example, it can be obtained based on the same data source, and it can be generated from the reference dataset.

[0085] For example, each piece of training data includes input sample data and output sample data. The input sample data includes multidimensional medical data of the corresponding patient who has reached the end of life at the corresponding time point, and the output sample data includes the survival label of the corresponding patient and reference explanatory text.

[0086] Furthermore, for each piece of training data, it is necessary to determine a similar multidimensional medical data training dataset from the training dataset. Each similar multidimensional medical data point and its corresponding survival time in this similar training dataset can serve as a prompt associated with the training data during the training process. The specific determination method can be similar to the method described above for determining similar datasets of multidimensional medical data for the current cancer patient. The current training data and the multidimensional medical data from its similar training datasets are combined and used as input to the large language model being fine-tuned, allowing the large language model to output prediction results.

[0087] Optionally, at least 1,500 patient-related data points for each tumor type are used to form the training data to cover sufficient case diversity, reduce the risk of overfitting, and ensure the predictive ability of the large language model for various tumor types.

[0088] Therefore, when training a large language model, the input sample data for each training dataset (multidimensional medical data and similar datasets) can be fed into the large language model, along with prediction task instructions. Based on the difference between the large language model's predicted output (including predicted survival time and predicted explanatory text) and the output sample data of the training dataset (actual survival time and reference explanatory text on which the prediction is based), the parameters of the large model can be fine-tuned. For example, iterative fine-tuning can be implemented, monitoring model performance through techniques such as cross-validation, and continuously adjusting the training strategy until the model achieves satisfactory prediction accuracy and explanatory power on the validation set.

[0089] To enable large language models to accurately predict survival and generate reasonable interpretations based on similar patient data, this application proposes a "Multi-task Retrieval-Augmented Instruction Tuning" method as a specific example. This method introduces special input construction and a hybrid loss function during the training phase.

[0090] 1. Prompt Engineering:

[0091] For each training sample, construct the input sequence (Context) with the following structure:

[0092] [System Instruction]: You are an oncology expert. Based on the clinical characteristics of the target patient and referring to the outcomes of similar historical cases, predict the survival time (in months) of the target patient and explain the basis for your prediction.

[0093] Reference Context: Retrieved Top-K similar patient information. Formatted as: "Similar Case 1: Features {...}, Survival: X months; Similar Case 2: Features {...}, Survival: Y months...".

[0094] [Target Patient]: Multidimensional medical data of the target patient.

[0095] [Reasoning Trigger]: First analyze the similarities and differences between the target patient and similar cases, then provide a prediction.

[0096] 2. Hybrid Loss Function Design:

[0097] Traditional language models only optimize the predicted probability of the next word (Cross-Entropy Loss), which is not accurate enough for numerical regression tasks (survival prediction). This application introduces a hybrid loss function L during fine-tuning. total =L CLM +λ1⋅L Reg +λ2⋅L Rank

[0098] L CLM (Causal Language Modeling Loss): Related to the fluency and logic of explanatory text, used to optimize the fluency and logic of the explanatory text generated by the model, ensuring that the model can generate explanations such as "because the patient has feature X, and is slightly better than similar case A..."

[0099] L Reg (Numerical Regression Loss): The numerical regression loss associated with the discrepancy between predicted and actual survival times is calculated by adding a linear regression head to the model's output layer or extracting the numerical portion from the output text. This forces the model to focus on the accuracy of time prediction.

[0100] L Rank(Ranking Consistency Loss): This loss constructs pairwise samples by associating them with the true survival ranking relationship between paired samples. If patient A's true survival is longer than patient B's, the model is forced to predict a higher value for A than for B. This helps the model learn the relative ranking relationship of disease severity and improves generalization ability.

[0101] 3. Training process:

[0102] Low-rank adaptive (LoRA) technique was employed to fine-tune the attention weight matrix of the large language model. During training, most parameters of the pre-trained model were frozen, and only the LoRA adapter parameters and regression head parameters were updated. This not only reduced GPU memory requirements but also preserved the general medical knowledge of the base model, preventing catastrophic amnesia on small-sample tumor data.

[0103] Through the above training, the model has learned an "analogical reasoning" ability: that is, instead of memorizing survival data, it has learned a reasoning pattern of "adjusting based on the benchmark of similar cases and combined with the individual differences of the target patient (such as complications and gene mutations)".

[0104] In some embodiments, the current cancer patient may also have corresponding multidimensional medical data from previous time points. Therefore, the current cancer patient's multidimensional medical data at the current time point, along with all previous multidimensional medical data, can be used as input to the large language model. In this case, when constructing training data, the corresponding time point for each patient and all previous multidimensional medical data can also be used as input to the large language model during the training / fine-tuning phase, enabling the model to perform analysis based on the patient's entire disease course.

[0105] Furthermore, as mentioned earlier, any data ultimately input into the large language model (e.g., input sample data for training, actual multidimensional medical data of current cancer patients, and actual integrated data of similar data from similar datasets) is generally a feature representation (e.g., embedding vector) obtained after encoding or other similar processing methods. Methods for encoding and processing the input text are widely used in this field, and therefore will not be described in detail here to avoid obscuring the focus of this application.

[0106] As can be seen, the proposed method for predicting the survival of cancer patients utilizes a customized large language model to analyze processed and integrated multidimensional medical data, predicting the survival time of cancer patients after a specified time point with accuracy down to the month and generating explanatory text. This explanatory text clearly points out the basis behind the prediction, enhancing its transparency and reliability. Furthermore, by constructing a reference dataset, multidimensional information from patients with different cancer types from various sources can be comprehensively analyzed. This allows for continuous training and fine-tuning of the large language model, enabling it to deeply understand the language of oncology and the complexity of disease progression, thus providing more accurate survival predictions. Additionally, when predicting survival, similar datasets matched with similar patients are used as prompts, providing even more precise survival predictions and explanations, thereby enhancing the model's generalization ability and predictive accuracy.

[0107] The following combination Figure 3 An example process is described for a scheme to predict the survival of cancer patients according to embodiments of this application.

[0108] Step 1 (S1): Obtain and process information about cancer patients.

[0109] For example, databases of cancer patient information, similar patient matching modules, and customized large-scale language models can be seamlessly integrated into a unified platform (e.g., Figure 1 On the server), and the platform can provide an interactive user interface (e.g., on...). Figure 1 (either at the server or via end-user devices), allowing doctors to easily input information about current cancer patients. For example, the information could be: "A 65-year-old male lung cancer patient, no smoking history, in the advanced stage, has received two cycles of chemotherapy. The most recent CT scan (time point: May 2023) shows that although the tumor has shrunk slightly, it is accompanied by mild pulmonary edema."

[0110] Step 2 (S2): Matching Similar Patients

[0111] For example, the platform can automatically trigger the similar patient matching module to identify the most relevant similar cases from the database (the specific process can be found in the previous description). In this example, the system matched cases A, B, and C, with the following details:

[0112] Similar patient case A: Male, 69 years old, non-smoker, advanced lung cancer, his condition stabilized after three cycles of chemotherapy, and his survival period was 24 months from the end of chemotherapy.

[0113] Similar patient case B: Female, 60 years old, with a light smoking history, advanced lung cancer, her condition stabilized after two rounds of chemotherapy, and her survival period was 18 months from the end of chemotherapy, but brain metastases were found in the 12th month.

[0114] Similar patient case C: Male, 68 years old, non-smoker, advanced lung cancer, responded well to chemotherapy, and survived for 30 months after the end of chemotherapy.

[0115] Step 3 (S3): Data Integration and Construction of Large Model Input

[0116] The platform combines information about current cancer patients with information about similar patient cases as input data for the large language model, as follows:

[0117] Patient Information: A 65-year-old male lung cancer patient with no smoking history, in an advanced stage, who has received two cycles of chemotherapy. The most recent CT scan (May 2023) showed slight tumor shrinkage, but was accompanied by mild pulmonary edema. This information was compiled and entered into the system.

[0118] Similar patient information:

[0119] A: Male, 69 years old, non-smoker, with advanced lung cancer. His condition stabilized after three cycles of chemotherapy. His survival period from the diagnosis of advanced lung cancer was 24 months.

[0120] B: Female, 60 years old, with a light smoking history, late-stage lung cancer. Her condition stabilized after two rounds of chemotherapy. Her survival period was 18 months from the diagnosis of late-stage lung cancer, but brain metastases were discovered in the 12th month.

[0121] C: Male, 68 years old, non-smoker, advanced lung cancer, good response to chemotherapy, survival period reached 30 months from the diagnosis of advanced lung cancer.

[0122] Step 4 (S4): Survival prediction and interpretation generation

[0123] Once the input is constructed, the platform or its associated encoding module can encode the input data and perform similar processing, and then invoke a customized large language model for predictive analysis. This model not only predicts the patient's survival of 22 months from the diagnosis of advanced lung cancer based on the input data, but also automatically generates explanatory text detailing the basis for the prediction, such as the patient's age, non-smoking history, positive response to chemotherapy, and existing complications (pulmonary edema). Simultaneously, the model references survival data from similar cases A and C, especially their good chemotherapy response in a non-smoking background, and synthesizes this information to provide the prediction result.

[0124] Predicted survival: The patient's predicted survival from the end of the second chemotherapy cycle is 22 months.

[0125] Explanatory Text: Based on the input patient information, the model analyzed several key factors: patient age, smoking history, treatment response, and complication status. In particular, it referenced several similar cases, such as cases A and C, both non-smokers who responded well to chemotherapy, with survival periods of 24 months and 30 months respectively from the end of chemotherapy. Considering your patient's tumor shrank after chemotherapy but has the potential complication of pulmonary edema, the model's comprehensive assessment predicts a survival of 22 months. This prediction integrates positive factors (such as chemotherapy efficacy) and potential risks (pulmonary edema), and specifically considers patient cases with similar conditions.

[0126] exist Figure 3 The training / fine-tuning process of the large language model can be referred to the previous description, and will not be repeated here.

[0127] According to another aspect of this application, a device for predicting the survival of cancer patients is also provided.

[0128] Figure 4 A structural block diagram of a device 400 for predicting the survival of cancer patients according to an embodiment of this application is shown. The device 400 may be... Figure 1 The server shown may include the server or be a part of it.

[0129] like Figure 4 As shown, the device 400 may include an acquisition module 410, a determination module 420, and a prediction module 430.

[0130] The acquisition module 410 can be used to acquire multidimensional medical data related to the condition of cancer patients, and the multidimensional medical data includes data in at least two dimensions.

[0131] The determination module 420 can be used to determine similar datasets of the multidimensional medical data from a reference dataset, wherein the reference dataset includes reference multidimensional medical data of each of a plurality of reference patients who have reached the end of life at different time points and their corresponding survival periods.

[0132] The prediction module 430 can be used to obtain a prediction result of the survival of the cancer patient based on the multidimensional medical data and the similar dataset using a finely tuned large language model, wherein the prediction result indicates the survival of the cancer patient and explanatory text of the information on which the prediction result is based.

[0133] Optionally, the device 400 may further include a training module 440 for training the large language model. For example, the training module 440 may be used to acquire a training dataset associated with the reference dataset, wherein each training data point includes input sample data and output sample data. The input sample data includes multidimensional medical data of a corresponding patient who has reached the end of life at a corresponding time point, and the output sample data includes the survival label of the corresponding patient and explanatory text. For each training data point, a similar training dataset is determined that is similar to the multidimensional medical data of the corresponding patient at the corresponding time point, wherein each similar multidimensional medical data point and its corresponding survival in the similar training dataset serve as a prompt associated with the training data during the training process. The input sample data of each training data point is then input into the large language model, and the large language model is fine-tuned based on the difference between the predicted output of the large language model and the corresponding output sample data.

[0134] Optionally, when determining similar datasets, the determining module 420 can be configured to: map the multidimensional medical data into a multimodal feature space, the multimodal feature space including hard constraint features, numerical clinical features, and unstructured text features corresponding to different dimensions; based on the hard constraint features, use an inverted index or Boolean logic to filter a pool of candidate patients from the reference dataset, wherein the reference patients in the candidate patient pool match the tumor patients on the hard constraint features; and for each reference patient in the candidate patient pool, determine similar datasets of the multidimensional medical data based on the unstructured text features and numerical clinical features corresponding to the reference patients and the numerical clinical features and unstructured text features corresponding to the tumor patients.

[0135] Alternatively, when determining similar datasets, the determining module 420 can be configured to: determine multidimensional medical data in the reference dataset that are identical or similar to the multidimensional medical data in terms of core attributes, thus obtaining a filtered dataset; generate corresponding reference feature vectors for each reference multidimensional medical data in the filtered dataset based on one or more selected dimensions; and use a similarity algorithm to determine a preset number of reference multidimensional medical data that are most similar to the multidimensional medical data of the tumor patient from the filtered dataset based on the feature vectors of the multidimensional medical data and each reference feature vector, thus obtaining the similar dataset of the multidimensional medical data. Alternatively, the determining module 420 may not perform pre-screening based on core attributes, but instead directly calculate the similarity between the multidimensional medical data and each reference multidimensional medical data in the reference dataset according to one or more selected dimensions to determine similar datasets.

[0136] Optionally, the reference dataset can also be obtained in the following ways (e.g., by...) Figure 1 The data processing and integration apparatus shown may be constructed as part of or independent of apparatus 400: acquiring multi-dimensional medical data for different patients at different time points from different data sources; filtering out multiple reference patients who have reached the end of life from the different patients, and determining the time point of the end of life for each reference patient; and for each of the multiple reference patients, generating corresponding multi-dimensional medical data at different time points from the multi-dimensional medical data belonging to the reference patient at different time points, and determining the corresponding survival period of the reference patient at the different time points based on the time point of the end of life of the reference patient.

[0137] Furthermore, the device 400 may also include a training module for: acquiring a training dataset associated with the reference dataset, wherein each training data includes input sample data and output sample data, the input sample data including multidimensional medical data of a corresponding patient who has reached the end of life at a corresponding time point, and the output sample data including the survival label of the corresponding patient and explanatory text; for each training data, determining a similar training dataset similar to the multidimensional medical data of the corresponding patient at the corresponding time point corresponding to the training data, each similar multidimensional medical data and corresponding survival in the similar training dataset serving as prompt information associated with the training data during the training process; and inputting the input sample data of each training data into the large language model, and fine-tuning the large language model based on the difference between the predicted output of the large language model and the corresponding output sample data. In some embodiments, the input sample data is an input sequence with a preset structure, the input sequence including one or more of system instructions, reference context, corresponding patient, and explanatory triggers. Furthermore, the large language model is fine-tuned, including by using a hybrid loss function, wherein the hybrid loss function is constructed based on a causal language modeling loss associated with the fluency and logicality of the explanatory text, a numerical regression loss associated with differences in survival prediction, and a ranking consistency loss associated with the true survival ranking relationship between paired samples.

[0138] Figure 4 More details about the various modules in device 400 can be found in the previous text. Figure 2 The description will not be repeated here.

[0139] In addition, although Figure 4The modules and sub-modules described above are illustrated by way of example. However, it should be understood that the device 400 may be divided into more or fewer modules depending on different functions, or each module may be divided into further more or fewer sub-modules. For example, the first identification module 410 may be the same module as the second identification module 420, and / or the evaluation module 430 may also include a determination sub-module and a division sub-module, etc. In some example embodiments, the modules or their sub-modules may be implemented using electronic hardware (e.g., general-purpose processors, DSPs, ASICs, FPGAs or other programmable logic devices, discrete gate or transistor logic, discrete hardware components, etc.), computer software (e.g., which may be stored in random access memory (RAM), flash memory, read-only memory (ROM), erasable programmable ROM (EPROM), etc.), or a combination of both.

[0140] refer to Figure 4 The device described for predicting the survival of cancer patients utilizes a customized large language model to analyze processed and integrated multidimensional medical data. It predicts the survival time of cancer patients after a specified time point, accurate to the month, and generates explanatory text that clearly explains the basis of the prediction, enhancing its transparency and reliability. Furthermore, by constructing a reference dataset, it can comprehensively analyze multi-dimensional information from patients with different cancer types from various sources, continuously training and fine-tuning the large language model. This allows the model to deeply understand the language of oncology and the complexity of disease progression, thus providing more accurate survival predictions. Additionally, when predicting survival, it uses similar datasets matched with similar patients as prompts, enabling more precise survival predictions and explanations, thereby enhancing the model's generalization ability and predictive accuracy.

[0141] Figure 5 A structural block diagram of a computing device according to an embodiment of this application is shown.

[0142] See Figure 5 The computing device 500 may include one or more processors 501 and one or more memories 502 connected to the processors 501. Both the processors 501 and the memories 502 can be connected via a bus 503. The computing device 500 can be any type of portable device (such as a smart camera, smartphone, tablet, etc.) or any type of stationary device (such as a desktop computer, server, etc.). For example, the computing device may be... Figure 1 The server shown.

[0143] Processor 501 can perform various actions and processes according to the computer program and computer instruction set stored in memory 502. Specifically, processor 501 can be an integrated circuit chip with signal processing capabilities. The processor can be a general-purpose processor, digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), or other programmable logic device, discrete gate or transistor logic device, or discrete hardware component. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this invention. The general-purpose processor can be a microprocessor or any conventional processor, and can be based on x85 architecture or ARM architecture.

[0144] Memory 502 stores computer-executable instructions that, when executed by processor 501, implement the method described above for predicting the survival of cancer patients. Memory 502 may be volatile memory or non-volatile memory, or may include both. Non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory may be random access memory (RAM) used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous linked dynamic random access memory (SLDRAM), and direct memory bus random access memory (DR RAM). It should be noted that the memory of the methods described herein is intended to include, but is not limited to, these and any other suitable categories of memory.

[0145] Furthermore, the method for predicting the survival of cancer patients according to the present invention can be recorded in a non-transitory computer-readable recording medium. Specifically, according to the present invention, a non-transitory computer-readable recording medium storing computer-executable instructions or a computer program can be provided, which, when executed by a processor, causes the processor to perform the method for predicting the survival of cancer patients as described above.

[0146] It should be noted that the flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, program segment, or portion of code containing at least one executable instruction for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0147] In general, various exemplary embodiments of the present invention can be implemented in hardware or dedicated circuitry, software, firmware, logic, or any combination thereof. Some aspects can be implemented in hardware, while others can be implemented in firmware or software that can be executed by a controller, microprocessor, or other computing device. When aspects of embodiments of the present invention are illustrated or described as block diagrams, flowcharts, or represented using certain other images, it will be understood that the blocks, apparatuses, systems, techniques, or methods described herein can be implemented as non-limiting examples in hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or controllers or other computing devices, or certain combinations thereof.

[0148] Unless otherwise defined, all terms used herein (including technical and scientific terms) shall have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. It should also be understood that terms such as those defined in a common dictionary shall be interpreted as having a meaning consistent with their meaning in the context of the relevant art, and not as having an idealized or highly formalized meaning, unless expressly defined herein.

[0149] The foregoing description is intended to illustrate the invention and should not be construed as limiting it. Although several exemplary embodiments of the invention have been described, those skilled in the art will readily understand that many modifications can be made to the exemplary embodiments without departing from the novel teachings and advantages of the invention. Therefore, all such modifications are intended to be included within the scope of the invention as defined in the claims. It should be understood that the foregoing description is intended to illustrate the invention and should not be construed as limiting it to the specific embodiments disclosed, and modifications to the disclosed embodiments and other embodiments are intended to be included within the scope of the appended claims. The invention is defined by the claims and their equivalents.

Claims

1. A method for predicting the survival of cancer patients, comprising: Acquire multidimensional medical data related to the condition of cancer patients, wherein the multidimensional medical data includes data from at least two dimensions. A similar dataset of the multidimensional medical data is determined from a reference dataset, wherein the reference dataset includes reference multidimensional medical data and corresponding survival times for each of multiple reference patients who have reached the end of their lives at different time points; and Using a finely tuned large language model based on the multidimensional medical data and the similar dataset, a prediction result for the survival of the cancer patient is obtained, wherein the prediction result indicates the survival of the cancer patient and an explanatory text of the information on which the prediction result is based.

2. The method according to claim 1, wherein, Identifying similar datasets of the multidimensional medical data from the reference dataset includes: The multidimensional medical data is mapped into a multimodal feature space, which includes hard constraint features, numerical clinical features, and unstructured text features, and corresponds to different dimensions. Based on the hard constraint features, a candidate patient pool is selected from the reference dataset using an inverted index or Boolean logic, wherein the reference patients in the candidate patient pool match the tumor patients on the hard constraint features; and For each reference patient in the candidate patient pool, a similar dataset of the multidimensional medical data is determined based on the unstructured text features and numerical clinical features corresponding to the reference patient and the numerical clinical features and unstructured text features corresponding to the tumor patient.

3. The method according to claim 2, wherein, Each numerical clinical feature corresponding to the different dimensions was assigned a corresponding weight. The similar datasets for determining the multidimensional medical data include: Based on the semantic similarity between the unstructured text features corresponding to the reference patient and the unstructured text features corresponding to the tumor patient, and the weighted difference between the numerical clinical features corresponding to the reference patient and the numerical clinical features corresponding to the tumor patient, a weighted mixed distance with the tumor patient is calculated; and Based on a dynamic threshold strategy, reference multidimensional data corresponding to reference patients whose weighted mixing distance is less than a preset threshold are retained as the similar dataset.

4. The method according to claim 1, wherein, Identifying similar datasets of the multidimensional medical data from the reference dataset includes: The multidimensional medical data that are the same as or similar to the multidimensional medical data in terms of core attributes are determined from the reference dataset to obtain the filtered dataset; Generate a corresponding reference feature vector for each piece of reference multidimensional medical data in the filtered dataset based on one or more selected dimensions; and Using a similarity algorithm based on the feature vectors of the multidimensional medical data and each reference feature vector, a preset number of reference multidimensional medical data that are most similar to the multidimensional medical data of the tumor patient are determined from the selected dataset, thus obtaining the similar dataset of the multidimensional medical data.

5. The method according to claim 1, wherein, The large language model is trained using the following method: Obtain a training dataset associated with the reference dataset, wherein each training dataset includes input sample data and output sample data. The input sample data includes multidimensional medical data of the corresponding patient who has reached the end of life at the corresponding time point, and the output sample data includes the survival label of the corresponding patient and reference explanatory text. For each piece of training data, a similar training dataset is determined that is similar to the multidimensional medical data in the training data. Each piece of similar multidimensional medical data and its corresponding survival time in the similar training dataset are used as prompt information associated with the training data during the training process. as well as The input sample data of each training data is input into the large language model, and the large language model is fine-tuned based on the difference between the predicted output of the large language model and the corresponding output sample data.

6. The method according to claim 5, wherein, The input sample data is an input sequence with a preset structure, which includes one or more of the following: system instructions, reference context, corresponding patient, and interpretive triggers.

7. The method according to claim 5, wherein, Fine-tuning the large language model includes using a mixture loss function to fine-tune the large language model. The hybrid loss function is constructed based on causal language modeling loss associated with the fluency and logicality of the explanatory text, numerical regression loss associated with differences in survival prediction, and ranking consistency loss associated with the true survival ranking relationship between paired samples.

8. The method according to claim 1, wherein, The reference dataset is constructed in the following way: Obtain multi-dimensional medical data from different data sources for different patients at different time points; From the different patients, select the multiple reference patients who have reached the end of life, and determine the time point of the end of life for each reference patient; as well as For each of the multiple reference patients, multidimensional medical data belonging to the reference patient at different time points is generated to form corresponding multidimensional medical data at different time points, and the corresponding survival period of the reference patient at different time points is determined based on the time point of the reference patient's end of life.

9. The method according to claim 8, wherein, Each piece of medical data acquired includes a timestamp, and the multidimensional medical data includes medical data with preset dimensions. Specifically, when the medical data of a reference patient at a first time point is missing the first dimension of medical data compared to the preset dimension, the medical data of the first dimension at the first time point is determined based on the first dimension of medical data of the reference patient at other time points.

10. The method according to claim 1, wherein, Using a finely tuned large language model based on the multidimensional medical data and the similar dataset, predictions for the survival of the cancer patients are obtained, including: The multidimensional medical data and the similar data are integrated to generate integrated data with a predetermined input format; The integrated data is encoded using an encoder associated with the fine-tuned large language model to obtain model input data; and Using the finely tuned large language model based on the model input data, the predicted survival time of the cancer patient is obtained.

11. The method according to claim 10, wherein, The encoder separately encodes the multidimensional medical data and the similar datasets in the integrated data, so that the encoded information of the similar datasets serves as prompt information for the large language model during prediction.

12. The method according to claim 8, wherein, Obtain multi-dimensional medical data from different data sources for different patients at different time points, including: Obtain text data from different data sources for different patients at different time points; and Natural language processing techniques are used to extract key information from the acquired text data, and structured, multi-dimensional medical data is generated based on the extracted key information.

13. The method according to claim 8, wherein, The different data sources include at least two of the following: electronic health records from different hospitals, laboratory test reports, imaging data, gene sequencing data, and clinical trial databases.

14. A device for predicting the survival of cancer patients, comprising: The acquisition module is used to acquire multidimensional medical data related to the condition of cancer patients, wherein the multidimensional medical data includes data from at least two dimensions. A determination module is used to determine similar datasets of the multidimensional medical data from a reference dataset, wherein the reference dataset includes reference multidimensional medical data and corresponding survival times for each of multiple reference patients who have reached the end of their lives at different time points; and The prediction module is used to obtain a prediction result of the survival time of the cancer patient based on the multidimensional medical data and the similar dataset using a fine-tuned large language model, wherein the prediction result indicates the survival time of the cancer patient and explanatory text of the information on which the prediction result is based.

15. The apparatus according to claim 14, wherein, When determining similar datasets of the multidimensional medical data from the reference dataset, the determination module is configured to: The multidimensional medical data is mapped into a multimodal feature space, which includes hard constraint features, numerical clinical features, and unstructured text features, and corresponds to different dimensions. Based on the hard constraint features, a candidate patient pool is selected from the reference dataset using an inverted index or Boolean logic, wherein the reference patients in the candidate patient pool match the tumor patients in terms of the hard constraint features. For each reference patient in the candidate patient pool, a similar dataset of the multidimensional medical data is determined based on the unstructured text features and numerical clinical features corresponding to the reference patient and the numerical clinical features and unstructured text features corresponding to the tumor patient.

16. The apparatus of claim 14, further comprising a training module for: Obtain a training dataset associated with the reference dataset, wherein each training dataset includes input sample data and output sample data. The input sample data includes multidimensional medical data of the corresponding patient who has reached the end of life at the corresponding time point, and the output sample data includes the survival label of the corresponding patient and explanatory text. For each piece of training data, a similar training dataset is determined that is similar to the multidimensional medical data of the corresponding patient at the corresponding time point. Each piece of similar multidimensional medical data and the corresponding survival time in the similar training dataset are used as prompt information associated with the training data during the training process. as well as The input sample data of each training data is input into the large language model, and the large language model is fine-tuned based on the difference between the predicted output of the large language model and the corresponding output sample data.

17. The apparatus according to claim 16, wherein, The input sample data is an input sequence with a preset structure, which includes one or more of the following: system instructions, reference context, corresponding patient, and interpretive triggers.

18. The apparatus according to claim 16, wherein, Fine-tuning the large language model includes using a mixture loss function to fine-tune the large language model. The hybrid loss function is constructed based on causal language modeling loss associated with the fluency and logicality of the explanatory text, numerical regression loss associated with differences in survival prediction, and ranking consistency loss associated with the true survival ranking relationship between paired samples.

19. A computing device, comprising: One or more processors; as well as One or more memories storing a computer program or instruction set that, when executed by the one or more processors, causes the one or more processors to perform the method as described in any one of claims 1-13.

20. A computer-readable storage medium storing a computer program or instruction set thereon, which, when executed by the one or more processors, causes the one or more processors to perform the method as described in any one of claims 1-13.