A tumor recurrence risk prediction method and system based on electronic medical record data

By constructing a tumor recurrence risk level mapping model based on graph neural networks and integrating electronic medical record data, individualized assessment and dynamic intervention of tumor recurrence risk are achieved, which solves the problems of data integration and intervention fusion in existing technologies and improves the accuracy of tumor recurrence risk prediction and clinical application efficiency.

CN120544908BActive Publication Date: 2025-10-17CHENGDU MILITARY GENERAL HOSPITAL OF PLA
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511030178.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-25
Publication Date
2025-10-17
Estimated Expiration
2045-07-25

AI Technical Summary

Technical Problem

Existing tumor recurrence risk prediction cannot effectively integrate structured, semi-structured, and unstructured data in electronic medical records, resulting in insufficient cross-modal association analysis and failure to effectively integrate prediction results with clinical intervention operations, limiting it to the retrospective research stage.

Method used

By collecting and preprocessing electronic medical record data of patients of multiple age groups, classifying and mapping them, a tumor recurrence risk level mapping model based on graph neural network is constructed, and dynamic intervention is carried out in combination with tumor marker parameters to achieve personalized risk assessment and treatment intervention.

Benefits of technology

It achieves accurate risk assessment of multi-dimensional features, improves the accuracy of identifying high-risk patients, reduces the risk of over-medicalization, improves the efficiency and accuracy of clinical intervention, and forms closed-loop management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120544908B_ABST
    Figure CN120544908B_ABST
Patent Text Reader

Abstract

The application discloses a tumor recurrence risk prediction method and system based on electronic medical record data, relates to the field of electronic medical record data processing and analysis, and can realize individualized recurrence risk assessment based on multidimensional characteristics, significantly reduces subjective judgment errors, and establishes a recurrence risk mapping model through structured processing of historical electronic medical record data; a risk grade stratification mechanism can automatically distinguish patients who need emergency intervention from patients who need routine monitoring, avoids excessive medical treatment or delayed treatment, and is especially suitable for chronic diseases such as tumors that need long-term management; through periodic marker measurement and risk grade feedback, a closed-loop management of 'assessment-intervention-reassessment' is formed, meeting the needs of continuous optimization of clinical diagnosis and treatment; in addition, through preprocessing, multidimensional characteristics after preprocessing can be directly called in the subsequent process, and repeated data cleaning work is avoided.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the field of electronic medical record data processing and analysis, in particular, relates to a tumor recurrence risk prediction method and system based on electronic medical record data. BACKGROUND

[0002] The existing tumor recurrence risk prediction cannot effectively integrate the structured test indicators, semi-structured text descriptions (such as medical history records) and unstructured image data in the electronic medical record, and then construct a unified feature space to realize cross-modal correlation analysis; in addition, the existing prediction model does not fuse the prediction results with the clinical actual intervention operation, resulting in that the model only stays in the retrospective research stage. SUMMARY

[0003] In view of the problems in the related art, the application provides a tumor recurrence risk prediction method and system based on electronic medical record data to overcome the above technical problems existing in the prior art.

[0004] To solve the above technical problems, the application is realized by the following technical scheme:

[0005] The application is a tumor recurrence risk prediction method based on electronic medical record data, comprising the following steps:

[0006] S1, collecting historical electronic medical record data of a plurality of patients corresponding to different age groups and pre-processing the data;

[0007] S2, classifying the pre-processed data in S1 and the tumor recurrence data of the corresponding patients according to age groups and tumor types and setting corresponding recurrence risk grades to obtain a historical recurrence risk grade classification data set and a historical pre-processed electronic medical record classification data set;

[0008] S3, constructing a tumor recurrence risk grade mapping model using the historical recurrence risk grade classification data set and the historical tumor recurrence classification data set;

[0009] S4, inputting the pre-processed electronic medical record data of the current patient to be predicted into the corresponding tumor recurrence risk grade mapping model for mapping to obtain current recurrence risk grade data;

[0010] S5, if the current recurrence risk grade data is 0, tumor marker parameter measurement is performed, if the measurement result corresponds to recurrence risk grade data that is always 0, no intervention is needed; otherwise, the corresponding intervention process is selected; if the current recurrence risk grade data is not 0, the corresponding intervention process is directly selected;

[0011] S6, treating and intervening the current patient to be predicted according to the intervention process selected in S5 until the recurrence risk grade data is 0.

[0012] Preferably, the S1 comprises the following steps:

[0013] S11, set a plurality of age stages of tumor treatment population to obtain a treatment population age stage set; according to the treatment population age stage set, collect historical electronic medical record data of a plurality of patients corresponding to each age stage and tumor types to obtain a historical electronic medical record data set and a historical tumor type data set;

[0014] S12, pre-process the historical electronic medical record data set to obtain a historical processed electronic medical record data set;

[0015] By integrating heterogeneous electronic medical record data by age stage, the analysis bias caused by age characteristics in traditional research is solved, and the risk prediction model can more accurately capture the tumor evolution law at different life cycle stages.

[0016] Preferably, the S12 comprises the following steps:

[0017] S121, data classification is performed on the historical electronic medical record data set to obtain a historical electronic medical record classification data set; the classification basis of the data classification includes classification according to structured, semi-structured and unstructured data;

[0018] S122, pre-process the data in each classification in the historical electronic medical record classification data set; after pre-processing, a historical processed electronic medical record data set is obtained;

[0019] The term mapping and unit unification of structured and semi-structured data eliminate the data heterogeneity among medical institutions, so that electronic medical records of different sources have direct comparability; by extracting clinical entities and association relations, free text is converted into an analyzable semantic network, supporting the construction of a subsequent tumor recurrence risk prediction model.

[0020] Preferably, the S2 comprises the following steps:

[0021] S21, cooperate with the historical tumor type data set to collect whether each patient corresponding to the historical electronic medical record data set has tumor recurrence and the recurrence time after obtaining the historical electronic medical record data of each patient to obtain a historical tumor recurrence data set; S22, according to the historical tumor type data set and the treatment population age stage set, classify the tumor recurrence data and electronic medical record data corresponding to different types of tumors and different age stages of patients in the historical tumor recurrence data set and the historical processed electronic medical record data set to obtain a historical tumor recurrence classification data set and a historical processed electronic medical record classification data set;

[0022] S23, according to the historical tumor recurrence classification data set, the recurrence risk level corresponding to each patient in S11 is set to obtain a historical recurrence risk level classification data set;

[0023] The data set classified based on the age range and the tumor type provides targeted data for subsequent construction of a specific age range and tumor type, which improves the accuracy compared to a general model.

[0024] Preferably, the S23 comprises the following steps:

[0025] S231, a plurality of recurrence time interval intervals corresponding to each type of tumor and a plurality of risk levels corresponding to each type of tumor are set to obtain a tumor recurrence time interval interval set and a tumor recurrence risk level set;

[0026] S232, if the recurrence flag data in the historical tumor recurrence classification data set is no, the corresponding recurrence risk is set to 0; otherwise, in cooperation with the tumor recurrence time interval interval set and the tumor recurrence risk level set, the recurrence time data in the historical tumor recurrence classification data set is classified and the corresponding tumor recurrence risk level data is obtained to obtain a historical recurrence risk level classification data set;

[0027] Based on the mapping relationship between the time interval interval and the risk level, the discrete recurrence data is converted into a quantifiable and comparable risk level, which supports the clinical rapid identification of high-risk patients who need priority intervention; the classification processing of non-recurrence patients and recurrence patients avoids data noise interference and improves the model training efficiency.

[0028] Preferably, the S3 comprises the following steps:

[0029] S31, in cooperation with the historical recurrence risk level classification data set and the historical processed electronic medical record classification data set, a mapping model between the tumor recurrence risk level data and the electronic medical record data corresponding to different types of tumors and different age groups of patients is respectively constructed to obtain a tumor recurrence risk level mapping model set;

[0030] Based on the difference modeling of tumor types and age stratification, the problem of insufficient prediction accuracy of traditional single model for different biological characteristics of breast cancer, prostate cancer and other tumors is effectively solved, and the identification accuracy of high-risk patients is improved. Secondly, through the dynamic mapping of electronic medical record data and recurrence risk level, multiple key indicators such as pathological grading, molecular typing and treatment response can be integrated to provide real-time risk warning for clinicians, which improves the efficiency compared to traditional manual assessment.

[0031] Preferably, the mapping model in the tumor recurrence risk level mapping model set in S31 adopts a graph neural network model;

[0032] Through the relationship modeling of nodes and edges, the structured data, unstructured text and image features in the electronic medical record can be processed simultaneously, the automatic extraction and interaction of cross-modal features are realized, the patient age, tumor type and other attributes are supported as node features, the treatment records and complications are supported as edge relationships, the dynamic heterogeneous graph is constructed, the time sequence dependence of the patient's historical medical records and the topological structure features of tumor evolution are captured through the graph convolution layer or the graph attention network, the discrete electronic medical record codes are converted into continuous vectors through the graph embedding technology, the causal logic between clinical events is retained, and the recurrence risk propagation law of similar patient groups is simulated through the message passing mechanism, and the prediction robustness of small sample high-risk medical records is improved.

[0033] Preferably, the S4 comprises the following steps:

[0034] S41, cooperate with the tumor recurrence risk level set, the treatment population age set and the historical tumor type data set, respectively set the treatment intervention process for patients of different age groups of different types of tumors under different recurrence risk levels, and obtain a tumor treatment intervention process set;

[0035] S42, obtain the electronic medical record data of the current patient to be predicted, denoted as current electronic medical record data; the current electronic medical record data is preprocessed in the manner of S12 to obtain current processed electronic medical record data; the current processed electronic medical record data is input into the tumor recurrence risk level mapping model corresponding to the age stage of the current patient to be predicted and the tumor type in the tumor recurrence risk level mapping model set to obtain current recurrence risk level data;

[0036] Through the graph neural network modeling of multi-dimensional data, the dynamic quantitative evaluation of recurrence risk is realized, and the static limitations of the traditional TNM staging system are overcome; a three-dimensional intervention prediction matching mechanism of risk level-tumor type-age is established, and the tumor recurrence risk level is mapped out so as to further match the corresponding intervention process in the subsequent.

[0037] Preferably, the S5 comprises the following steps:

[0038] S51, set the mapping relationship between the current prediction result observation period, the current tumor marker parameter and the parameter value interval and the tumor recurrence risk level data, and obtain a current parameter interval level mapping relationship set; when the current recurrence risk level data is 0, enter S52; otherwise, enter S53;

[0039] S52, measure each parameter of the tumor marker of the current patient to be predicted according to the current prediction result observation period and the current tumor marker parameter type; cooperate with the current parameter interval level mapping relationship set, if the tumor recurrence risk level data corresponding to the parameter of the tumor marker of the current patient to be predicted is always 0, the prediction is ended, and no intervention is needed;

[0040] Otherwise, add the maximum tumor recurrence risk level data corresponding to the parameter of the tumor marker of the current patient to be predicted and the current processed electronic medical record data into the historical recurrence risk level classification data set and the historical processed electronic medical record classification data set respectively, and repeat S31;

[0041] Then, according to the maximum tumor recurrence risk level data corresponding to the parameter of the tumor marker of the current patient to be predicted, the age stage of the current patient to be predicted and the tumor type, select the corresponding tumor treatment intervention process from the tumor treatment intervention process set to obtain a first current tumor treatment intervention process, and enter S61;

[0042] S53, according to the current recurrence risk level data, the age stage of the current patient to be predicted and the tumor type, select the corresponding tumor treatment intervention process from the tumor treatment intervention process set to obtain a second current tumor treatment intervention process, and enter S62;

[0043] By establishing the quantitative mapping relationship between the tumor marker parameter interval and the recurrence risk level, the transition from static single detection to dynamic continuous evaluation is realized; when the marker parameter breaks through the threshold, the historical data is automatically traced back, combined with multi-dimensional data such as age and tumor type to optimize risk stratification (re-optimize data, re-train and test the model), avoid the lag of traditional TNM staging system; according to the risk level, the differentiated treatment path is automatically matched to ensure that low-risk patients avoid over-treatment and high-risk patients start timely precise intervention.

[0044] Preferably, the S6 comprises the following steps:

[0045] S61, set a first current intervention deadline; use the first current tumor treatment intervention process to intervene in the tumor treatment of the current patient to be predicted within the first current intervention deadline; after the intervention is completed, collect the tumor marker parameters of the current patient to be predicted and obtain the corresponding tumor recurrence risk level data to obtain the first current intervention after recurrence risk level data;

[0046] When the first current intervention after recurrence risk level data is greater than 0, according to the first current intervention after recurrence risk level data and the tumor treatment intervention process set, the corresponding tumor treatment intervention process is obtained and the above treatment intervention process in S61 is repeated until the first current intervention after recurrence risk level data is equal to 0.

[0047] S62, set a second current intervention deadline; adopt a second current tumor treatment intervention process to perform tumor treatment intervention on the current to-be-predicted patient within the second current intervention deadline; after the intervention is completed, collect the tumor marker parameters of the current to-be-predicted patient and obtain corresponding tumor recurrence risk level data to obtain second current post-intervention recurrence risk level data;

[0048] When the second current post-intervention recurrence risk level data is greater than 0, a corresponding tumor treatment intervention process is obtained according to the second current post-intervention recurrence risk level data and a tumor treatment intervention process set, and the above-mentioned treatment intervention process in S62 is repeated until the second current post-intervention recurrence risk level data is equal to 0;

[0049] According to the initial risk level, the intervention process is automatically matched, over-treatment of low-risk patients is avoided, and at the same time, timely intensive intervention of high-risk patients is ensured; through real-time feedback of the recurrence risk level data, dynamic adjustment of the treatment scheme is realized until the biological indicators are completely relieved; when the marker parameter breaks through the safety threshold, the process iteration is triggered immediately, the tumor marker parameter is used as the core evaluation index of the intervention effect, and the simple image evaluation is replaced, and the microscopic lesion change is captured earlier.

[0050] A tumor recurrence risk prediction system based on electronic medical record data, comprising a historical electronic medical record data acquisition and processing module, a historical patient tumor data classification module, a historical recurrence risk level setting module, a recurrence risk level mapping model construction module, a current patient recurrence risk level mapping module, a current patient treatment intervention process determination module and a current patient treatment intervention module.

[0051] The present application has the following beneficial effects:

[0052] 1. In the present application, the historical electronic medical record data is processed by structuring and the recurrence risk mapping model is established, individualized recurrence risk assessment can be realized based on multi-dimensional characteristics, subjective judgment errors are significantly reduced, the risk level stratification mechanism can automatically distinguish patients who need emergency intervention and routine monitoring, over-treatment or delayed treatment is avoided, and the present application is especially suitable for chronic diseases such as tumors which need long-term management; through periodic marker measurement and risk level feedback, a closed-loop management of "evaluation-intervention-re-evaluation" is formed, and the needs of continuous optimization of clinical diagnosis and treatment are met.

[0053] 2. In the present application, by integrating heterogeneous electronic medical record data by age, the analysis bias caused by the mixing of age characteristics in traditional research is solved, so that the risk prediction model can more accurately capture the tumor evolution law at different life cycle stages; the preprocessing procedure can eliminate the data heterogeneity among medical institutions and improve the generalization ability of subsequent modeling; in addition, through preprocessing, the preprocessed multi-dimensional features can be directly called for subsequent work, avoiding repeated data cleaning work.

[0054] 3. In the present application, the differential modeling based on tumor type and age stratification effectively solves the problem of insufficient prediction accuracy of traditional single model for different biological characteristics of breast cancer, prostate cancer and other tumors, and improves the identification accuracy of high-risk patients; secondly, through the dynamic mapping of electronic medical record data and recurrence risk level, multiple key indicators such as pathological grading, molecular typing and treatment response can be integrated to provide real-time risk warning for clinicians, which improves the efficiency compared with traditional manual assessment; finally, the ensemble learning architecture of the model set can adaptively optimize the prediction logic, keeping high sensitivity while controlling the false positive rate below a reasonable range, significantly reducing the risk of over-treatment.

[0055] Of course, implementing any product of the present application does not necessarily require all the advantages described above. BRIEF DESCRIPTION OF DRAWINGS

[0056] In order to more clearly illustrate the technical solutions of the embodiments of the application, the following will briefly introduce the drawings needed to be used in the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the application, and for those skilled in the art, other drawings can also be obtained without creative labor based on these drawings.

[0057] Figure 1 The overall flowchart of the present application is shown in the flowchart of the present application.

[0058] Figure 2 The flowchart of the present application is shown in the flowchart of the present application.

[0059] Figure 3 The flowchart of the present application is shown in the flowchart of the present application.

[0060] Figure 4 The flowchart of the present application is shown in the flowchart of the present application.

[0061] Figure 5 The module diagram of the present application is shown in the module diagram of the present application. DETAILED DESCRIPTION

[0062] The technical solutions in the embodiments of the application will be described clearly and completely below with reference to the drawings in the embodiments of the application. Obviously, the described embodiments are only part of the embodiments of the application, rather than all the embodiments of the application. Based on the embodiments of the application, all other embodiments obtained by a person of ordinary skill in the art without creative effort belong to the protection scope of the application.

[0063] Embodiment one: refer to Figures 1-4 The embodiment is a tumor recurrence risk prediction method based on electronic medical record data, comprising the following steps.

[0064] S1, collecting historical electronic medical record data of a plurality of patients corresponding to a plurality of age groups and performing preprocessing;

[0065] The S1 comprises the following steps.

[0066] S11, setting a plurality of age groups of tumor treatment populations to obtain a treatment population age group set; according to the treatment population age group set, collecting historical electronic medical record data of a plurality of patients corresponding to each age group and tumor types to obtain a historical electronic medical record data set and a historical tumor type data set;

[0067] S12, preprocessing the historical electronic medical record data set to obtain a historical preprocessed electronic medical record data set;

[0068] The S12 comprises the following steps.

[0069] S121, performing data classification on the historical electronic medical record data set to obtain a historical electronic medical record classification data set; the classification basis of the data classification comprises classifying according to structured, semi-structured and unstructured data;

[0070] S122, preprocessing the data in each classification in the historical electronic medical record classification data set; after preprocessing is completed, a historical preprocessed electronic medical record data set is obtained;

[0071] Exemplarily, the specific process of the preprocessing is as follows.

[0072] For structured and semi-structured data: fixed template fields are matched through regular expression (such as “pathological diagnosis: {conclusion}”), diagnosis conclusions and classification information are extracted, JSON / XML format inspection reports (such as endoscopic description) are converted into two-dimensional table structure, descriptive terms are mapped to standard codes (such as “CA” is unified as “cancer”, “Tumor” is mapped to ICD-O code), non-standard units are converted (such as “5 thousand / mm³”→“5×10 9 / L”) and multiple inspection records of the same patient are associated to construct a time sequence chain (such as comparison of postoperative pathology and review report).

[0073] It can be handled by developing special-purpose parsers (e.g., breast cancer pathology templates ≠ lung cancer templates);

[0074] For unstructured data (free text, images, audio, etc.):

[0075] (1) Text data (medical records, chief complaints, etc.)

[0076] Extract clinical entities (e.g., symptoms "persistent chest pain," drugs "aspirin") and establish entity associations (e.g., "diabetes" -> "insulin therapy") through models like BERT; eliminate meaningless symbols (e.g., formatting errors), and merge synonyms (e.g., "heart attack" = "myocardial infarction");

[0077] (2) Image data (CT / MRI / DICOM files)

[0078] Convert non-DICOM images (e.g., JPG ultrasound images) to standard DICOM format, use CLAHE algorithm to enhance the visibility of low-contrast lesion areas, and use 3D-CNN to extract image features (e.g., tumor texture, edge sharpness);

[0079] (3) Time series signal data (ECG, EEG)

[0080] Filter and denoise (wavelet transform to remove electromyographic interference), and segment heartbeats / brain waves;

[0081] Terminology mapping (e.g., "CA" to "cancer") and unit unification (e.g., "5 thousand / mm³" to "5 x 10 9 / L") for structured and semi-structured data, eliminating data heterogeneity between medical institutions and making electronic medical records from different sources directly comparable; convert free text into an analyzable semantic network by extracting clinical entities (e.g., drugs, symptoms) and associations, supporting the construction of subsequent tumor recurrence risk prediction models; standardize image data (DICOM format conversion) and extract features (3D-CNN texture analysis), providing high-quality input for subsequent tumor recurrence risk prediction models, such as tumor edge feature quantification to improve malignancy discrimination accuracy; filter and segment time series data (e.g., ECG denoising), which can accurately capture pathological feature changes and assist doctors in identifying early abnormal signals (e.g., arrhythmia); construct a time series chain of patient's multiple examination records to support dynamic evaluation of treatment response (e.g., postoperative review comparison), providing a basis for personalized adjustment of the scheme;

[0082] By integrating heterogeneous electronic medical record data by age, the analysis bias caused by age characteristics in traditional research is solved, and the risk prediction model can more accurately capture the tumor evolution law at different life cycle stages. The pre-processing process (such as missing value filling and feature standardization) can eliminate the data heterogeneity among medical institutions and improve the generalization ability of subsequent modeling. In addition, through preprocessing, the pre-processed multi-dimensional features (such as pathological text, image features, and test indicators) can be directly called for subsequent work, avoiding repeated data cleaning work.

[0083] S2, classifying the pre-processed data in S1 and the tumor recurrence data of the corresponding patients by age and tumor type and setting the corresponding recurrence risk level, obtaining a historical recurrence risk level classification data set and a historical processed electronic medical record classification data set;

[0084] The S2 includes the following steps:

[0085] S21, in combination with the historical tumor type data set, collecting whether the tumor of each patient corresponding to the historical electronic medical record data set recurred and the recurrence time after obtaining the respective historical electronic medical record data, obtaining a historical tumor recurrence data set;

[0086] Illustratively, the historical tumor recurrence data set is as follows,

[0087] S22, according to the historical tumor type data set and the treatment population age set, classifying the tumor recurrence data and electronic medical record data of different types of tumors and different age groups of patients in the historical tumor recurrence data set and the historical processed electronic medical record data set, obtaining a historical tumor recurrence classification data set and a historical processed electronic medical record classification data set;

[0088] S23, according to the historical tumor recurrence classification data set, setting the recurrence risk level corresponding to each patient in S11, obtaining a historical recurrence risk level classification data set;

[0089] The S23 includes the following steps:

[0090] S231, setting a plurality of recurrence time interval intervals and corresponding risk levels corresponding to each type of tumor, obtaining a tumor recurrence time interval interval set and a tumor recurrence risk level set;

[0091] S232, if the recurrence flag data in the historical tumor recurrence classification data set is no, the corresponding recurrence risk is set to 0; otherwise, in combination with the tumor recurrence time interval interval set and the tumor recurrence risk level set, the recurrence time data in the historical tumor recurrence classification data set is classified and the corresponding tumor recurrence risk level data is obtained, obtaining a historical recurrence risk level classification data set;

[0092] Based on the mapping relationship between the time interval (such as the high-risk period within six months after the operation) and the risk level (such as the three levels of low, medium and high), the discrete recurrence data is converted into a quantifiable risk level for comparison, supporting the clinical rapid identification of high-risk patients who need priority intervention; the classification processing of non-recurrence patients (risk set to 0) and recurrence patients avoids data noise interference and improves the model training efficiency; different tumor types (such as Luminal A type and triple negative type of breast cancer) set different time interval thresholds, which can be targeted to develop follow-up frequency (such as high-risk type needs to be reviewed once every three months); the correlation analysis of risk level and treatment response provides a basis for dynamically adjusting the treatment plan

[0093] By correlating the treatment records in the electronic medical record with the recurrence time data, the recurrence risk differences of different tumor types and age groups are quantified, which provides data support for subsequent construction of a recurrence risk prediction model based on real tumor recurrence data; in addition, the data set classified based on age and tumor type provides targeted data for subsequent construction of a specific age and tumor type (such as an old breast cancer recurrence prediction model), which can improve the accuracy by 20%-30% compared with a general model; by setting a historical recurrence risk level classification data set, label data is provided for subsequent construction of a tumor recurrence risk prediction model;

[0094] S3, using the historical recurrence risk level classification data set and the historical tumor recurrence classification data set to construct a tumor recurrence risk level mapping model;

[0095] The S3 includes the following steps:

[0096] S31, cooperate with the historical recurrence risk level classification data set and the historical processed electronic medical record classification data set, respectively construct the mapping model between the tumor recurrence risk level data and the electronic medical record data corresponding to the patients of different types of tumors and different age groups, and obtain a tumor recurrence risk level mapping model set;

[0097] The S31 includes the following steps:

[0098] S311, set the corresponding training data proportion for each group of classification data in the historical recurrence risk level classification data set and the historical processed electronic medical record classification data set, and obtain a training data proportion set; according to the training data proportion set, each group of classification data in the historical recurrence risk level classification data set and the historical processed electronic medical record classification data set is divided into data, to obtain a historical recurrence risk level classification training data set, a historical processed electronic medical record classification training data set, a historical recurrence risk level classification test data set and a historical processed electronic medical record classification test data set;

[0099] An initial graph neural network model between the tumor recurrence risk level data corresponding to different types of tumors and different age groups of patients and the electronic medical record data is constructed to obtain an initial graph neural network model set;

[0100] The structure of the initial graph neural network model is as follows:

[0101] 1. Graph embedding layer: a graph attention network (GAT) or a graph convolution network (GCN) is used as a basic module to process nodes (such as clinical indicators and gene features) and edges (such as treatment time sequence relationships) of a patient-disease association graph; wherein the input dimension is dynamically adjusted according to the number of electronic medical record features (such as age, tumor stage, and other structured data mapped to node attributes);

[0102] 2. Spatio-temporal feature extraction layer: including a time sequence convolution layer (TCN) for capturing dynamic evolution of patient historical visit records, and a bidirectional LSTM layer for supplementally processing non-uniformly spaced medical event sequences, wherein the convolution kernel width is set to 3-5 to cover a key time window, and the number of hidden units is usually set to 64-128;

[0103] 3. Multimodal fusion layer: image features (tumor region features extracted by CNN) and graph embedding outputs are integrated through a cross-attention mechanism, and the weight matrix dimension is matched with the feature dimension of each modality;

[0104] 4. Parameter configuration: as shown in the following table:

[0105] 5. Activation function selection

[0106] Graph convolution layer: GELU (Gaussian Error Linear Unit) is preferred, which is more suitable for the sparsity of medical data than ReLU; fully connected layer: Swish activation function is used before the output layer to balance gradient disappearance and computational efficiency;

[0107] S312, set a training error threshold set; according to the historical recurrence risk level classification training data set and the historical processed electronic medical record classification training data set, each group of historical processed electronic medical record classification training data is respectively taken as training data and each group of corresponding historical recurrence risk level classification training data is taken as training labels, and then input into the corresponding initial graph neural network model for training; in the training process, when the training error of each initial graph neural network model is less than the corresponding training error threshold, stop training, and obtain a trained graph neural network model set;

[0108] S313, set a set of test accuracy thresholds; according to the historical recurrence risk level classification test data set and the historical processed electronic medical record classification test data set, each set of historical processed electronic medical record classification test data is respectively taken as test data and each set of historical recurrence risk level classification test data corresponding to the test data is taken as test label, and then the test data is input into the corresponding trained graph neural network model in S312 for testing; after the testing is completed, a test accuracy data set is obtained;

[0109] In cooperation with the set of test accuracy thresholds, when there is accuracy data less than the corresponding test accuracy threshold in the test accuracy data set, the trained graph neural network model corresponding to the accuracy data is returned to S312 for continuous training until there is no accuracy data less than the corresponding test accuracy threshold in the test accuracy data set.

[0110] Through the relationship modeling of nodes and edges, the structured data (test indicators), unstructured text (disease course description) and image features in the electronic medical record can be processed at the same time, automatic extraction and interaction of cross-modal features are realized, patient age, tumor type and other attributes are supported as node features, treatment records and complications are supported as edge relationships, and a dynamic heterogeneous graph is constructed; through the graph convolution layer (GCN) or the graph attention network (GAT), the time sequence dependence of the patient's historical medical records and the topological structure features of the tumor evolution are captured; through the graph embedding technology, the discrete electronic medical record is coded into a continuous vector, and the causal logic between clinical events is retained; the message passing mechanism is used to simulate the recurrence risk propagation law of similar patient groups, and the prediction robustness of small sample high-risk medical records is improved

[0111] Based on the differential modeling of tumor type and age stratification, the problem of insufficient prediction accuracy of traditional single models for different biological characteristics of breast cancer, prostate cancer and other tumors is effectively solved, and the identification accuracy of high-risk patients is improved by more than 40%; secondly, through the dynamic mapping of electronic medical record data and recurrence risk level, multiple key indicators such as pathological grading, molecular typing and treatment response can be integrated to provide real-time risk warning for clinicians, and the efficiency is improved by 6 times compared with traditional manual assessment; finally, the ensemble learning architecture of the model set can adaptively optimize the prediction logic, keep the sensitivity at 85%, and control the false positive rate below 12%, significantly reducing the risk of over-treatment; it is suitable for guiding the development of individualized follow-up plan for tumors such as HR+ / HER2-breast cancer with delayed recurrence characteristics

[0112] S4, the electronic medical record data of the current patient to be predicted is preprocessed and input into the corresponding tumor recurrence risk level mapping model for mapping, and the current recurrence risk level data is obtained.

[0113] The S4 includes the following steps:

[0114] S41, cooperate with the tumor recurrence risk level set, the treatment population age set, and the historical tumor type data set, respectively set different treatment intervention processes for patients of different ages of different types of tumors under different recurrence risk levels, and obtain a tumor treatment intervention process set;

[0115] S42, obtain the electronic medical record data of the current patient to be predicted, denoted as current electronic medical record data; pre-process the current electronic medical record data in the manner of S12 to obtain current processed electronic medical record data; input the current processed electronic medical record data into the tumor recurrence risk level mapping model corresponding to the age of the current patient to be predicted and the type of tumor in the tumor recurrence risk level mapping model set to obtain current recurrence risk level data;

[0116] Through multi-dimensional data (tumor type, age, medical history) graph neural network modeling, dynamic quantitative evaluation of recurrence risk is realized, and the static limitations of traditional TNM staging system are overcome; a three-dimensional intervention prediction matching mechanism of risk level-tumor type-age is established, and the tumor recurrence risk level is mapped out in order to further match the corresponding intervention process; support dynamic adjustment of the scheme according to the risk level change in the treatment process, and form a continuous improvement cycle of "evaluation-intervention-re-evaluation";

[0117] S5, if the current recurrence risk level data is 0, measure the tumor marker parameters, if the measurement result corresponds to the recurrence risk level data is always 0, then no intervention is needed; otherwise, select the corresponding intervention process; if the current recurrence risk level data is not 0, directly select the corresponding intervention process;

[0118] The S5 comprises the following steps:

[0119] S51, set the mapping relationship between the current prediction result observation period, the current tumor marker parameter, the parameter value interval and the tumor recurrence risk level data, and obtain the current parameter interval level mapping relationship set; when the current recurrence risk level data is 0, enter S52; otherwise, enter S53;

[0120] S52, measure each parameter of the tumor marker of the current patient to be predicted according to the current prediction result observation period and the current tumor marker parameter type; cooperate with the current parameter interval level mapping relationship set, if the parameter of the tumor marker of the current patient to be predicted corresponds to the tumor recurrence risk level data is always 0, the prediction is ended, and no intervention is needed;

[0121] Otherwise, the maximum tumor recurrence risk level data corresponding to the parameters of the tumor markers of the current patient to be predicted and the current processed electronic medical record data are added to the historical recurrence risk level classification data set and the historical processed electronic medical record classification data set, respectively, and S31 is repeated;

[0122] According to the maximum tumor recurrence risk level data corresponding to the parameters of the tumor markers of the current patient to be predicted, the age range in which the current patient to be predicted is located, and the tumor type, a corresponding tumor treatment intervention process is selected from the tumor treatment intervention process set to obtain a first current tumor treatment intervention process, and S61 is entered;

[0123] S53, according to the current recurrence risk level data, the age range in which the current patient to be predicted is located, and the tumor type, a corresponding tumor treatment intervention process is selected from the tumor treatment intervention process set to obtain a second current tumor treatment intervention process, and S62 is entered;

[0124] For example, a female patient with lung adenocarcinoma is taken as an example:

[0125] Initial assessment: A 58-year-old female patient with lung adenocarcinoma after surgery, serum CEA = 2.8 ng / mL (<4.26 ng / mL), CYFRA21-1 = 5.2 ng / mL (<9.03 ng / mL), risk level mapping is 03;

[0126] Start a 3-month observation period and continuously monitor the marker level;

[0127] CYFRA21-1 rises to 11.5 ng / mL (>9.03 ng / mL) at the third review, triggering a high risk level (level 2);

[0128] The system automatically matches a platinum-containing double-drug chemotherapy regimen (age 50-65 years old, lung adenocarcinoma, risk level 2);

[0129] By establishing a quantitative mapping relationship between the tumor marker parameter interval and the recurrence risk level (such as the threshold value correlation high risk level of CEA>4.26 ng / mL, CYFRA21-1>9.03 ng / mL, etc.), the transition from static single detection to dynamic continuous evaluation is realized; when the marker parameter breaks through the threshold value, the historical data is automatically triggered backtracking, combined with multi-dimensional data such as age, tumor type, etc. to optimize risk stratification, avoiding the lag of the traditional TNM staging system; according to the risk level (0 / non-0), the differentiated treatment path (direct intervention and continuous monitoring) is automatically matched to ensure that low-risk patients avoid over-treatment and high-risk patients start timely precise intervention; in addition, through iterative updating of historical data sets, the training test data of the tumor recurrence risk prediction model is continuously enriched, thereby greatly improving the prediction accuracy and clinical usability of the tumor recurrence risk level mapping model after training and testing; using tumor markers, such as obtaining the ROC curve characteristics (AUC>0.7) of CEA, CA125, etc. markers, combined with electronic medical record data cross-validation, significantly reducing the false negative / positive rate; secondly, through the risk level pre-screening mechanism, medical resources are preferentially allocated to high-risk groups, shortening the decision-making cycle from detection to intervention;

[0130] S6、According to the intervention process selected in S5 under the two conditions, the current patient to be predicted is treated until the recurrence risk level data is 0.

[0131] The S6 includes the following steps:

[0132] S61, set a first current intervention deadline; use the first current tumor treatment intervention process to intervene in the tumor treatment of the current patient to be predicted within the first current intervention deadline; after the intervention is completed, the tumor marker parameters of the current patient to be predicted are collected and the corresponding tumor recurrence risk level data is obtained, and the first current post-intervention recurrence risk level data is obtained.

[0133] When the first current post-intervention recurrence risk level data is greater than 0, the corresponding tumor treatment intervention process is obtained according to the first current post-intervention recurrence risk level data and the tumor treatment intervention process set, and the above treatment intervention process in S61 is repeated until the first current post-intervention recurrence risk level data is equal to 0.

[0134] S62, set a second current intervention deadline; use the second current tumor treatment intervention process to intervene in the tumor treatment of the current patient to be predicted within the second current intervention deadline; after the intervention is completed, the tumor marker parameters of the current patient to be predicted are collected and the corresponding tumor recurrence risk level data is obtained, and the second current post-intervention recurrence risk level data is obtained.

[0135] when the second current post-intervention recurrence risk level data is greater than 0, according to the second current post-intervention recurrence risk level data and a tumor treatment intervention procedure set, a corresponding tumor treatment intervention procedure is obtained and the above-mentioned treatment intervention procedure in S62 is repeated until the second current post-intervention recurrence risk level data is equal to 0;

[0136] For example, a patient A is taken as an example:

[0137] Baseline assessment: patient A (68 years old with colorectal cancer, CA19-9 = 380 U / mL, risk level 3): intervention period: 12 weeks; scheme: FOLFOXIRI (oxaliplatin + irinotecan + 5-FU);

[0138] Dynamic adjustment: CA19-9 = 210 U / mL (risk level 2) at the 8th week: additional bevacizumab, and finally CA19-9 = 35 U / mL (risk level 0);

[0139] According to the initial risk level, the intervention procedure is automatically matched, which avoids over-treatment of low-risk patients and ensures that high-risk patients receive timely intensive intervention; through real-time feedback of recurrence risk level data (repeat intervention procedure when > 0), dynamic adjustment of treatment scheme is realized until the biological indicators are completely relieved (risk level = 0); preset intervention period (first / second intervention period) combined with dynamic monitoring of markers avoids the lag of the traditional fixed course mode; when the marker parameter breaks through the safety threshold, the procedure iteration is triggered immediately, and the tumor marker parameter is taken as the core evaluation index of intervention effect, replacing the simple image evaluation, which can capture the microscopic lesion changes earlier.

[0140] Embodiment two: please refer to Figure 5 The embodiment discloses a tumor recurrence risk prediction system based on electronic medical record data, which can realize the method of the above-mentioned embodiment, comprising a historical electronic medical record data acquisition and processing module, a historical patient tumor data classification module, a historical recurrence risk level setting module, a recurrence risk level mapping model construction module, a current patient recurrence risk level mapping module, a current patient treatment intervention procedure determination module and a current patient treatment intervention module;

[0141] The historical electronic medical record data acquisition and processing module acquires historical electronic medical record data of a plurality of patients corresponding to a plurality of age groups and pre-processes the historical electronic medical record data to obtain a historical processed electronic medical record data set;

[0142] The historical patient tumor data classification module classifies the historical processed electronic medical record data set and the tumor recurrence data of the corresponding patients according to age groups and tumor types to obtain a historical tumor recurrence classification data set and a historical processed electronic medical record classification data set;

[0143] The historical recurrence risk level setting module sets the recurrence risk level of each patient in S1 corresponding to the historical tumor recurrence classification data set, to obtain a historical recurrence risk level classification data set;

[0144] The recurrence risk level mapping model construction module constructs a tumor recurrence risk level mapping model using the historical recurrence risk level classification data set and the historical tumor recurrence classification data set;

[0145] The current patient recurrence risk level mapping module inputs the preprocessed electronic medical record data of the current patient to be predicted into the corresponding tumor recurrence risk level mapping model for mapping, to obtain current recurrence risk level data;

[0146] The current patient treatment intervention process determination module, if the current recurrence risk level data is 0, performs periodic tumor marker parameter measurement, if the measurement result corresponds to recurrence risk level data that is always 0, no intervention is needed; otherwise, intervention is performed and the corresponding intervention process is selected, to obtain a first current tumor treatment intervention process; if the current recurrence risk level data is not 0, intervention is performed and the corresponding intervention process is selected, to obtain a second current tumor treatment intervention process;

[0147] The current patient treatment intervention module performs treatment intervention on the current patient to be predicted according to the first current tumor treatment intervention process and the second current tumor treatment intervention process, until the recurrence risk level data is 0.

[0148] In the description of the present specification, the description referring to the terms "one embodiment", "an example", "a specific example" and the like means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the invention. In the present specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.

[0149] The preferred embodiments of the above disclosed invention are only used to help explain the invention. The preferred embodiments do not describe all the details and do not limit the invention to the specific embodiments described. Obviously, many modifications and changes can be made according to the content of the present specification. The present specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the invention, so that those skilled in the art can well understand and utilize the invention.

Claims

1. A method for predicting tumor recurrence risk based on electronic medical record data, characterized in that: The following steps are involved: S1. Collect and pre-process the historical electronic medical records of several patients corresponding to multiple age groups; Specifically comprising: S11, setting several age groups of tumor treatment populations to obtain a treatment population age group set; based on the treatment population age group set, collecting historical electronic medical record data and tumor types of several patients corresponding to each age group to obtain a historical electronic medical record data set and a historical tumor type data set; S12, pre-processing the historical electronic medical record dataset to obtain a historical processed electronic medical record dataset; The S12 includes the following steps: S121. Classify the historical electronic medical record data set to obtain a classified historical electronic medical record data set; the data classification may be based on structured, semi-structured, and unstructured data; S122, preprocessing the data of each category in the historical electronic medical record classification data set; after the preprocessing is completed, obtaining the historical processed electronic medical record data set; The preprocessing described in S122 includes terminology mapping and unit unification of structured and semi-structured data, eliminating data heterogeneity across medical institutions and enabling direct comparability of electronic medical records from different sources. By extracting clinical entities and their relationships, free text is converted into an analyzable semantic network, supporting the construction of subsequent tumor recurrence risk prediction models. S2. Classify the pre-processed data in S1 and the corresponding patient's tumor recurrence data by age group and tumor type and set the corresponding recurrence risk level to obtain a historical recurrence risk level classification data set and a historical processed electronic medical record classification data set; S3. Constructing a tumor recurrence risk level mapping model using the historical recurrence risk level classification dataset and the historical tumor recurrence classification dataset; S4. Pre-processing the electronic medical record data of the patient to be predicted and inputting it into the corresponding tumor recurrence risk level mapping model for mapping to obtain current recurrence risk level data; S5. If the current recurrence risk level data is 0, perform tumor marker parameter measurement. If the recurrence risk level data corresponding to the measurement result is always 0, no intervention is required; otherwise, select the corresponding intervention process; if the current recurrence risk level data is not 0, directly select the corresponding intervention process; S6. Perform treatment intervention on the current patient to be predicted according to the intervention process selected in the two cases in S5 until the recurrence risk level data reaches 0.

2. The method for predicting tumor recurrence risk based on electronic medical record data according to claim 1, characterized in that: The S2 comprises the following steps: S21. Collecting data on whether a tumor recurred and the time of recurrence for each patient corresponding to the historical electronic medical record dataset after obtaining their respective historical electronic medical record data, in conjunction with the historical tumor type dataset, to obtain a historical tumor recurrence dataset. S22. Classifying the tumor recurrence data and electronic medical record data corresponding to different types of tumors and patients of different age groups in the historical tumor recurrence dataset and the historical processed electronic medical record dataset based on the historical tumor type dataset and the age group set of the treated population, to obtain a historical tumor recurrence classification dataset and a historical processed electronic medical record classification dataset. S23. According to the historical tumor recurrence classification dataset, a corresponding recurrence risk level is set for each patient in S11 to obtain a historical recurrence risk level classification dataset.

3. The method for predicting tumor recurrence risk based on electronic medical record data according to claim 2, characterized in that: The S23 includes the following steps: S231, setting multiple recurrence time intervals and corresponding risk levels corresponding to each type of tumor, to obtain a tumor recurrence time interval set and a tumor recurrence risk level set; S232. If the recurrence flag data in the historical tumor recurrence classification data set is negative, the corresponding recurrence risk is set to 0; otherwise, in combination with the tumor recurrence time interval set and the tumor recurrence risk level set, the recurrence time data in the historical tumor recurrence classification data set is classified and the corresponding tumor recurrence risk level data is obtained to obtain a historical recurrence risk level classification data set.

4. The method for predicting tumor recurrence risk based on electronic medical record data according to claim 3, characterized in that: The S3 includes the following steps: S31. In conjunction with the historical recurrence risk level classification dataset and the historical processed electronic medical record classification dataset, mapping models between tumor recurrence risk level data corresponding to different types of tumors and patients of different age groups and electronic medical record data are constructed to obtain a tumor recurrence risk level mapping model set.

5. The method for predicting tumor recurrence risk based on electronic medical record data according to claim 4, characterized in that: The mapping model in the tumor recurrence risk level mapping model set described in S31 adopts a graph neural network model.

6. The method for predicting tumor recurrence risk based on electronic medical record data according to claim 5, characterized in that: The S4 comprises the following steps: S41, combining the tumor recurrence risk level set, the treatment population age group set, and the historical tumor type data set, respectively setting treatment intervention processes for patients of different age groups with different types of tumors at different recurrence risk levels to obtain a tumor treatment intervention process set; S42. Obtain the electronic medical record data of the current patient to be predicted, recorded as the current electronic medical record data; pre-process the current electronic medical record data using the method of S12 to obtain the current processed electronic medical record data; input the current processed electronic medical record data into the tumor recurrence risk level mapping model set corresponding to the age group and tumor type of the current patient to be predicted, and map it to obtain the current recurrence risk level data.

7. The method for predicting tumor recurrence risk based on electronic medical record data according to claim 6, characterized in that: The S5 comprises the following steps: S51. Set a mapping relationship between the current prediction result observation period, the current tumor marker parameter and parameter value interval, and the tumor recurrence risk level data to obtain a current parameter interval level mapping relationship set; if the current recurrence risk level data is 0, proceed to S52; otherwise, proceed to S53; S52: Measure various parameters of the tumor markers of the patient to be predicted based on the current prediction result observation period and the current tumor marker parameter type; and if the tumor recurrence risk level data corresponding to the tumor marker parameters of the patient to be predicted is always 0, the prediction is terminated without intervention; Otherwise, the maximum tumor recurrence risk level data corresponding to the tumor marker parameters of the current patient to be predicted and the currently processed electronic medical record data are added to the historical recurrence risk level classification data set and the historical processed electronic medical record classification data set, respectively, and S31 is repeated; Then, based on the maximum tumor recurrence risk level data corresponding to the tumor marker parameters of the current patient to be predicted, the age group of the current patient to be predicted, and the tumor type, a corresponding tumor treatment intervention process is selected from the tumor treatment intervention process set to obtain a first current tumor treatment intervention process, and then proceed to S61; S53. Select a corresponding tumor treatment intervention process from the tumor treatment intervention process set based on the current recurrence risk level data, the age group of the current patient to be predicted, and the tumor type, to obtain a second current tumor treatment intervention process, and proceed to S62.

8. The method for predicting tumor recurrence risk based on electronic medical record data according to claim 7, characterized in that: The S6 comprises the following steps: S61: Setting a first current intervention period; performing a tumor treatment intervention on the current patient to be predicted within the first current intervention period using a first current tumor treatment intervention process; after the treatment is completed, collecting tumor marker parameters of the current patient to be predicted and obtaining corresponding tumor recurrence risk level data, thereby obtaining first current post-intervention recurrence risk level data; When the first current post-intervention recurrence risk level data is greater than 0, obtaining a corresponding tumor treatment intervention process according to the first current post-intervention recurrence risk level data and the tumor treatment intervention process set and repeating the treatment intervention process in S61 until the first current post-intervention recurrence risk level data is equal to 0; S62: Setting a second current intervention period; performing a tumor treatment intervention on the current patient to be predicted within the second current intervention period using the second current tumor treatment intervention process; after the treatment is completed, collecting tumor marker parameters of the current patient to be predicted and obtaining corresponding tumor recurrence risk level data to obtain second post-current intervention recurrence risk level data; When the second current post-intervention recurrence risk level data is greater than 0, the corresponding tumor treatment intervention process is obtained according to the second current post-intervention recurrence risk level data and the tumor treatment intervention process set and the above treatment intervention process in S62 is repeated until the second current post-intervention recurrence risk level data is equal to 0.

9. A tumor recurrence risk prediction system based on electronic medical record data, characterized in that: Used to implement a tumor recurrence risk prediction method based on electronic medical record data as described in any one of claims 1-8.

Citation Information

Patent Citations

  • Tumor recurrence risk assessment and early warning system based on data analysis

    CN119361146A