Method and system for predicting symptom cluster trajectory after colorectal cancer surgery based on multi-source data

By using a multi-source data method to predict the trajectory of postoperative symptom clusters in colorectal cancer, and constructing a trajectory prediction model using a generalized dynamic Bayesian network and a genetic algorithm, the problem of continuous management of postoperative symptom clusters in colorectal cancer is solved, and the ability to describe the recovery process and the level of intelligence of intervention measures are improved.

CN121034640BActive Publication Date: 2026-02-17SICHUAN ACADEMY OF MEDICAL SCI SICHUAN PROVINCIAL PEOPLES HOSPITAL
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511554934.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-29
Publication Date
2026-02-17
Estimated Expiration
2045-10-29

Smart Images

  • Figure CN121034640B_ABST
    Figure CN121034640B_ABST
Patent Text Reader

Abstract

The application discloses a postoperative symptom group trajectory prediction method and system based on multi-source data, relates to the technical field of postoperative management of colorectal cancer, and comprises the following steps: obtaining a unique identity identifier and static physiological characteristics of a target patient, querying a historical evolution trend and a historical intervention trajectory of a postoperative symptom group of the target patient according to the unique identity identifier; obtaining a target intervention measure; and inputting the static physiological characteristics, the historical evolution trend, the historical intervention trajectory and the target intervention measure into a pre-trained trajectory prediction model to generate an expected evolution trend of a target symptom, so that a technical transformation from single-time risk prediction to continuous trajectory prediction for auxiliary decision-making is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of postoperative management technology for colorectal cancer, specifically to a method and system for predicting postoperative symptom cluster trajectories for colorectal cancer based on multi-source data. Background Technology

[0002] Colorectal cancer is one of the most common malignant tumors of the digestive system. Patients often experience various symptoms during postoperative recovery, such as abdominal pain, bloating, constipation, fatigue, and loss of appetite. Accurate prediction of the occurrence and evolution of postoperative syndromes is beneficial for assessing the patient's recovery process and customizing intervention plans.

[0003] Existing postoperative symptom prediction systems primarily predict the probability of postoperative symptoms based on patients' physiological characteristics. For example, the method and system for predicting anastomotic leakage after colorectal cancer surgery (application number CN202411463054.1) uses multivariate regression analysis to screen input features for the prediction model and constructs a probability prediction model for anastomotic leakage based on a logistic regression model. However, for clinical rehabilitation management, patients and medical staff need to pay attention to the evolution of symptoms throughout the entire recovery period. Symptom risk at a single moment is insufficient to provide effective support for the continuous management of postoperative symptom clusters. Summary of the Invention

[0004] The purpose of this invention is to solve the technical problem that existing technologies, which predict symptom risk at a single moment, cannot effectively support the continuous management of postoperative symptom clusters. This invention provides a method and system for predicting the trajectory of postoperative symptom clusters in colorectal cancer based on multi-source data, and provides prediction results of the trajectory of postoperative symptom cluster changes in colorectal cancer patients, so as to guide individualized rehabilitation decision-making.

[0005] According to a first aspect of the present invention, the present invention claims protection for a method for predicting the trajectory of postoperative symptom clusters of colorectal cancer based on multi-source data, comprising: obtaining a unique identifier and static physiological characteristics of a target patient, and querying the historical evolution trend and historical intervention trajectory of the target patient's postoperative symptom clusters based on the unique identifier;

[0006] To obtain the target intervention, static physiological characteristics, historical evolution trends, historical intervention trajectories, and the target intervention are input into a pre-trained trajectory prediction model to generate the expected evolution trend of the target symptom.

[0007] Preferably, the step of generating the desired evolution trend of the target symptom further includes:

[0008] Based on a generalized dynamic Bayesian network, the patient's postoperative stage is mapped to a hidden state, the symptom evolution trend is mapped to an observation vector, the historical intervention trajectory is used as a dynamic intervention variable emitted from the hidden state to the observation value, and the static physiological characteristics are used as a static intervention variable.

[0009] The model parameters of the trajectory prediction model are estimated using the expectation-maximization algorithm;

[0010] The trajectory prediction model is based on the recursive input of the target patient's static physiological characteristics, historical evolution trends, historical intervention trajectories, and target intervention measures into the trajectory prediction model.

[0011] The expected evolution trend is obtained based on the observed vectors of all symptoms at each time point.

[0012] Preferably, the acquisition of target intervention measures also includes:

[0013] In response to input from candidate interventions, determine the number of candidate interventions;

[0014] If the number of candidate interventions is greater than 1, all candidate interventions are encoded into initial gene sequences. The initial gene sequences are then used to construct an initial population through crossover and mutation. The risk value of each individual in the initial population corresponding to the intervention is calculated for the target patient. Based on the risk value, the corresponding fitness is generated. Individuals with higher fitness are preferentially selected to enter the next generation population. The remaining individuals are generated through crossover and mutation. The new population replaces the initial population and the iteration continues until the number of iterations reaches a threshold. The intervention corresponding to the individual with the highest fitness is then selected as the target intervention.

[0015] Preferably, the step of calculating the individual risk value further includes:

[0016] Based on the target patient's static physiological characteristics, historical evolution trend, historical intervention trajectory, and individual corresponding intervention measures, the probability of occurrence and expected recovery time of candidate symptom clusters are generated respectively.

[0017] A risk component is generated based on the expected recovery time for each symptom; the risk components are weighted and summed using the probability of symptom occurrence as a weighting coefficient to generate the risk value for the corresponding individual.

[0018] Preferably, the step of generating an individual's risk component further includes:

[0019] Based on the patient profile of the target patient, the average recovery time of the patient group for the corresponding symptoms is matched, and the corresponding risk component is obtained based on the difference between the expected recovery time and the corresponding average recovery time.

[0020] Preferably, the risk component of each symptom is obtained by the ratio between the duration difference and the preset maximum difference. The duration difference is the preset maximum difference only when the symptom type is irreversible. When the symptom type is reversible, the corresponding duration difference is less than the preset maximum difference.

[0021] Preferably, in the step of weighted summation of all risk components, all candidate symptoms with an occurrence probability not less than a probability threshold are selected to generate a target symptom cluster, and all risk components corresponding to the target symptom cluster are weighted summation.

[0022] According to a second aspect of the present invention, the present invention claims protection for a postoperative symptom cluster trajectory prediction system for colorectal cancer based on multi-source data, comprising:

[0023] The acquisition module is used to acquire the unique identification, static physiological characteristics, and target intervention measures of the target patient;

[0024] The query module is used to query the historical evolution trend of postoperative symptom clusters and historical intervention trajectory of target patients based on unique identification identifiers;

[0025] The generation module is used to input static physiological characteristics, historical evolution trends, historical intervention trajectories, and target intervention measures into a pre-trained trajectory prediction model to generate the expected evolution trend of the target symptoms.

[0026] Preferably, the generation module further includes:

[0027] Based on a generalized dynamic Bayesian network, the patient's postoperative stage is mapped to a hidden state, the symptom evolution trend is mapped to an observation vector, the historical intervention trajectory is used as a dynamic intervention variable emitted from the hidden state to the observation value, and the static physiological characteristics are used as a static intervention variable.

[0028] The model parameters of the trajectory prediction model are estimated using the expectation-maximization algorithm;

[0029] The trajectory prediction model is based on the recursive input of the target patient's static physiological characteristics, historical evolution trends, historical intervention trajectories, and target intervention measures into the trajectory prediction model.

[0030] The expected evolution trend is obtained based on the observed vectors of all symptoms at each time point.

[0031] Preferably, the acquisition module further includes:

[0032] In response to input from candidate interventions, determine the number of candidate interventions;

[0033] If the number of candidate interventions is greater than 1, all candidate interventions are encoded into initial gene sequences. The initial gene sequences are then used to construct an initial population through crossover and mutation. The risk value of each individual in the initial population corresponding to the intervention is calculated for the target patient. Based on the risk value, the corresponding fitness is generated. Individuals with higher fitness are preferentially selected to enter the next generation population. The remaining individuals are generated through crossover and mutation. The new population replaces the initial population and the iteration continues until the number of iterations reaches a threshold. The intervention corresponding to the individual with the highest fitness is then selected as the target intervention.

[0034] Preferably, the acquisition module further includes:

[0035] Based on the target patient's static physiological characteristics, historical evolution trend, historical intervention trajectory, and individual corresponding intervention measures, the probability of occurrence and expected recovery time of candidate symptom clusters are generated respectively.

[0036] A risk component is generated based on the expected recovery time for each symptom; the risk components are weighted and summed using the probability of symptom occurrence as a weighting coefficient to generate the risk value for the corresponding individual.

[0037] Preferably, the acquisition module further includes:

[0038] Based on the patient profile of the target patient, the average recovery time of the patient group for the corresponding symptoms is matched, and the corresponding risk component is obtained based on the difference between the expected recovery time of each symptom and the corresponding average recovery time.

[0039] Preferably, the risk component of each symptom is obtained by the ratio between the duration difference and the preset maximum difference. The duration difference is the preset maximum difference only when the symptom type is irreversible. When the symptom type is reversible, the corresponding duration difference is less than the preset maximum difference.

[0040] Preferably, the acquisition module further includes generating a target symptom cluster by taking all candidate symptoms with an occurrence probability not less than a probability threshold, and weighted summing all risk components corresponding to the target symptom cluster.

[0041] According to a third aspect of the present invention, the present invention claims protection for a device for predicting the trajectory of postoperative symptom clusters of colorectal cancer based on multi-source data, comprising a processor and a memory, the memory storing computer-readable instructions that, when executed by the processor, perform the steps of the method described in the first aspect above.

[0042] This application has the following beneficial effects:

[0043] 1. By using unique identifiers and static physiological characteristics, the model ensures that the predicted input accurately corresponds to the individual characteristics of the target patient. By introducing historical symptom evolution trends and historical intervention trajectories, the model provides time-series contextual information, enabling it to capture the dynamic patterns of symptom clusters at different recovery stages. By modeling static physiological characteristics, target intervention measures, and symptom evolution trends as explicit features, the model achieves a technological shift from single-moment risk prediction to continuous trajectory prediction, thereby enhancing its ability to assist in the continuous management of postoperative symptom clusters in colorectal cancer.

[0044] 2. The trajectory prediction model is based on a generalized dynamic Bayesian network modeling approach, enabling it to simultaneously describe the latent changes in the postoperative recovery stage and the mechanism of action of target interventions under multi-source data conditions. However, the symptom clusters of colorectal cancer patients after surgery exhibit dynamic characteristics, and traditional time series prediction struggles to characterize the implicit patterns of different recovery stages, resulting in insufficient interpretability. This method, by introducing latent state nodes and directed conditional dependency structures, integrates the patient's recovery stage, interventions, and symptom manifestations into a unified probabilistic modeling framework, achieving a dynamic representation of "state transition - intervention response - symptom emission." This allows for the prediction of future symptom trends in target patients while also simulating the recovery evolution path, providing interpretable evidence for clinical decision-making. Furthermore, the model estimates model parameters using the EM algorithm, maintaining stable convergence even with limited sample size, demonstrating high robustness and significantly improving the descriptive ability of the postoperative recovery process and its practical application value in clinical scenarios.

[0045] 3. By introducing the adaptive optimization mechanism of genetic algorithm, this method can automatically search for the global optimal or near-optimal solution among multiple potential intervention options, reduce the bias caused by human experience, improve the intelligence level of multi-option intervention selection, and enable intervention measures to automatically optimize the configuration of intervention variables according to individual patient differences, assisting medical staff in judging the best or near-optimal symptom recovery trajectory of patients after surgery.

[0046] 4. By incorporating expected recovery time into the fitness calculation method, unified risk quantification across symptom clusters of patients is achieved. For symptoms such as pain, constipation, diarrhea, or fatigue, the system can use recovery speed to evaluate severity, which is then converted into a numerical risk component that can be compared across symptoms. This accurately reflects the patient's overall symptom compliance and intervention response status, solving the problem of comparison differences caused by the dimensional differences between different symptoms. This allows risk assessment and intervention optimization to be performed more accurately under cross-symptom cluster conditions.

[0047] 5. By introducing the average recovery time of the patient group as a benchmark, the expected recovery time of the target patient is compared with the average recovery time of the group. Compared with the absolute expected recovery time, the system can more accurately identify abnormal recovery of symptoms, providing more accurate cross-symptom severity comparison results for the genetic algorithm.

[0048] 6. For irreversible symptoms, the expected recovery time and average recovery time are usually considered as +∞, which causes the fitness calculation to fail. Setting a preset maximum difference for irreversible symptoms allows the genetic algorithm to optimize intervention measures based on the occurrence of irreversible symptoms, expanding application scenarios and improving practicality.

[0049] 7. By using a probability threshold screening mechanism, only candidate symptoms with an occurrence probability of not less than a preset threshold are included in the risk value calculation. This effectively reduces low-probability symptoms caused by data noise, model uncertainty, or abnormal input, guides medical staff to prioritize high-incidence symptoms, and improves the robustness of intervention decisions and trajectory prediction. Attached Figure Description

[0050] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.

[0051] Figure 1 This is a flowchart of a method for predicting postoperative symptom clusters of colorectal cancer based on multi-source data, which is an embodiment of this application.

[0052] Figure 2 This is a flowchart illustrating the steps involved in generating the desired evolution trend in an embodiment of this application.

[0053] Figure 3 This is a flowchart illustrating the process of obtaining the target intervention measures involved in the embodiments of this application;

[0054] Figure 4 This is a schematic diagram of the structure of the colorectal cancer postoperative symptom cluster trajectory prediction system based on multi-source data involved in the embodiments of this application;

[0055] Figure 5 This is a schematic diagram of the electronic device structure involved in the embodiments of this application. Detailed Implementation

[0056] This invention provides a method and system for predicting the trajectory of postoperative symptom clusters in colorectal cancer based on multi-source data. To make the above-mentioned objectives, features, and advantages of this application more apparent and understandable, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. The components of the embodiments of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, referring to terms such as "an embodiment," "some embodiments," "implementation," "embodiment," "illustrative embodiment," "example," "specific example," or "some examples," the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely indicates that the specific features, structures, or characteristics described in connection with that embodiment or example are included in at least one embodiment or example of the invention. Moreover, the specific features, structures, or characteristics described can be combined in any suitable manner in one or more embodiments or examples. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.

[0057] It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. Furthermore, in the description of this application, relational terms such as "first," "second," "S1," and "S2" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations.

[0058] According to a first aspect of the present invention, the present invention claims protection for a method for predicting postoperative symptom cluster trajectories of colorectal cancer based on multi-source data, with reference to the appendix. Figure 1 As shown, it includes:

[0059] S1: Obtain the unique identifier of the target patient.

[0060] It should be noted that the unique identifier of the target patient can be automatically obtained through a hospital information system (HIS), electronic medical record system, or follow-up management platform. This unique identifier is used to correspond and associate patient information across different data sources. The unique identifier can be a patient number, medical record number, hospitalization number, or other feasible encrypted identifier.

[0061] S2: Based on the identity identifier, query the evolution trend of the target patient's postoperative symptom cluster to obtain the historical evolution trend. The historical evolution trend refers to the sequence of changes in symptom cluster characteristics at different postoperative time points. Based on the identity identifier, query the evolution trend of the target patient's postoperative intervention measures to obtain the historical intervention trajectory. The historical intervention trajectory refers to the sequence of changes in the intervention measures used at different postoperative time points.

[0062] It should be noted that the historical evolution trend includes several historical evolution components, each corresponding to the evolution trend of a characteristic of a postoperative symptom. For example, for pain symptoms, the historical evolution trend can be obtained through time series recording of pain levels; for constipation symptoms, the historical evolution trend can be obtained through time series recording of defecation frequency, defecation duration, etc.; for fatigue symptoms, it can be obtained through time series recording based on the assessment results of the Brief Fatigue Scale (BFI), but it is not limited to these.

[0063] S3: Obtain the static physiological characteristics of the target patient, such as gender, age, past medical history, family medical history, disease name, etc., but it is not limited to these.

[0064] S4: Obtain the target intervention. The target intervention can be manually entered by healthcare professionals, generated by an existing intervention recommendation system, or obtained through other feasible implementation methods.

[0065] It should be noted that the target intervention includes several intervention variables, which represent specific intervention content. For example, the target intervention for the first patient includes a first intervention variable, a second intervention variable, and a third intervention variable. The first intervention variable could be the use of a first medication, 2 tablets three times a day; the second intervention variable could be the use of a second medication, 30 ml twice a day; and the third intervention variable could be the use of a first device once a day for one hour. The corresponding target intervention is obtained by combining the first, second, and third intervention variables. The types of intervention variables include drug intervention, device intervention, nutritional support, and exercise intervention, etc., but are not limited to these.

[0066] S5: Encode and perform feature fusion operations on static physiological characteristics, historical evolution trends of postoperative symptom clusters, historical intervention trajectories, and target intervention measures, and input them into a pre-trained trajectory prediction model to generate the expected evolution trend of the target symptoms.

[0067] It should be noted that the trajectory prediction model corresponding to each postoperative symptom can be constructed by models such as recurrent neural networks and convolutional neural networks, using retrospectively collected historical patient data as the corresponding dataset, and trained through supervised learning and other methods.

[0068] In this embodiment, a trajectory prediction model is constructed based on a Generalized Dynamic Bayesian Network (GDBN). The postoperative process is modeled as a discrete-time state transition model, where the latent state (recovery phase) drives observations (symptom change trends). State transitions are modulated by intervention measures and static physiological characteristics. Parameters are learned using the Expectation-Maximization (EM) algorithm based on historical data. Given a target intervention, the model predicts the future pain level distribution based on the inferred current latent state and outputs the expected evolution trend corresponding to the pain level. A detailed description is provided using the postoperative pain level change trend as an example, refer to the appendix. Figure 2 As shown, step S5 also includes the following sub-steps:

[0069] S51. Obtain multi-source datasets of colorectal cancer patients after surgery from historical medical records, and preprocess the multi-source datasets.

[0070] It should be noted that the multi-source dataset includes at least static physiological characteristics, time series of symptom observations, and time series of intervention measures, but it is not limited to these.

[0071] In this embodiment, time steps are denoted as t, t=1, ..., T. Here, T represents the total number of time steps up to the present. The time difference between two adjacent time steps is the same, which can be obtained through a preset method, for example, in days. When time step t=1, it represents the first day after surgery. The patient's recovery stage is defined as a hidden state variable, denoted as S, and the patient's recovery stage at time t is denoted as S0. t The patient's characteristics at time t are denoted as {Y}. t A t}. Among them, Y t A represents the patient's pain level at time t. t This indicates the intervention measures taken by the patient at time t.

[0072] It should be noted that, for time t, the patient's condition is represented by the value Y. t The corresponding pain level is based on intervention plan A used at or before time t-1. t-1 .

[0073] It should be noted that the latent state is obtained through pre-setting and represents the target patient's postoperative stage, including the acute inflammatory phase, early recovery phase, stable phase, metabolic imbalance phase, complication phase, deterioration phase, relapse phase, and palliative phase. Of course, the phase may not be limited to these. Specifically, the acute inflammatory phase indicates that the inflammatory response is dominant immediately or early after surgery, manifested by tissue stress and elevated inflammatory markers; the early recovery phase indicates that inflammation weakens and nutrition and tissue repair begin, manifested by gradual improvement in symptoms and physiological indicators; the stable phase indicates that the patient has entered a relatively stable recovery stage, manifested by mild symptoms or near baseline; the metabolic imbalance phase indicates metabolic abnormalities, such as insufficient nutritional support, affecting wound healing and functional recovery; the complication phase indicates the occurrence of postoperative infections or anastomotic leakage and other complications; the deterioration phase indicates the occurrence of organ dysfunction or systemic complications, with the condition progressing towards a more severe state; the relapse phase indicates the worsening of symptoms related to the primary disease, recurrence, or metastasis; and the palliative phase indicates that the disease has entered a stage primarily focused on symptom relief and quality of life maintenance, such as home-based follow-up treatment.

[0074] In this embodiment, the static physiological characteristics are denoted as Z. For continuous static physiological characteristics such as age, z-score standardization is performed, while for discrete static physiological characteristics such as gender and past medical history, one-hot encoding is used. Pain levels are mapped to a rating of 0-10, with higher ratings indicating more severe pain. The categories of intervention measures are one-hot encoded, and the specific daily dosage is converted into standardized continuous values, thereby constructing the intervention feature vector.

[0075] S52. Randomly initialize the hidden state distribution P(S) for each patient. t ), and randomly assign values ​​to the model parameters.

[0076] It should be noted that the model parameters include initial state distribution parameters, state transition parameters, and observation emission parameters. The initial state distribution parameters represent the probability distribution of the patient's initial recovery phase after surgery for a patient with static physiological characteristics Z, i.e., the probability of the patient being in each hidden state on the first day after surgery; the state transition parameters represent the probability of the patient being in each hidden state when performing intervention A. t In the case of hidden state S t Transition to the next hidden state S t+1 The probability; the observed emission parameters represent the hidden state S. t The pain level Y is generated. t The conditional probability.

[0077] S53. Given the current parameter values, calculate the posterior probability distribution for each patient in each hidden state, thereby completing the soft classification of the hidden states. That is, given the observed pain level and intervention measures, calculate the probability of each hidden state occurring at each time step, where s is the probability of the i-th hidden state.i The posterior probability can be expressed as:

[0078] ;

[0079] Among them, Y 1:T This represents the pain level observation sequence from day 1 post-surgery to day T; A 1:T-1 This indicates the sequence of interventions used from postoperative day 1 to day T-1. This represents the parameter estimate for the k-th iteration.

[0080] It should be noted that the molecule is represented in the hidden state s i The probability of observing the current pain trajectory is given by the formula, where the denominator represents the total probability of that observed trajectory across all hidden states. This probability can be calculated using forward, backward, or forward-backward algorithms.

[0081] S54. Update the initial state distribution parameters by maximizing the expected likelihood.

[0082] In this embodiment, because different patients have different physical conditions after surgery, they exhibit different recovery stages initially, such as varying recovery speeds, early complications, or deterioration of the condition. Therefore, the initial state distribution parameter π(Z) is modeled using a softmax form, i.e.:

[0083] ;

[0084] Where, π i (Z) represents the value of the initial state distribution parameter in the i-th hidden state, that is, the probability that the patient is in the i-th hidden state on the first day after surgery; exp represents the exponential function; and The transpose of the weight coefficients of static physiological characteristics is represented by K, where K represents the number of hidden states.

[0085] The update is transformed into a weighted multinomial logistic regression, that is:

[0086] ;

[0087] Where, ω i The weighting coefficients represent static physiological characteristics; N represents the number of patients in the multi-source dataset. This indicates that the nth patient is in the hidden state s at the initial time. i The posterior probability; This represents the static physiological characteristics of the nth patient.

[0088] It should be noted that the weight coefficients of the static physiological features corresponding to the maximum expected likelihood can be obtained by methods such as gradient ascent or standard softmax regression, and then the specific values ​​of the updated initial state distribution parameters can be obtained by modeling and calculation.

[0089] S55. Update the transition probability parameter by maximizing the expected likelihood.

[0090] In this implementation, to account for the intervention effect and individual differences, the transition probability is modeled as follows:

[0091] ;

[0092] Where φ represents the joint feature vector of the intervention and static physiological characteristics, β ij Indicates from state s i Transfer to s j The weight vector.

[0093] In this embodiment, the state transition parameters are updated based on maximizing the expected likelihood, that is:

[0094] ;

[0095] in, This indicates that the nth patient starts from state s at time t. i Transition to state s j The posterior probability; This represents the intervention for the nth patient at time t.

[0096] It should be noted that the weight coefficients of the state transition corresponding to maximizing the expected likelihood can be obtained through methods such as gradient ascent or standard softmax regression, and then the specific values ​​of the updated state transition parameters can be obtained through modeling calculation.

[0097] S56. Update the observed launch parameters.

[0098] In this embodiment, pain level is preset as a continuous variable and modeled using a Gaussian distribution. Simultaneously, to reduce the increase in parameter dimensionality due to repeated modeling of static physiological characteristics, regardless of the values ​​of the static physiological characteristics, when patients are in the same latent state, their pain manifestations are modeled as similarly distributed; that is, the pain level is modeled as follows:

[0099] ;

[0100] ;

[0101] ;

[0102] in, Let represent the expectation of the j-th hidden state. Let represent the variance of the j-th hidden state.

[0103] S57. Determine whether the current iteration meets the stopping rule. If it does, complete the parameter training of the trajectory prediction model. If it does not, return to step S53 to continue execution.

[0104] It should be noted that the stopping rule can be obtained through pre-setting, such as when the parameter changes are all below a preset change threshold or the number of iterations is greater than a preset iteration threshold, but it is not limited to these.

[0105] S58. Input the static physiological characteristics of the target patient, the historical evolution trajectory of pain level, the historical intervention trajectory, and the target intervention measures into the trajectory prediction model to output the expected value of pain level. and the corresponding hidden state Furthermore, by recursively calculating, the expected value and latent state corresponding to each moment within a future preset time period are obtained, thus obtaining the sequence data of the pain level changing over time within the future preset time period, i.e., the expected evolution trend.

[0106] It should be noted that the evolution trend of other symptoms is similar to that of pain level. After the symptom evolution is quantified and defined, the same steps as in step S5 can be followed, and will not be described in detail again.

[0107] In one feasible implementation, refer to the appendix. Figure 3 As shown, in step S4, the method further includes:

[0108] S41. In response to the input operation of candidate intervention measures, determine the number of candidate intervention measures; if the number of candidate intervention measures is 1, continue to execute step S5; if the number of candidate intervention measures is greater than 1, continue from step S42.

[0109] In this embodiment, candidate intervention measures can be manually input by medical staff based on the characteristics of the target patient. The target intervention measure is obtained after judging and / or optimizing the candidate intervention measures input by the medical staff.

[0110] S42. Encode all candidate interventions into initial gene sequences.

[0111] In this implementation, all selectable intervention variables are used as standard variables, and initial gene sequences are generated based on the daily implementation amount of each standard variable in each candidate measure. The initial gene sequences are denoted as G, where G = [g1, ..., g2]. i , ..., g M] Where M represents the number of standard variables, gi This represents the daily implementation amount of the candidate intervention in the i-th standard variable.

[0112] It should be noted that the standard variables can be pre-defined based on all possible intervention variables that may be used after colorectal cancer patients.

[0113] S43. Construct the initial population by crossover and mutation of the initial gene sequence.

[0114] It should be noted that the crossover operation is used to exchange partial gene segments between two initial genes, and can be single-point crossover, multi-point crossover, or uniform crossover, etc., but it is not limited to these. The mutation operation is used to randomly change certain gene segments in the initial gene sequence within a preset value range, and can be positional mutation or uniform mutation, etc., but it is not limited to these. The gene sequences generated by the crossover and mutation operations, together with the initial gene sequence, form the initial population.

[0115] S44. Calculate the risk value of the intervention for each individual in the initial population for the target patient.

[0116] It should be noted that the risk value can be obtained based on the probability of symptom worsening and / or adverse reaction score. For example, it can be obtained by predicting the probability of symptom worsening after the patient uses the intervention based on a pre-built neural network model, or by matching whether the patient will have adverse reactions, the probability of occurrence, and the severity of adverse reactions based on preset medication safety rules, and then obtaining the risk value through a linear weighting operation.

[0117] S45. Generate the corresponding fitness based on the risk value. The higher the risk value, the lower the fitness; the lower the risk value, the higher the fitness. This can be achieved using linearly monotonically decreasing functions, exponentially decaying functions, logarithmic functions, reciprocal functions, etc.

[0118] S46. Based on the selection rules, individuals with higher fitness are preferentially selected from all the initial populations to enter the next generation population. Examples include roulette wheel selection strategies or tournament selection strategies, but these are not limited to these.

[0119] S47. Generate remaining individuals through crossover and mutation, replace the initial population with the new population, and continue iterating until the number of iterations reaches a threshold. The threshold for the number of iterations can be preset.

[0120] S8. Obtain the maximum fitness value in the current population and take the corresponding intervention as the target intervention.

[0121] In one feasible implementation, step S44 further includes:

[0122] S441. Based on the target patient's static physiological characteristics, historical evolution trend, historical intervention trajectory, and individual corresponding intervention measures, generate the probability of occurrence of candidate symptom clusters.

[0123] It should be noted that the candidate symptom cluster includes a set of various postoperative symptoms that patients may experience under the current intervention plan. For example, for patients after colorectal cancer surgery, the candidate symptom cluster may include diarrhea, constipation, abdominal pain, fatigue, and bowel dysfunction, but it is not limited to these. The candidate symptom cluster can be obtained by pre-setting all possible postoperative symptoms, matching them based on the records of patients who have visited the hospital in the past, or through other feasible implementation methods.

[0124] In this embodiment, an incidence prediction model for each symptom can be constructed based on a long short-term memory neural network model. The hidden layers capture the time-series dependencies, outputting the probability of occurrence for each candidate symptom. This incidence prediction model can be trained using data from historical patients' medical records.

[0125] S442. Based on the target patient's static physiological characteristics, historical evolution trend, historical intervention trajectory, and individual corresponding intervention measures, generate the expected recovery time of each candidate symptom cluster.

[0126] In this embodiment, a trajectory prediction model is constructed based on a generalized Bayesian network to obtain recovery evaluation rules for each candidate symptom. For example, for pain level, when the pain level is less than 3, the patient's symptoms can be considered recovered; when the pain level is not less than 3, the patient's symptoms are considered not recovered. Then, the patient's features are input into the trajectory prediction model in a recursive manner, and the output predicted value is compared with the recovery evaluation rules. If the recovery evaluation rules are met, the time step of the recursive operation, i.e., the number of iterations, is recorded and used as the expected recovery time. If the recovery evaluation rules are not met, the recursion continues, and the prediction result and the input features are combined to generate new input features, which are then re-input into the trajectory prediction model.

[0127] It should be noted that the specific implementation method of the trajectory prediction model is as described in step S5, and will not be described in detail here.

[0128] S44. Generate a risk component based on the expected recovery time for each symptom. The longer the expected recovery time, the larger the corresponding risk component value; the shorter the expected recovery time, the smaller the corresponding risk component value.

[0129] It should be noted that the expected recovery time can be directly used as the risk component of the symptom, or the normalized mapping result of the expected recovery time can be used as the component.

[0130] S45. Using the probability of symptom occurrence as a weighting coefficient, sum all risk components to generate the risk value for the corresponding individual.

[0131] In this embodiment, the risk value R is calculated as follows:

[0132] ;

[0133] Where m represents the number of candidate symptoms; α i r represents the incidence rate of the i-th candidate symptom; i denoted as the expected recovery time for the i-th candidate symptom; f represents the mapping function, where a larger input value results in a larger mapping result, and a smaller input value results in a smaller mapping result; R1 represents the penalty value, used to improve the safety of intervention measures. Adverse reaction matching rules are set according to clinical guidelines and other guiding rules. When a patient experiences an adverse reaction after using the corresponding treatment plan, R1 can accumulate the corresponding penalty value.

[0134] In one feasible implementation, the risk value The calculation method can also be:

[0135] ;

[0136] in, This represents the average recovery time for the target patient group at the i-th candidate symptom, which can be calculated based on historical medical records. The larger the difference between the expected recovery time and the average recovery time, the greater the corresponding risk component; the smaller the difference, the smaller the corresponding risk component.

[0137] In one feasible implementation, the risk component of each symptom The value is obtained by the ratio between the duration difference and the preset maximum difference. The duration difference is the preset maximum difference only when the symptom type is irreversible; when the symptom type is reversible, the corresponding duration difference is less than the preset maximum difference. That is:

[0138] ;

[0139] Among them, R max This indicates that the preset maximum difference can be obtained through a pre-setting method.

[0140] In one feasible implementation, step S45 further includes generating a target symptom cluster by taking all candidate symptoms with an occurrence probability not less than a probability threshold, and weighting and summing all risk components corresponding to the target symptom cluster.

[0141] It should be noted that the probability threshold can be obtained by presetting, and the probability thresholds for different candidate symptoms can be the same or different.

[0142] According to a second aspect of the present invention, the present invention claims protection for a postoperative symptom cluster trajectory prediction system for colorectal cancer based on multi-source data, as detailed in the appendix. Figure 4 As shown, it includes:

[0143] The acquisition module is used to acquire the unique identification, static physiological characteristics, and target intervention measures of the target patient;

[0144] The query module is used to query the historical evolution trend of postoperative symptom clusters and historical intervention trajectory of target patients based on their identity identifiers;

[0145] The generation module is used to input static physiological characteristics, historical evolution trends, historical intervention trajectories, and target intervention measures into a pre-trained trajectory prediction model to generate the expected evolution trend of the target symptoms.

[0146] It should be noted that the acquisition module further includes an identity acquisition submodule, a feature acquisition submodule, and an intervention acquisition submodule. Specifically, the identity acquisition submodule is used to acquire the unique identity identifier of the target patient, the feature acquisition submodule is used to acquire the static physiological characteristics of the target patient, and the intervention acquisition submodule is used to acquire the target intervention measures for the target patient.

[0147] In one feasible implementation, the generation module further includes:

[0148] Based on a generalized dynamic Bayesian network, the patient's postoperative stage is mapped to a hidden state, the symptom evolution trend is mapped to an observation vector, the historical intervention trajectory is used as a dynamic intervention variable emitted from the hidden state to the observation value, and the static physiological characteristics are used as a static intervention variable.

[0149] The model parameters of the trajectory prediction model are estimated using the expectation-maximization algorithm;

[0150] The trajectory prediction model is based on the recursive input of the target patient's static physiological characteristics, historical evolution trends, historical intervention trajectories, and target intervention measures into the trajectory prediction model.

[0151] The expected evolution trend is obtained based on the observed vectors of all symptoms at each time point.

[0152] Preferably, the intervention acquisition submodule further includes:

[0153] In response to input from candidate interventions, determine the number of candidate interventions;

[0154] If the number of candidate interventions is greater than 1, all candidate interventions are encoded into initial gene sequences. The initial gene sequences are then used to construct an initial population through crossover and mutation. The risk value of each individual in the initial population corresponding to the intervention is calculated for the target patient. Based on the risk value, the corresponding fitness is generated. Individuals with higher fitness are preferentially selected to enter the next generation population. The remaining individuals are generated through crossover and mutation. The new population replaces the initial population and the iteration continues until the number of iterations reaches a threshold. The intervention corresponding to the individual with the highest fitness is then selected as the target intervention.

[0155] Preferably, the intervention acquisition submodule further includes:

[0156] Based on the target patient's static physiological characteristics, historical evolution trend, historical intervention trajectory, and individual corresponding intervention measures, the probability of occurrence and expected recovery time of candidate symptom clusters are generated respectively.

[0157] A risk component is generated based on the expected recovery time for each symptom; the risk components are weighted and summed using the probability of symptom occurrence as a weighting coefficient to generate the risk value for the corresponding individual.

[0158] Preferably, the intervention acquisition submodule further includes:

[0159] Based on the patient profile of the target patient, the average recovery time of the patient group for the corresponding symptoms is matched, and the corresponding risk component is obtained based on the difference between the expected recovery time of each symptom and the corresponding average recovery time.

[0160] Preferably, the risk component of each symptom is obtained by the ratio between the duration difference and the preset maximum difference. The duration difference is the preset maximum difference only when the symptom type is irreversible. When the symptom type is reversible, the corresponding duration difference is less than the preset maximum difference.

[0161] Preferably, the intervention acquisition submodule further includes generating a target symptom cluster by taking all candidate symptoms with an occurrence probability not less than a probability threshold, and weighted summing all risk components corresponding to the target symptom cluster.

[0162] See attached document Figure 5 As shown, this application provides an electronic device including a processor and a memory. The processor and the memory are interconnected and communicate with each other via a communication bus and / or other forms of connection mechanism (not shown). The memory stores a computer program executable by the processor. When the computing device is running, the processor executes the computer program to perform a system in any of the optional implementations of the above embodiments.

[0163] This application provides a storage medium in which, when the computer program is executed by a processor, it performs a system according to any optional implementation of the above embodiments. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0164] It should be understood that the disclosed system can be implemented in other ways, as illustrated in the embodiments provided in this application. The system embodiments described above are merely illustrative. For example, the division of modules is only a logical functional division, and in actual implementation, it can be divided in other ways. Furthermore, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Additionally, the coupling or direct coupling or communication connection shown or discussed can be through some communication interfaces; the indirect coupling or communication connection between systems or units can be electrical, mechanical, or other forms.

[0165] Furthermore, the units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0166] Furthermore, the functional modules in the various embodiments of this application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.

[0167] Flowcharts are used herein to illustrate the steps of the methods according to embodiments of this disclosure. It should be understood that the preceding or following steps are not necessarily performed in exact order. Instead, the steps can be evaluated in reverse order or simultaneously. Furthermore, other operations can be added to these processes.

[0168] Unless otherwise defined, all terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains. It should also be understood that terms such as those defined in a common dictionary should be interpreted as having a meaning consistent with their meaning in the context of the relevant art, and not as having an idealized or highly formalized meaning, unless expressly defined herein.

[0169] The above provides a detailed description of the method and system for predicting postoperative symptom clusters of colorectal cancer based on multi-source data. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are merely examples of this application and are intended only to help understand the method and system for predicting postoperative symptom clusters of colorectal cancer based on multi-source data. They are not intended to limit the scope of protection of this application. Furthermore, various modifications and variations can be made to this application by those skilled in the art. Any modifications or equivalent substitutions made within the spirit and principles of this application should be included within the scope of protection of this application.

Claims

1. A method for predicting the trajectory of postoperative symptom clusters of colorectal cancer based on multi-source data, characterized in that, The method comprises the following steps: Obtain the unique identity and static physiological characteristics of the target patient, and query the historical evolution trend and historical intervention trajectory of the postoperative symptom group of the target patient according to the unique identity; Obtain the target intervention measure, and input the static physiological characteristics, historical evolution trend, historical intervention trajectory and target intervention measure into the pre-trained trajectory prediction model to generate the expected evolution trend of the target symptom; In the step of obtaining the target intervention measure, the method further comprises the following steps: In response to the input operation of the candidate intervention measure, the number of candidate intervention measures is determined; If the number of candidate intervention measures is greater than 1, all candidate intervention measures are coded into an initial gene sequence, the initial gene sequence is constructed into an initial population through crossover and mutation, the risk value of the target patient using each individual in the initial population is calculated, the corresponding fitness is generated according to the risk value, the individual with higher fitness is preferentially selected into the next generation population, the remaining individuals are generated through crossover and mutation, and the new population is replaced by the initial population to continue iteration until the iteration times reach a threshold value, and the intervention measure corresponding to the individual with the highest fitness is selected as the target intervention measure; In the step of calculating the risk value of the individual, the method further comprises the following steps: According to the static physiological characteristics, historical evolution trend, historical intervention trajectory and intervention measure corresponding to the individual of the target patient, the occurrence probability and expected recovery time length of the candidate symptom group are generated respectively; The risk component is generated according to the expected recovery time length of each symptom, the risk components are weighted and summed with the symptom occurrence probability as the weighting coefficient, and the penalty value is combined to generate the risk value of the corresponding individual, and when the patient uses the corresponding treatment scheme, the penalty value corresponding to the adverse reaction is accumulated; The calculation method of the risk value R is as follows: ; wherein m represents the number of candidate symptoms; a i represents the incidence of the i-th candidate symptom; r i represents the expected recovery duration of the i-th candidate symptom; f represents a mapping function; R1 represents a penalty value; In the step of weighting and summing all risk components, the target symptom group is generated by taking all candidate symptoms with an occurrence probability not less than a probability threshold value, and all risk components corresponding to the target symptom group are weighted and summed. 2.The method of claim 1, wherein, In the step of generating the expected evolution trend of the target symptom, the method further comprises the following steps: Based on the generalized dynamic Bayesian network, the stage of the patient after the operation is mapped to the hidden state, the symptom evolution trend is mapped to the observation vector, the historical intervention trajectory is taken as the dynamic intervention variable emitted from the hidden state to the observation value, and the static physiological characteristics are taken as the static intervention variable; The model parameters of the trajectory prediction model are estimated according to the expectation maximization algorithm; The static physiological characteristics, historical evolution trend, historical intervention trajectory and target intervention measure of the target patient are input into the trajectory prediction model based on the recursive formula; The expected evolution trend is obtained according to the observation vectors of all symptoms at each time. 3.The method of claim 2, wherein, In the step of generating the risk component of the individual, the method further comprises the following steps: The average recovery time length of the corresponding symptom of the patient group is matched according to the patient portrait of the target patient, and the corresponding risk component is obtained according to the time length difference between the expected recovery time length and the corresponding average recovery time length. 4.The method of claim 3, wherein, The value of the risk component of each symptom is obtained by the ratio between the time length difference and a preset maximum difference value, and the time length difference is the preset maximum difference value only when the symptom type is irreversible, and the corresponding time length difference is less than the preset maximum difference value when the symptom type is reversible.

5. A system for predicting trajectories of postoperative symptom clusters in colorectal cancer based on multi-source data, characterized by, The method comprises the following steps: The acquisition module is configured to acquire a unique identity of a target patient, a static physiological feature, and a target intervention measure; The query module is configured to query a historical evolution trend of a postoperative symptom group and a historical intervention trajectory of the target patient according to the unique identity; The generation module is configured to input the static physiological feature, the historical evolution trend, the historical intervention trajectory, and the target intervention measure into a pre-trained trajectory prediction model to generate an expected evolution trend of a target symptom; The acquisition module further includes: In response to an input operation of the candidate intervention measure, the number of candidate intervention measures is determined; If the number of candidate intervention measures is greater than 1, all candidate intervention measures are encoded into an initial gene sequence, the initial gene sequence is constructed into an initial population through crossover and mutation, a risk value of the target patient using each individual in the initial population corresponding to the intervention measure is calculated, a corresponding fitness is generated according to the risk value, individuals with higher fitness are preferentially selected into a next generation population, the remaining individuals are generated through crossover and mutation, the initial population is replaced by the new population, and the iteration is continued until the iteration times reach a threshold value, and the intervention measure corresponding to the individual with the highest fitness is selected as the target intervention measure; The acquisition module further includes: According to the static physiological feature, the historical evolution trend, the historical intervention trajectory, and the individual corresponding intervention measure of the target patient, the occurrence probability and the expected recovery time length of the candidate symptom group are generated respectively; A risk component is generated according to the expected recovery time length of each symptom, the risk components are weighted and summed with the occurrence probability of the symptom as a weighting coefficient, and a penalty value is combined to generate a risk value of the corresponding individual, and when the patient uses the corresponding treatment scheme, an adverse reaction will occur, and the corresponding penalty value is accumulated; In the step of weighting and summing all risk components, the target symptom group is generated from all candidate symptoms with an occurrence probability not less than a probability threshold, and all risk components corresponding to the target symptom group are weighted and summed.

Citation Information

Patent Citations

  • Prediction method and system for occurrence of anastomotic fistula after colorectal cancer operation

    CN119541825A

  • Decision-making method and device based on causal effect estimation model and related equipment

    CN117744808A

  • Predictive risk assessment in patient and health modeling

    US20220230759A1

  • Automated treatment selection method

    US6063028A