A Precision Nutrition Decision-Making System and Method Based on Offline Reinforcement Learning
By using an offline reinforcement learning-based precision nutrition decision-making system, a model is built using patient data to dynamically adjust the nutrition intervention plan. This solves the problem of the lack of individualization and precision in nutrition intervention in existing technologies, and achieves more efficient use of medical resources and improvement of patients' nutritional status.
Patent Information
- Application Number
- CN202411830804.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-12
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2044-12-12
AI Technical Summary
Existing nutritional intervention programs lack individualization and precision, leading to a waste of medical resources and poor treatment outcomes. Current nutritional intervention tools cannot be dynamically adjusted according to the patient's specific nutritional indicators.
A precision nutrition decision-making system based on offline reinforcement learning is adopted. By collecting patient data, an offline reinforcement learning model is constructed, and an intelligent agent is trained using offline data to dynamically adjust the nutrition intervention plan, thereby achieving individualized and precise nutrition intervention.
It improved the precision and flexibility of nutritional intervention, reduced the waste of medical resources, significantly improved patients' nutritional status and clinical outcomes, and enhanced treatment effectiveness.
Smart Images

Figure CN119833107B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of medical informatics and relates to a precision nutrition decision-making system and method based on offline reinforcement learning. Background Technology
[0002] Malnutrition is a common disease with complex etiologies, caused by decreased nutrient intake or absorption, disease-related inflammation, or other mechanisms. The incidence of malnutrition is high among hospitalized patients, leading to adverse consequences such as decreased treatment effectiveness, increased complication rates, reduced quality of life, and worsening long-term prognosis. Therefore, early identification of target populations that can benefit from malnutrition and implementing precise, individualized nutritional interventions targeting reversible factors in the development of malnutrition are crucial for improving the prognosis of hospitalized patients.
[0003] Clinical nutritional intervention methods for hospitalized patients mainly include nutritional education, enteral nutrition, and parenteral nutrition. Enteral and parenteral nutrition interventions using nutritional preparations are inherently complex, for example: 1) diverse intervention routes, including oral, tube feeding, and parenteral nutrition; 2) different dosage forms, dosages, and formulations of nutritional preparations; 3) dynamic adjustments to the intervention plan based on changes in the patient's nutrition-related efficacy indicators during treatment; 4) separate intervention plans need to be developed for different phenotypes of malnutrition; 5) potential limitations of existing nutritional risk screening or assessment tools, leading to some patients being considered unable to benefit from nutritional intervention when they actually can; 6) some patients who do not actually benefit receive unnecessary nutritional interventions, resulting in a waste of medical resources and even negative treatment effects.
[0004] Current nutritional intervention programs are primarily selected based on the personal experience of healthcare professionals, as well as nutritional intervention guidelines, standards, and expert consensus. Because these methods mainly consider standardization at the population level, they sacrifice individual precision and cannot allow for more accurate intervention selection based on the specific values of each patient's core nutritional indicators. Conventional nutritional intervention programs led by healthcare professionals cannot fully guarantee benefits in core nutritional indicators such as weight. Therefore, a decision-making system that intelligently and dynamically selects nutritional intervention programs based on patients' different nutritional phenotypes would provide greater flexibility, help improve treatment outcomes, and promote the rational allocation of medical resources.
[0005] Reinforcement learning, a major branch of machine learning, has become a hot topic in artificial intelligence research in recent years and is increasingly being applied in the medical field. Due to its close resemblance to human learning, it is considered one of the most promising machine learning algorithms for achieving general-purpose artificial intelligence. Unlike supervised learning, which primarily addresses intelligent cognition, the core idea of reinforcement learning is that an agent takes actions within an environment and learns and improves its policy based on the reward information from the environment, thereby achieving specific intelligent decision-making goals. This is similar to the classic scenario in nutritional intervention: the patient can be viewed as a Markovian system exhibiting a dynamically changing nutritional phenotype. Healthcare professionals apply nutritional interventions to change this state and adjust the intervention strategy based on changes in the patient's measurable physical indicators, thereby maximizing medical benefits. Offline reinforcement learning is a form of reinforcement learning where the agent learns from pre-collected historical data (usually collected by other strategies or manually) without needing real-time interaction with the environment. Previous studies have shown that intelligent systems based on offline reinforcement learning do indeed demonstrate significant advantages comparable to or surpassing humans in disease treatment decisions, such as the treatment of sepsis and tumors. However, similar technologies have been less applied to the development of nutritional intervention decision support systems.
[0006] Using offline reinforcement learning to learn effective nutritional intervention strategies from a large amount of collected human experience and construct a decision-making system can assist healthcare professionals in selecting intervention programs. The deployable and automated nature of machine learning models can also improve the efficiency of healthcare services and promote the rational allocation of medical resources, thus possessing high application value. Therefore, the inventors have developed a precision nutrition decision-making system and method based on offline reinforcement learning. Summary of the Invention
[0007] In view of this, the purpose of this invention is to provide a precision nutrition decision-making system and method based on offline reinforcement learning, thereby serving as an artificial intelligence-assisted decision-making tool to support healthcare professionals in formulating and adjusting patient nutrition intervention plans in real time.
[0008] To achieve the above objectives, the present invention provides the following technical solution:
[0009] A precision nutrition decision-making method based on offline reinforcement learning includes the following steps:
[0010] Step 1: Collect medical record data of hospitalized patients, including general patient information, disease information, treatment information, nutrition-related anthropometric indicators, hematological indicators, changes in nutrition-related indicators before and after nutritional intervention, specific nutritional intervention plan, and short- and long-term clinical outcomes;
[0011] Step 2: Preprocess the collected medical record data, including removing outliers, excluding samples with missing key information, imputing missing data, performing one-hot encoding on multi-class data, standardizing continuous variables, removing variables with the same value for all samples, and adding timestamp identification information based on whether the interaction segment between the agent and the environment has ended.
[0012] Step 3: Randomly divide the preprocessed medical record data into training and test sets according to a certain ratio;
[0013] Step 4: Define the patient's nutritional status index as the initial state space, and label this index as the initial state S in the training and testing data;
[0014] Step 5: Define the type of the actual nutritional intervention program as the agent action space, and label it as action A in the training and testing data;
[0015] Step 6: Define the nutritional status benefit after nutritional intervention as the environmental reward, and select one or more nutrition-related indicators as the basis for calculating the environmental reward function. Express the reward result in numerical form and label it as reward R in the training and testing data.
[0016] Step 7: Use the same nutrient state indices as the initial state space as the new state space, and label it as the new state S_new in the training and testing data;
[0017] Step 8: Set the hyperparameters for offline reinforcement learning, including but not limited to learning rate, discount factor, greed coefficient, number of hidden layers in the neural network, loss function, and number of training epochs;
[0018] Step 9: Based on the training set data, according to the parameter settings in Steps 4 to 8, use the offline reinforcement learning algorithm to build the model, and perform the interaction between the agent and the environment based on the offline data in the training set until the model converges or reaches the preset number of iterations.
[0019] Step 10: Evaluate model performance on the test dataset. Group the test data according to whether the nutritional intervention plan in the test dataset is consistent with the model's recommended strategy, and compare whether there are differences in weight, limb skeletal muscle index, nutritional risk screening score, nutritional assessment score, and 30-day mortality rate before and after nutritional intervention between the inconsistent group and the consistent group.
[0020] Furthermore, the offline reinforcement learning algorithm is Q-learning, SARSA, or batch-constrained Q-learning.
[0021] Furthermore, the patient's nutritional status indicators include, but are not limited to, weight, limb skeletal muscle index, calf circumference, triceps skinfold thickness, inflammation status score, nutritional risk screening score, nutritional assessment score, malnutrition diagnosis and severity classification.
[0022] Furthermore, the types of nutritional intervention programs actually implemented include, but are not limited to, parenteral nutrition intervention, enteral nutrition intervention, combined parenteral and enteral nutrition intervention, nutritional intervention not performed in accordance with guidelines, and no nutritional intervention at all.
[0023] Furthermore, the nutritional status benefit after the nutritional intervention is calculated based on the changes in one or more nutrition-related indicators before and after a certain intervention period, the length of which includes, but is not limited to, 24 hours, one week, or one month.
[0024] Furthermore, the calculation method for the environmental reward is as follows: R = post-intervention indicator - pre-intervention indicator. If the indicator increases, the reward is positive; if the indicator decreases, the reward is negative; if there is no change, the reward is zero.
[0025] Furthermore, the collection of medical record data for hospitalized patients in step one also includes information on the patient's surgical treatment and medication treatment.
[0026] Furthermore, in step two, multiple interpolation is used to interpolate the missing data.
[0027] Furthermore, the evaluation methods used in step ten include independent samples t-test, rank-sum test, and chi-square test.
[0028] A precision nutrition decision-making system based on offline reinforcement learning includes:
[0029] Data acquisition module: used to collect patients' electronic medical record data, including general patient information, disease information, treatment information, nutrition-related anthropometric indicators, hematological indicators, changes in nutrition-related indicators before and after nutritional intervention, specific nutritional intervention plans, and short- and long-term clinical outcomes;
[0030] The data analysis module includes a reinforcement learning modeling submodule and an individual recommendation data calculation submodule. The reinforcement learning modeling submodule is used to perform offline reinforcement learning-based modeling based on the data obtained by the data acquisition module, so that the agent can learn to obtain the best nutritional intervention strategy under different nutritional statuses. The individual recommendation data calculation submodule is used to find or calculate the best nutritional intervention plan under the current state based on the intervention strategy obtained by the reinforcement learning modeling submodule and the input individual recommendation data.
[0031] The strategy output module includes a strategy display submodule and a personalized recommendation submodule. The strategy display submodule displays some nutritional intervention plans corresponding to different nutritional states on the screen based on the reinforcement learning modeling results of the reinforcement learning modeling submodule. The personalized recommendation submodule calculates the results of the individual recommendation data calculation submodule and the input patient status, and displays the input patient status and the obtained best nutritional intervention plan.
[0032] The beneficial effects of this invention are as follows:
[0033] (1) Compared to conventional nutritional intervention programs based on the personal experience of healthcare professionals, this invention is based on past human experience, possessing the characteristics of objectivity, prior knowledge, and effect orientation. Furthermore, the offline reinforcement learning algorithm used in this invention can utilize existing data, unlike online reinforcement learning, eliminating the need for real-time data acquisition for modeling. This is because the interaction between intelligent agents and the environment in medical settings carries certain risks, while offline learning offers superior safety and cost savings.
[0034] (2) Compared with nutritional intervention guidelines for specific populations that sacrifice some accuracy, the present invention can theoretically set an infinitely divisible continuous state space, thereby achieving high-granularity segmentation of the population and obtaining more accurate nutritional intervention plans; the present invention can achieve different treatment objectives by replacing different nutritional status indicators and environmental rewards, thus possessing high flexibility; importantly, conventional patient grouping analysis methods may result in a small number of cases that do not conform to experience or guidelines, but the effect of potential intervention plans with significant efficacy is masked, while the present invention is expected to discover such effective intervention strategies.
[0035] (3) This invention does not specify a particular offline reinforcement learning algorithm. Therefore, the specific algorithm can be reasonably selected according to different disease populations, disease stages, data types, and sample sizes. This improves the flexibility of this invention in application and can meet the need to save computing costs under different hardware conditions. Therefore, this model can be easily applied to clinical healthcare scenarios with different conditions and disease complexity.
[0036] Other advantages, objectives, and features of the invention will be set forth in part in the description which follows, and in part will be apparent to those skilled in the art from the following examination, or may be learned from practice of the invention. The objectives and other advantages of the invention can be realized and obtained through the following description. Attached Figure Description
[0037] To make the objectives, technical solutions, and advantages of the present invention clearer, the preferred embodiments of the present invention will be described in detail below with reference to the accompanying drawings, wherein:
[0038] Figure 1 This is a flowchart illustrating the technical process of the present invention.
[0039] Figure 2 This is a schematic diagram of offline reinforcement learning modeling in Embodiment 1 of the present invention and a three-dimensional visualization of the Q-table generated based on the modeling results;
[0040] Figure 3 This is a bar chart showing the changes in body mass index and limb skeletal muscle index before and after intervention and in the cross-group of strategy consistency in Embodiment 1 of the present invention.
[0041] Figure 4 This is a schematic diagram of offline reinforcement learning modeling in Embodiment 2 of the present invention and a three-dimensional visualization of the Q-table generated based on the modeling results;
[0042] Figure 5 This is a bar chart showing the changes in body mass index and limb skeletal muscle index before and after intervention and in the cross-grouping of the strategy consistency in Embodiment 2 of the present invention. Detailed Implementation
[0043] The following specific examples illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of the present invention. Unless otherwise specified, the following embodiments and features can be combined with each other.
[0044] The accompanying drawings are for illustrative purposes only and are schematic diagrams, not actual pictures. They should not be construed as limiting the invention. To better illustrate the embodiments of the invention, some parts in the drawings may be omitted, enlarged, or reduced, and do not represent the actual product dimensions. It is understandable to those skilled in the art that some well-known structures and their descriptions may be omitted in the drawings.
[0045] In the accompanying drawings of the embodiments of the present invention, the same or similar reference numerals correspond to the same or similar components. In the description of the present invention, it should be understood that if terms such as "upper," "lower," "left," "right," "front," and "rear" indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, they are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, the terms used to describe positional relationships in the drawings are only for illustrative purposes and should not be construed as limiting the present invention. For those skilled in the art, the specific meaning of the above terms can be understood according to the specific circumstances.
[0046] In a first aspect, the present invention provides a method for constructing a precision nutrition decision-making model based on offline reinforcement learning, comprising the following steps:
[0047] 1) Selection of the modeling population
[0048] The study selected hospitalized patients and collected variable information from sources such as the medical electronic medical record system. This included, but was not limited to, general patient information, disease information, surgical treatment information, drug treatment information, weight, height, limb skeletal muscle index, calf circumference, triceps skinfold thickness, serum albumin, serum transferrin, C-reactive protein, specific nutritional intervention plan, type of nutritional preparation, specifications of nutritional preparation, dosage of nutritional preparation, route of administration of nutritional preparation, weight changes before and after one month of nutritional intervention, changes in limb skeletal muscle index before and after one month of nutritional intervention, and short- and long-term clinical outcomes as the dataset sample.
[0049] 2) Sample data preprocessing:
[0050] The collected sample datasets are preprocessed, including but not limited to removing outliers, excluding samples with missing key modeling information, imputing missing data using multiple imputation, one-hot encoding of multi-class data, standardizing continuous variables using the z-score method, removing variables with the same value for all samples, and adding timestamp identification information based on whether the interaction segment between the agent and the environment has ended.
[0051] 3) Divide the dataset into training and testing sets:
[0052] The preprocessed sample dataset from step 2) is randomly divided into a training set and a test set according to a certain proportion. The training set comprises a larger proportion and is used for the reinforcement learning model to learn strategies, while the test set comprises a smaller proportion and is used to evaluate the effectiveness of the strategies obtained from the training set.
[0053] 4) Define the range of values for the initial state space (State, S):
[0054] The patient's status is modeled as an environmental status. Specifically, the patient's nutritional status is modeled as a partially observable environmental status. An indicator, set of indicators, or status score that reflects the patient's nutritional status before the implementation of nutritional intervention is used as the initial state space value, and the variable name is labeled as the initial state S in the training and testing data. The selection of indicators includes, but is not limited to, weight, limb skeletal muscle index, calf circumference, triceps skinfold thickness, inflammation status score, nutritional risk screening score, nutritional assessment score, malnutrition diagnosis and severity classification.
[0055] 5) Define the range of values for the agent's action space (Action, A):
[0056] The physician or nutritionist performing the nutritional intervention is modeled as an agent. The type of nutritional intervention plan actually implemented is used as the range of values in the action space. The variable name is labeled as action A in the training and testing data. For algorithms that require the next action to update the model parameters or Q table, the variable name is labeled as A_new. The selection of action variables includes, but is not limited to, the coarse classification of parenteral and enteral nutrition intervention, the daily dose of parenteral nutrition, the number of days of parenteral nutrition intervention, the daily dose of enteral nutrition corrected in kilograms of body weight, and the total dose of parenteral and enteral nutrition actually implemented by the patient calculated based on the individual's energy and protein consumption.
[0057] 6) Define the range of values for the environmental reward (Reward, R):
[0058] The nutritional status benefit of patients after nutritional intervention is modeled as an environmental reward. Changes in one or more nutrition-related indicators before and after a certain intervention period are used as the basis for calculating the environmental reward function. The length of the intervention period includes, but is not limited to, 24 hours, one week, and one month. The selection of nutrition-related indicators depends on one or more indicators that the treatment aims to improve. The reward result is expressed numerically, and the reward value is determined based on whether the patient's nutritional status improves, remains unchanged, or worsens after the nutritional intervention. The variable name is "reward" in the training and testing data. Nutrition-related indicators include, but are not limited to, body weight, limb skeletal muscle index, nutritional risk screening score, and nutritional assessment score. The variable name is "reward R" in the training and testing data.
[0059] 7) Define the range of values for the new state space (New state, S_new):
[0060] The same indicators, indicator sets, or state scores that reflect the nutritional status of patients after the implementation of nutritional intervention as those in step 4) are used as the values of the new state space, and the variable name is marked as the new state S_new in the training and testing data.
[0061] 8) Set the model hyperparameters:
[0062] Pre-set the hyperparameters for offline reinforcement learning, including but not limited to the learning rate α, discount factor γ, greed coefficient ε, number of hidden layers n_layer, loss function loss_function, and training episode;
[0063] 9) Construct an offline reinforcement learning model
[0064] Based on the training data in step 3), and according to the parameter settings in steps 4) to 8), an offline reinforcement learning model is constructed. First, the agent's behavior policy is initialized. Then, based on the offline data in the training set, the agent begins interacting with the environment. Each interaction includes: sampling a patient instance from the dataset and recording its initial state; the agent taking an action; updating the state; calculating the reward based on the state update; and updating the model parameters or Q-table state. This interaction process is repeated until the number of iterations predefined in step 8) is completed or the model converges, generating the agent's behavior policy. The choice of reinforcement learning algorithm includes, but is not limited to, Q-learning, SARSA (State-Action-Reward-State-Action), and Batch Constrained Q-learning (BCQ).
[0065] 10) Evaluate model performance on the test dataset.
[0066] Based on the agent behavior strategy output in step 9), the model performance is evaluated using the test data in step 3). The test indicators include, but are not limited to: grouping patients according to whether the strategy used in the test data is the model-recommended strategy; using independent samples t-test or rank-sum test to evaluate whether there is a significant difference in weight before and after nutritional intervention between the inconsistent group and the consistent group; using independent samples t-test or rank-sum test to evaluate whether there is a significant difference in weight before and after nutritional intervention within the inconsistent group and within the consistent group; using independent samples t-test or rank-sum test to evaluate whether there is a significant difference in limb skeletal muscle index before and after nutritional intervention between the inconsistent group and the consistent group; using independent samples t-test or rank-sum test to evaluate whether there is a significant difference in limb skeletal muscle index before and after nutritional intervention within the inconsistent group and within the consistent group; and using chi-square test to compare whether there is a difference in 30-day mortality between the inconsistent group and the consistent group.
[0067] In step 1), specific disease subgroups are selected for stratified modeling to reduce computational load and obtain more targeted models.
[0068] In steps 4) and 7), the patient's body mass index before and after nutritional intervention is used as continuous and discrete state spaces for modeling, respectively.
[0069] In steps 4) and 7), the skeletal muscle index of the limbs before and after the patient's nutritional intervention is used as the continuous and discrete state spaces for modeling, respectively.
[0070] In steps 4) and 7), the quantitative or qualitative results of nutritional screening, assessment or diagnosis before and after nutritional intervention are used as continuous or discrete state spaces for modeling.
[0071] In steps 4) and 7), cluster analysis is performed on various nutrition-related indicators of the patients, and the clustering results are used as a discrete state space for modeling.
[0072] In step 5), only whether the nutritional intervention plan is carried out in accordance with the guidelines is distinguished, thereby simplifying the action space values and speeding up the model convergence.
[0073] In step 5), cluster analysis is performed on complex intervention schemes, and the results of the cluster analysis are used as values for the discrete action space, thereby accelerating the model convergence speed.
[0074] In step 6), if the patient's nutritional status changes from malnutrition to good nutrition, the positive reward value is multiplied by a coefficient greater than 1 to enhance the reward effect on the agent.
[0075] In step 6), if the patient's nutritional status changes from good nutrition to poor nutrition, the negative reward value is multiplied by a coefficient greater than 1 to enhance the punishment effect on the agent.
[0076] In step 8), the model is optimized for parameters, which is a mesh parameter search in a predefined parameter space;
[0077] In step 9), the experience replay technique is used to sample the non-independent and identically distributed time series data in the training set to form an experience base containing several records. The agent learns by sampling from the experience base.
[0078] In step 10), the evaluation method used is to compare whether there is a difference in the improvement of patients' weight when the human intervention plan and the reinforcement learning model strategy are the same or different.
[0079] In step 10), the evaluation method used is to compare whether there is a difference in the length of hospital stay for patients when the human intervention plan and the reinforcement learning model strategy are the same or different.
[0080] Secondly, this invention provides a precision nutrition decision-making system based on offline reinforcement learning, which mainly includes the following modules:
[0081] Data acquisition module:
[0082] 1) Model training data input submodule: used to collect patients' electronic medical record data, including general information, disease information, treatment information, nutrition-related anthropometric indicators, hematological indicators, changes in nutrition-related indicators before and after nutritional intervention, specific nutritional intervention plans, and short- and long-term clinical outcomes as input data;
[0083] 2) Individual Recommendation Data Input Submodule: Receives single or batch values of specific nutritional status indicators of patients as input data;
[0084] Data Analysis Module:
[0085] 1) Reinforcement Learning Modeling Submodule: Using the data obtained from input module 1), this module performs the offline reinforcement learning-based modeling, enabling the agent to learn the optimal nutritional intervention strategy under different nutritional statuses. This module uses offline reinforcement learning technology to simulate the process of a doctor or nutritionist providing nutritional intervention to a patient, thereby intelligently learning and improving the nutritional intervention strategy. Based on core indicators of clinical nutritional diagnosis and treatment, the nutritional intervention treatment process is modeled as a Markov decision process, the doctor is modeled as an agent, the patient's nutritional status is modeled as a partially observable environment, nutritional intervention is defined as an action, and changes in indicators reflecting the effect of nutritional intervention treatment are defined as reward functions. Through model learning on offline data collection, a precise nutritional intervention strategy that maximizes therapeutic efficacy is obtained, and a decision-making system is constructed.
[0086] 2) Individual Recommendation Data Calculation Submodule: Based on the intervention strategy obtained by the reinforcement learning modeling submodule of the data analysis module 1), the input data obtained by the individual recommendation data input submodule of the input module 2) is used to perform search or function calculation to obtain the best nutrition intervention plan under the current state;
[0087] Strategy output module:
[0088] 1) Strategy Display Submodule: Based on the reinforcement learning modeling results of the data analysis module 1), this module displays some nutritional intervention plans corresponding to different nutritional states on the screen. If there are too many nutritional states to list them all, several examples are displayed.
[0089] Personalized recommendation submodule: Based on the calculation results of the data analysis module 2), it simultaneously displays the patient status input in the input module 2) and the optimal nutritional intervention plan obtained by the data analysis module 2);
[0090] Example 1:
[0091] In this embodiment, see Figure 1 The method for constructing a precision nutrition decision-making model based on offline reinforcement learning includes the following steps:
[0092] 1) Establish a sample set:
[0093] Data from 1210 hospitalized patients with chronic non-communicable diseases who met the research requirements were selected, including general information about the patient population, disease information, weight changes before and after nutritional intervention, and specific nutritional intervention plans, as the dataset sample.
[0094] 2) Sample data preprocessing:
[0095] The collected sample datasets are preprocessed, including removing outliers, excluding samples with missing key modeling information, and imputing missing data.
[0096] 3) Divide the dataset into training and testing sets:
[0097] The preprocessed sample dataset from step 2) is randomly divided, with 80% used as the training set and saved in the experience replay pool, and the remaining 20% used as the test set, which are used for training and performance evaluation of the offline reinforcement learning model, respectively.
[0098] 4) Define the range of values for the initial state space (State, S):
[0099] See Figure 2 The initial discrete state space S is the patient's body mass index (BMI) before nutritional intervention, rounded to one decimal place. The BMI is calculated as: weight (kg) ÷ height (m). 2 ;
[0100] 5) Define the values of the agent's action space (Action, A):
[0101] See Figure 2 The actual nutritional intervention program is taken as the action space A of the intelligent agent, including the following categories: parenteral nutrition intervention performed according to the guidelines, enteral nutrition intervention performed according to the guidelines, combined parenteral and enteral nutrition intervention performed according to the guidelines, nutritional intervention not performed according to the guidelines, and no nutritional intervention performed.
[0102] 6) Define the values for the environmental reward (Reward, R):
[0103] See Figure 2 The difference in body mass index (BMI) before and after nutritional intervention is used as the reward value R, i.e., R = BMI. 干预后 -BMI 干预前 If the body mass index increases, the reward is positive; if the body mass index decreases, the reward is negative; if there is no change, the reward is zero.
[0104] 7) Define the values of the changed state space (New state, S_new):
[0105] See Figure 2 The body mass index after intervention is calculated using the same method as in step 4) and used as the value S_new in the changed discrete state space;
[0106] 8) Set the model hyperparameters:
[0107] The hyperparameters for offline reinforcement learning are set as follows: learning rate α = 0.1, which means that new information accounts for 10% of the Q-value update in each update; discount factor γ = 0.9, which determines the degree of influence of future rewards on the current decision; greed coefficient ε = 0.1, which represents the probability of the agent randomly choosing an action; and the number of learning iterations episode = 50000, which represents the total number of times the agent interacts with the environment.
[0108] 9) Construct an offline reinforcement learning model;
[0109] Based on the training data from step 3), a model is constructed using the Q-learning algorithm in a Python 3.9.11 environment. The steps include: initializing the Q-table using the NumPy library. Rows in the table represent different values in the state space, and columns represent different values in the action space. The initial value of the table is set to 0. An ε-greedy policy is used to select actions. Probabilities are set based on the greedy coefficient ε: if a random number between 0 and 1 is less than ε, an action is randomly selected (exploration). Otherwise, the action with the largest Q-value corresponding to the current state in the Q-table is selected (exploitation). A for loop is used to define 50,000 training rounds. In each loop, an if statement is used to control the agent to take actions. In each training round, the agent selects an action based on the current state, interacts with the environment, receives a reward, and updates the Q-table; based on the current state, the reward R obtained after the action is executed, and the next state S_new. To calculate the Q-value update, for each state-action pair, the Bellman equation is used to update the Q-value: Q[S,A]=Q[S,A]+α×(R+γ×argmax(Q[S_new])-Q[S,A]). In each training round, the agent moves to the next state (S_new) based on the current action and environmental feedback, and updates the round counter. This process is repeated until the training round ends. After training, each row in the Q-table represents the Q-value of different actions in the current state. The agent's optimal strategy is to select the action with the maximum Q-value in each state, i.e., select the column index corresponding to the maximum Q-value in each row as the best nutritional intervention plan learned by the agent. See also Figure 2 .
[0110] 10) Analyze the test dataset
[0111] Based on the agent behavior strategy table output in step 9), a nutritional intervention effect analysis is performed on the test dataset in step 3). The test data is grouped according to whether the nutritional intervention plan in the test dataset is consistent with the optimal plan in the Q-table. Based on the grouping information, inter-group comparisons are made between the body mass index and limb skeletal muscle index before and after the intervention. (See [link to relevant documentation]). Figure 3 .
[0112] The results show that the nutritional adjustment strategy obtained by combining offline data with Q-learning reinforcement learning agents achieved better results than conventional clinical strategies, significantly improving patients' weight gain and effectively improving the atrophy of limb skeletal muscle, with statistically significant differences. Therefore, this invention has the potential to achieve dynamic and intelligent adjustment of nutritional intervention strategies for hospitalized patients, and is expected to significantly improve their weight and maintain improved skeletal muscle levels, thereby improving nutrition-related short- and long-term clinical outcomes.
[0113] Example 2:
[0114] This embodiment is largely the same as Embodiment 1. See [link / reference] Figure 4 However, different state spaces and environmental rewards were defined, and different algorithms were used for modeling. The method for constructing a precise nutrition decision-making model based on offline reinforcement learning includes the following steps:
[0115] 1) Establish a sample set:
[0116] Data from 647 hospitalized patients with chronic non-communicable diseases who met the research requirements were selected, including general information about the patient population, disease information, changes in limb skeletal muscle index before and after nutritional intervention, and specific nutritional intervention plans, as the dataset sample.
[0117] 2) Sample data preprocessing:
[0118] The collected sample datasets are preprocessed, including removing outliers, excluding samples with missing key modeling information, and imputing missing data.
[0119] 3) Divide the dataset into training and testing sets:
[0120] The preprocessed sample dataset from step 2) is randomly divided, with 80% used as the training set and the remaining 20% as the test set, for training and performance evaluation of the offline reinforcement learning model, respectively.
[0121] 4) Define the initial state space (State, S) values:
[0122] See Figure 4 The initial discrete state space was set with the appendicular skeletal muscle mass index (ASMI) of the patient before nutritional intervention, keeping the integer value. The ASMI was calculated as: skeletal muscle mass of the limbs (kg) obtained by bioelectrical impedance analysis ÷ height (m). 2 ;
[0123] 5) Define the values of the agent's action space (Action, A):
[0124] See Figure 4 The actual nutritional intervention plan is used as the agent's action space A, and the ε-greedy strategy selects the next action A_new. The values of the action space include: parenteral nutrition intervention performed according to the guidelines, enteral nutrition intervention performed according to the guidelines, combined parenteral and enteral nutrition intervention performed according to the guidelines, nutritional intervention not performed according to the guidelines, and no nutritional intervention performed.
[0125] 6) Define the values for the environmental reward (Reward, R):
[0126] See Figure 4 The difference in body mass index (BMI) before and after nutritional intervention is used as the reward value R, i.e., R = ASMI. 干预后 -ASMI 干预前 If the skeletal muscle index of the limbs increases, the reward is positive; if the skeletal muscle index of the limbs decreases, the reward is negative; if there is no change, the reward is zero.
[0127] 7) Define the values of the changed state space (New state, S_new):
[0128] See Figure 4 The skeletal muscle index of the limbs after intervention is calculated using the same method as in step 4) and used as the value S_new of the changed discrete state space.
[0129] 8) Set the model hyperparameters
[0130] The hyperparameters for offline reinforcement learning are set as follows: learning rate α = 0.1, which means that new information accounts for 10% of the Q-value update in each update; discount factor γ = 0.9, which determines the degree of influence of future rewards on the current decision; greed coefficient ε = 0.1, which represents the probability of the agent randomly choosing an action; and the number of learning iterations episode = 50000, which represents the total number of times the agent interacts with the environment.
[0131] 9) Construct an offline reinforcement learning model
[0132] Based on the training data from step 3), a model is constructed using the SARSA (State-Action-Reward-State-Action) algorithm in a Python 3.9.11 environment. The steps include: initializing the Q-table using the NumPy library. Rows in the table represent different values in the state space, and columns represent different values in the action space. The initial value of the table is set to 0. An ε-greedy policy is used to select actions. Probabilities are set based on the greedy coefficient ε: if a random number between 0 and 1 is less than ε, an action is randomly selected (exploration). Otherwise, the action with the largest Q-value corresponding to the current state in the Q-table is selected (exploitation). A for loop is used to define 50,000 training rounds. In each loop, an if statement is used to control the agent to take actions. In each training round, the agent selects an action based on the current state, interacts with the environment, receives a reward, and updates the Q-table; based on the current state, the reward R obtained after the action is executed, and the next state S_new. To calculate the Q-value update, SARSA uses the Q-values of the current state and current action, as well as the Q-values of the next state S_new and the next action A_new, using the formula: Q[S,A]=Q[S,A]+α×(R+γ×argmax(Q[S_new,A_new])-Q[S,A]). In each training round, the agent selects an action based on the current state, receives a reward, and moves to the next state S_new. Then, in the new state, the agent selects the next action A_new according to an ε-greedy policy. The corresponding values in the Q-table are updated according to the above Q-value update formula. This process is repeated until the training round ends. After training, each row in the Q-table represents the Q-value of different actions in the current state. The agent's optimal policy is to select the action with the maximum Q-value in each state, i.e., to select the column index corresponding to the maximum Q-value in each row, as the best nutritional intervention plan learned by the agent. See [link to relevant documentation]. Figure 4 .
[0133] 10) Analyze the test dataset
[0134] Based on the agent behavior strategy table output in step 9), a nutritional intervention effect analysis is performed on the test dataset in step 3). The test data is grouped according to whether the nutritional intervention plan in the test dataset is consistent with the optimal plan in the Q-table. Based on the grouping information, inter-group comparisons are made between the body mass index and limb skeletal muscle index before and after the intervention. (See [link to relevant documentation]). Figure 5 .
[0135] The results show that the nutritional adjustment strategy learned by the present invention using offline data combined with SARSA reinforcement learning agent achieves better results than conventional clinical strategies, significantly improving patients' weight and skeletal muscle benefits. Therefore, this invention holds promise for achieving dynamic and intelligent adjustment of nutritional intervention strategies for hospitalized patients, potentially significantly improving their weight and skeletal muscle levels, thereby improving nutrition-related short- and long-term clinical outcomes. Notably, analysis of the results compared to Example 2 shows that the present invention can achieve benefits for different nutritional indicators by combining different state spaces and reward functions, demonstrating good flexibility and adaptability to different clinical diagnostic and treatment purposes and scenarios.
[0136] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A precise nutrition decision-making method based on offline reinforcement learning, characterized by: Includes the following steps: Step 1: Collect medical record data of hospitalized patients, including general patient information, disease information, treatment information, nutrition-related anthropometric indicators, hematological indicators, changes in nutrition-related indicators before and after nutritional intervention, specific nutritional intervention plan, and short- and long-term clinical outcomes; Step 2: Preprocess the collected medical record data, including removing outliers, excluding samples with missing key information, imputing missing data, performing one-hot encoding on multi-class data, standardizing continuous variables, removing variables with the same value for all samples, and adding timestamp identification information based on whether the interaction segment between the agent and the environment has ended. Step 3: Randomly divide the preprocessed medical record data into training and test sets according to a certain ratio; Step 4: Define the patient's nutritional status index as the initial state space, and label this index as the initial state S in the training and testing data; Step 5: Define the actual nutritional intervention program type as the agent action space and label it as action A in the training and testing data; Step 6: Define the nutritional status benefit after nutritional intervention as the environmental reward, and select one or more nutrition-related indicators as the basis for calculating the environmental reward function. Express the reward result in numerical form and label it as reward R in the training and testing data. Step 7: Use the same nutrient state indices as the initial state space as the new state space, and label it as the new state S_new in the training and testing data; Step 8: Set the hyperparameters for offline reinforcement learning, including but not limited to learning rate, discount factor, greed coefficient, number of hidden layers in the neural network, loss function, and number of training epochs; Step 9: Based on the training set data, according to the parameter settings in Steps 4 to 8, use the offline reinforcement learning algorithm to build the model, and perform the interaction between the agent and the environment based on the offline data in the training set until the model converges or reaches the preset number of iterations. Step 10: Evaluate model performance on the test dataset. Group the test data according to whether the nutritional intervention plan in the test dataset is consistent with the model's recommended strategy, and compare whether there are differences in weight, limb skeletal muscle index, nutritional risk screening score, nutritional assessment score, and 30-day mortality rate before and after nutritional intervention between the inconsistent group and the consistent group.
2. The precise nutrition decision-making method based on offline reinforcement learning according to claim 1, characterized in that: The offline reinforcement learning algorithm is Q-learning, SARSA, or batch-constrained Q-learning.
3. The precise nutrition decision-making method based on offline reinforcement learning according to claim 1, characterized in that: The patient's nutritional status indicators include, but are not limited to, weight, limb skeletal muscle index, calf circumference, triceps skinfold thickness, inflammation status score, nutritional risk screening score, nutritional assessment score, malnutrition diagnosis and severity classification.
4. The precise nutrition decision-making method based on offline reinforcement learning according to claim 1, characterized in that: The types of nutritional interventions actually implemented include, but are not limited to, parenteral nutrition intervention, enteral nutrition intervention, combined parenteral and enteral nutrition intervention, nutritional intervention not performed in accordance with guidelines, and no nutritional intervention at all.
5. The precise nutrition decision-making method based on offline reinforcement learning according to claim 1, characterized in that: The nutritional status benefit after the nutritional intervention is calculated based on the changes in one or more nutrition-related indicators before and after a certain intervention period, the length of which includes, but is not limited to, 24 hours, one week, or one month.
6. The precise nutrition decision-making method based on offline reinforcement learning according to claim 1, characterized in that: The calculation method for the environmental reward is as follows: R = post-intervention index - pre-intervention index. If the index increases, the reward is positive; if the index decreases, the reward is negative; if there is no change, the reward is zero.
7. The precise nutrition decision-making method based on offline reinforcement learning according to claim 1, characterized in that: The collection of medical record data for hospitalized patients in step one also includes information on surgical treatment and medication treatment.
8. The precise nutrition decision-making method based on offline reinforcement learning according to claim 1, characterized in that: In step two, multiple interpolation is used to impute the missing data.
9. The precise nutrition decision-making method based on offline reinforcement learning according to claim 1, characterized in that: The evaluation methods used in step ten include independent samples t-test, rank-sum test, and chi-square test.
10. A precision nutrition decision-making system based on offline reinforcement learning, characterized in that: include: Data acquisition module: used to collect patients' electronic medical record data, including general patient information, disease information, treatment information, nutrition-related anthropometric indicators, hematological indicators, changes in nutrition-related indicators before and after nutritional intervention, specific nutritional intervention plans, and short- and long-term clinical outcomes; The data analysis module includes a reinforcement learning modeling submodule and an individual recommendation data calculation submodule. The reinforcement learning modeling submodule is used to perform offline reinforcement learning-based modeling based on the data obtained by the data acquisition module, so that the agent can learn to obtain the best nutritional intervention strategy under different nutritional statuses. The individual recommendation data calculation submodule is used to find or calculate the best nutritional intervention plan under the current state based on the intervention strategy obtained by the reinforcement learning modeling submodule and the input individual recommendation data. The strategy output module includes a strategy display submodule and a personalized recommendation submodule. The strategy display submodule displays some nutritional intervention plans corresponding to different nutritional states on the screen based on the reinforcement learning modeling results of the reinforcement learning modeling submodule. The personalized recommendation submodule calculates the results of the individual recommendation data calculation submodule and the input patient status, and displays the input patient status and the obtained best nutritional intervention plan.
Citation Information
Patent Citations
Hemodialysis patient dry weight auxiliary adjusting system based on deep reinforcement learning
CN114496235A
Nasopharynx cancer patient nutrition risk screening system
CN117457202A