Reinforcement learning-based nail and breast nursing intervention optimization method and system
By using reinforcement learning-based time-to-event prediction and state transition prediction models, the problems of delayed reward and censoring characteristics in breast and thyroid care were solved, achieving stable optimization and reliable recommendation of nursing intervention strategies.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-31
- Publication Date
- 2026-04-10
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In existing technologies, breast and thyroid care outcome events have delayed reward and censoring characteristics, making it difficult to form stable feedback signals that can be used for time-step optimization. Furthermore, they lack the ability to evaluate and compare the effects of nursing intervention sequence actions, which affects the operability and reliability of nursing intervention optimization decisions.
By employing a reinforcement learning-based approach, a state sequence and a nursing intervention action sequence are constructed through joint training of a time-to-event prediction model and a state transition prediction model. The result is an output risk rate sequence and an instant reward sequence, thereby optimizing nursing intervention strategies.
Transforming delayed outcomes into learnable feedback signals at each time step improves the stability and assessability of nursing intervention strategies, quantifies the impact of nursing actions on outcome risk, enhances the reliability of recommendation decisions, and reduces the risk of adverse events.
Smart Images

Figure CN121839016A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of medical care informatization, and in particular to a nursing intervention optimization method and system based on reinforcement learning. BACKGROUND
[0002] The nursing management of thyroid and breast surgery patients during the perioperative period and after discharge usually involves wound / drainage tube care, analgesic care, functional exercise guidance, retest arrangement and follow-up reminders, and other types of nursing interventions. With the popularization of hospital informatization and nursing electronic medical records and mobile follow-up platforms, nursing process data (vital signs, laboratory tests, pain scores, drainage volume, incision assessment, etc.) and nursing action records are gradually stored in a structured and time-stamped manner, providing a foundation for data-driven nursing decision support. In the prior art, nursing intervention decisions mainly rely on clinical pathways, guidelines and experience rules; at the same time, there have also been risk assessment methods based on statistics or machine learning, such as the construction of Logistic regression, Cox survival model or tree model / depth model for risk prediction of outcomes such as incision infection, low calcium-related adverse events, subcutaneous effusion, lymphedema, unplanned readmission, etc., to prompt high-risk patients and assist in developing nursing plans. In recent years, sequential decision algorithms such as reinforcement learning have also been used to explore dynamic treatment or management strategies in medical settings in order to optimize long-term outcomes over multiple time steps.
[0003] However, the above prior art still has the following problems:
[0004] 1. The nursing outcomes are mostly complications or readmission, etc. "time-to-event" indicators, which have the characteristics of delayed return and censoring, making it difficult to directly form a stable feedback signal that can be used for time-step-by-time-step optimization, thereby limiting the effective learning and evaluation of nursing intervention strategies;
[0005] 2. Existing risk prediction methods mostly focus on the "patient state-outcome" association, and often fail to capture the conditional influence of nursing intervention actions on risk rates / event occurrence, and lack the ability to evaluate and compare the effects of different nursing action sequences;
[0006] 3. Many methods lack a predictable model of "state-action-state" dynamic evolution, making it difficult to perform forward-looking simulation and planning evaluation of candidate nursing intervention action sequences before recommendation, thereby affecting the operability and reliability of nursing intervention optimization decisions.
[0007] Therefore, a nursing intervention optimization method and system that can solve the above problems of the prior art is needed to solve the problems of those skilled in the art. SUMMARY
[0008] One purpose of the present application is to propose a method for optimizing nursing intervention based on reinforcement learning, aiming at the problems in the prior art that the nursing outcome events (complications / rehospitalization) have the characteristics of delayed reward and deletion, it is difficult to form learnable feedback and guide the sequential decision of nursing action, and the technical scheme based on "time-to-event prediction model + model-assisted reinforcement learning" is proposed: the nursing process data is uniformly aligned and missing processed in time steps, the state sequence and the nursing intervention action sequence aligned therewith are constructed; the action-conditioned time-to-event prediction model is trained to output the risk rate sequence, and the cumulative risk sequence is obtained by integrating within a preset time window, and then the immediate reward sequence is constructed by adjacent difference to realize reward remodeling; the state transition prediction model is trained and the predicted state sequence is generated by rolling, the predicted state sequence is input into the time-to-event prediction model to obtain the predicted risk rate sequence, the consistency loss is calculated with the real risk rate sequence, and the two models are trained jointly; based on the immediate reward and combined with the model after joint training, the candidate nursing intervention action sequence is planned and evaluated to maximize the cumulative reward, the nursing intervention strategy is trained, and the recommended action is output. The present application has the technical effects of converting the delayed outcome into step-by-step optimizable reward, improving the reliability of candidate nursing intervention sequence evaluation and recommendation, and reducing the risk of adverse events and unplanned rehospitalization.
[0009] The present application provides a method for optimizing nursing intervention based on reinforcement learning, comprising:
[0010] S1. Collect nursing process data and nursing outcome event data for patients with breast and thyroid cancer; S2. Time-align the nursing process data and process missing data to obtain a state sequence and a nursing intervention action sequence aligned with it. Based on the nursing outcome event data, determine the nursing outcome event type, occurrence time, and censoring markers corresponding to the state sequence; S3. Train a time-to-event prediction model based on the state sequence, nursing intervention action sequence, nursing outcome event type, occurrence time, and censoring markers, enabling it to output a hazard rate sequence based on the state sequence and nursing intervention action sequence; S4. Integrate the hazard rate sequence within a preset time window to obtain a cumulative risk sequence, and differ the cumulative risks of adjacent time steps to obtain an immediate reward sequence; S5. Train a state transition prediction model based on the state sequence and nursing intervention action sequence, enabling it to respond to the current time step state and the current... S6. When performing a nursing intervention action at a previous time step, output the predicted state value for the next time step and obtain the predicted state sequence through rolling prediction; S7. Input the predicted state sequence and the nursing intervention action sequence into the time-to-event prediction model to obtain the predicted hazard rate sequence and calculate the consistency loss with the hazard rate sequence. Combine the state prediction loss of the state transition prediction model with the state prediction loss of the state transition prediction model to obtain the joint loss. Update the parameters of the time-to-event prediction model and the state transition prediction model based on the joint loss to obtain the jointly trained time-to-event prediction model and the state transition prediction model; S8. Use the immediate reward sequence and train the nursing intervention strategy model with the goal of maximizing the cumulative reward; S9. Collect the current state of the target patient and input it into the trained nursing intervention strategy model to output at least one candidate nursing intervention action sequence. Select the first nursing intervention action in the candidate nursing intervention action sequence with the best cumulative reward as the recommended nursing intervention action.
[0011] Optionally, S1 includes:
[0012] Nursing process data and nursing outcome event data corresponding to each breast and thyroid patient are collected. The nursing process data includes time-stamped patient status data and time-stamped nursing intervention data.
[0013] The patient status data includes at least one or more of the following: vital signs data, laboratory test data, drainage-related data, pain score data, and incision assessment data;
[0014] The nursing intervention action data at least includes one or more of a wound nursing action, a drainage tube nursing action, an analgesic nursing action, a functional exercise guidance action, a retest arrangement action, and a follow-up reminding action, and each piece of nursing intervention action data at least includes a nursing intervention action type and a nursing intervention action parameter; the nursing outcome event data at least includes a nursing outcome event type, a nursing outcome event occurrence time, and a censoring label, wherein the nursing outcome event type at least includes one or more of a low calcium related adverse event, an incision infection, a subcutaneous effusion, lymphedema, and unplanned readmission, the nursing outcome event occurrence time is the occurrence time of the event corresponding to the nursing outcome event type, and the censoring label is used to represent whether the occurrence of the event corresponding to the nursing outcome event type is observed within a preset observation time length.
[0015] Optionally, S2 includes:
[0016] determining a time step of a uniform time step and an observation start time and an observation end time, and taking the observation start time as a time zero point; discretizing the nursing process data according to the uniform time step to obtain patient state data of a corresponding time step in each uniform time step, so as to form a state sequence, wherein when there are multiple pieces of patient state data in the same uniform time step, the same index is averaged or the last recorded value is taken; performing missing data processing on the missing patient state data in the nursing process data to obtain a state sequence with patient state data in each uniform time step, the missing data processing at least includes one of a previous value filling, a linear interpolation, and a preset value filling; merging the nursing intervention action data according to the uniform time step to determine the nursing intervention action data of the corresponding time step in each uniform time step, so as to form a nursing intervention action sequence corresponding to the state sequence step by step, wherein when there are multiple pieces of nursing intervention action data in the same uniform time step, the multiple pieces of nursing intervention action data are merged into the nursing intervention action data corresponding to the uniform time step; determining a nursing outcome event type, a nursing outcome event occurrence time, and a censoring label based on the nursing outcome event data, wherein the nursing outcome event occurrence time is a time interval of the event corresponding to the nursing outcome event type relative to the observation start time, and the censoring label is used to represent whether the occurrence of the event corresponding to the nursing outcome event type is observed before the observation end time.
[0017] Optionally, S3 includes:
[0018] The state sequence, the nursing intervention action sequence, the nursing outcome event type, the nursing outcome event occurrence time, and the deletion mark are taken as training samples to supervise training of a time-to-event prediction model, wherein the time-to-event prediction model comprises an action encoding layer and a hazard rate output layer; the action encoding layer encodes the nursing intervention action sequence to obtain an action encoding sequence corresponding to the state sequence at each time step, wherein the action encoding layer generates the action encoding sequence based on at least a nursing intervention action type and a nursing intervention action parameter; the hazard rate output layer takes patient state data at each time step and action encoding at the corresponding time step as input, and outputs a hazard rate at each time step corresponding to the nursing outcome event type, thereby obtaining a hazard rate sequence; an event occurrence time step is determined based on the nursing outcome event occurrence time, and the training samples are subjected to deletion processing based on the deletion mark, to construct a loss function of the time-to-event prediction model, wherein for a training sample that is not deleted, the loss function comprises an event likelihood item corresponding to the hazard rate of the event occurrence time step and a survival likelihood item of each time step before the event occurrence time step, and for a deleted training sample, the loss function comprises a survival likelihood item of each time step before the deletion time step; the loss function is used to update parameters of the time-to-event prediction model, to obtain a trained time-to-event prediction model, and the state sequence and the nursing intervention action sequence are calculated forward by using the trained time-to-event prediction model, to output the hazard rate sequence.
[0019] Optionally, S4 comprises:
[0020] A time window length corresponding to a preset time window and a time step length of a uniform time step are determined; the hazard rate sequence is subjected to sliding integral operation according to the time window length, to obtain a cumulative risk sequence, wherein at discrete time steps, the cumulative risk of any time step is the sum of the product of the hazard rate of each time step from the time step to the last time step corresponding to the time window length and the time step length; the cumulative risk sequence is subjected to difference operation between adjacent time steps, to construct an instant return sequence, wherein the instant return of any time step is the cumulative risk of the previous time step minus the cumulative risk of the current time step, so that the instant return corresponding to the decrease of the cumulative risk is a positive value; when the nursing outcome event type comprises multiple event types, the cumulative risk sequence corresponding to each event type is calculated respectively, and the cumulative risk sequence corresponding to each event type is weighted and summed according to a preset event weight, to obtain the cumulative risk sequence for difference operation.
[0021] Optionally, S5 comprises:
[0022] The state sequence and the nursing intervention action sequence are input, the state sequence is divided into state pairs of adjacent time steps according to uniform time steps, a state of a current time step and a nursing intervention action of the current time step are combined to form a model input sample, and a state of a next time step is taken as a model output label; a state transition prediction model is trained based on the model input sample and the model output label, so that the state transition prediction model outputs a predicted value of a state of a next time step corresponding to a current time step when inputting a state of the current time step and a nursing intervention action of the current time step; a state prediction loss is constructed based on a difference between the predicted value of the state of the next time step and the state of the next time step, and the state transition prediction model is updated based on the state prediction loss to obtain a trained state transition prediction model; the state sequence and the nursing intervention action sequence are predicted by using the trained state transition prediction model, wherein at any time step, a state of the time step and a nursing intervention action of the time step are input into the trained state transition prediction model to obtain a predicted value of a state of a next time step, and the predicted value of the state of the next time step is taken as a state of the next time step to continue the prediction, and a predicted state sequence is output.
[0023] Optionally, S6 includes:
[0024] After the predicted state sequence and the nursing intervention action sequence are aligned according to uniform time steps, the trained time-to-event prediction model is input, and a predicted risk rate sequence is output; a consistency loss is calculated based on the predicted risk rate sequence and the risk rate sequence, wherein the consistency loss includes a mean square error loss or a divergence loss between the predicted risk rate sequence and the risk rate sequence; a state prediction loss is obtained, and the consistency loss and the state prediction loss are combined into a joint loss according to a preset weight coefficient group; the trained time-to-event prediction model and the trained state transition prediction model are updated based on the joint loss to obtain a jointly trained time-to-event prediction model and a jointly trained state transition prediction model.
[0025] Optionally, S7 includes:
[0026] Based on state sequences and nursing intervention action sequences, state transition samples are constructed for each time step, such that each state transition sample includes the current time step state, the current time step nursing intervention action, the corresponding time step immediate reward, and the next time step state. These state transition samples are used as reinforcement learning training samples to initialize the nursing intervention strategy model. During training, for any current time step state, the nursing intervention strategy model outputs at least one candidate nursing intervention action sequence. The candidate nursing intervention action sequence and the current time step state are input into a jointly trained state transition prediction model, and rolling prediction is performed for a preset number of prediction steps to obtain a candidate predicted state sequence corresponding to the candidate nursing intervention action sequence. The candidate predicted state sequence is input into a jointly trained time-to-event prediction model to output a candidate risk rate sequence, and the cumulative reward corresponding to the candidate nursing intervention action sequence is calculated according to the construction method of the cumulative risk sequence and immediate reward sequence. An optimization objective function for the nursing intervention strategy model is constructed with the goal of maximizing the cumulative reward, and the parameters of the nursing intervention strategy model are updated based on the optimization objective function to obtain the trained nursing intervention strategy model.
[0027] Optionally, S8 includes:
[0028] The current state of the target patient is collected and input into a trained nursing intervention strategy model, which outputs at least one candidate nursing intervention action sequence. For each candidate nursing intervention action sequence, the candidate nursing intervention action sequence and the current state are input into a jointly trained state transition prediction model, which performs a rolling prediction of a preset number of prediction steps, and outputs a candidate predicted state sequence corresponding to the candidate nursing intervention action sequence. The candidate predicted state sequence is input into a jointly trained time-to-event prediction model, which outputs a candidate risk rate sequence, and calculates the cumulative reward corresponding to the candidate nursing intervention action sequence according to the construction method of the cumulative risk sequence and the immediate reward sequence. The first nursing intervention action in the candidate nursing intervention action sequence with the best cumulative reward is determined as the recommended nursing intervention action, and the recommended nursing intervention action is output.
[0029] On the other hand, the present invention also provides a reinforcement learning-based system for optimizing breast and nail care interventions, comprising:
[0030] The data processing module is used to collect nursing process data and nursing outcome event data of patients with breast and thyroid diseases, and to perform time alignment and missing data processing to generate state sequences, nursing intervention action sequences aligned with them, and nursing outcome event types, event occurrence times and censoring marks corresponding to the state sequences.
[0031] a time-to-event prediction module configured to train a time-to-event prediction model based on the state sequence, the sequence of nursing intervention actions, the time of event occurrence, and the censoring label, and output a sequence of hazard rates;
[0032] a return construction module configured to integrate the sequence of hazard rates within a preset time window to obtain a sequence of cumulative risks, and generate an immediate return sequence through adjacent difference;
[0033] a state transition prediction module configured to train a state transition prediction model to predict a next time step state, and roll a predicted state sequence;
[0034] a joint training module configured to obtain a predicted sequence of hazard rates based on the predicted state sequence, and calculate a consistency loss with the sequence of hazard rates, combine the consistency loss and a state prediction loss into a joint loss to jointly update the time-to-event prediction model and the state transition prediction model;
[0035] a policy training module configured to train a nursing intervention policy model using the immediate return sequence, and update the nursing intervention policy model based on the jointly trained time-to-event prediction model and state transition prediction model to maximize the cumulative return of the candidate sequence of nursing intervention actions;
[0036] an intervention recommendation module configured to input a current state of a target patient into the trained nursing intervention policy model to obtain at least one candidate sequence of nursing intervention actions, and select a first nursing intervention action in the candidate sequence of nursing intervention actions with the optimal cumulative return as a recommended nursing intervention action.
[0037] The present application has the following advantages:
[0038] 1. The sequence of hazard rates is output by the time-to-event prediction model, the cumulative risk is formed by integrating within a preset time window, and the immediate return is constructed through difference, so as to convert the delayed outcomes such as complications and readmission into a time step-learnable feedback signal, thereby improving the stability and evaluability of the nursing intervention policy training;
[0039] 2. The action-conditioned hazard rate output structure is adopted, so that the hazard rate is dynamically updated with the change of the nursing intervention action, the influence of different nursing action types and parameters on the outcome risk can be quantified, and the discrimination ability and recommendation specificity of the candidate sequence of nursing intervention actions are improved;
[0040] 3. The state transition prediction model is introduced and jointly trained with the time-to-event prediction model, the forward-looking simulation and planning evaluation of the "state-action-state-risk" link are realized, the cumulative return of different candidate action sequences can be compared before recommendation, so as to improve the reliability of the recommended decision and reduce the risk of adverse events. BRIEF DESCRIPTION OF DRAWINGS
[0041] The accompanying drawings are included to provide a further understanding of the application and are incorporated in and constitute a part of this specification, illustrate embodiments of the application and together with the description serve to explain the application. In the drawings:
[0042] Figure 1 A flowchart of a reinforcement learning-based nursing intervention optimization method and system for breast cancer is provided. DETAILED DESCRIPTION
[0043] The application will now be described in further detail with reference to the drawings. These drawings show only the essential features of the application and are therefore to be regarded only as being schematic illustrations. In the drawings:
[0044] Reference Figure 1 A reinforcement learning-based nursing intervention optimization method, comprising:
[0045] S1, collecting nursing process data and nursing outcome event data of breast cancer patients; S2, performing time alignment on the nursing process data and missing data processing to obtain a state sequence and a nursing intervention action sequence aligned therewith, determining a nursing outcome event type corresponding to the state sequence, a nursing outcome event occurrence time, and a deletion mark based on the nursing outcome event data; S3, training a time-to-event prediction model based on the state sequence, the nursing intervention action sequence, the nursing outcome event type, the nursing outcome event occurrence time, and the deletion mark, so that the time-to-event prediction model outputs a risk rate sequence based on the state sequence and the nursing intervention action sequence; S4, integrating the risk rate sequence within a preset time window to obtain a cumulative risk sequence, and differentiating the cumulative risks of adjacent time steps to obtain an instantaneous return sequence; S5, training a state transition prediction model based on the state sequence and the nursing intervention action sequence, so that the state transition prediction model outputs a next time step state prediction value when inputting a current time step state and a current time step nursing intervention action, and obtains a predicted state sequence through rolling prediction; S6, inputting the predicted state sequence and the nursing intervention action sequence into the time-to-event prediction model to obtain a predicted risk rate sequence and calculating a consistency loss with the risk rate sequence, combining the state prediction loss of the state transition prediction model into a joint loss, updating the parameters of the time-to-event prediction model and the state transition prediction model based on the joint loss, and obtaining the time-to-event prediction model and the state transition prediction model after joint training; S7, training a nursing intervention strategy model using the instantaneous return sequence and aiming to maximize the cumulative return; S8, inputting the current state of a target patient into the trained nursing intervention strategy model to output at least one candidate nursing intervention action sequence, and selecting the first nursing intervention action in the candidate nursing intervention action sequence with the optimal cumulative return as the recommended nursing intervention action.
[0046] In this specific embodiment, S1 comprises:
[0047] A same patient identifier mapping relationship is established between the nursing process data and the nursing outcome event data with the patient as the minimum data unit, wherein the patient identifier is denoted as , and all records of the same patient within the perioperative period to the follow-up period are collected and de-identified;
[0048] The nursing process data includes time-stamped patient status data and time-stamped nursing intervention action data, the patient status data at least selects one or more of vital sign data, laboratory examination data, drainage related data, pain score data and incision evaluation data as a status indicator, and a field dictionary and unit specification are established for each indicator to achieve consistency across sources, such as temperature unified in Celsius, drainage volume unified in milliliters, pain score unified in integer level from 0 to 10, incision evaluation unified in grading or enumeration coding, and time stamp unified to the minute and using the same time zone;
[0049] The nursing intervention action data at least includes one or more of wound care action, drainage tube care action, analgesic care action, functional exercise guidance action, retest arrangement action and follow-up reminder action, and each nursing intervention action data at least contains nursing intervention action type and nursing intervention action parameter, wherein the nursing intervention action type is used to represent the action category, and the nursing intervention action parameter is used to represent the executable details and intensity information of the category action, and the nursing intervention action parameter is saved in structured key-value pairs for subsequent coding, such as recording drug name, administration route, single dose, administration frequency for analgesic care action, recording whether to flush the tube, flushing fluid volume, negative pressure maintenance state for drainage tube care action, recording retest items and planned time for retest arrangement action;
[0050] The nursing outcome event data at least includes nursing outcome event type, nursing outcome event occurrence time and deletion marker, the nursing outcome event type at least includes one or more of low calcium related adverse events, incision infection, subcutaneous fluid collection, lymphedema and unplanned readmission, the nursing outcome event occurrence time is the occurrence time of the corresponding event in the actual world, and the deletion marker is used to represent whether the occurrence of the corresponding event is observed within the preset observation time, which can be configured by the hospital process as 30 days or 90 days after discharge and remains consistent for the same batch of data;
[0051] To ensure data availability, after collection is completed, each record is subjected to integrity check and time stamp rationality check, and obviously impossible abnormal values are removed or truncated according to preset rules, such as temperature less than or greater than is considered abnormal and marked for review, and when laboratory indicators exceed the upper limit of the device range, the upper limit of the range is truncated and the original value is retained for reference.
[0052] In the specific embodiment, S2 comprises:
[0053] determining the time step of the uniform time step, the observation start time and the observation end time, wherein the observation start time is denoted as , the observation end time is denoted as , the time step of the uniform time step is denoted as , and all subsequent relative times are calculated based on the same reference system with as the time zero point, and the uniform time step boundary is generated according to the following formula:
[0054] ;
[0055] wherein denotes the start time of the th uniform time step, denotes the uniform time step number, denotes the fixed time interval between adjacent uniform time steps, denotes the total number of uniform time steps required from to , and denotes the observation end time, denotes the rounding up to ensure coverage of the end time;
[0056] wherein is determined by the nursing service rhythm and data density, for example, the in-hospital perioperative period can be set to 1 hour or 4 hours, and the post-discharge follow-up period can be set to 1 day. If a patient has both in-hospital and follow-up data, the smallest common granularity is taken as and is kept consistent through the aggregation rule;
[0057] After the uniform time step is determined, the patient state data is discretized and aggregated according to the time stamp falling into the corresponding uniform time step. If there are multiple records of the same state indicator in the same uniform time step, the "latest record" is preferred to retain the state at the adjacent decision-making time. If the indicator is a high-frequency continuous monitoring with large noise, the "mean value" is adopted to enhance robustness. Moreover, a fixed aggregation rule is adopted for each state indicator and written into the configuration table to ensure reproducibility.
[0058] In terms of state sequence generation, for each patient Patient status data corresponding to each time step is output according to a unified time step and spliced into a status sequence according to a predefined index order. Then, missing data processing is performed on the status sequence to ensure that patient status data is available at each unified time step. Missing data processing is performed according to the following priorities and a fixed strategy is used for each index to avoid information leakage: when an index is missing at a time step and there is an observed value in the previous time step, the previous value is used to fill it; when consecutive missing values occur and there are observed values at both ends of the missing interval, linear interpolation is used; when the missing value occurs at the beginning of the sequence or is missing throughout, the preset value is used to fill it. The preset value is determined by the median of the clinical norm or the hospital standard reference range and is fixed in the configuration table.
[0059] In terms of generating nursing intervention action sequences, nursing intervention action data is merged according to a unified time step. If there are multiple nursing intervention action data within a unified time step, they are merged into the nursing intervention action data corresponding to that unified time step. The merged result includes at least the set of nursing intervention action types that occurred in that time step and the set of nursing intervention action parameters corresponding to each type of action. When the same nursing intervention action type appears multiple times in the same time step, it is merged according to business interpretable rules, such as taking the total amount for dosage parameters, taking the maximum value for frequency parameters, and taking a logical OR for yes / no parameters, while retaining the original details for traceability.
[0060] Regarding nursing outcome event alignment, nursing outcome event types, occurrence times, and censoring markers are determined based on nursing outcome event data, and the occurrence times of nursing outcome events are converted relative to the observation start time. Time interval:
[0061] ;
[0062] in Indicates the patient event types The corresponding relative time interval between events Indicates the patient event types The corresponding actual time of the event, This indicates the start time of observation for this patient and the event type. Used to differentiate between different nursing outcomes such as hypocalcemia-related adverse events, surgical site infection, subcutaneous effusion, lymphedema, and unplanned readmission;
[0063] Censoring markers are generated and bound to the state sequence patient by patient according to the following rules: if at the end of observation time Previously observed event types If it occurs, it is recorded as not censored, and the event occurrence time step is determined as the unified time step into which the event occurrence time falls. type of event that has not been observed before is recorded as censored and the censored time step is determined as the unified time step that the fall into, thus obtaining a triple of nursing outcome event type, relative event occurrence time and censoring label corresponding to the state sequence, which can be used for subsequent time-to-event modeling.
[0064] In the specific embodiment, S3 comprises:
[0065] taking the state sequence and the nursing intervention action sequence aligned therewith as input, and taking the nursing outcome event type, the nursing outcome event occurrence time and the censoring label as the supervision signal to construct the training sample, and supervising the training of the time-to-event prediction model to output the hazard rate sequence;
[0066] wherein for any training sample identified as , the state sequence is denoted as , the nursing intervention action sequence is denoted as , the event type is denoted as , the event occurrence time step is denoted as , the censoring label is denoted as , and is the total number of unified time steps, is the unified time step number, the determination rule of is to take the event type corresponding to the unified time step that the nursing outcome event occurrence time falls into as the event occurrence time step, if the event type has not been observed before the observation termination time, then take the unified time step that the observation termination time falls into as the censored time step and let equal to the censored time step, denotes that the event type has been observed and is not censored, denotes that the event type has not been observed and is censored;
[0067] The time-to-event prediction model comprises an action encoding layer and a hazard rate output layer, wherein the action encoding layer is configured to encode the nursing intervention action sequence into an action encoding sequence corresponding to the state sequence step by step, specifically, for the nursing intervention action at each time step is at least split into a nursing intervention action type and a nursing intervention action parameter and is represented in a structured manner, the nursing intervention action type is denoted as and takes a value from a preset action type set, and the nursing intervention action parameter is denoted as and is a parameter vector matched with the action type; for the nursing intervention action type The trainable embedding table is used for encoding, and the embedding dimension is taken as In order to control the model complexity while maintaining the expression ability, the nursing intervention action parameters First, the numerical value is processed according to the field type, and the continuous type parameter is standardized with zero mean and unit variance, and the enumeration type parameter is one-hot encoded, and the missing parameter is represented by a preset default value and a missing flag bit to avoid mistaking the missing value as a real value. Subsequently, the normalized parameter vector is input into a two-layer multilayer perception for encoding, and the hidden layer width of the multilayer perception is taken as 64 and 32 in turn, the activation function is taken as ReLU, and the output dimension is taken as ;
[0068] The nursing intervention action type embedding vector and the nursing intervention action parameter encoding vector are spliced according to the feature dimension and obtained through a linear mapping to obtain an action encoding vector , the dimension of the action encoding vector is taken as , thereby forming an action encoding sequence ;
[0069] The risk rate output layer is used for outputting the risk rate of each time step based on the patient state data of each time step and the action encoding output and event type of the corresponding time step , specifically, the state vector and the action encoding vector of each time step are spliced into a time sequence input vector, and input into a single-layer gated recurrent unit to extract time sequence dependent information, the hidden state dimension of the gated recurrent unit is taken as and 0.1 is adopted during training to suppress overfitting, thereby obtaining a time sequence feature vector of each time step, and then the event type corresponding linear output head and Sigmoid function map the time sequence feature vector to a discrete time risk rate, the discrete time risk rate represents the probability of occurrence of the event type within the time step under the condition that the patient has survived to the start time of the time step , and the calculation is as follows:
[0070] ;
[0071] wherein represents the risk rate of the patient to the event type at the time step , represents the Sigmoid function for compressing the output to the interval of 0 to 1, represents the linear output head weight vector of the event type , denotes a transpose operation, denotes a time step output by the gated recurrent unit , denotes an event type , outputs a hazard rate sequence ;
[0072] In terms of loss function construction, based on the event occurrence time step and the censoring indicator , the training samples are censored and the discrete-time survival likelihood is used to construct the supervision signal, the event likelihood term is introduced at the event occurrence time step for uncensored samples and the survival likelihood term is introduced before the event occurrence time step, and only the survival likelihood term is introduced before the censoring time step for censored samples, and the negative log-likelihood loss of a single sample and a single event type is defined as:
[0073] ;
[0074] wherein denotes a patient , the loss value for an event type , denotes a natural logarithm function, denotes a summation operation, denotes a censoring indicator and is used to enable the event likelihood term when uncensored and disable the event likelihood term when censored, denotes an event occurrence time step or a censoring time step, denotes a time step number to be accumulated, denotes a hazard rate at the event occurrence time step , denotes a conditional survival probability at the time step , and when the upper bound of the summation becomes to only accumulate the survival likelihood before the event occurrence, and when the upper bound of the summation becomes to accumulate the survival likelihood to the censoring time step;
[0075] When modeling multiple types of nursing outcome events, an independent output head is set for each event type, and the total loss is obtained by weighting and summing the loss of each event type according to the preset event weight for end-to-end training, and the model parameters are updated using small batch gradient descent during training, the optimizer is Adam and the learning rate is , and the weight decay is The batch size is set to 128, the maximum number of training rounds is set to 50, and early stopping is based on the validation set loss with a patience round of 5. At the same time, the gradient norm is pruned to 1.0 to improve training stability. After training is completed, the state sequence and nursing intervention action sequence are input into the event prediction model after training to perform forward calculation to output the risk rate sequence of the corresponding event type.
[0076] In this specific embodiment, S4 includes:
[0077] Constructing an immediate reward sequence for reinforcement learning based on a hazard rate sequence involves first determining the time window length corresponding to a preset time window and the time step size of a uniform time step, where the time window length is denoted as... The time step of the unified time step is denoted as The time window length is discretized into time steps. ,in Indicates the number of uniform time steps covered by the preset time window, round This indicates rounding to the nearest integer to ensure that the time window length is aligned with the discrete time step. Maintain consistency with step S2 and keep it fixed within the same batch of training data;
[0078] For any patient, the identifier is Samples and any nursing outcome event type The hazard rate sequence corresponding to this event type is obtained from the time-to-event prediction model. ,in Indicates the patient In the Event types occurring within a unified time step Discrete-time hazard rate, Indicates a unified time step sequence number. The total number of time steps is represented by a constant. Then, a sliding integral is performed on the hazard rate sequence within a preset time window to obtain the cumulative risk sequence. A numerical integration approximation using "summation multiplied by step size" is applied at discrete time steps. The cumulative risk is calculated using the following formula:
[0079] ;
[0080] in Indicates the patient At time step Starting time facing the preset time window Weighted cumulative risk This represents the time step number that is accumulated by integration within the time window. This indicates that a smaller value is used to handle the case where the time window at the end of the sequence goes out of bounds. denotes a set of concurrent outcome event types being modeled, denotes an event type with preset event weights and satisfies and can be configured as to achieve a comparable scale of different event types risk, if only a single event type is modeled, it makes and so that degrades to the cumulative risk of the event type;
[0081] In implementation, the cumulative risk sequence is calculated in a sliding window accumulation manner to reduce complexity, first calculate the unweighted cumulative risk sequence for each event type respectively and cache, then do linear weighting according to the preset event weights to get , and can be truncated to avoid numerical anomalies, limit to interval to keep consistent with the "window cumulative conditional occurrence probability upper bound";
[0082] After obtaining the cumulative risk sequence, the cumulative risks of adjacent time steps are differenced to construct the immediate return sequence, so that the cumulative risk decline corresponds to a positive return, and the immediate return is defined as follows:
[0083] ;
[0084] wherein denotes the immediate return of the patient at time step , denotes the cumulative risk corresponding to the previous time step, denotes the cumulative risk corresponding to the current time step;
[0085] To handle the boundary conditions of the difference, initialize to 0 or equivalently let to ensure that the length of the return sequence is aligned with the state sequence and the nursing intervention action sequence step by step, and the final output immediate return sequence as the step-by-step return input of the reinforcement learning training sample in step S7, so as to realize the conversion of the risk change of delayed outcomes such as complications or readmission into step-by-step learnable and comparable immediate feedback signals.
[0086] In the specific embodiment, S5 includes:
[0087] Train the state transition prediction model based on the state sequence and the nursing intervention action sequence to depict the dynamic evolution relationship of "current state + current nursing intervention action" to "next time step state" and generate a predicted state sequence;
[0088] Specifically, for any patient identified as , the state sequence is denoted as , and the nursing intervention action sequence aligned with it is denoted as , where is the uniform time step number and consistent with step S2, is the total number of uniform time steps, the adjacent time step states are first constructed into state pairs ( ) according to the uniform time steps, and the current time step state and the current time step nursing intervention action form the model input sample, and the next time step state is taken as the model output label;
[0089] To ensure the model trainability and cross-index scale consistency, the continuous index in the state vector is standardized according to the training set statistics with zero mean and unit variance, the enumeration or hierarchical index is encoded by one-hot encoding or trainable embedding encoding, and the encoded indexes are spliced according to the pre-defined order to form a fixed dimension state vector, and the state vector dimension is denoted as Meanwhile, to avoid the difficulty of directly inputting the discrete type and structured parameters of the nursing intervention action, the action encoding method consistent with step S3 is used to encode the nursing intervention action into an action encoding vector , and the action encoding vector dimension is denoted as , where the action encoding is generated based on at least the type and parameters of the nursing intervention action, and the default value and missing flag are used to represent the missing parameters to avoid ambiguity;
[0090] The state transition prediction model is implemented by using a multi-layer perception structure, denoted as , where is the model parameter set including the weight matrix and bias vector of each layer, the model input is the spliced vector of with the input dimension of , the network layer number is 3, the hidden layer width is 256 and 128 in turn, the activation function is ReLU, the output layer is a linear layer, and the output dimension is to obtain the prediction value with the same dimension as the next time step state , and the single-step prediction relationship is given by the following formula:
[0091] ;
[0092] where represents the patient at time step The predicted state at the next time step. Indicates by parameters Parameterized state transition prediction model Indicates the patient At time step The current time step state, Indicates the patient At time step The current time step nursing intervention action encoding vector;
[0093] During training, the state prediction loss is constructed based on the difference between the predicted value and the label. The mean squared error is used as the state prediction loss and averaged over mini-batch operations.
[0094] ;
[0095] in This represents the state prediction loss. This represents the set of patient-time step index pairs sampled within a small batch. Represents a set The number of elements, Represents the L2 norm, This represents the predicted state value at the next time step. Indicates the actual state at the next time step;
[0096] The optimizer uses Adam and the learning rate is... Batch size is 256, maximum number of training rounds is 50, and training is based on the validation set. Early stopping with a patience round count of 5 is implemented, while the gradient norm is pruned to 1.0 to improve training stability and to apply a modulo operation to the model weights. Weight decay is used to suppress overfitting;
[0097] After training, the trained state transition prediction model is used to perform rolling predictions to output a sequence of predicted states. Specifically, for each patient, the initial predicted state is set to... Subsequently to Gradually adjust the current predicted state Nursing intervention action codes corresponding to the time steps enter get and will The input for the next time step continues to roll until a complete predicted state sequence is obtained. wherein the continuous type index can be truncated within a clinically reasonable range during the rolling process to avoid numerical divergence caused by error accumulation, and the upper and lower bounds of the truncation are pre-configured by training data quantiles or clinical thresholds, thereby providing reproducible dynamic simulation inputs for the consistency joint training in step S6 and subsequent prospective evaluation of candidate nursing intervention action sequences.
[0098] In the present specific embodiment, S6 comprises:
[0099] consistency and stability of risk assessment under rolling prediction are improved through consistency joint training based on the time-to-event prediction model and the state transition prediction model;
[0100] wherein the parameter set of the time-to-event prediction model is denoted as , and the parameter set at least includes embedding table parameters of the action encoding layer, multi-layer perception parameters, time series modeling unit parameters, and risk rate output head parameters of each nursing outcome event type ; , and the parameter set of the state transition prediction model is denoted as ;
[0101] When training any mini-batch training sample set , first, a continuous time segment with a length of is extracted from each sample for rolling consistency constraints, wherein is a preset prediction step number and takes uniform time steps to balance the consistency constraint strength and error accumulation risk, then the starting real state of each time segment is taken as the rolling initial value, the aligned nursing intervention action sequence is input into the state transition prediction model to obtain the predicted state sequence , and after strictly aligning the predicted state sequence and the corresponding nursing intervention action sequence at uniform time steps, the predicted risk rate sequence is obtained by inputting the predicted state sequence and the corresponding nursing intervention action sequence into the time-to-event prediction model, wherein denotes the predicted state of the sample with patient identifier at time step , denotes the predicted risk rate of the event type based on the predicted state and the nursing intervention action;
[0102] In parallel, the real state sequence within the same time segment and the same nursing intervention action sequence are input into the time-to-event prediction model to obtain the reference risk rate sequence , wherein reference hazard rate based on the real state and the nursing intervention action;
[0103] To avoid the reference branch and the prediction branch chasing each other simultaneously, causing training instability, the reference hazard rate sequence is processed with a stop gradient when calculating the consistency loss, so that it only serves as a consistency supervision target and does not affect the intermediate calculation graph of the reference branch in reverse;
[0104] The consistency loss and the state prediction loss are combined into a joint loss according to a preset weight and are used to update and simultaneously, and the joint loss is defined as:
[0105]
[0106] wherein represents the joint loss, represents the state prediction loss and adopts the mean square error definition of step S5 and is obtained by subtracting and in the same time segment, represents the consistency loss and adopts the mean square error of the prediction hazard rate and the reference hazard rate, and respectively represent the state prediction loss weight and the consistency loss weight and take and to ensure that state dynamic learning is the main part and risk consistency is the auxiliary part, represents the number of small batches of samples and is equal to the number of time segments sampled in the set represents the set of nursing outcome event types, represents the number of event types, represents the rolling step sequence number in the time segment, represents the time segment index pair starting from the th uniform time step of the patient , and the difference between and
[0107] is used to depict the risk shift under rolling prediction; The Adam optimizer is used for joint backpropagation and synchronous update of the parameter sets , the batch size is 64, the number of joint training rounds is 20, and the gradient norm is clipped with a threshold of 1 to suppress the gradient explosion caused by the rolling link, so as to obtain the time-to-event prediction model after joint training and the state transition prediction model after joint training.
[0108] In the specific embodiment, S7 comprises:
[0109] based on the state sequence , the nursing intervention action sequence , and the immediate reward sequence , a state transition sample set corresponding to each time step is constructed for offline training of the nursing intervention strategy model;
[0110] wherein for any time step identified as , a state transition sample is constructed, wherein represents the current time step state vector of the patient at time step , and its index constitutes, the standardization manner and step S5 are consistent, represents the current time step nursing intervention action of the patient at time step , and its structured fields include the nursing intervention action type and the nursing intervention action parameters and are consistent with the action coding criteria of step S3, represents the immediate reward at time step and is obtained by accumulating the risk difference, represents the next time step real state vector of the patient at time step , and the training sample covers all time steps satisfying to ensure the existence of label ;
[0111] The nursing intervention strategy model is denoted as , wherein is a set of strategy model parameters, the strategy model input is the current time step state , and the output is at least one candidate nursing intervention action sequence , wherein is the preset prediction step number and is is the candidate sequence index and the number of candidates is to balance the search coverage and the amount of calculation;
[0112] In the embodiment, the policy model adopts the structure of "action type distribution + action parameter regression" to cover both discrete nursing intervention action types and continuous nursing intervention action parameters, specifically, a two-layer multi-layer perception is used as a shared backbone network with hidden layer widths of 128 and 128 in sequence, the activation function is ReLU, then two types of output heads are divided out, the first output head outputs the probability distribution of the nursing intervention action type by Softmax to sample the nursing intervention action type, and the second output head outputs the nursing intervention action parameter vector matched with the nursing intervention action type and adopts Softmax to scale the continuous parameters to the service allowable range of each parameter, and the Softmax sampling of the enumerated parameters obtains the enumerated values, thereby forming an executable nursing intervention action .
[0113] At each parameter update, for any current time step state , a candidate nursing intervention action sequence of length is generated, and each candidate nursing intervention action sequence is input into the state transition prediction model trained jointly . The state transition prediction model trained jointly is used for rolling prediction to obtain a candidate predicted state sequence , where is the parameter set of the state transition prediction model trained jointly and is consistent with step S6, initialized as to ensure consistency of the rolling link and subsequent planning evaluation; Subsequently, the candidate predicted state sequence and the candidate nursing intervention action sequence are input into the time-to-event prediction model trained jointly to obtain a candidate risk rate sequence
[0114] , where is the parameter set of the time-to-event prediction model trained jointly and is consistent with step S6, represents the predicted risk rate of the candidate sequence for the nursing outcome event type at time step . When evaluating the candidate sequence, the return construction method of step S4 is strictly reused to integrate the candidate risk rate sequence within a preset time window to obtain a candidate cumulative risk sequence and to obtain a candidate immediate return sequence
[0115] by adjacent difference, and the cumulative return of the candidate nursing intervention action sequence is calculated as:
[0116] .
[0117] wherein denotes the predicted cumulative return of a candidate sequence starting from state , denotes the predicted immediate return of a candidate sequence at time step , denotes the step index within a candidate sequence, denotes the preset prediction step number, denotes the discount factor and takes to emphasize recent risk improvement and take into account long-term effects;
[0118] The policy model is updated with the goal of maximizing the cumulative return, and the optimization objective function is defined as:
[0119] .
[0120] wherein denotes the expected objective function of the policy model, denotes the expectation of the sampled patient and time step index pair in the training sample, denotes the expectation of the candidate sequence of care intervention action distribution generated by the policy model under the given state, denotes the predicted cumulative return of the corresponding candidate sequence.
[0121] In the implementation, an elite sample-based policy improvement is adopted to ensure training stability. Specifically, in each iteration, the top candidate sequences with the largest cumulative return are selected from the candidate sequences as an elite set, and take 4, and the update of is completed by minimizing the negative log-likelihood of the elite set, which is equivalent to maximizing the generation probability of the elite set under , and the mean square error supervision of the elite sample is adopted for the action parameter regression head to converge the parameter output to the high return area;
[0122] The training adopts the Adam optimizer, the learning rate is taken as , the batch size is taken as 128, the training round is taken as 30, and the gradient norm of the policy network is clipped by 1.0 to avoid high-variance updates introduced by rolling evaluation, thereby obtaining the trained care intervention policy model used for online recommendation in step S8.
[0123] In the specific embodiment, S8 includes:
[0124] For the target patient, the identification is In the online recommendation scenario, the current state of the target patient at the current unified time step is first collected, and a current state vector is constructed according to the same caliber as in the training stage , wherein the index set and the field order are consistent with steps S2 to S5, the continuous index is standardized to zero mean and unit variance using the training set statistics, the enumeration or hierarchical index is encoded using one-hot encoding or embedding encoding, and the missing index is filled with the previous value, linear interpolation or a preset value according to the fixed strategy in the training stage, and the missing flag is retained to avoid mistaking the missing value as a real value;
[0125] The current state is input into the trained nursing intervention strategy model to generate at least one candidate nursing intervention action sequence, and in the embodiment, the number of candidates is set to and the length of each candidate nursing intervention action sequence is a preset prediction step number , wherein is the parameter set of the nursing intervention strategy model, the candidate nursing intervention action sequence is denoted as is the candidate sequence index, and is the step index in the candidate sequence, and the nursing intervention action each contains a nursing intervention action type and a nursing intervention action parameter, and the action type is valued in a preset action type set, and the action parameter is constrained within the business allowed range;
[0126] After the candidate sequence is generated, optional safety filtering is performed to remove candidate nursing intervention action sequences that do not conform to the nursing specification or conflict with the patient's contraindication information, and the contraindication information at least includes allergy history, postoperative contraindicated body position, drug interaction and medical order conflict markers. The safety filtering rules are fixed in the form of a configurable table and maintained by nursing management personnel when the system is deployed;
[0127] For each retained candidate nursing intervention action sequence , the candidate nursing intervention action sequence and the current state are input into the jointly trained state transition prediction model to perform rolling prediction to obtain a candidate predicted state sequence , wherein is the parameter set of the jointly trained state transition prediction model, is initialized as , and and are input into to obtain to ensure that the rolling link is consistent with the training stage;
[0128] inputting the candidate prediction state sequence and the candidate care intervention action sequence into the jointly trained time-to-event prediction model at a uniform time step outputting a candidate risk rate sequence, wherein the candidate risk rate sequence covers a set of care outcome event types for the jointly trained time-to-event prediction model parameter set and outputs a time-step-wise risk rate for each event type;
[0129] Subsequently, strictly reuse the cumulative risk and immediate return construction mode of step S4, and construct the cumulative return of the candidate sequence at a preset time window length and a uniform time step length Slide the candidate risk rate sequence to obtain a candidate cumulative risk sequence, and perform adjacent difference to obtain a candidate immediate return sequence , wherein represents the predicted immediate return of the candidate sequence at time step ;
[0130] On this basis, the predicted cumulative return of each candidate care intervention action sequence is calculated:
[0131] ;
[0132] , wherein represents the predicted cumulative return of the target patient at the current time step executing the candidate sequence , represents a discount factor and takes to take into account short-term and long-term risk improvement, represents a summation operation on the return of each step in the candidate sequence;
[0133] After obtaining the predicted cumulative return of all candidate sequences, the candidate care intervention action sequence with the optimal cumulative return is selected and its first care intervention action is output as the recommended care intervention action, and the selection rule is:
[0134] and ;
[0135] , wherein represents the optimal candidate sequence index, represents an index operation that takes the maximum value of the objective function, represents the predicted cumulative return of the candidate sequence , represents the recommended care intervention action output to the care provider or care information system, represents a first nursing intervention action of an optimal candidate nursing intervention action sequence at a current time step.
[0136] The above merely describes a preferred specific embodiment of the present application, but the protection scope of the present application is not limited thereto, and any person skilled in the art can make equivalent replacements or changes to the technical solution and the inventive concept of the present application within the technical scope disclosed by the present application, which should be covered within the protection scope of the present application.
[0137] The present application models the nursing intervention optimization problem as a sequential decision problem, and in view of the characteristics of nursing outcomes such as complications and unplanned readmission, i.e., "delayed return, censored, and difficult to evaluate step by step", a time-to-event prediction model is used to output a risk rate sequence at each time step, and then the cumulative risk is obtained by integrating within a preset time window, and the immediate return is constructed by adjacent difference, so that the influence of nursing action on future outcome risk can be converted into a return signal that can be accumulated and compared step by step; On this basis, a state transition prediction model is introduced to rollingly predict the dynamic evolution of "state-action-state", and the candidate nursing intervention action sequence is evaluated in a model-assisted manner, so as to realize the effect quantification and optimization selection of different nursing intervention sequences, improve the evaluability and executability of recommended intervention decision, and thus achieve the technical effects of reducing the risk of adverse events and improving the optimization efficiency of nursing intervention.
[0138] The present application improves the algorithm structure in view of the above technical problems: firstly, the cumulative risk based on risk rate integration and differential return remodeling is used to move the delayed outcome signal forward and refine it to each time step, which alleviates the problems of sparse return and evaluation lag in reinforcement learning training; secondly, the time-to-event prediction model sets action coding and realizes action-conditioned risk rate output, so that the model can depict the conditional influence of nursing intervention action type and parameter on risk rate, and improve the ability to distinguish the advantages and disadvantages of candidate action sequence; thirdly, the predicted risk rate is calculated by the predicted state sequence generated by the world model, and the risk rate obtained based on the real state sequence is introduced into the consistency loss, and the state prediction loss is jointly trained to form a joint loss, which enhances the consistency and stability of state prediction and risk prediction in the rolling planning process, so as to more reliably support the prediction evaluation and recommended output based on the candidate action sequence.
Claims
1. A method for optimizing breast and nail care intervention based on reinforcement learning, characterized in that, include: S1. Collect nursing process data and nursing outcome event data for patients with breast and thyroid cancer; S2. Time-align the nursing process data and process missing data to obtain the state sequence and the nursing intervention action sequence aligned with it. Based on the nursing outcome event data, determine the nursing outcome event type, the occurrence time of the nursing outcome event and the censoring mark corresponding to the state sequence. S3. Train the time-to-event prediction model based on the state sequence, nursing intervention action sequence, nursing outcome event type, nursing outcome event occurrence time, and censoring markers, so that it outputs a hazard rate sequence based on the state sequence and nursing intervention action sequence; S4. Integrate the hazard rate sequence within a preset time window to obtain a cumulative risk sequence, and subtract the cumulative risks of adjacent time steps to obtain an immediate reward sequence; S5. Train the state transition prediction model based on the state sequence and nursing intervention action sequence, so that it outputs the predicted state value of the next time step when the current time step state and the current time step nursing intervention action are input, and obtains the predicted state sequence through rolling prediction; S6. Input the predicted state sequence and nursing intervention action sequence into the time-to-event prediction model to obtain the predicted hazard rate sequence, calculate the consistency loss with the hazard rate sequence, combine it with the state prediction loss of the state transition prediction model to form a joint loss, update the parameters of the time-to-event prediction model and the state transition prediction model based on the joint loss, and obtain the jointly trained time-to-event prediction model and the state transition prediction model; S7. Train the nursing intervention strategy model using the immediate reward sequence and with the goal of maximizing the cumulative reward; S8. Collect the current state of the target patient and input the trained nursing intervention strategy model, output at least one candidate nursing intervention action sequence, and select the first nursing intervention action in the candidate nursing intervention action sequence with the best cumulative return as the recommended nursing intervention action.
2. The method for optimizing breast and thyroid care intervention based on reinforcement learning according to claim 1, characterized in that, S1 includes: Nursing process data and nursing outcome event data corresponding to each breast and thyroid patient are collected. The nursing process data includes time-stamped patient status data and time-stamped nursing intervention data. The patient status data includes at least one or more of the following: vital signs data, laboratory test data, drainage-related data, pain score data, and incision assessment data; The nursing intervention data includes at least one or more of the following: wound care actions, drainage tube care actions, analgesia care actions, functional exercise guidance actions, retesting arrangement actions, and follow-up reminder actions. Each nursing intervention data item includes at least the nursing intervention action type and nursing intervention action parameters. The nursing outcome event data includes at least the nursing outcome event type, the nursing outcome event occurrence time, and censoring markers. The nursing outcome event type includes at least one or more of the following: hypocalcemia-related adverse events, incision infection, subcutaneous effusion, lymphedema, and unplanned readmission. The nursing outcome event occurrence time is the time when the event corresponding to the nursing outcome event type occurs. The censoring markers are used to indicate whether the occurrence of the event corresponding to the nursing outcome event type was observed within the preset observation period.
3. The method for optimizing breast and thyroid care intervention based on reinforcement learning according to claim 1, characterized in that, S2 include: Determine the time step size of the unified time step, as well as the observation start time and observation end time, and use the observation start time as the time zero point; The nursing process data is discretized according to the unified time step to obtain the patient status data for each unified time step, thus forming a status sequence. When multiple patient status data exist within the same unified time step, the same indicator is summarized by averaging or taking the most recent record. Missing patient status data in the nursing process data is processed to obtain a status sequence with patient status data at each unified time step. The missing data processing includes at least one of the following: previous value imputation, linear interpolation, and preset value imputation. Nursing intervention action data is merged according to the unified time step to obtain the patient status data for each unified time step. Nursing intervention action data for corresponding time steps are determined to form a nursing intervention action sequence corresponding to the state sequence for each time step. When multiple nursing intervention action data exist within the same unified time step, the multiple nursing intervention action data are merged into the nursing intervention action data corresponding to that unified time step. Nursing outcome event type, nursing outcome event occurrence time, and censoring mark are determined based on nursing outcome event data. The nursing outcome event occurrence time is the time interval between the event corresponding to the nursing outcome event type and the observation start time. The censoring mark is used to indicate whether the occurrence of the event corresponding to the nursing outcome event type was observed before the observation end time.
4. The method for optimizing breast and thyroid care intervention based on reinforcement learning according to claim 1, characterized in that, S3 include: A supervised training method is used to train a time-to-event prediction model, employing state sequences, nursing intervention action sequences, nursing outcome event types, nursing outcome event occurrence times, and censoring markers. The model includes an action encoding layer and a hazard output layer. The action encoding layer encodes the nursing intervention action sequences to obtain action encoding sequences corresponding to each time step of the state sequence. This action encoding layer generates the action encoding sequences based at least on the nursing intervention action type and nursing intervention action parameters. The hazard output layer takes the patient state data and the corresponding action codes at each time step as input and outputs the hazard rates for each time step corresponding to the nursing outcome event type, thereby obtaining a hazard rate sequence. The process involves determining the event occurrence time step based on the occurrence time of the nursing outcome event, and censoring the training samples based on the censoring markers to construct a loss function for the time-to-event prediction model. For uncensored training samples, the loss function includes the event likelihood term corresponding to the hazard rate at the event occurrence time step and the survival likelihood term for each time step before the event occurrence time step. For censored training samples, the loss function includes the survival likelihood term for each time step before the censored time step. The loss function is then used to update the parameters of the time-to-event prediction model, resulting in a trained time-to-event prediction model. This trained model is then used to perform forward calculations on the state sequence and the nursing intervention action sequence to output a hazard rate sequence.
5. The method for optimizing breast and thyroid care intervention based on reinforcement learning according to claim 1, characterized in that, S4 include: The time window length corresponding to the preset time window and the time step length of the unified time step are determined. A sliding integral operation is performed on the hazard rate sequence according to the time window length to obtain the cumulative risk sequence. In discrete time steps, the cumulative risk for any time step is the sum of the products of the hazard rate of each time step from that time step to the last time step corresponding to the time window length and the time step length. The cumulative risk sequence is then subjected to a difference operation between adjacent time steps to construct an instantaneous return sequence. The instantaneous return for any time step is the cumulative risk of the previous time step minus the cumulative risk of the current time step, with the instantaneous return corresponding to a decrease in cumulative risk being a positive value. When the nursing outcome event type includes multiple event types, the cumulative risk sequence corresponding to each event type is calculated separately, and the cumulative risk sequences corresponding to each event type are weighted and summed according to preset event weights to obtain the cumulative risk sequence used for the difference operation.
6. The method for optimizing breast and thyroid care intervention based on reinforcement learning according to claim 1, characterized in that, S5 include: Using a state sequence and a nursing intervention sequence as input, the state sequence is divided into state pairs of adjacent time steps according to a unified time step. The state of the current time step and the nursing intervention action of the current time step are used as the model input sample, and the state of the next time step is used as the model output label. A state transition prediction model is trained based on the model input sample and the model output label, so that when the state of the current time step and the nursing intervention action of the current time step are input, the state transition prediction model outputs the predicted value of the state of the next time step corresponding to the current time step. The difference between the predicted value of the state of the next time step and the state of the next time step is used to construct a state prediction loss, and the parameters of the state transition prediction model are updated based on the state prediction loss to obtain the trained state transition prediction model. The trained state transition prediction model is used to perform rolling prediction on the state sequence and the nursing intervention sequence. Specifically, at any time step, the state of the time step and the nursing intervention action of the time step are input into the trained state transition prediction model to obtain the predicted value of the state of the next time step, and the predicted value of the state of the next time step is used as the state of the next time step to continue rolling prediction, outputting the predicted state sequence.
7. The method for optimizing breast and thyroid care intervention based on reinforcement learning according to claim 1, characterized in that, S6 include: After aligning the predicted state sequence and the nursing intervention sequence at a unified time step, input them into the trained time-to-event prediction model and output the predicted hazard rate sequence. Calculate the consistency loss based on the predicted hazard rate sequence and the hazard rate sequence, where the consistency loss includes the mean square error loss or divergence loss between the predicted hazard rate sequence and the hazard rate sequence. Obtain the state prediction loss and combine the consistency loss and the state prediction loss according to preset weight coefficients to form a joint loss. Update the parameters of the trained time-to-event prediction model and the trained state transition prediction model based on the joint loss to obtain a jointly trained time-to-event prediction model and a jointly trained state transition prediction model.
8. The method for optimizing breast and thyroid care intervention based on reinforcement learning according to claim 1, characterized in that, S7 includes: Based on the state sequence and nursing intervention action sequence, construct state transition samples corresponding to each time step, so that each state transition sample includes the current time step state, the current time step nursing intervention action, the corresponding time step immediate report, and the next time step state. Using the state transition samples as reinforcement learning training samples, a nursing intervention strategy model is initialized. During training, for any current time step state, the nursing intervention strategy model outputs at least one candidate nursing intervention action sequence. The candidate nursing intervention action sequence and the current time step state are input into a jointly trained state transition prediction model, and a preset number of prediction steps are performed for rolling prediction to obtain a candidate predicted state sequence corresponding to the candidate nursing intervention action sequence. The candidate predicted state sequence is input into a jointly trained time-to-event prediction model to output a candidate risk rate sequence, and the cumulative reward corresponding to the candidate nursing intervention action sequence is calculated according to the construction method of the cumulative risk sequence and the immediate reward sequence. An optimization objective function for the nursing intervention strategy model is constructed with the goal of maximizing the cumulative return. The parameters of the nursing intervention strategy model are then updated based on the optimization objective function to obtain the trained nursing intervention strategy model.
9. The method for optimizing breast and thyroid care intervention based on reinforcement learning according to claim 1, characterized in that, S8 includes: The current state of the target patient is collected and input into a trained nursing intervention strategy model, which outputs at least one candidate nursing intervention action sequence. For each candidate nursing intervention action sequence, the candidate nursing intervention action sequence and the current state are input into a jointly trained state transition prediction model, which performs a rolling prediction of a preset number of prediction steps, and outputs a candidate predicted state sequence corresponding to the candidate nursing intervention action sequence. The candidate predicted state sequence is input into a jointly trained time-to-event prediction model, which outputs a candidate risk rate sequence, and calculates the cumulative reward corresponding to the candidate nursing intervention action sequence according to the construction method of the cumulative risk sequence and the immediate reward sequence. The first nursing intervention action in the candidate nursing intervention action sequence with the best cumulative reward is determined as the recommended nursing intervention action, and the recommended nursing intervention action is output.
10. A reinforcement learning-based nail and breast care intervention optimization system, used to execute the reinforcement learning-based nail and breast care intervention optimization method according to any one of claims 1 to 9, comprising: The data processing module is used to collect nursing process data and nursing outcome event data of patients with breast and thyroid diseases, and to perform time alignment and missing data processing to generate state sequences, nursing intervention action sequences aligned with them, and nursing outcome event types, event occurrence times and censoring marks corresponding to the state sequences. The time-to-event prediction module is used to train a time-to-event prediction model based on the state sequence, nursing intervention action sequence, event occurrence time and censoring markers, and output a hazard rate sequence. The reward construction module is used to integrate the hazard rate sequence within a preset time window to obtain a cumulative risk sequence, and generate an instant reward sequence through adjacent differences. The state transition prediction module is used to train the state transition prediction model to predict the state at the next time step and to generate a sequence of predicted states in a rolling manner. The joint training module is used to obtain a predicted hazard rate sequence based on the predicted state sequence, calculate a consistency loss with the hazard rate sequence, and combine the consistency loss with the state prediction loss into a joint loss to jointly update the time-to-event prediction model and the state transition prediction model. The strategy training module is used to train the nursing intervention strategy model using the immediate reward sequence, and evaluate the cumulative reward of the candidate nursing intervention action sequence based on the jointly trained time-to-event prediction model and state transition prediction model, so as to update the nursing intervention strategy model with the goal of maximizing the cumulative reward. The intervention recommendation module is used to input the current state of the target patient into the trained nursing intervention strategy model to obtain at least one candidate nursing intervention action sequence, and select the first nursing intervention action in the candidate nursing intervention action sequence with the best cumulative return as the recommended nursing intervention action output.