Multi-agent reinforcement learning joint decision method and system for whole course collaborative diagnosis and treatment
Patent Information
- Application Number
- CN202611032545.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-13
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2046-07-13
AI Technical Summary
这种震荡和冲突使得生成的综合管理策略在过渡期无法平滑衔接,最终导致护理人员无法执行一致的全病程托管方案,延误患者恢复进程
[0018]Compared to existing technologies, the advantages of this invention are as follows: By constructing a phased state observation tensor and performing disease stage switching detection based on Kullback-Leibler divergence, this invention effectively solves the problem of blurred stage boundaries caused by non-stationary changes in multi-source heterogeneous data during the transition period across disease stages. This provides a stable time reference for subsequent multi-agent collaborative decision-making and avoids false alarms caused by single abnormal sampling values. Furthermore, the phased state observation tensor is decoupled into local observation sub-tensors according to the agent's responsibility domain, allowing nutritionists, sports rehabilitation therapists, and health managers to independently output action probability distributions. A dynamic weight vector is then obtained through the coupling calculation of the intra-stage action consistency matrix and the confidence level of the disease stage offset. This ensures that the weight distribution is nearly uniform in weak switching scenarios and concentrated towards agents with stable strategies in the new stage during strong switching scenarios, thus mechanistically suppressing conflicts and oscillations in intervention parameters during the transition period. Moreover, the phased experience replay buffer is used to store transformed tuples in disease stage partitions, ensuring that the gradient update of the strategy network is always aligned with the current stage state distribution, eliminating parameter update lag and cumulative errors caused by cross-stage sample mixing. Simultaneously, the dynamic weight vector is recalculated synchronously after the strategy parameters are updated, forming a closed feedback loop from strategy update to weight response within a single time step. This ensures that the joint decision-making actions obtain a weight benchmark consistent with the latest strategy at the stage switching moment. The synergistic effect of these mechanisms enables smooth connection and adaptive adjustment of multidimensional intervention strategies in the whole course of disease management, improving the decision consistency and clinical feasibility of complex disease management solutions in the field of healthcare informatics.
Smart Images

Figure CN122552185B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of medical intelligent decision-making technology, and in particular to a multi-agent reinforcement learning joint decision-making method and system for collaborative diagnosis and treatment throughout the entire disease course. Background Technology
[0002] With the development and application of artificial intelligence technology in the healthcare field, intelligent decision-making systems based on full-course disease management are increasingly being used to assist in generating intervention plans for patients at each stage from hospitalization to rehabilitation. This system utilizes a single decision model, receiving real-time physiological state data and daily activity data input by the user. By analyzing multi-dimensional features, it outputs intervention parameters, including nutritional intake and exercise rehabilitation, to achieve continuous monitoring and dynamic management of the patient's condition. This model typically employs time-series prediction methods, calculating and updating various management parameters based on the current data distribution characteristics to generate a unified full-course disease management plan.
[0003] In practical applications, when patients transition from hospitalization to home rehabilitation, healthcare professionals or operators have observed inconsistent system responses to the same patient's physiological data within the same timeframe. For example, in response to the gradually increasing trend of heart rate variability, the model's nutritional intake parameters recommend increasing protein intake, while exercise rehabilitation parameters suggest reducing activity levels. These two parameters cancel each other out within the same timeframe, resulting in increased tube feeding nutrition but decreased rehabilitation training intensity during actual implementation. With continuous data input, output parameters exhibit repeated oscillations, meaning today's recommended nutritional ratios differ significantly from yesterday's plan, and exercise guidance parameters frequently change. This oscillation and conflict prevent a smooth transition in the generated integrated management strategy, ultimately leading to an inability of nursing staff to implement a consistent full-course care plan, delaying the patient's recovery. Therefore, reducing output conflicts of intervention parameters during the transition period in AI-based full-course management is a pressing issue that needs to be addressed in this field. Summary of the Invention
[0004] To address the shortcomings of existing technologies, the present invention aims to provide a multi-agent reinforcement learning joint decision-making method and system for collaborative diagnosis and treatment throughout the entire disease course. This method involves acquiring multi-dimensional physiological state data of the patient throughout the entire disease course to construct a phased state observation tensor, performing disease stage switching detection to generate a disease stage label sequence, extracting local observation sub-tensors for each agent and independently outputting action probability distribution vectors, using a cross-agent action state circular queue to calculate the action consistency matrix and dynamic weight vector within each stage, weighted aggregation to generate a joint decision action distribution, and storing it in a phased experience replay buffer for policy network updates. Iterative decision-making based on stage exit conditions enables dynamic adaptation to changes in disease stages through multi-agent collaborative decision-making, improving the accuracy and robustness of time-series decision-making throughout the entire disease course.
[0005] To achieve the above objectives, the present invention adopts the following technical solution:
[0006] In a first aspect, the present invention provides a multi-agent reinforcement learning joint decision-making method for collaborative diagnosis and treatment throughout the entire disease course, comprising: acquiring multidimensional physiological state data of a patient throughout the entire disease course and constructing a phased state observation tensor; performing disease course phase switching detection on the phased state observation tensor to generate a disease course phase label sequence; the multidimensional physiological state data throughout the entire disease course includes real-time physiological indicator data and disease state data; extracting local observation sub-tensors of each agent from the phased state observation tensor according to the disease course phase label sequence; using the local observation sub-tensors as input, driving each agent to independently output an action probability distribution vector, and writing the action probability distribution vectors of each agent into a cross-agent action state circular queue; and reading the action probabilities of each agent from the cross-agent action state circular queue. The distribution vector is used to calculate the action consistency matrix within the stage; based on the action consistency matrix within the stage and the confidence level of the disease stage offset, the dynamic weight vector of each agent is calculated; the action probability distribution vector of each agent is weighted and aggregated using the dynamic weight vector to generate a joint decision action distribution; the execution action sequence of the current time step is sampled according to the joint decision action distribution, and the execution action sequence and the patient's state feedback for the next time step are combined to form a transformation tuple and written into the staged experience replay buffer; the strategy network execution parameters of each agent are updated according to the transformation tuple in the staged experience replay buffer; the dynamic weight vector is updated according to the new round of action probability distribution vector after parameter update, and a determination is made based on the stage exit condition, outputting the full disease course decision record table and continuing the sequential decision loop.
[0007] Furthermore, the acquisition of multidimensional physiological state data of the patient throughout the entire disease course and the construction of a phased state observation tensor includes: continuously acquiring the patient's real-time physiological indicator data and disease state data; aligning the acquired real-time physiological indicator data and disease state data according to the acquisition timestamp, and arranging them along three dimensions: time axis, indicator category axis, and data modality axis, to construct a phased state observation tensor.
[0008] Further, the step of performing disease stage switching detection on the phased state observation tensor and generating a disease stage labeling sequence includes: extracting the statistical distribution features of each indicator within the sliding window step by step along the time axis of the phased state observation tensor; calculating the Kullback-Leibler divergence of the statistical distribution features of each indicator between adjacent non-overlapping sliding windows to obtain a step-by-step distribution offset sequence; performing disease stage switching detection based on the relationship between the step-by-step distribution offset sequence and a preset stage switching judgment threshold; if the distribution offset of a certain time step is less than the stage switching judgment threshold, then the time step is marked as a continuation frame of the current disease stage; if the distribution offset of a certain time step is greater than or equal to the stage switching judgment threshold, then the time step is marked as a disease stage switching trigger frame, and a new disease stage number is reassigned to the time steps after the switching trigger frame; and summarizing the labeling results of all frames in time step order to generate a disease stage labeling sequence.
[0009] Further, the calculation of the Kullback-Leibler divergence of the statistical distribution characteristics of each index between adjacent non-overlapping sliding windows to obtain the time-step distribution offset sequence includes: for each index item within the sliding window sub-tensor, collecting all values of the index item within the sliding window along the time axis, dividing it into equal-width histogram intervals according to a preset number of bins, counting the frequency of values in each interval and normalizing it to obtain the statistical distribution feature vector of the index item under the current sliding window; calculating the Kullback-Leibler divergence of the statistical distribution feature vector of the same index item under the sliding window at the corresponding starting position of the previous non-overlapping observation period for the current time step; averaging the Kullback-Leibler divergence of all index items along the index category axis to obtain the distribution offset of the current time step; and outputting the time-step distribution offset sequence after traversing all time steps of the phased state observation tensor along the time axis.
[0010] Further, the step of extracting the local observation sub-tensor of each Agent from the staged state observation tensor according to the disease stage marker sequence includes: dividing the staged state observation tensor along the indicator category axis according to the Agent's responsibility domain based on the disease stage marker sequence; enabling the nutritionist Agent to read nutrition-related indicator slices, the exercise rehabilitation therapist Agent to read exercise function-related indicator slices, and the health manager Agent to read chronic disease comprehensive management-related indicator slices; and recording each slice as the local observation sub-tensor of the corresponding Agent.
[0011] Furthermore, the step of writing the action probability distribution vector of each Agent into the cross-Agent action state circular queue includes: when the disease stage marking sequence detects a disease stage switching trigger frame at the current time step, clearing the cross-Agent action state circular queue and resetting the slot counter, so that the cross-Agent action state circular queue only retains the action probability distribution vector within the current new stage; when the disease stage marking sequence is marked as a continuation frame at the current time step, appending the action probability distribution vector of each Agent to the next slot of the cross-Agent action state circular queue.
[0012] Furthermore, the calculation of the intra-stage action consistency matrix includes: reading the action probability distribution vectors of all written slots in the current stage from the cross-Agent action state circular queue; calculating the pairwise cosine similarity of the action probability distribution vectors of the same Agent at each time step in the current stage to construct the intra-stage action self-consistency score of the Agent; calculating the cross-Agent action cosine similarity of the action probability distribution vectors of any two different Agents; and arranging the intra-stage action self-consistency scores and cross-Agent action cosine similarities of each Agent into a square matrix to obtain the intra-stage action consistency matrix.
[0013] Furthermore, the step of calculating the dynamic weight vector of each Agent based on the intra-stage action consistency matrix and the disease stage offset confidence includes: taking the distribution offset corresponding to the current disease stage switching trigger frame, subtracting it from the stage switching judgment threshold, and then normalizing it to obtain the disease stage offset confidence; using the sum of each row of the intra-stage action consistency matrix as the basic weight vector, multiplying the basic weight vector with the disease stage offset confidence, and then performing softmax normalization to obtain the dynamic weight vector of each Agent.
[0014] Furthermore, the step of weighting and aggregating the action probability distribution vectors of each Agent using the dynamic weight vector to generate a joint decision action distribution includes: pre-setting a conservative baseline distribution vector within the corresponding Agent's responsibility domain, performing weighted interpolation fusion to obtain a weighted action probability distribution vector; performing argmax sampling on the weighted action probability distribution vector of each Agent, and taking the action option number corresponding to the maximum probability as the action to be executed by that Agent at the current time step; arranging the execution actions of the three Agents in order of Agent number to obtain the execution action sequence at the current time step; the joint decision action distribution is a composite vector formed by concatenating the three weighted action probability distribution vectors.
[0015] Furthermore, the step of writing the execution action sequence and the patient's state feedback at the next time step into a transformation tuple and then writing it into the phased experience replay buffer includes: constructing a phased experience replay buffer, wherein the phased experience replay buffer is stored in partitions according to the disease stage number, and each partition independently maintains a first-in-first-out queue; packaging the phased state observation tensor of the current time step, the execution action sequence, the patient's state feedback at the next time step, and the current value of the disease stage marker sequence into a transformation tuple; reading the disease stage number corresponding to the current time step from the disease stage marker sequence, and appending the transformation tuple to the partition in the phased experience replay buffer corresponding to the current disease stage number.
[0016] Furthermore, the process of updating the dynamic weight vector based on the new round of action probability distribution vector after parameter updates, determining the exit condition, outputting the full-process decision record table, and continuing the sequential decision loop includes: after each Agent policy network completes parameter updates, it re-executes forward inference to obtain the updated action probability distribution vector; it replaces the original action probability distribution vector in the current slot of the cross-Agent action state circular queue with the updated action probability distribution vector, and re-executes the action consistency matrix calculation and dynamic weight vector update within the stage; if the stage exit condition is met, it writes the final dynamic weight vector, joint decision action distribution, and stage number of the current stage into the full-process decision record table, and freezes the parameter updates and dynamic weight recalculation of each Agent policy network; if the stage exit condition is not met, it marks the current time step as a continuation frame and returns to the previous step to continue executing the sequential decision of the next time step within the current stage.
[0017] Secondly, this invention provides a multi-agent reinforcement learning joint decision-making system for collaborative diagnosis and treatment throughout the entire disease course, comprising: a tensor construction and stage division module, used to acquire multidimensional physiological state data of the patient throughout the entire disease course and construct a staged state observation tensor; perform disease course stage switching detection on the staged state observation tensor to generate a disease course stage label sequence; the multidimensional physiological state data throughout the entire disease course includes real-time physiological indicator data and disease state data; a local inference and queue writing module, used to extract local observation sub-tensors of each agent from the staged state observation tensor according to the disease course stage label sequence; drive each agent to independently output an action probability distribution vector using the local observation sub-tensors as input, and write the action probability distribution vectors of each agent into a cross-Agent action state circular queue; and a consistency analysis and weight allocation module, used to read each agent from the cross-Agent action state circular queue. The system calculates the action probability distribution vector of each agent and the action consistency matrix within each phase. Based on the action consistency matrix and the confidence level of the disease phase offset, it calculates the dynamic weight vector for each agent. The action aggregation and experience replay module performs weighted aggregation on the action probability distribution vector of each agent using the dynamic weight vector to generate a joint decision action distribution. It samples the action sequence for the current time step based on the joint decision action distribution and writes the action sequence and the patient's state feedback for the next time step into a transformation tuple in the phased experience replay buffer. It updates the strategy network execution parameters for each agent based on the transformation tuple in the phased experience replay buffer. The loop feedback and output archiving module updates the dynamic weight vector based on the updated action probability distribution vector, determines the exit condition based on the phase, outputs a full-process decision record table, and continues the sequential decision loop.
[0018] Compared to existing technologies, the advantages of this invention are as follows: By constructing a phased state observation tensor and performing disease stage switching detection based on Kullback-Leibler divergence, this invention effectively solves the problem of blurred stage boundaries caused by non-stationary changes in multi-source heterogeneous data during the transition period across disease stages. This provides a stable time reference for subsequent multi-agent collaborative decision-making and avoids false alarms caused by single abnormal sampling values. Furthermore, the phased state observation tensor is decoupled into local observation sub-tensors according to the agent's responsibility domain, allowing nutritionists, sports rehabilitation therapists, and health managers to independently output action probability distributions. A dynamic weight vector is then obtained through the coupling calculation of the intra-stage action consistency matrix and the confidence level of the disease stage offset. This ensures that the weight distribution is nearly uniform in weak switching scenarios and concentrated towards agents with stable strategies in the new stage during strong switching scenarios, thus mechanistically suppressing conflicts and oscillations in intervention parameters during the transition period. Moreover, the phased experience replay buffer is used to store transformed tuples in disease stage partitions, ensuring that the gradient update of the strategy network is always aligned with the current stage state distribution, eliminating parameter update lag and cumulative errors caused by cross-stage sample mixing. Simultaneously, the dynamic weight vector is recalculated synchronously after the strategy parameters are updated, forming a closed feedback loop from strategy update to weight response within a single time step. This ensures that the joint decision-making actions obtain a weight benchmark consistent with the latest strategy at the stage switching moment. The synergistic effect of these mechanisms enables smooth connection and adaptive adjustment of multidimensional intervention strategies in the whole course of disease management, improving the decision consistency and clinical feasibility of complex disease management solutions in the field of healthcare informatics. Attached Figure Description
[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0020] Figure 1 This is a flowchart of the multi-agent reinforcement learning joint decision-making method for collaborative diagnosis and treatment throughout the entire disease course in this invention;
[0021] Figure 2 This is a schematic diagram of the phased state observation tensor construction in an embodiment of the present invention;
[0022] Figure 3 This is a schematic diagram of sliding window statistical distribution feature extraction in an embodiment of the present invention;
[0023] Figure 4 This is a schematic diagram of the disease stage marker sequence in an embodiment of the present invention;
[0024] Figure 5This is a schematic diagram of local observation sub-tensor slice extraction in an embodiment of the present invention;
[0025] Figure 6 This is a schematic diagram of a circular queue of cross-Agent action states in an embodiment of the present invention;
[0026] Figure 7 This is a schematic diagram of the action consistency matrix within a stage in an embodiment of the present invention;
[0027] Figure 8 This is a schematic diagram of the phased experience playback buffer structure in an embodiment of the present invention;
[0028] Figure 9 This is a functional block diagram of the multi-agent reinforcement learning joint decision-making system for collaborative diagnosis and treatment throughout the entire disease course in this invention. Detailed Implementation
[0029] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0030] Example 1:
[0031] Please see Figure 1 As shown, this embodiment provides a multi-agent reinforcement learning joint decision-making method for collaborative diagnosis and treatment throughout the entire disease course, including:
[0032] S1: Obtain multidimensional physiological state data of the patient throughout the entire course of the disease, and construct a staged state observation tensor; perform disease stage switching detection on the staged state observation tensor to generate a disease stage marker sequence; the multidimensional physiological state data throughout the entire course of the disease includes real-time physiological indicator data and disease state data.
[0033] This step addresses the issue of non-stationary distribution changes in real-time physiological status and daily activity data during patient transitions across disease stages (such as the shift from hospitalization to home rehabilitation). At these transition points, the statistical characteristics of multi-source heterogeneous data undergo abrupt changes. Without a clear mechanism for locating stage boundaries, subsequent local observation partitioning and dynamic weight allocation by each agent will lack a stable time reference, leading to lag in parameter updates and accumulated temporal errors at the stage transition point. This step constructs a stage-specific state observation tensor to uniformly organize multimodal data and automatically detects disease stage transition boundaries based on distribution offset metrics, providing a consistent data foundation and transition signals for subsequent steps.
[0034] Further, step S1 includes:
[0035] S11: Continuously acquire real-time physiological indicator data and disease status data of patients. After aligning the acquired real-time physiological indicator data and disease status data according to the acquisition timestamp, arrange them along three dimensions: time axis, indicator category axis, and data modality axis to construct a phased state observation tensor.
[0036] This step aims to unify the multidimensional physiological state data from different sources, with different sampling rates, and different modalities throughout the entire disease course into a structured tensor that can be directly indexed by subsequent steps. If discrete data streams are directly used in the subsequent agent's policy inference, timestamp misalignment and modality confusion will lead to the failure of semantic alignment of different indicators at the same time step, thus making it impossible to stably execute stage switching detection and local observation partitioning.
[0037] Specifically, real-time physiological indicator data is obtained from the patient monitoring interface of the in-hospital monitoring system and the data reporting interface of the wearable device. The real-time physiological indicator data includes vital sign time-series signals, biochemical test value sequences, and motor function assessment score sequences. The vital sign time-series signals include heart rate, blood pressure, blood oxygen saturation, body temperature, and respiratory rate, which are reported by the wearable device at a sampling rate of 1Hz. The biochemical test value sequences include fasting blood glucose, glycated hemoglobin, four blood lipids, and liver and kidney function indicators, which are synchronously pushed by the in-hospital laboratory information system when the tests are completed. The motor function assessment score sequences include 6-minute walking distance, grip strength test values, and Berg balance scale scores, which are entered by the rehabilitation assessment terminal after each assessment.
[0038] Disease status data is obtained from the structured data interface of the hospital's electronic medical record system. The disease status data includes structured fields of medical records and diagnostic coding sequences. The structured fields of medical records include past medical history markers, present medical history summary tags, and comorbidity markers. The diagnostic coding sequences adopt the ICD-10 encoding format.
[0039] Real-time physiological indicator data and disease status data are aligned according to the collection timestamp: using a preset alignment time step Δt as the baseline granularity, multiple samples of each data source within the Δt range are averaged, and missing samples within the Δt range are filled using a forward imputation strategy, ensuring that all indicators have a unique value at the same time step. The alignment time step Δt is determined based on the clinical decision-making cycle of full-course disease management; for example, Δt can be set to 5 minutes.
[0040] The aligned data is arranged along three dimensions: the time axis length equals the number of time steps collected throughout the entire disease course (T); the indicator category axis length equals the total number of indicator items in all real-time physiological indicator data and disease state data (N); and the data modality axis length equals the total number of data source modalities (M). The resulting phased state observation tensor has a data structure of a three-dimensional floating-point tensor of size T×N×M. (See also...) Figure 2This is a schematic diagram of the phased state observation tensor construction provided in an embodiment of this application. For example... Figure 2 As shown, the left side establishes an alignment association between the acquired real-time physiological indicator data and disease state data according to the acquisition timestamp, and maps a three-dimensional data cube to the right based on the alignment time step. This data cube is expanded along three orthogonal directions and labeled as the time axis, indicator category axis, and data modality axis, respectively. After arrangement and combination, it forms the staged state observation tensor labeled on the right. Multidimensional physiological state data from different sources, with different sampling rates, and different modalities throughout the entire disease course are uniformly presented in this tensor organization structure. This method can uniformly organize multimodal data by constructing a staged state observation tensor, eliminating the risk of semantic alignment failure of different indicators at the same time step due to timestamp misalignment and modality confusion. It establishes semantically consistent data alignment coordinates and a stable time reference benchmark for subsequent local observation division within each agent stage and stage switching detection.
[0041] S12: Extract the statistical distribution characteristics of each index within the sliding window step by step along the time axis from the phased state observation tensor, calculate the Kullback-Leibler divergence of the statistical distribution characteristics of each index between adjacent non-overlapping sliding windows, and obtain the time-step distribution offset sequence.
[0042] This step aims to transform the non-stationary distribution of the phased state observation tensor along the time axis into a quantifiable scalar sequence, providing a single benchmark for determining the phase switch in S13. If the difference between the original observations is used directly as the switch signal, occasional fluctuations in single abnormal sampling values will be confused with the actual phase switch, causing the false alarm rate of switch detection to increase linearly with the sampling frequency.
[0043] Specifically, along the time axis, with a sliding window length W as the window range and a unit time step as the sliding step size, the sliding window sub-tensor corresponding to each time step is extracted from the staged state observation tensor; the sliding window length W is determined according to the ratio of the duration of the typical stage in the whole course of disease management to the aligned time step size Δt. For example, W can be set to 288, corresponding to the amount of data for one day when Δt = 5 minutes.
[0044] For each index term within the sliding window subtensor, all values of that index term within the sliding window are collected along the time axis. These values are then divided into equal-width histogram intervals according to a preset binning number B. The frequency of values within each interval is counted and normalized to a frequency, yielding the statistical distribution feature vector of that index term within the current sliding window. The preset binning number B is determined by rounding up the square root of the sliding window length W; for example, B can be set to 17. (See also...) Figure 3 This is a schematic diagram of the sliding window statistical distribution feature extraction provided in an embodiment of this application. For example... Figure 3As shown, the upper part displays the extraction of corresponding sliding window sub-tensors along the time direction from the dashed-line tensor spanning the time axis, i.e., the staged state observation tensor. The lower part transforms and reorganizes the extracted data into histogram intervals, establishing multiple sets of equal-width histogram intervals. Black solid frequency fill blocks represent the frequencies collected in the distribution, reflecting the normalized frequency according to a preset bin number, thus obtaining the statistical distribution feature vector of the index under the current sliding window. This graphical structure intuitively describes the entire process of transforming continuous time-separated slices of the tensor into discrete-frequency statistical distribution feature vectors. This structure transforms the non-stationary distribution changes of the staged state observation tensor along the time axis into a quantifiable scalar sequence, providing a consistent and single comparison benchmark. This allows the difference between the original observation values to be used as a switching signal benchmark during the Kullback-Leibler divergence feature extraction stage comparison, making the stage switching detection structurally insensitive to occasional fluctuations in single abnormal sampling values and effectively reducing the false alarm rate.
[0045] Calculate the Kullback-Leibler divergence for the statistical distribution feature vectors of the same index term under the sliding window of the current time step t and the sliding window of time step tW (i.e., the corresponding starting position of the previous non-overlapping observation period):
[0046]
[0047] in and These represent the frequencies of the i-th bin interval for this indicator item within the sliding window corresponding to time step t and tW, respectively. The two sliding windows do not overlap on the time axis. To prevent the smoothing constant from reaching zero logarithm, take By replacing adjacent time step comparisons with non-overlapping window comparisons, the dynamic range of KL divergence is made to cover the entire interval from stable changes within a stage to abrupt changes in distribution across stages, in line with the stage switching threshold τ and the upper bound of the distribution offset saturation. The magnitude setting is consistent with itself.
[0048] The Kullback-Leibler divergence of all indicators is averaged along the indicator category axis to obtain the distribution offset at time step t. After traversing all time steps of the phased state observation tensor along the time axis, the time-step distribution offset sequence is output. For the initial segment where time step t < W, the sliding window of the current time step t is compared with the sliding window of the 0th time step (filling the insufficient part with zero values) to ensure that the length of the output sequence is equal to T. The data structure of the time-step distribution offset sequence is a one-dimensional floating-point number sequence with a length of T.
[0049] S13: Based on the relationship between the distribution offset sequence of each time step and the preset stage switching judgment threshold τ, perform disease stage switching detection; if the distribution offset of a certain time step is less than the stage switching judgment threshold τ, then mark the time step as a continuation frame of the current disease stage; if the distribution offset of a certain time step is greater than or equal to the stage switching judgment threshold τ, then mark the time step as a disease stage switching trigger frame, and reassign a new disease stage number to the time steps after the switching trigger frame; summarize the marking results of all frames in time step order to generate a disease stage marking sequence.
[0050] This step aims to transform the time-step distribution offset sequence output by S12 into a discrete disease stage number sequence, so that subsequent steps S2 to S5 can perform data isolation and weight allocation based on clear stage boundaries.
[0051] Specifically, the method for determining the stage switching threshold τ is as follows: take the distribution offset samples corresponding to all known stage switching times in the historical full-course data, and the distribution offset samples corresponding to all non-switching times, and solve the optimal separation threshold as τ according to the criterion of maximizing the F1 score of the two classes of samples; for example, the stage switching threshold τ can be set to 0.25.
[0052] Initialize the disease stage number to 1, and traverse the time-step distribution offset sequence in time step order. For each time step t, if the value of the time-step distribution offset sequence at time step t is less than the stage switching judgment threshold τ, then time step t is marked as a continuation frame and assigned the current disease stage number. If the value of the time-step distribution offset sequence at time step t is greater than or equal to the stage switching judgment threshold τ, then time step t is marked as a disease stage switching trigger frame, the disease stage number is incremented by 1 and assigned to time step t, and the incremented disease stage number is used as the default number for all time steps after time step t.
[0053] Summarize the frame tag type and disease stage number of all time steps in time step order, and output the disease stage tag sequence. The data structure of the disease stage tag sequence is a sequence of tuples of length T, where each tuple contains the frame tag type (continuation frame or switch trigger frame) and the disease stage number. See also Figure 4 This is a schematic diagram of the disease stage marker sequence provided in the embodiments of this application. For example... Figure 4As shown, the upper part displays the time-step distributed offset sequence. A horizontal dashed line is used to mark the stage switching threshold. Different paths extend based on whether the sequence value crosses this threshold: non-marked black squares with values below the threshold connect to the continuation frame indicator label, while sequence elements that cross the threshold and are marked with black squares connect to the switching trigger frame indicator label. Based on different frame attribute mapping classifications, the disease stage marking sequence generated in the lower part is divided into two parts by different lines and outlines, with new default numbers assigned by the range labels of the disease stage numbers. This demonstrates the logical process of automatically converting the time-step distributed offset sequence into a discretely calibrated disease stage marking sequence based on the threshold. This mechanism automatically detects the disease stage switching boundary based on the distributed offset metric and isolates the time-step signal during the transition between disease stages, enabling agents in subsequent stages to perform data processing and policy determination based on the clearly defined and uniformly isolated and identified disease stage number boundaries.
[0054] Using Kullback-Leibler divergence as a measure of the offset between statistical distribution feature vectors makes stage switching detection structurally insensitive to occasional fluctuations in single anomalous sampling values. In the context of full-course disease management, occasional data loss from wearable devices or extreme values in single biochemical tests often cause pulse-like jumps in the original observations at a single time step. However, these pulse-like jumps only affect the frequency values of a single bin in the sliding window histogram, without substantially affecting the frequency distribution of the remaining B-1 bins. This keeps the overall shape of the statistical distribution feature vector stable, and the Kullback-Leibler divergence value remains below the stage switching threshold τ. Only when the values of most sampling points within a non-overlapping observation period synchronously shift towards the new distribution interval, causing the frequency values of multiple bins to migrate in the same direction, will the Kullback-Leibler divergence continuously exceed the stage switching threshold τ, triggering a disease stage switching marker. This provides a stable stage boundary location for the subsequent weight allocation of each agent.
[0055] S2: Based on the disease stage marker sequence, extract the local observation sub-tensor of each Agent from the stage state observation tensor; use the local observation sub-tensor as input to drive each Agent to independently output the action probability distribution vector, and write the action probability distribution vector of each Agent into the cross-Agent action state circular queue.
[0056] This step addresses the problem of single-decision models lacking real-time information interaction and game theory mechanisms between different professional management dimensions in long-term sequential decision-making processes. In whole-course disease management, the key indicators of the three sub-tasks—nutritional assessment, exercise rehabilitation, and comprehensive chronic disease management—are physically independent but temporally coupled. If a single model processes all indicators simultaneously, it will be impossible to independently optimize the strategy outputs of different sub-tasks. This step divides the phased state observation tensor according to the agent's responsibility domain, allowing each agent to reason independently within its own domain. Then, a circular queue of cross-agent action states aggregates the action outputs of all agents, providing a data foundation for S3 consistency measurement and weight allocation.
[0057] Further, step S2 includes:
[0058] S21: Based on the disease stage marker sequence, the staged state observation tensor is divided along the index category axis according to the Agent's responsibility domain. The nutritionist Agent reads the nutrition-related index slices, the exercise rehabilitation agent reads the exercise function-related index slices, and the health manager Agent reads the chronic disease comprehensive management-related index slices. Each slice is recorded as the local observation sub-tensor of the corresponding Agent.
[0059] This step aims to decouple the phased state observation tensor along the index category axis into three independent local observation sub-tensors, ensuring that each agent's policy inference is based solely on indices within its own responsibility domain. If decoupling is not performed and all N index terms are input into each agent's policy network, each agent will consume policy network capacity on indices outside its responsibility domain, causing the policy gradient update direction to be interfered with by irrelevant indices.
[0060] Specifically, an Agent responsibility domain index table is pre-established. The Agent responsibility domain index table is a mapping table with Agent ID as the key and set of indicator item IDs as the value. The initial labeling method of the Agent responsibility domain index table is as follows: senior physicians from the three specialties of clinical nutrition, rehabilitation medicine and chronic disease management independently label all N indicator items to the Agent, and the indicator items that are consistently labeled by the three physicians are taken as the responsibility domain indicators of the corresponding Agent.
[0061] The responsibility domain indicators for a nutritionist agent include weight-related indicators in the vital signs time series signal, fasting blood glucose, glycated hemoglobin, and blood lipids in the biochemical test numerical sequence, and dietary preference markers in the structured fields of the medical record; the responsibility domain indicators for a sports rehabilitation therapist agent include all indicators in the exercise function assessment scoring sequence and heart rate and blood oxygen saturation in the vital signs time series signal; the responsibility domain indicators for a health manager agent include past medical history markers, comorbidity markers, and diagnostic coding sequences in the structured fields of the medical record.
[0062] Based on the disease stage marker sequence, the time step range corresponding to the current disease stage number is extracted from the staged state observation tensor. Then, according to the Agent responsibility domain index table, index slices corresponding to each Agent are extracted along the index category axis, resulting in local observation sub-tensors for each of the three Agents. The data structure of each local observation sub-tensor is of shape [shape missing]. A three-dimensional floating-point tensor, in which This represents the accumulated time steps at the current stage of the disease course. This represents the number of metrics within the Agent's responsibility domain being traversed. See also... Figure 5 This is a schematic diagram of local observation sub-tensor slice extraction provided in an embodiment of this application. For example... Figure 5 As shown, the large three-dimensional cuboid structure on the left represents the phased state observation tensor, where two parallel thick dashed lines are used for region isolation and positioning along the bottom indicator category axis. The output on the right features vertically arranged container frames with three-dimensional local block identifiers representing the health manager agent, sports rehabilitation agent, and nutritionist agent. The three sets of indicator slices extracted from the left are injected into the corresponding three agents according to the horizontal association lines with arrows, reading their respective local observation sub-tensors according to their respective responsibilities. This demonstrates the process of decoupling and slice allocation of the phased state observation tensor along the corresponding indicators based on professional functional domains. This method isolates and extracts the overall tensor slices along the indicator category axis according to the agent's responsibility domain, allowing each agent's policy network to reason independently within its own local observation sub-tensor. This avoids the situation where inputting all indicator items into each agent would consume policy network capacity on indicators outside their responsibility domain, completely eliminating the defect of policy gradient update direction being interfered with by irrelevant indicators at different time steps.
[0063] S22: Each Agent takes its own local observation sub-tensor as input, and outputs the action probability distribution vector at the current time step through its own policy network forward reasoning; each dimension of the action probability distribution vector corresponds to an intervention action option within the Agent's domain of responsibility.
[0064] This step aims to enable each agent to independently generate the probability distribution of intervention actions based on its own local observation sub-tensor, providing the original probability input for the consistency measurement in S3 and the weighted aggregation in S4.
[0065] Specifically, each agent's policy network uses the same network architecture but different learnable parameters. The input layer of the policy network receives a slice of the local observation sub-tensor corresponding to the currently traversed agent at the current time step, and flattens the slice to a length equal to... The input vector is multiplied by the first fully connected hidden layer. The first fully connected hidden layer multiplies the input vector by the first weight matrix and adds the first bias vector, then processes it through the ReLU activation function to output the first hidden feature vector, which has a dimension of 128. The second fully connected hidden layer performs a linear transformation on the first hidden feature vector, multiplies it by the second weight matrix, adds the second bias vector, and then processes it through the ReLU activation function to output the second hidden feature vector, which has a dimension of 64. The output layer has two parallel branches: the policy branch multiplies the second hidden feature vector by the policy output weight matrix and adds the policy output bias vector, then normalizes it using the Softmax function to output the action probability distribution vector; the value branch multiplies the second hidden feature vector by the value output weight matrix and adds the value output bias vector to output a scalar-form state value estimate. This is used for calculating the time-series difference error in S43. The first weight matrix, the second weight matrix, the policy output weight matrix, the value output weight matrix, and the corresponding bias vector are all learnable parameters obtained through training data.
[0066] Each dimension of the action probability distribution vector corresponds to an intervention action option within the current Agent's domain of responsibility: the nutritionist Agent's action options include increasing the proportion of protein intake, decreasing the proportion of carbohydrate intake, increasing dietary fiber intake, and maintaining the current dietary structure, etc. The exercise rehabilitation agent offers a variety of options, including increasing aerobic exercise intensity, increasing resistance training intensity, reducing total exercise duration, and maintaining the current exercise plan. The health management agent's action options include increasing the frequency of blood glucose monitoring, adjusting the frequency of medication reminders, arranging remote follow-ups, and maintaining the current management frequency. The sum of all components of the action probability distribution vector of each Agent is always 1.
[0067] S23: Construct a cross-Agent action state circular queue. The number of slots in the cross-Agent action state circular queue is equal to the number of time steps accumulated in the current stage of the disease stage marker sequence. Write the action probability distribution vector of each Agent at the current time step into the current slot of the cross-Agent action state circular queue in Agent number order.
[0068] This step aims to aggregate the action probability distribution vectors of all time steps and all agents within the current disease stage using a circular queue structure, enabling S3 to centrally read the data required for consistency measurement within the stage domain.
[0069] Further, step S23 includes:
[0070] S231: When the disease stage marker sequence detects a disease stage switching trigger frame at the current time step, clear the cross-Agent action state circular queue and reset the slot counter so that the cross-Agent action state circular queue only retains the action probability distribution vector within the current new stage and does not mix in the historical action records of the previous stage.
[0071] This step aims to ensure that the content of the cross-Agent action state circular queue strictly corresponds to the current disease stage, eliminating the residual influence of the action probability distribution vector from the previous stage during the weight aggregation process in the new stage. During the transition period across disease stages, the action probability distribution vector from the previous stage reflects the policy preferences of each agent under the state distribution of the previous stage, and no longer matches the state distribution of the new stage.
[0072] Specifically, the slot counter is initialized to 0, and the data structure of the cross-Agent action status circular queue has a length equal to the preset maximum number of slots. The circular buffer stores a composite vector in each slot, which is formed by concatenating the action probability distribution vectors of the three agents in order of their agent numbers; the preset maximum number of slots... The value is determined based on the ratio of the duration of a typical stage in the whole course of disease management to the aligned time step Δt. For example, It can be set to 2016.
[0073] When the frame label type of the disease stage marker sequence at the current time step t is a disease stage switching trigger frame, a zero-value clearing operation is performed on all slots in the cross-Agent action state circular queue, resetting the slot counter value to 0; subsequently, the action probability distribution vector of each Agent at the current time step t is written into slot 0 in Agent number order, and the slot counter value is incremented by 1. See also Figure 6 This is a schematic diagram of a circular queue of cross-Agent action states provided in an embodiment of this application. Figure 6 As shown, the top three-segment composite vector block is formed by splicing the action probability distribution vectors of the three agents in numerical order. After an append write operation, it is sequentially written to each slot of the circular queue under the index control of the slot counter. When the disease stage marker sequence detects a switching trigger frame, the clear flag on the left reaches the circular structure through the broken arrow, performs a zero-value clear operation on all slots and resets the slot counter, ensuring that the circular queue only retains the action probability distribution vector within the current new stage.
[0074] S232: When the disease stage marker sequence is marked as a continuation frame at the current time step, the action probability distribution vector of each Agent is appended to the next slot of the cross-Agent action state circular queue.
[0075] Specifically, when the frame label type of the disease stage marker sequence at the current time step t is a continuation frame, the current value of the slot counter is read. As the write location index, the action probability distribution vectors of each Agent at the current time step t are concatenated into a composite vector according to the Agent number order and then written into the cross-Agent action state circular queue. For slot number 1, increment the slot counter value by 1; if the slot counter value reaches the preset maximum number of slots... If the write position is reversed to slot 0 according to the circular overwrite strategy, the next write position will be rolled back.
[0076] S231 directly binds the switching trigger frame of the disease stage marker sequence to the clearing operation of the cross-Agent action state circular queue, causing the cross-Agent action state circular queue to be forcibly reset at the moment of stage switching. The probability distribution vector of the previous stage action is physically cleared in the first time step of the new stage. This binding mechanism stems from the synchronization relationship between the cross-Agent action state circular queue and the disease stage number in the data organization dimension: the cross-Agent action state circular queue uses the time step as the slot index and the current disease stage as the effective scope. The disease stage switching trigger frame is essentially a boundary marker of the effective scope. Using the boundary marker directly as the clearing trigger signal eliminates the need for additional judgment logic to update the effective scope. The original design goal of this mechanism was solely to isolate the interference of the probability distribution vector of actions in the preceding stage on the calculation of the action consistency matrix in stage S3. However, the content written to the circular queue of cross-Agent action states is also the source data for the transformation tuples in the experience replay buffer of stage S4. The forced clearing operation ensures that the gradient update direction of each agent's policy network in S43 is aligned with the state distribution of the new stage in the first time step after the stage switch, eliminating the hidden defect that "the policy network parameters fail to reflect the state distribution of the new stage in a timely manner within a short time window after the stage switch". Technicians usually attribute the lag in policy network parameter updates to the inherent convergence speed of gradient descent, without associating it with the timing of clearing the circular queue of cross-Agent action states. The binding mechanism of S231 structurally avoids this hidden defect without modifying the policy network structure or learning rate scheduling strategy.
[0077] S3: Read the action probability distribution vector of each agent from the cross-Agent action state circular queue, and calculate the action consistency matrix within the stage; based on the action consistency matrix within the stage and the confidence of the disease stage offset, calculate the dynamic weight vector of each agent.
[0078] This step addresses the problem that single-decision models cannot dynamically calculate and allocate priority weights for different vertical management subtasks when faced with multi-source heterogeneous physiological data input. The magnitude of weight allocation should adaptively adjust to scenarios with varying phase switching intensities: in strong switching scenarios, weights need to be significantly concentrated on agents with stable strategies in the new phase; in weak switching scenarios, a near-uniform weight allocation should be maintained to avoid decision oscillations. This step uses a coupled calculation of the intra-phase action consistency matrix and the confidence level of the disease stage offset to simultaneously use the strategy stability of each agent and the phase switching intensity as the basis for weight allocation.
[0079] Further, step S3 includes:
[0080] S31: Read the action probability distribution vectors of all written slots in the current stage from the cross-Agent action state circular queue. Calculate the pairwise cosine similarity of the action probability distribution vectors of the same Agent at each time step in the current stage to construct the in-stage action self-consistency score of the Agent. Calculate the cross-Agent action cosine similarity for the action probability distribution vectors of any two different Agents. Arrange the in-stage action self-consistency scores and cross-Agent action cosine similarities of each Agent into a square matrix to obtain the in-stage action consistency matrix.
[0081] This step aims to simultaneously characterize the temporal stability of each agent's own policy output and the consistency between policy outputs of different agents in a matrix format. During the cross-stage transition period, the self-consistency score can identify which agents have formed stable policies under the new stage state distribution, while the cross-agent cosine similarity reflects whether there are potential conflicts in policy direction between different agents.
[0082] Specifically, read all composite vectors from slot 0 to the slot corresponding to the current value of the slot counter minus 1 in the cross-Agent action state circular queue, and decompose the composite vectors into a time series set of action probability distribution vectors for each Agent according to the Agent number.
[0083] For Agent number For each Agent, it iterates through all pairwise combinations of its action probability distribution vectors in the time series set, calculates the cosine similarity between the two action probability distribution vectors in each combination, and takes the arithmetic mean of the cosine similarities for all combinations, using this as the Agent's ID. Intra-phase action self-consistency score .
[0084] For Agent number With Agent number For two different agents, calculate the mean vector of the action probability distribution vector of each agent at each time step in the current stage, and then calculate the cosine similarity between the two mean vectors as the cross-agent action cosine similarity. .
[0085] Action self-consistency scoring within each stage Arranged on the diagonal of the intra-stage action consistency matrix, the cosine similarity of actions across agents is calculated. Arranged on the off-diagonal positions of the intra-phase motion consistency matrix, the intra-phase motion consistency matrix is obtained. The data structure of the action consistency matrix within the aforementioned stage is a 3×3 symmetric floating-point matrix, with matrix elements ranging from -1 to 1. See also... Figure 7 This is a schematic diagram of the consistency matrix of actions within a stage provided in an embodiment of this application. For example... Figure 7 As shown, the central area features a 3x3 matrix grid, with the grid cells divided into two categories based on their internal fill type. Cells located on the upper left to lower right diagonal with a dark gray background are labeled with solid-background connecting lines, representing intra-stage action self-consistency scores. Cells located off-diagonal with diagonal striped backgrounds are labeled with cross-Agent action cosine similarity scores, connected by another stripe connecting lines. These two types of extracted and calculated dissimilar score matrices are arranged in a common matrix module to form the intra-stage action consistency matrix structure. This composite matrix represents the process of reconstructing the temporal set of action probability distribution vectors into an output form for evaluating multi-Agent synergy. It integrates, in a matrix structure, both a self-score reflecting the agent's intra-stage temporal stability and a component determining whether there are potential conflict deviations in the cross-dimensional relationships between their mutual actions. It mathematically represents the policy stability and coordination of each agent within a single matrix dimension as the basis for consistency measurement data output.
[0086] S32: Take the distribution offset corresponding to the current disease stage switching trigger frame in step S1, subtract it from the stage switching judgment threshold τ, and normalize it to obtain the disease stage offset confidence level; the disease stage offset confidence level reflects the intensity of the current stage switching.
[0087] This step aims to transform the distribution offset that triggers the disease stage switching in S1 into a standardized indicator that can be directly used as a weight allocation amplitude adjustment coefficient, so that the intensity of weight allocation strictly corresponds to the actual intensity of stage switching.
[0088] Specifically, the switching trigger frame time step corresponding to the current disease stage number is retrieved from the disease stage marker sequence output by S13. Read the time-step distribution offset sequence output by S12 at time step The value of Calculate the confidence level of the disease stage offset using the following expression. :
[0089]
[0090] in The threshold for stage switching determination in S13, This is a preset upper bound for the distribution offset saturation. Determined based on the 99th percentile of the distribution offset sample in historical full-course data, for example, It can be set to 2.0.
[0091] The design principle of the confidence level for the disease stage offset is: molecule This represents the magnitude by which the distribution offset exceeds the judgment threshold at the current handover trigger moment; the greater the magnitude of the exceedance, the higher the handover intensity. The denominator represents... The excess amplitude is normalized to the [0,1] interval to ensure comparability of the disease stage offset confidence scores across different patients and different stages of switch events. A disease stage offset confidence score close to 0 indicates a weak switch, while a disease stage offset confidence score close to 1 indicates a strong switch. If the value is greater than 1, it is truncated to 1.
[0092] S33: Using the sum of each row of the action consistency matrix within a stage as the base weight vector, multiply the base weight vector by the confidence level of the disease stage offset and then perform softmax normalization to obtain the dynamic weight vector of each agent.
[0093] This step aims to ensure that weight allocation responds to both the policy stability of each agent in the new phase and the overall intensity of phase switching, avoiding the mismatch problem of the fixed-ratio update mechanism under different switching intensity scenarios.
[0094] Specifically, the intra-stage action consistency matrix output by S31 Summing by row yields a basic weight vector of length 3. Each component of the basic weight vector is equal to the sum of the intra-stage action self-consistency score of the corresponding Agent and the cross-Agent action cosine similarity of the Agent to other Agents.
[0095] The basic weight vector Confidence level of each component's offset from the disease stage Multiplying them together yields the coupling weight vector. Each component of the coupling weight vector is calculated using the following expression:
[0096]
[0097] For the coupling weight vector Performing softmax normalization yields the dynamic weight vector. :
[0098]
[0099] The design principle of the dynamic weight vector calculation expression is as follows: the basic weight vector reflects the policy stability and cross-agent coordination of each agent in the new stage. Agents with high self-consistency scores and consistent with the direction of other agents receive higher basic weights. The confidence of the disease stage offset is used as a multiplicative coefficient, which amplifies the component differences of the basic weight vector in strong switching scenarios and suppresses them in weak switching scenarios. When the confidence of the disease stage offset approaches 0, the components of the coupled weight vector approach 0, the softmax normalization output approaches a uniform distribution, and the dynamic weight vector degenerates into a uniform weight close to 1 / 3. When the confidence of the disease stage offset approaches 1, the difference of each component of the coupled weight vector is equal to the difference of the basic weight vector, and the softmax normalization significantly concentrates the weights on agents with high self-consistency scores in the stage.
[0100] For example, suppose the action consistency matrix in the current stage is: Then the basic weight vector If the confidence level of the disease stage shift Then the coupling weight vector After softmax normalization, the dynamic weight vector is approximately If the confidence level of the disease stage shift Then the coupling weight vector After softmax normalization, the dynamic weight vector is approximately It is nearly uniformly distributed.
[0101] The intra-stage action consistency matrix and the confidence level of the disease stage offset form a structural coupling in the calculation of the dynamic weight vector. The physical quantities reflected by the two are independent in dimension but complementary in weight allocation mechanism. When the intra-stage action consistency matrix acts alone, it can only reflect the relative stability of each agent's current strategy and cannot determine the magnitude of the weight allocation. When the confidence level of the disease stage offset acts alone, it can only reflect the overall intensity of the stage switch and cannot determine which agent should receive the weight. If weight allocation is based solely on the intra-stage action consistency matrix and the confidence level of the disease stage offset is ignored, in a weak switching scenario, the softmax normalization will still output a non-uniform dynamic weight vector because the sum of each row of the intra-stage action consistency matrix always has natural differences. This will cause unnecessary agent priority flipping in the joint decision-making action during the stable disease evolution stage. If weight allocation is based solely on the confidence level of the disease stage offset and the intra-stage action consistency matrix is ignored, in a strong switching scenario, the weights cannot be accurately guided to agents with stable strategies in the new stage, and the weight allocation will degenerate into uniform amplification or uniform compression. The coupling form of multiplication between the two allows the component differences of the base weight vector in weak handover scenarios to be minimized. The multiplicative coefficients are suppressed to near 0, and the softmax normalization naturally results in a nearly uniform output distribution; the component differences of the basic weight vector in strong switching scenarios are... The multiplicative coefficients are fully preserved, and softmax normalization significantly concentrates the weights towards agents with high self-consistency scores. This coupling mechanism ensures that the joint decision action distribution obtained by weighted aggregation in S4 maintains a balanced combination of agent outputs in weak switching scenarios and accurately gravitates towards agents with strong adaptability to the new stage in strong switching scenarios. The two response modes are mathematically coupled, and the output of the other becomes meaningless in the corresponding scenario if one is missing.
[0102] S4: Perform weighted aggregation on the action probability distribution vector of each Agent using dynamic weight vectors to generate a joint decision action distribution; sample the execution action sequence of the current time step based on the joint decision action distribution, and write the execution action sequence and the patient's state feedback of the next time step into a transformation tuple written to the phased experience playback buffer; update the policy network execution parameters of each Agent based on the transformation tuple in the phased experience playback buffer.
[0103] This step addresses the problem of policy gradient update direction being dominated by data from previous stages in multi-agent collaborative decision-making using a single decision model. In multi-agent reinforcement learning, the sample distribution of the experience replay buffer directly determines the policy gradient update direction. If the sample source is not isolated across stages, cross-stage sample mixing will cause the policy network parameter update in the new stage to be dragged down by the state distribution of the previous stage. This step, through a staged experience replay buffer partitioned storage according to the disease stage number, ensures that the policy gradient of each agent is always updated within the state distribution of the current stage.
[0104] Further, step S4 includes:
[0105] S41: For each Agent, a conservative baseline distribution vector is preset within the Agent's domain of responsibility. This conservative baseline distribution vector is a One-Hot vector where the position corresponding to the "Maintain Current Plan" action option number is 1, and the other positions are 0. The vector is then weighted according to the corresponding weight components in the dynamic weight vector. As the retention ratio of the probability distribution vector of the agent's actions, As the retention ratio of the conservative baseline distribution vector, weighted interpolation fusion is performed to obtain the weighted action probability distribution vector:
[0106]
[0107] in Agent number is The action probability distribution vector, The corresponding conservative baseline distribution vector; the weighted action probability distribution vector for each Agent. argmax sampling is performed, and the action option number corresponding to the highest probability is taken as the action to be executed by the Agent at the current time step. This mechanism makes Agents with higher dynamic weights (with stable policies and high coordination with other Agents) more likely to output the active intervention actions that they actually infer, while the weighted action probability distribution vector of Agents with lower dynamic weights is pulled toward a conservative baseline, and argmax sampling tends to output "maintain the current plan", thereby realizing the dynamic adjustment of the intervention intensity of each Agent according to the policy confidence.
[0108] The execution actions of the three agents are arranged in order of agent number to obtain the execution action sequence at the current time step. The data structure of the execution action sequence is an integer tuple of length 3, where each integer corresponds to the index of the specific intervention instruction of the current agent within its own domain of responsibility. The data structure of the joint decision action distribution is composed of three weighted action probability distribution vectors. A composite vector formed by concatenating Agent IDs in sequence.
[0109] S42: Construct a phased experience replay buffer. The phased experience replay buffer is partitioned and stored according to the disease stage number. Each partition maintains an independent first-in-first-out queue. Pack the phased state observation tensor of the current time step, the sequence of executed actions, the patient's state feedback for the next time step, and the current value of the disease stage marker sequence into a transformation tuple and write it into the partition of the phased experience replay buffer corresponding to the current disease stage number.
[0110] This step aims to provide a source of samples that strictly corresponds to the stage number for updating the parameters of the policy network, thereby avoiding policy gradient pollution caused by cross-stage sample mixing.
[0111] Specifically, the data structure of the phased experience playback buffer is a mapping table with the disease stage number as the key and the first-in-first-out queue as the value, and the maximum capacity of each first-in-first-out queue is [missing information]. The maximum capacity of the first-in-first-out queue. The value is determined based on the ratio of the duration of a typical stage in the whole course of disease management to the aligned time step Δt. For example, It can be set to 5000.
[0112] The patient's status feedback for the next time step is obtained from the patient monitoring interface of the in-hospital monitoring system and the data reporting interface of the wearable device at the end of the next aligned time step Δt. The data structure of the status feedback is the same as the slice structure of the phased status observation tensor on a single time step, and is supplemented with an immediate reward value calculated by the clinical reward function. The output of the clinical reward function is a weighted combination of scalar values. The inputs include the distance of the blood glucose level from the target value, the distance of the blood pressure from the target value, the change in the exercise function assessment score, and the comorbidity markers for the next time step. The weighting coefficients are pre-calibrated by the clinical physician team.
[0113] The slice of the phased state observation tensor at the current time step, the sequence of actions performed, the patient's state feedback for the next time step, and the current value of the disease stage marker sequence are packaged into a transformation tuple. The disease stage number corresponding to the current time step is read from the disease stage marker sequence, and the transformation tuple is appended to the partition corresponding to the current disease stage number in the phased experience playback buffer. If the first-in-first-out queue length of the partition reaches... If the first-in-first-out (FIFO) strategy is applied, the earliest transformed tuple written to the head of the queue is discarded.
[0114] See Figure 8 This is a schematic diagram of the phased experience playback buffer structure provided in an embodiment of this application. Figure 8As shown, the main structure consists of a multi-dimensional parallel matrix representing a phased experience replay buffer. Within the structure, three channels with labeled containers are arranged parallel to each other from top to bottom. The leftmost column of each channel contains a data mapping table with labeled disease stage numbers and their different partition descriptions. To the right of the boxes, horizontally connected are pipes with open ends representing a first-in-first-out (FIFO) queue. The channel contains rows of closely spaced small black squares representing transformation tuples of independent sequence units. A group of breakpoint arrows pointing outwards at the end of the queue indicates the dynamic effect of tuples being discarded due to length accumulation reaching a threshold. The logic of transformation tuples being assigned to the corresponding FIFO partition queue based on their numbers is demonstrated. Its partitioned storage design avoids cross-stage mixing and cross-referencing of transformation tuples, using only the experience within the current disease stage partition to drive the gradient descent strategy of each agent network to obtain update strategies. This prevents the parameters of the new stage from being misled by the experience of the previous stage during the same decision-making process, thus ensuring that the sources of experience replay samples are all aligned with a single, pure gradient update direction in the new disease stage.
[0115] S43: Sample several transformation tuples from the current disease stage partition of the phased experience replay buffer, and calculate the temporal difference error of each agent; perform gradient descent parameter update on the policy network of each agent according to the temporal difference error; during the parameter update process, only transformation tuples within the current disease stage partition are used, and cross-partition mixing sampling is not performed.
[0116] This step aims to drive the update of the network parameters of each agent's policy using the transformation tuples within the current disease stage partition as a single sample source, so that the gradient descent direction is strictly aligned with the current stage state distribution.
[0117] Specifically, the partition corresponding to the current disease stage number is retrieved from the phased experience replay buffer, and a batch of [size missing] is drawn from the first-in-first-out queue of that partition using a uniform random sampling method. The batch size of the converted tuple samples; The hidden layer dimension of the policy network and the available computing resources are determined, for example, It can be set to 32.
[0118] For each transformed tuple in the sample batch, perform the following calculations sequentially: Extract the local observation sub-tensor slices of the current Agent from the phased state observation tensor slices in the transformed tuples according to the Agent responsibility domain index table in S21, and input them into the policy network of the current Agent to obtain the action probability distribution vector and state value estimate at the current time step. Extract the local observation sub-tensor slices from the state feedback in the transformed tuples in the same way to the next time step, and input them into the policy network of the current Agent to obtain the state value estimate for the next time step. Calculate the timing difference error using the following expression. :
[0119]
[0120] in To transform the instant reward value in the tuple, and These are the state value estimates corresponding to the local observation sub-tensor slices at the current time step and the next time step, respectively (given by the value output head of the policy network). The discount factor is determined based on the long-term reward trade-offs required for full-course disease management. For example, It can be set to 0.95.
[0121] The mean squared value of the temporal difference error of all transformed tuples within the sample batch is used as the loss function. The gradient descent parameters of the Adam optimizer are updated on the first weight matrix, second weight matrix, policy output weight matrix, value output weight matrix, and corresponding bias vector in the policy network of the current Agent. The initial learning rate is set to... The learning rate is adjusted according to the cosine annealing strategy.
[0122] The design of storing the phased experience replay buffer in partitions according to the disease stage number, along with the clearing mechanism of the cross-Agent action state circular queue in S231, provides dual protection at the data isolation level: the cross-Agent action state circular queue ensures the purity of the input for calculating the action consistency matrix at the current time step, while the phased experience replay buffer ensures the purity of the training samples for updating the policy network parameters. This two-layer data isolation mechanism ensures that the parameter update direction of each agent's policy network in a new disease stage is always dominated by the state distribution of the new stage, and is not dragged down by the transformation tuples of previous stages.
[0123] S5: Update the dynamic weight vector based on the probability distribution vector of the new round of actions after parameter update, make a judgment based on the stage exit condition, output the full course decision record table and continue the sequential decision loop.
[0124] This step addresses the hidden flaw of a time-step lag between policy network parameter updates and weight allocation. In conventional implementations of multi-agent reinforcement learning, weight allocation is typically calculated based on the probability distribution vector of the old actions before the parameter update, resulting in a one-time-step lag between the weight response and the policy update. This step, through a closed feedback loop from policy update to weight update within a single time step and a stage exit condition determination, ensures that the weight response and policy update occur synchronously, and archives the decision results of this stage after stage stabilization.
[0125] Further, step S5 includes:
[0126] S51: After each Agent policy network completes parameter updates, forward inference is re-executed with the local observation sub-tensor of the current time step as input to obtain the updated action probability distribution vector; the updated action probability distribution vector replaces the original action probability distribution vector in the current slot of the cross-Agent action state circular queue, and the intra-stage action consistency matrix calculation and dynamic weight vector update of step S3 are re-executed.
[0127] This step aims to ensure that the policy network parameter update results are perceived by the weight allocation mechanism within the same time step, eliminating the time step lag between parameter update and weight response.
[0128] Specifically, after S43 completes the parameter update of the policy network of each Agent, it immediately takes the local observation sub-tensor slice of the current time step t as input and re-executes forward inference on the policy network of each Agent (using the same network structure and calculation process as S22) to output the updated action probability distribution vector of each Agent at the current time step t.
[0129] Read the current value of the slot counter in the cross-agent action state circular queue. locating to the The original composite vector of the action probability distribution vector written by S22 in slot number 1 (i.e., the write slot corresponding to this time step) is completely replaced by a new composite vector formed by concatenating the updated action probability distribution vectors in Agent number order.
[0130] Using the replaced cross-Agent action state circular queue as input, the entire calculation process from S31 to S33 is re-executed: the action consistency matrix within the stage is recalculated, multiplied again with the confidence of the disease stage offset, and softmax normalization is performed to obtain the updated dynamic weight vector at the current time step t.
[0131] S51 writes the updated action probability distribution vector after the policy network parameters are updated back to the current slot in the cross-Agent action state circular queue, and triggers the synchronous recalculation of the intra-phase action consistency matrix and dynamic weight vector, so that the weight allocation result forms a closed feedback loop from policy update to weight update within a single time step. The original design goal of this closed feedback loop was only to ensure that the dynamic weight vector responds to the policy changes of each agent immediately after each policy network parameter update, without waiting for the next time step to reflect the weight changes. Those skilled in the art usually attribute the weight response lag to the inherent training order of multi-agent reinforcement learning, regarding it as an unavoidable delay, without associating this lag with the timing of writing to the cross-Agent action state circular queue. S51 achieves this by writing back the updated action probability distribution vector to the current slot of the cross-Agent action state circular queue—a simple operation that does not add any extra modules. This allows the downstream intra-stage action consistency matrix to automatically use the updated policy as the measurement object, thereby enabling the dynamic weight vector and policy network parameters to converge synchronously within the same time step. This eliminates the hidden defect that "the weight allocation within a short window after the disease stage switch still uses the estimation results of the previous policy network parameters," allowing the joint decision action distribution to obtain a weight allocation benchmark consistent with the latest policy network parameters as soon as the stage switch occurs.
[0132] S52: Calculate the mean of the time-step distribution offset sequence of several consecutive time steps within the current disease stage. If the mean is lower than the stage switching judgment threshold τ and the cumulative number of time steps in the current stage exceeds the preset minimum stage dwell time, then the current disease stage is determined to meet the stage exit condition.
[0133] This step aims to provide stable exit judgment logic for archiving the current disease stage, and avoid the archiving operation being mistakenly triggered due to a temporary drop in stage switching detection.
[0134] Specifically, the current time step t is read from the time-step distribution offset sequence output by S12 and backtracked. The distribution offset values within a time step range are used to calculate the arithmetic mean of the backtracking range as the exit decision mean; the exit decision window length The value is determined based on the ratio of the duration of stable disease progression in the whole course of management to the aligned time step Δt. For example, It can be set to 12, corresponding to 1 hour of continuous observation at Δt=5 minutes.
[0135] Read the cumulative time steps corresponding to the current disease stage number in the disease stage marker sequence. The minimum number of stage dwell steps The value is determined based on the ratio of the shortest clinically significant duration of each stage in the whole course of disease management to the aligned time step Δt. For example, It can be set to 144, corresponding to the shortest stay of 12 hours at Δt=5 minutes.
[0136] If the average exit judgment value is less than the stage switching judgment threshold τ, and the cumulative time steps are... Greater than If the current disease stage meets the stage exit conditions, the stage exit flag is output as true; otherwise, the stage exit flag is output as false.
[0137] S53: If the stage exit condition is met, the final dynamic weight vector, joint decision action distribution, and stage number of the current disease stage are written into the full disease decision record table, and the parameter updates and dynamic weight recalculation of each agent policy network are frozen. Before the end of the current disease stage, each agent policy network only performs forward inference and does not perform the gradient descent parameter update in S43. The execution action sequence of each time step is directly output using the policy parameters corresponding to the archived joint decision action distribution, ensuring that the patient continues to receive intervention instructions during the stable period of the disease. When step S1 detects that the distribution offset exceeds the stage switching judgment threshold τ in the time-step distribution offset sequence of subsequent time steps, the parameter update freeze state is lifted, and a new round of disease stage label sequence generation and policy network parameter update process begins. If the stage exit condition is not met, the current time step is marked as a continuation frame and the process returns to step S2 to continue executing the sequential decision of the next time step within the current disease stage.
[0138] This step aims to branch between the archiving path and the continued loop path based on the stage exit flag, so that the decision-making results of the whole course of the disease can be archived in stages, while ensuring the continuous advancement of sequential decision-making in unstable stages, and ensuring that patients continue to receive intervention instructions based on the archiving strategy in stable stages.
[0139] Specifically, the data structure of the full course decision record table is an ordered mapping table with the course stage number as the key and the triples (final dynamic weight vector, joint decision action distribution, and stage cumulative time steps) as the values.
[0140] If the stage exit flag output by S52 is true, then update the dynamic weight vector of S51 corresponding to the current disease stage number, the joint decision action distribution output by S41, and the cumulative time steps of the current stage. Write the entire disease course decision record table and freeze the parameter updates and dynamic weight recalculation of each Agent policy network; before the end of the current disease course stage (i.e. before S1 detects a new disease course stage switching trigger frame), each Agent policy network only performs forward inference and does not perform gradient descent parameter updates in S43. The execution action sequence of each time step is directly output with the policy parameters corresponding to the archived joint decision action distribution, ensuring that the patient continues to receive intervention instructions during the stable period of the disease; when the S1 step detects that the distribution offset exceeds the stage switching judgment threshold τ in the time step distribution offset sequence of subsequent time steps, the parameter update freeze state is lifted, and a new round of disease course stage label sequence generation and policy network parameter update process begins.
[0141] If the stage exit flag output by S52 is false, then the frame label type of the current time step in the disease stage label sequence is set to a continuation frame, the disease stage number remains unchanged, and the process jumps back to step S2. The process continues to execute the complete sequential decision loop, which includes local observation sub-tensor extraction, policy network forward inference, cross-Agent action state circular queue writing, intra-stage action consistency matrix calculation, dynamic weight vector update, joint decision action distribution generation, transformation tuple writing, policy network parameter update, and weight writeback.
[0142] The full-course decision record table is continuously accumulated until the patient's full-course management ends. The final dynamic weight vector of each stage arranged in the order of the disease stage number in the full-course decision record table, together with the distribution of joint decision actions, constitutes the full-course management plan. The full-course management plan is distributed to the patient's health management terminal through the plan distribution interface of the full-course management platform. The health management terminal generates specific nutritional intake instructions, exercise rehabilitation instructions, and daily health management instructions according to the execution action sequence of each time step, and pushes them to the nutrition management subsystem, rehabilitation training subsystem, and chronic disease follow-up subsystem for execution, respectively.
[0143] Example 2:
[0144] This embodiment, based on Embodiment 1, provides a multi-agent reinforcement learning joint decision-making system for collaborative diagnosis and treatment throughout the entire disease course, such as... Figure 9 As shown, it includes:
[0145] The tensor construction and stage division module is used to acquire multidimensional physiological state data of patients throughout the entire course of the disease and construct staged state observation tensors; perform disease stage switching detection on the staged state observation tensors and generate disease stage marker sequences; the multidimensional physiological state data throughout the entire course of the disease includes real-time physiological indicator data and disease state data;
[0146] The local inference and queue writing module is used to extract the local observation sub-tensor of each Agent from the staged state observation tensor according to the disease stage marker sequence; using the local observation sub-tensor as input, drive each Agent to independently output the action probability distribution vector, and write the action probability distribution vector of each Agent into the cross-Agent action state circular queue.
[0147] The consistency analysis and weight allocation module is used to read the action probability distribution vector of each agent from the cross-Agent action state circular queue, calculate the action consistency matrix within the stage, and calculate the dynamic weight vector of each agent based on the action consistency matrix within the stage and the confidence of the disease stage offset.
[0148] The action aggregation and experience replay module is used to perform weighted aggregation on the action probability distribution vector of each Agent using the dynamic weight vector to generate a joint decision action distribution; sample the execution action sequence of the current time step according to the joint decision action distribution, and write the execution action sequence and the patient's state feedback of the next time step into a transformation tuple written to the phased experience replay buffer; update the policy network execution parameters of each Agent according to the transformation tuple in the phased experience replay buffer;
[0149] The loop feedback and output archiving module is used to update the dynamic weight vector based on the probability distribution vector of the new round of actions after parameter updates, make a judgment based on the stage exit condition, output the full course decision record table, and continue the sequential decision loop.
Claims
1. A multi-agent reinforcement learning joint decision-making method for collaborative diagnosis and treatment throughout the entire disease course, characterized in that, include: Acquire multidimensional physiological state data of patients throughout the entire course of their illness, and construct a phased state observation tensor; The multidimensional physiological state data throughout the entire disease course includes real-time physiological indicator data and disease status data; The acquisition of multidimensional physiological state data throughout the patient's entire disease course and the construction of a phased state observation tensor include: Continuously acquire patients' real-time physiological indicators and disease status data; After aligning the acquired real-time physiological index data and disease state data according to the collection timestamp, they are arranged along three dimensions: time axis, index category axis, and data modality axis to construct a phased state observation tensor. Perform disease stage switching detection on the phased state observation tensor to generate a disease stage marker sequence; the process of performing disease stage switching detection on the phased state observation tensor to generate a disease stage marker sequence includes: The statistical distribution characteristics of each index within the sliding window are extracted step by step along the time axis from the phased state observation tensor. The Kullback-Leibler divergence of the statistical distribution characteristics of each index between adjacent non-overlapping sliding windows is calculated to obtain the time-step distribution offset sequence. Based on the relationship between the time-step distribution offset sequence and the preset stage switching determination threshold, disease stage switching detection is performed; If the distribution offset of a certain time step is less than the stage switching determination threshold, then the time step is marked as a continuation frame of the current disease stage; if the distribution offset of a certain time step is greater than or equal to the stage switching determination threshold, then the time step is marked as a disease stage switching trigger frame, and a new disease stage number is reassigned to the time steps after the switching trigger frame. The labeling results of all frames are summarized in time step order to generate a disease stage labeling sequence; Based on the disease stage marker sequence, extract the local observation sub-tensor of each Agent from the staged state observation tensor; use the local observation sub-tensor as input to drive each Agent to independently output the action probability distribution vector, and write the action probability distribution vector of each Agent into the cross-Agent action state circular queue. Read the action probability distribution vector of each agent from the cross-Agent action state circular queue, and calculate the action consistency matrix within the stage; calculate the dynamic weight vector of each agent based on the action consistency matrix within the stage and the confidence of the disease stage offset. The dynamic weight vector is used to perform weighted aggregation on the action probability distribution vector of each Agent to generate a joint decision action distribution; the execution action sequence of the current time step is sampled according to the joint decision action distribution, and the execution action sequence and the patient's state feedback of the next time step are combined to form a transformation tuple and written into the phased experience replay buffer; the policy network execution parameters of each Agent are updated according to the transformation tuple in the phased experience replay buffer. The dynamic weight vector is updated based on the probability distribution vector of the new round of actions after parameter updates. The decision is made based on the stage exit condition, and the full course decision record table is output and the sequential decision loop continues.
2. The multi-agent reinforcement learning joint decision-making method for collaborative diagnosis and treatment throughout the entire disease course as described in claim 1, characterized in that, The calculation of the Kullback-Leibler divergence of the statistical distribution characteristics of each index between adjacent non-overlapping sliding windows, to obtain the time-step distribution offset sequence includes: For each index item within the subtensor of the sliding window, collect all values of the index item within the sliding window along the time axis, divide it into equal-width histogram intervals according to a preset number of bins, count the frequency of values in each interval and normalize it to the frequency, and obtain the statistical distribution feature vector of the index item under the current sliding window. Calculate the Kullback-Leibler divergence for the statistical distribution feature vector of the same index item under the corresponding starting position of the sliding window at the current time step and the sliding window at the corresponding starting position of the previous non-overlapping observation period. The distribution offset of the current time step is obtained by averaging the Kullback-Leibler divergence of all index items along the index category axis; after traversing all time steps of the phased state observation tensor along the time axis, the time step-by-time distribution offset sequence is output.
3. The multi-agent reinforcement learning joint decision-making method for collaborative diagnosis and treatment throughout the entire disease course as described in claim 1, characterized in that, The step of extracting local observation sub-tensors for each agent from the staged state observation tensor based on the disease stage marker sequence includes: Based on the disease stage marker sequence, the staged state observation tensor is divided along the index category axis according to the Agent responsibility domain; The nutritionist agent reads nutrition-related indicator slices, the sports rehabilitation therapist agent reads exercise function-related indicator slices, and the health manager agent reads chronic disease comprehensive management-related indicator slices. Each slice is denoted as a local observation sub-tensor of the corresponding Agent.
4. The multi-agent reinforcement learning joint decision-making method for collaborative diagnosis and treatment throughout the entire disease course as described in claim 1, characterized in that, The step of writing the action probability distribution vector of each Agent into the cross-Agent action state circular queue includes: When the disease stage marker sequence detects a disease stage switching trigger frame at the current time step, the cross-Agent action state circular queue is cleared and the slot counter is reset so that the cross-Agent action state circular queue only retains the action probability distribution vector within the current new stage. When the disease stage marker sequence is marked as a continuation frame at the current time step, the action probability distribution vector of each Agent is appended to the next slot of the cross-Agent action state circular queue.
5. The multi-agent reinforcement learning joint decision-making method for collaborative diagnosis and treatment throughout the entire disease course as described in claim 1, characterized in that, The action consistency matrix during the calculation phase includes: Read the action probability distribution vectors of all written slots in the current stage from the cross-Agent action state circular queue, calculate the pairwise cosine similarity of the action probability distribution vectors of the same Agent at each time step in the current stage, and construct the action self-consistency score of the Agent in the current stage. Calculate the cross-Agent action cosine similarity for the action probability distribution vectors of any two different Agents; The intra-stage action consistency matrix is obtained by arranging the intra-stage action self-consistency scores and cross-Agent action cosine similarity scores of each Agent into a square matrix.
6. The multi-agent reinforcement learning joint decision-making method for collaborative diagnosis and treatment throughout the entire disease course as described in claim 1, characterized in that, The calculation of the dynamic weight vector for each agent based on the consistency matrix of actions within the stage and the confidence level of the disease stage offset includes: Take the distribution offset corresponding to the current disease stage switching trigger frame, subtract it from the stage switching judgment threshold, and then normalize it to obtain the disease stage offset confidence level. Using the sum of each row of the action consistency matrix within a stage as the base weight vector, the base weight vector is multiplied by the confidence level of the disease stage offset and then subjected to softmax normalization to obtain the dynamic weight vector of each agent.
7. The multi-agent reinforcement learning joint decision-making method for collaborative diagnosis and treatment throughout the entire disease course as described in claim 1, characterized in that, The step of performing weighted aggregation on the action probability distribution vectors of each Agent using the dynamic weight vector to generate a joint decision action distribution includes: A conservative baseline distribution vector within the corresponding Agent's responsibility domain is preset, and weighted interpolation fusion is performed to obtain a weighted action probability distribution vector; For each Agent, perform argmax sampling on the weighted action probability distribution vector, and take the action option number corresponding to the maximum probability as the action to be executed by the Agent at the current time step; The execution actions of the three agents are arranged in order of agent number to obtain the execution action sequence at the current time step; the joint decision action distribution is a composite vector formed by concatenating three weighted action probability distribution vectors.
8. The multi-agent reinforcement learning joint decision-making method for collaborative diagnosis and treatment throughout the entire disease course as described in claim 1, characterized in that, The step of writing the sequence of executed actions and the patient's state feedback at the next time step into a transformation tuple and then writing it into the phased experience playback buffer includes: A phased experience replay buffer is constructed, which is stored in partitions according to the disease stage number, and each partition maintains an independent first-in-first-out queue. The current time step's phased state observation tensor, the sequence of executed actions, the patient's state feedback for the next time step, and the current value of the disease stage marker sequence are packaged into a transformation tuple; Read the disease stage number corresponding to the current time step from the disease stage marker sequence, and append the transformed tuple to the partition corresponding to the current disease stage number in the staged experience playback buffer.
9. The multi-agent reinforcement learning joint decision-making method for collaborative diagnosis and treatment throughout the entire disease course as described in claim 1, characterized in that, The process of updating the dynamic weight vector based on the new round of action probability distribution vector after parameter updates, determining the exit condition based on the stage, outputting the full course decision record table, and continuing the sequential decision loop includes: After each Agent policy network completes parameter updates, it re-executes forward inference to obtain the updated action probability distribution vector. Replace the original action probability distribution vector in the current slot of the cross-Agent action state circular queue with the updated action probability distribution vector, and re-execute the intra-phase action consistency matrix calculation and dynamic weight vector update. If the exit condition is met, the final dynamic weight vector, joint decision action distribution, and disease stage number of the current disease stage are written into the full disease decision record table, and the parameter updates and dynamic weight recalculation of each Agent strategy network are frozen. If the stage exit condition is not met, the current time step is marked as a continuation frame and the previous step is returned to continue executing the sequential decision for the next time step within the current disease stage.
10. A multi-agent reinforcement learning joint decision-making system for collaborative diagnosis and treatment throughout the entire disease course, characterized in that, The multi-agent reinforcement learning joint decision-making method for collaborative diagnosis and treatment throughout the entire disease course, as described in any one of claims 1-9, includes: The tensor construction and stage division module is used to acquire multidimensional physiological state data of patients throughout the entire course of the disease and construct staged state observation tensors; perform disease stage switching detection on the staged state observation tensors and generate disease stage marker sequences; the multidimensional physiological state data throughout the entire course of the disease includes real-time physiological indicator data and disease state data; The local inference and queue writing module is used to extract the local observation sub-tensor of each Agent from the staged state observation tensor according to the disease stage marker sequence; using the local observation sub-tensor as input, drive each Agent to independently output the action probability distribution vector, and write the action probability distribution vector of each Agent into the cross-Agent action state circular queue. The consistency analysis and weight allocation module is used to read the action probability distribution vector of each agent from the cross-Agent action state circular queue, calculate the action consistency matrix within the stage, and calculate the dynamic weight vector of each agent based on the action consistency matrix within the stage and the confidence of the disease stage offset. The action aggregation and experience replay module is used to perform weighted aggregation on the action probability distribution vector of each Agent using the dynamic weight vector to generate a joint decision action distribution; sample the execution action sequence of the current time step according to the joint decision action distribution, and write the execution action sequence and the patient's state feedback of the next time step into a transformation tuple written to the phased experience replay buffer; update the policy network execution parameters of each Agent according to the transformation tuple in the phased experience replay buffer; The loop feedback and output archiving module is used to update the dynamic weight vector based on the probability distribution vector of the new round of actions after parameter updates, make a judgment based on the stage exit condition, output the full course decision record table, and continue the sequential decision loop.
Citation Information
Patent Citations
Abnormity analysis method and device based on multi-agent cooperation, equipment and medium
CN120929133A
Diagnosis and treatment auxiliary framework forming method and system based on multi-agent collaboration
CN121460123A