Patient behavior reminding method based on federated learning and multi-modal time series anomaly detection

By constructing a drug knowledge base and a multimodal temporal anomaly detection model, and combining federated learning and reinforcement learning, the patient behavior judgment benchmark is dynamically adjusted to generate personalized intervention strategies. This solves the problems of static judgment and data silos in existing technologies, and realizes personalized, dynamic, and adaptive patient behavior management.

CN122369988APending Publication Date: 2026-07-10YUNCHANG (BEIJING) DIGITAL TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610507061.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-04-16
Publication Date
2026-07-10

AI Technical Summary

Technical Problem

Existing patient behavior management systems suffer from several problems, including deficiencies in static benchmark judgment, lack of adaptive adjustment mechanisms, insufficient identification of medication requirements, lack of analysis on the correlation between behavior and medication, data silos leading to limited model training, low accuracy of behavior analysis, and limited effectiveness of intervention strategies.

Method used

We employ a federated learning and multimodal temporal anomaly detection approach to construct a drug knowledge base, collect multimodal health behavior data, train a multimodal temporal anomaly detection model locally using a federated learning architecture, identify patients’ historical behavioral deviations, dynamically adjust judgment criteria, generate personalized intervention strategies, and push reminder content.

Benefits of technology

It enables personalized, dynamic, and adaptive patient behavior management, improves the accuracy of behavior analysis and the effectiveness of intervention strategies, solves the data silo problem, ensures privacy protection, and reduces the misjudgment rate and invalid alerts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122369988A_ABST
    Figure CN122369988A_ABST
Patent Text Reader

Abstract

This disclosure provides a patient behavior alert method based on federated learning and multimodal temporal anomaly detection, applicable to the field of healthcare technology. The method includes constructing a drug knowledge base and extracting patient multimodal behavioral features; analyzing the matching degree and correlation between medication requirements and multimodal behavioral features, identifying historical behavioral deviations, and establishing personalized behavioral benchmarks accordingly; training a multimodal temporal anomaly detection model; dynamically adjusting subsequent behavioral judgment benchmarks based on the deviation between the patient's most recent medication information and the personalized behavioral benchmarks; inputting current multimodal behavioral features into the model, combining the adjusted benchmarks to identify real-time behavioral deviations and generate intervention criteria; and generating personalized intervention strategies based on the intervention criteria and the patient's historical feedback and pushing them to the patient. This approach addresses the problems of personalization, accuracy, and interpretability in existing solutions, achieving intelligent, personalized, and adaptive management of patient behavior intervention alerts.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of artificial intelligence technology in healthcare, and in particular to a patient behavior alerting method based on federated learning and multimodal temporal anomaly detection. Background Technology

[0002] In patient behavior management, particularly medication adherence management, non-adherence to medication remains one of the most persistent and costly challenges in healthcare. Up to 25% of patients never begin prescription treatment, and approximately half do not take medication as directed. This directly impacts treatment outcomes and healthcare costs, and healthcare institutions face direct economic penalties due to their inability to meet quality indicators. Existing technologies primarily employ patient behavior reminders based on pre-defined rule bases, such as scheduled medication reminders and regular exercise encouragement. These methods rely on static judgment benchmarks and fixed rule logic, failing to dynamically adjust based on patients' real-time medication usage and behavioral characteristics. For example, when a patient delays medication due to business trips, holidays, or illness, subsequent medication decisions still rely on fixed benchmark intervals, leading to inappropriate reminder timing, high misjudgment rates, and a lack of flexibility in intervention strategies. Furthermore, statistical regression models and rule-based systems often fail to capture the high-dimensional, temporal evolution, and probabilistic characteristics inherent in medication behavior trajectories, severely limiting their effectiveness in precision medicine and strategy simulation scenarios.

[0003] In terms of drug information utilization, existing technologies generally lack the ability to finely identify and manage drug administration requirements. Different drugs vary significantly in terms of frequency, dosage, whether to take them before or after meals, and the interval between doses. Although these requirements are detailed in the drug instructions, existing systems have failed to effectively parse and utilize this unstructured textual information, resulting in generic and unspecific reminders. Some researchers have attempted to extract medical entities from drug instructions using pre-trained models, but due to the significant heterogeneity of medical entities in drug instructions for treating different diseases, model training requires a large number of labeled samples, making it technically challenging. Regarding patient behavior analysis, existing technologies mostly employ single data sources and simple machine learning methods, lacking the ability to fuse multimodal data and detect temporal anomalies. This makes it difficult to comprehensively grasp patients' behavioral characteristics across multiple dimensions (medication, exercise, diet, sleep, physiological indicators, etc.) and accurately identify complex behavioral deviation patterns. Traditional methods are often limited to specific modalities when integrating different types of clinical data, requiring extensive manual feature engineering and being difficult to generalize to different clinical environments.

[0004] At the data collaboration level, current patient health management faces a severe data silo problem. Medical and health data is primarily stored within different levels of medical institutions, and cross-institutional data collaboration has become a prominent weakness in the industry. Patient data from various medical institutions, community health service centers, and family doctor teams are isolated, preventing the formation of economies of scale for in-depth analysis. The limited data volume of individual institutions leads to poor training results for machine learning models, making it difficult to accurately identify patient behavior patterns. Simultaneously, data sharing faces privacy and compliance challenges, and traditional centralized data processing methods pose security risks. Federated learning, as a distributed training paradigm where "data is usable but not visible," provides a technical direction for resolving this contradiction. Research teams have already used federated learning to achieve superior model performance compared to single-center models in medical image analysis and disease risk assessment, but its systematic application in patient behavior intervention remains a gap. Furthermore, existing intervention strategies generally lack dynamic optimization mechanisms and interpretability. Reminders are difficult to adjust in real time based on patient feedback and behavioral changes, the intervention decision-making process is opaque, and patients and medical staff struggle to understand the basis of intervention recommendations, resulting in low clinical credibility and patient acceptance. In the cross-application of knowledge graphs and reinforcement learning, existing research has explored the use of knowledge graphs to enhance the interpretability of reinforcement learning, achieving superior performance compared to traditional methods in drug sensitivity prediction and adverse drug reaction detection. However, a systematic technical solution has not yet been formed in the field of patient behavior intervention reminders.

[0005] In summary, existing technologies have significant shortcomings in terms of the adaptability of behavioral judgment, the refined utilization of drug information, cross-institutional data collaboration, the accuracy of multimodal behavioral analysis, and the interpretability and dynamic optimization of intervention strategies. There is an urgent need for a patient behavior management solution that can integrate advanced artificial intelligence technology, balance privacy protection and cross-institutional collaboration, and achieve personalized adaptive intervention. Summary of the Invention

[0006] This disclosure provides a patient behavior reminder method based on federated learning and multimodal temporal anomaly detection, which solves the technical problems existing in the current patient behavior management system, such as static benchmark judgment defects, lack of adaptive adjustment mechanism, insufficient identification of medication requirements, lack of behavior-drug correlation analysis, data silos leading to limited model training, low accuracy of behavior analysis, and limited effectiveness of intervention strategies.

[0007] According to a first aspect of this disclosure, a patient behavior alerting method based on federated learning and multimodal temporal anomaly detection is provided. The method includes: A drug knowledge base was constructed, and patients' multimodal health behavior data was collected to extract multimodal behavioral features; The matching degree and correlation between the drug administration requirements information in the drug knowledge base and the multimodal behavioral characteristics are analyzed to identify the patient's historical behavioral deviations and establish personalized behavioral benchmarks for the patient based on the historical behavioral deviations. A federated learning architecture is adopted to collaboratively train a multimodal temporal anomaly detection model without leaving the local storage environment of the local patient data of multiple participants. Obtain the patient's most recent medication information, and dynamically adjust the benchmark for subsequent behavior judgment based on the deviation between the most recent medication information and the personalized behavior benchmark, thereby generating a dynamically adjusted judgment benchmark. The patient's current multimodal behavioral characteristics are input into the trained multimodal temporal anomaly detection model, and the real-time multimodal temporal behavioral deviations are identified in combination with the dynamically adjusted judgment criteria. Based on the preset medical knowledge graph, the identification results are used to generate intervention basis. Using reinforcement learning algorithms, personalized intervention strategies are generated based on the intervention criteria and patient feedback data on historical interventions, and personalized reminders are then pushed to patients accordingly.

[0008] In addition to the aspects and any possible implementations described above, a further implementation is provided in which the construction of the drug knowledge base and the collection of patients' multimodal health behavior data to extract multimodal behavioral features include: Obtain the raw text data of the drug instruction manual and preprocess the raw text data; Using a named entity recognition model trained based on deep learning, key entities are extracted from preprocessed text data. The key entities include drug name, active ingredient, indications, usage and dosage description, and precautions description. The relationship between the key entities is identified using a relation extraction model. The relationship includes: the dosage relationship between the drug and the single dose, the frequency relationship between the drug and the number of times it is taken per day, the time relationship between the drug and the time point of administration, the interval relationship between the drug and the duration of the administration interval, and the special requirement relationship between the drug and the instructions to take it before or after meals. The extracted key entities and their relationships are structured and stored in the form of triples to construct the drug knowledge base; The system continuously collects patients' health behavior data through multiple sensors, and preprocesses and extracts features from the health behavior data to generate multimodal behavioral features. The multimodal behavioral features include: medication time features, medication dosage features, medication frequency features, meal time features, exercise intensity features, exercise duration features, sleep duration features, and physiological indicator features. The health behavior data includes at least medication record data, dietary record data, exercise monitoring data, sleep monitoring data, and physiological indicator monitoring data.

[0009] In addition to the aspects and any possible implementations described above, a further implementation is provided in which the analysis of the matching degree and correlation between the drug administration requirement information in the drug knowledge base and the multimodal behavioral characteristics to identify the patient's historical behavioral deviations, and the establishment of personalized behavioral benchmarks for the patient based on the historical behavioral deviations, includes: Based on the name of the medication the patient is currently taking, the corresponding medication usage requirements are retrieved from the drug knowledge base. These requirements include the prescribed time of administration, the prescribed dosage, the prescribed frequency of administration, the prescribed interval between administrations, and any special requirements. The extracted medication time feature, medication dosage feature, and medication frequency feature from the multimodal behavioral features are matched with the prescribed medication time, prescribed medication dosage, and prescribed medication frequency in the medication administration requirement information, and the matching deviation values ​​for each feature are calculated. Cross-correlation analysis was performed on the patient's medication time data series and meal time data series to determine the lag time relationship between the two. Granger causality tests were performed on the time series of patients' medication behavior and the time series of physiological indicators to analyze whether changes in medication behavior constituted a statistical predictive cause of changes in physiological indicators. Clustering algorithms were used to perform cluster analysis on the historical multimodal behavioral characteristics of patients to identify the typical behavioral pattern categories of patients; An association rule mining algorithm was used to extract association rules between behavioral patterns and medication requirements from patients' historical behavioral data. Based on the results of the above cross-correlation analysis, Granger causality test, cluster analysis, and association rule mining, historical behavioral deviations of patients' actual medication behavior relative to the prescribed medication requirements are identified; wherein, the historical behavioral deviations include medication time deviations, medication dosage deviations, medication frequency deviations, and deviations from special medication requirements; Based on the preset risk assessment rules, assess the health risk level corresponding to each historical behavioral deviation; Based on the identified types of historical behavioral deviations and their corresponding health risk levels, personalized behavioral benchmarks are established for patients; wherein, the personalized behavioral benchmarks include benchmark medication time, benchmark medication dosage, benchmark medication frequency, benchmark medication interval, and the allowable deviation range for each benchmark item.

[0010] In addition to the aspects and any possible implementations described above, a further implementation is provided, wherein the federated learning architecture, under the condition that the local patient data of multiple participants does not leave their respective local storage environments, collaboratively trains a multimodal temporal anomaly detection model, including: A three-layer federated learning architecture is constructed, consisting of an edge node layer, a regional aggregator layer, and a central server layer. Each medical institution acts as an edge node, a medical consortium or medical group acts as a regional aggregator, and the central server is responsible for the initialization, distribution, and aggregation of global model parameters. The central server initializes the model parameters of the global multimodal temporal anomaly detection model and encrypts the global model parameters before sending them to each edge node. After receiving the global model parameters, each edge node uses the locally stored patient multimodal health behavior data to train a copy of the locally held multimodal temporal anomaly detection model. During the training process, a differential privacy protection mechanism is applied to generate local model gradients, and the local model gradients are encrypted and then uploaded to the corresponding regional aggregator. Each regional aggregator collects the encrypted model gradients uploaded by the edge nodes under its jurisdiction, performs weighted average aggregation using a federated averaging algorithm, generates regional model parameters, and uploads the regional model parameters to the central server. The central server collects the regional model parameters uploaded by each regional aggregator, performs global aggregation calculations to update the global model parameters, and then encrypts and sends the updated global model parameters to each edge node again. The above steps are executed iteratively until the global model converges or the preset number of communication rounds is reached; Each edge node uses local data to fine-tune the converged global model to generate a personalized multimodal temporal anomaly detection model that adapts to the distribution characteristics of local patient data.

[0011] In addition to the aspects and any possible implementations described above, a further implementation is provided, wherein obtaining the patient's most recent medication information and dynamically adjusting the benchmark for subsequent behavior judgment based on the deviation between the most recent medication information and the personalized behavioral benchmark, and generating the dynamically adjusted judgment benchmark, includes: Real-time monitoring of the patient's most recent medication time, dosage, and administration method, generating a recent medication record; Determine the baseline dosing interval based on the drug dosing interval requirements retrieved from the drug knowledge base; When it is detected that the actual time of a patient's medication is delayed compared to the baseline medication time in the personalized behavioral baseline, the judgment baseline time for the next medication is postponed by the same amount of time as the delay. When a patient's actual medication time is detected to be earlier than the baseline medication time, the safety risks of early medication are assessed, and it is checked whether the minimum safe dosing interval requirement for the drug is violated as obtained from the drug knowledge base. A time interval tolerance range is set for the dynamically adjusted judgment benchmark time. When the actual medication interval exceeds the tolerance range, it is determined to be an abnormal medication time and the abnormal event is recorded. The system continuously monitors the time interval deviation of multiple medication administrations and calculates the cumulative deviation value. When the cumulative deviation value exceeds a preset warning threshold, a warning signal is triggered, and the process of re-establishing the personalized behavioral benchmark is initiated. Acquire the patient's current meal time characteristics, exercise characteristics, and physiological indicators; adaptively adjust the baseline time for postprandial medication based on changes in meal time characteristics; assess the impact of changes in exercise intensity and duration on drug metabolism rate and adjust the baseline time for medication accordingly; assess the impact on medication safety based on significant changes in physiological indicators relative to the normal range and adjust the baseline time for medication accordingly.

[0012] In addition to the aspects and any possible implementations described above, a further implementation is provided, wherein the step of inputting the patient's current multimodal behavioral characteristics into the trained multimodal temporal anomaly detection model, and combining the dynamically adjusted judgment criterion to identify real-time multimodal temporal behavioral deviations includes: Load a multimodal temporal anomaly detection model obtained through federated collaborative training; wherein, the multimodal temporal anomaly detection model includes a multimodal encoder, a cross-modal attention fusion layer, and a reconstruction-based anomaly detection network; The patient's current multimodal behavioral features, collected and extracted in real time, are input into the multimodal temporal anomaly detection model, and the following processing is performed: The medication record encoder, motion data encoder, diet information encoder, and physiological indicator encoder in the multimodal encoder are used to perform temporal encoding on the medication time feature, medication dosage feature, medication frequency feature, exercise intensity feature, exercise duration feature, meal time feature, and physiological indicator feature in the multimodal behavioral features, respectively, to generate temporal feature vectors corresponding to each modality. The cross-modal attention fusion layer performs multi-head attention calculation on the temporal feature vectors of each modality to determine the correlation weight matrix between different modal features, and then performs weighted fusion on the temporal feature vectors of each modality according to the correlation weight matrix to generate a unified multimodal fusion feature representation. The multimodal fusion feature representation is input into the reconstruction-based anomaly detection network, and anomaly scores are generated by calculating the reconstruction error of the input features. Based on the abnormality score and the dynamic abnormality threshold set in the dynamically adjusted judgment criteria, it is determined whether the patient's current behavior is abnormal and the corresponding risk level is classified. A temporal pattern clustering algorithm is used to compare the similarity between the multimodal fusion feature representation corresponding to the current behavior and the feature representation of the patient's historical normal behavior pattern to identify specific real-time multimodal temporal behavior deviation categories; wherein, the behavior deviation categories include medication time deviation, medication dosage deviation, medication frequency deviation, and special medication requirement deviation.

[0013] In addition to the aspects and any possible implementations described above, a further implementation is provided, wherein the reasoning based on the recognition results using a preset medical knowledge graph to generate intervention criteria includes: Construct a medical knowledge graph; wherein the medical knowledge graph includes disease entity nodes, behavior entity nodes, intervention entity nodes, drug entity nodes, medication requirement entity nodes, and directed relation edges between the above nodes; the knowledge sources of the medical knowledge graph include at least one of clinical diagnosis and treatment guidelines, medical literature, expert experience knowledge, historical case data, and drug instructions; Based on the collected multimodal health behavior data and patients' electronic medical records, a patient profile is constructed; wherein, the patient profile includes the patient's basic demographic information, disease diagnosis information, health status information, historical behavioral characteristics, psychological status assessment information, social support information, and medication history information; Using the key features in the patient profile and the real-time multimodal temporal behavioral deviation category as query conditions, a graph traversal search is performed in the medical knowledge graph to query the disease-behavior-intervention path related to the current patient status and current behavioral deviation; Based on the retrieved graph path and graph traversal results, a reasoning operation is performed using rule-based reasoning or path-sorting-based reasoning algorithms to generate a health risk assessment report that may be caused by the patient's current behavior, output preliminary intervention suggestions, and combine the health risk assessment report and the preliminary intervention suggestions as the basis for intervention.

[0014] In addition to the aspects and any possible implementations described above, a further implementation is provided, wherein the use of a reinforcement learning algorithm to generate a personalized intervention strategy based on the intervention criteria and patient feedback data on historical interventions includes: Define the state space and action space of reinforcement learning; wherein, the state space includes the patient's physiological state parameters, behavioral state parameters, environmental state parameters, psychological state parameters, drug state parameters, and dynamic state parameters; the action space includes intervention type, reminder time parameter, reminder frequency parameter, reminder content template, and reminder tone parameter; A multi-objective reward function is designed based on the state space and the action space, and a hierarchical reinforcement learning architecture is constructed. The hierarchical reinforcement learning architecture includes a high-level policy network and a low-level policy network. The high-level policy network is used to output the intervention type in each decision cycle, and the low-level policy network is used to optimize the specific action parameters corresponding to the selected intervention type according to the output of the high-level policy network in the decision cycle. The strategy gradient algorithm is used to calculate the cumulative reward value based on the patient's real-time feedback data on historical interventions, and the parameters of the high-level strategy network and the low-level strategy network are updated based on the cumulative reward value. The intervention criteria are used as constraints or prior knowledge of the initial strategy to guide the strategy exploration and optimization direction of the strategy gradient algorithm, thereby generating an optimized personalized intervention strategy.

[0015] In addition to the aspects and any possible implementations described above, a further implementation is provided in which the generation of personalized reminder content and its push to the patient includes: The type of intervention to be performed and the corresponding reminder parameters are determined based on the personalized intervention strategy. Based on the patient's educational background, health knowledge level, and psychological state assessment information, determine the complexity level and expression method of the reminder content; Select the corresponding reminder template from the preset reminder template library according to the intervention type; wherein, the reminder template library includes medication reminder templates, dietary advice templates, exercise encouragement templates, and health monitoring reminder templates; Based on the patient's specific behavioral deviations and the dynamically adjusted judgment criteria, the variable fields in the reminder template are filled in to generate personalized reminder text. Based on the preset urgency level of the reminder content, the patient's preset push preferences, and the dynamically adjusted judgment benchmark time, a push channel is selected and the optimal push timing is determined; wherein, the push channel includes at least one of SMS push, APP message push, voice call push, and smart wearable device push.

[0016] In addition to the aspects and any possible implementations described above, a further implementation is provided, wherein the method further includes: Collect patient feedback data on push notifications; wherein, the feedback data includes medication adherence feedback data, patient satisfaction feedback data, health indicator change feedback data, and behavior pattern change feedback data; The collected feedback data is cleaned, labeled, and integrated to generate a standardized feedback dataset. Based on the standardized feedback dataset, the anomaly detection accuracy and recall of the multimodal temporal anomaly detection model, as well as the effectiveness metrics of the intervention strategy generated by the reinforcement learning algorithm, are evaluated; the effectiveness metrics include changes in medication adherence, behavioral deviation correction rate, and patient satisfaction score. When the evaluation results meet the preset model update trigger conditions, at least one of the following will be updated: the model parameters of the multimodal temporal anomaly detection model, the policy network parameters of the reinforcement learning algorithm, the benchmark parameters in the personalized behavior benchmark, and the entities and relationships in the medical knowledge graph. Through the federated learning architecture, the updated model parameters of each participant are encrypted, uploaded, and aggregated, and the aggregated global model parameters are then distributed back to each participant.

[0017] According to a second aspect of this disclosure, a patient behavior alerting device based on federated learning and multimodal temporal anomaly detection is provided. The device includes: The knowledge base construction and feature extraction module is used to build a drug knowledge base and collect patients' multimodal health behavior data to extract multimodal behavioral features; The behavior-drug correlation analysis module is used to analyze the matching degree and correlation between the drug administration requirements information in the drug knowledge base and the multimodal behavioral characteristics, so as to identify the patient's historical behavioral deviations and establish personalized behavioral benchmarks for the patient based on the historical behavioral deviations. The federated collaborative training module is used to collaboratively train a multimodal temporal anomaly detection model using a federated learning architecture, without leaving the local patient data of multiple participants in their respective local storage environments. The dynamic baseline adjustment module is used to obtain the patient's most recent medication information and dynamically adjust the baseline for subsequent behavior judgment based on the deviation between the most recent medication information and the personalized behavior baseline, thereby generating a dynamically adjusted judgment baseline. The anomaly detection and reasoning module is used to input the patient's current multimodal behavioral characteristics into the trained multimodal temporal anomaly detection model, identify real-time multimodal temporal behavioral deviations in combination with the dynamically adjusted judgment criteria, and generate intervention basis based on the identification results according to the preset medical knowledge graph. The strategy generation and push module is used to generate personalized intervention strategies based on the intervention criteria and patient feedback data on historical interventions using reinforcement learning algorithms, and to generate personalized reminder content to push to patients accordingly.

[0018] According to a third aspect of this disclosure, an electronic device is provided. The electronic device includes a memory and a processor, wherein the memory stores a computer program, and the processor executes the program to implement the method described above.

[0019] According to a fourth aspect of this disclosure, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the methods according to the first and / or second aspects of this disclosure.

[0020] First, by constructing a drug knowledge base and collecting multimodal health behavior data to extract features, discrete drug instruction information can be transformed into computable structured knowledge. At the same time, by integrating multi-source data such as medication, diet, exercise, sleep and physiological indicators, a comprehensive patient behavior profile can be provided for subsequent analysis.

[0021] Secondly, by analyzing the matching degree and correlation between medication requirements and behavioral characteristics and identifying historical behavioral deviations, a personalized behavioral benchmark is established for each patient. This process not only quantifies the differences between patients and standard protocols, but also reveals the association between hidden behavioral patterns and health risks through clustering and association rule mining. This makes the benchmark setting no longer a uniform and fixed threshold, but a dynamic reference system that varies from person to person and from medication to medication.

[0022] Based on this, a federated learning architecture is adopted for cross-institutional collaborative training. The local data of each participant can be used to jointly optimize the multimodal temporal anomaly detection model without leaving the storage environment. This not only solves the problem of poor model generalization ability caused by insufficient data from a single institution, but also strictly protects patient privacy through differential privacy and encrypted upload mechanisms, making large-scale, distributed patient behavior learning possible.

[0023] Furthermore, by obtaining the patient's most recent medication information and comparing it with personalized benchmarks, the system dynamically adjusts the benchmarks for subsequent behavior judgments. For example, if a patient's medication is delayed due to a business trip, the system automatically postpones the next judgment time and adjusts the reminder timing accordingly. This adaptive mechanism based on actual medication events completely changes the rigid mode of traditional fixed-interval judgments, and can adapt to real-world scenarios such as holidays, physical discomfort, and temporary changes in travel plans, significantly reducing the misjudgment rate and invalid reminders.

[0024] Next, the current multimodal behavioral features are input into the trained anomaly detection model, and real-time behavioral deviations are identified by combining the dynamically adjusted baseline. Then, intervention basis is generated through medical knowledge graph reasoning. This process utilizes cross-modal attention fusion to deeply encode multi-source time-series data, enabling the model to simultaneously perceive the complex interactions between medication time deviations, dietary changes, and fluctuations in physiological indicators. The graph traversal and reasoning of the knowledge graph provide traceable clinical logic support for intervention recommendations, solving the pain point of traditional "black box" models lacking interpretability.

[0025] Finally, reinforcement learning algorithms are used to dynamically optimize personalized intervention strategies based on intervention criteria and patient feedback data on historical interventions, and generate reminder content to push to patients. The intervention type is selected through a high-level strategy network, the reminder parameters are refined through a low-level strategy network, and a multi-objective reward function guides the strategy to evolve in the direction of improving compliance, satisfaction and health outcomes, thus forming a closed loop of "assessment-intervention-feedback-optimization".

[0026] In summary, this disclosure addresses the comprehensive shortcomings of existing technologies in terms of flexibility, accuracy, privacy protection, and interpretability through a series of interconnected steps, from data collection, knowledge construction, personalized benchmark establishment, federated collaborative training, dynamic adaptive judgment, multimodal anomaly detection, knowledge reasoning to reinforcement learning optimization. This enables intelligent, personalized, and dynamically adaptive management of patient behavioral intervention reminders.

[0027] It should be understood that the description in the Summary of the Invention is not intended to limit the key or essential features of the embodiments of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0028] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. The drawings are provided for a better understanding of the invention and are not intended to limit the scope of this disclosure. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein: Figure 1 A flowchart of a patient behavior alerting method based on federated learning and multimodal temporal anomaly detection provided by an embodiment of this disclosure is shown; Figure 2 A structural diagram of a patient behavior alert device based on federated learning and multimodal temporal anomaly detection provided by an embodiment of this disclosure is shown. Figure 3 A structural diagram of an exemplary electronic device capable of implementing embodiments of the present disclosure is shown. Detailed Implementation

[0029] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this disclosure, and not all embodiments. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.

[0030] Furthermore, the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.

[0031] First, by constructing a drug knowledge base and collecting multimodal health behavior data to extract features, discrete drug instruction information can be transformed into computable structured knowledge. At the same time, by integrating multi-source data such as medication, diet, exercise, sleep and physiological indicators, a comprehensive patient behavior profile can be provided for subsequent analysis.

[0032] Secondly, by analyzing the matching degree and correlation between medication requirements and behavioral characteristics and identifying historical behavioral deviations, a personalized behavioral benchmark is established for each patient. This process not only quantifies the differences between patients and standard protocols, but also reveals the association between hidden behavioral patterns and health risks through clustering and association rule mining. This makes the benchmark setting no longer a uniform and fixed threshold, but a dynamic reference system that varies from person to person and from medication to medication.

[0033] Based on this, a federated learning architecture is adopted for cross-institutional collaborative training. The local data of each participant can be used to jointly optimize the multimodal temporal anomaly detection model without leaving the storage environment. This not only solves the problem of poor model generalization ability caused by insufficient data from a single institution, but also strictly protects patient privacy through differential privacy and encrypted upload mechanisms, making large-scale, distributed patient behavior learning possible.

[0034] Furthermore, by obtaining the patient's most recent medication information and comparing it with personalized benchmarks, the system dynamically adjusts the benchmarks for subsequent behavior judgments. For example, if a patient's medication is delayed due to a business trip, the system automatically postpones the next judgment time and adjusts the reminder timing accordingly. This adaptive mechanism based on actual medication events completely changes the rigid mode of traditional fixed-interval judgments, and can adapt to real-world scenarios such as holidays, physical discomfort, and temporary changes in travel plans, significantly reducing the misjudgment rate and invalid reminders.

[0035] Next, the current multimodal behavioral features are input into the trained anomaly detection model, and real-time behavioral deviations are identified by combining the dynamically adjusted baseline. Then, intervention basis is generated through medical knowledge graph reasoning. This process utilizes cross-modal attention fusion to deeply encode multi-source time-series data, enabling the model to simultaneously perceive the complex interactions between medication time deviations, dietary changes, and fluctuations in physiological indicators. The graph traversal and reasoning of the knowledge graph provide traceable clinical logic support for intervention recommendations, solving the pain point of traditional "black box" models lacking interpretability.

[0036] Finally, reinforcement learning algorithms are used to dynamically optimize personalized intervention strategies based on intervention criteria and patient feedback data on historical interventions, and generate reminder content to push to patients. The intervention type is selected through a high-level strategy network, the reminder parameters are refined through a low-level strategy network, and a multi-objective reward function guides the strategy to evolve in the direction of improving compliance, satisfaction and health outcomes, thus forming a closed loop of "assessment-intervention-feedback-optimization".

[0037] In summary, this disclosure addresses the comprehensive shortcomings of existing technologies in terms of flexibility, accuracy, privacy protection, and interpretability through a series of interconnected steps, from data collection, knowledge construction, personalized benchmark establishment, federated collaborative training, dynamic adaptive judgment, multimodal anomaly detection, knowledge reasoning to reinforcement learning optimization. This enables intelligent, personalized, and dynamically adaptive management of patient behavioral intervention reminders.

[0038] Figure 1 A flowchart of a patient behavior alerting method based on federated learning and multimodal temporal anomaly detection provided by embodiments of this disclosure is shown, such as... Figure 1 As shown, the patient behavior reminder method 100 based on federated learning and multimodal temporal anomaly detection may include the following steps: S110: Construct a drug knowledge base and collect patients' multimodal health behavior data to extract multimodal behavioral features.

[0039] In some embodiments, constructing a drug knowledge base and collecting patients' multimodal health behavior data to extract multimodal behavioral features includes: Obtain the raw text data of the drug instruction manual and preprocess the raw text data; Using a named entity recognition model trained based on deep learning, key entities are extracted from preprocessed text data. The key entities include drug name, active ingredient, indications, usage and dosage description, and precautions description. The relationship between the key entities is identified using a relation extraction model. The relationship includes: the dosage relationship between the drug and the single dose, the frequency relationship between the drug and the number of times it is taken per day, the time relationship between the drug and the time point of administration, the interval relationship between the drug and the duration of the administration interval, and the special requirement relationship between the drug and the instructions to take it before or after meals. The extracted key entities and their relationships are structured and stored in the form of triples to construct the drug knowledge base; The system continuously collects patients' health behavior data through multiple sensors, and preprocesses and extracts features from the health behavior data to generate multimodal behavioral features. The multimodal behavioral features include: medication time features, medication dosage features, medication frequency features, meal time features, exercise intensity features, exercise duration features, sleep duration features, and physiological indicator features. The health behavior data includes at least medication record data, dietary record data, exercise monitoring data, sleep monitoring data, and physiological indicator monitoring data.

[0040] Specifically, the construction of a drug knowledge base requires obtaining the original text data of the drug instructions in practice. Electronic instructions can be read directly; for paper instructions, OCR (Optical Character Recognition) technology can be used for scanning and conversion. For example, open-source OCR engines such as EasyOCR and Tesseract can be used to extract the text from the paper document. Then, tools like olmOCR can be used for post-processing to remove noise such as headers, footers, and table borders, converting the image-based document into a clear, naturally readable plain text format. After obtaining the text data, preprocessing is also necessary, including removing irrelevant characters, blank lines, headers, footers, and other noise. The text is then segmented into sentences and words, and parts of speech are tagged as needed, laying the foundation for subsequent entity extraction and relation recognition.

[0041] Furthermore, for the extraction of key entities, deep learning-based Named Entity Recognition (NER) models can be used to automatically identify various key entities from the preprocessed text. Specifically, medical NER models fine-tuned from pre-trained language models such as RoBERTa or BERT can be employed. These models, trained on large-scale clinical and biomedical texts, can accurately identify key entities such as drug names, active ingredients, indications, usage and dosage descriptions, and precautions. Considering the significant heterogeneity of medical entities in drug instructions for different diseases, model training needs to cover a sufficiently large number of drug types to improve the model's generalization ability to different instruction texts.

[0042] Specifically, after extracting key entities, it is necessary to further identify the relationships between these entities. This can be achieved through relation extraction models, such as using a joint extraction architecture based on pre-trained language models, or utilizing traditional methods based on part-of-speech tagging and dependency parsing to identify dosage relationships between "drug - single dose", frequency relationships between "drug - number of daily doses", temporal relationships between "drug - time of administration", interval relationships between "drug - duration of dosing", and special requirements relationships between "drug - instructions for administration before or after meals". Through entity extraction and relation extraction, the originally unstructured instruction manual text can be transformed into structured knowledge units.

[0043] Subsequently, the extracted key entities and relationships are structured and stored in the form of triples to construct a drug knowledge base. For example, triples such as (drug name, dosage, 0.5 grams per dose) and (drug name, frequency, twice daily) can be stored in a graph database (such as Neo4j) or a relational database, and corresponding indexes can be created to support fast querying and retrieval. This drug knowledge base also needs to support dynamic updates. When a new drug is launched or the drug instructions are updated, a new parsing process should be automatically triggered, or the knowledge base should be supplemented and corrected through manual review, thereby maintaining the timeliness and accuracy of the knowledge base.

[0044] On the other hand, in the collection and feature extraction of multimodal behavioral data from patients, multi-source sensors can be used to continuously acquire patients' health behavior data. Specifically, medication record data (including drug name, medication time, dosage, etc.) can be collected using smart pillboxes or electronic medication record devices; activity monitoring data (steps, exercise duration, heart rate, etc.) and sleep monitoring data can be collected using wearable devices (such as smartwatches and wristbands); dietary record data (meal time, food type, etc.) can be collected using patient mobile health apps or smart tableware; and physiological indicator monitoring data can be collected using home health monitoring devices (blood glucose meters, blood pressure monitors, etc.). These data come from diverse sources and have varying sampling frequencies. After collection, they need to undergo a unified preprocessing procedure, including removing noise data and outliers, converting data of different formats to a unified format, aligning data with different timestamps, and standardizing numerical data to ensure consistency and comparability in subsequent feature extraction.

[0045] Specifically, after preprocessing, various behavioral characteristics are extracted from the above data, including medication time characteristics (the specific time of actual medication administration, the regularity of medication time, etc.), medication dosage characteristics (the degree to which the actual dosage conforms to the prescribed dosage, etc.), medication frequency characteristics (the actual number of times of medication administration, the stability of the medication interval, etc.), meal time characteristics (the specific time and regularity of breakfast, lunch, and dinner), exercise intensity characteristics and exercise duration characteristics, sleep duration characteristics, and physiological indicators such as blood glucose, blood pressure, and heart rate. These multimodal behavioral characteristics cover multiple dimensions of patients' daily life, including medication, diet, exercise, sleep, and health status, providing a rich data foundation for subsequent behavioral analysis, deviation identification, and personalized intervention. By correlating these characteristics with the medication requirements information in the drug knowledge base, personalized behavioral benchmarks can be established for each patient, enabling precise management of medication administration behavior and health behavior.

[0046] S120, Analyze the matching degree and correlation between the drug administration requirements information in the drug knowledge base and the multimodal behavioral characteristics to identify the patient's historical behavioral deviations, and establish personalized behavioral benchmarks for the patient based on the historical behavioral deviations.

[0047] In some embodiments, analyzing the matching degree and correlation between the drug administration requirements information in the drug knowledge base and the multimodal behavioral characteristics to identify the patient's historical behavioral deviations and establishing personalized behavioral benchmarks for the patient based on the historical behavioral deviations includes: Based on the name of the medication the patient is currently taking, the corresponding medication usage requirements are retrieved from the drug knowledge base. These requirements include the prescribed time of administration, the prescribed dosage, the prescribed frequency of administration, the prescribed interval between administrations, and any special requirements. The extracted medication time feature, medication dosage feature, and medication frequency feature from the multimodal behavioral features are matched with the prescribed medication time, prescribed medication dosage, and prescribed medication frequency in the medication administration requirement information, and the matching deviation values ​​for each feature are calculated. Cross-correlation analysis was performed on the patient's medication time data series and meal time data series to determine the lag time relationship between the two. Granger causality tests were performed on the time series of patients' medication behavior and the time series of physiological indicators to analyze whether changes in medication behavior constituted a statistical predictive cause of changes in physiological indicators. Clustering algorithms were used to perform cluster analysis on the historical multimodal behavioral characteristics of patients to identify the typical behavioral pattern categories of patients; An association rule mining algorithm was used to extract association rules between behavioral patterns and medication requirements from patients' historical behavioral data. Based on the results of the above cross-correlation analysis, Granger causality test, cluster analysis, and association rule mining, historical behavioral deviations of patients' actual medication behavior relative to the prescribed medication requirements are identified; wherein, the historical behavioral deviations include medication time deviations, medication dosage deviations, medication frequency deviations, and deviations from special medication requirements; Based on the preset risk assessment rules, assess the health risk level corresponding to each historical behavioral deviation; Based on the identified types of historical behavioral deviations and their corresponding health risk levels, personalized behavioral benchmarks are established for patients; wherein, the personalized behavioral benchmarks include benchmark medication time, benchmark medication dosage, benchmark medication frequency, benchmark medication interval, and the allowable deviation range for each benchmark item.

[0048] Specifically, based on the name of the medication the patient is currently taking (e.g., obtained from an electronic prescription or medication record), a search is performed in an established drug knowledge base to retrieve the corresponding dosage instructions. These instructions include at least the prescribed time of administration (e.g., "30 minutes after breakfast"), the prescribed dosage (e.g., "0.5 grams per dose"), the prescribed frequency of administration (e.g., "twice daily"), the prescribed interval of administration (e.g., "every 8 hours"), and any special instructions (e.g., "swallow whole," "chew," or "take with food"). The search can be performed quickly using query languages ​​for graph databases (e.g., Cypher) or SQL statements for relational databases, returning structured tuples of dosage instructions. Then, the extracted multimodal behavioral features—including medication time, dosage, and frequency—are matched item by item with the prescribed time, dosage, and frequency of administration, respectively. The matching process can calculate medication time deviation based on the absolute value of the timestamp difference (e.g., the minute difference between the actual medication time and the prescribed time), calculate dosage deviation by comparing the ratio or absolute difference between the actual dose and the prescribed dose, and calculate frequency deviation by comparing the difference between the actual number of doses and the prescribed number of doses. For each match, an initial tolerance range (e.g., ±30 minutes) can be set to distinguish between normal fluctuations and significant deviations, thereby obtaining a quantified matching deviation value.

[0049] Specifically, after completing the basic matching, deeper correlation analysis is needed to reveal the dynamic relationship between behavior and medication requirements. An important analysis involves cross-correlation analysis of the patient's medication time data series and meal time data series. The cross-correlation function can be calculated by converting the two event sequences into binary sequences sampled at fixed time intervals (e.g., 15 minutes) (1 indicates the event occurred, 0 indicates it did not occur), and then calculating the correlation coefficient at different lag times. The lag time with the largest correlation coefficient represents the typical delay time of medication behavior relative to meal behavior. For example, if the largest correlation coefficient occurs at a lag of 30 minutes, it indicates that the patient habitually takes medication half an hour after meals. This result can be compared with the "take after meals" requirement in the drug instructions to determine whether the prescribed post-meal medication window is met.

[0050] Another in-depth analysis involves performing Granger causality tests on the time series of patients' medication behavior and physiological indicators. In practice, the original discrete medication events need to be converted into equally spaced time series (e.g., hourly or daily medication adherence rates, mean medication time deviations, etc.), while simultaneously sampling physiological indicators (e.g., blood glucose, blood pressure) at the same time intervals. Then, two vector autoregressive models are established, one with a medication behavior history term and the other without. An F-test is used to determine whether adding the medication behavior history term significantly improves the predictive ability of the current physiological indicator. If the test result is significant (p-value less than 0.05), it indicates that the change in medication behavior statistically constitutes a predictive cause of the physiological indicator change; for example, two consecutive days of medication delay can predict elevated blood glucose on the third day. Identifying this causal relationship is crucial for assessing the actual health impact of behavioral deviations.

[0051] Furthermore, to gain a holistic understanding of patients' behavioral patterns, clustering algorithms can be used to perform cluster analysis on patients' historical multimodal behavioral characteristics. For example, daily feature vectors (including the average deviation of medication time, regularity of meal times, exercise duration, sleep duration, and average blood glucose levels) can be input into DBSCAN or K-means clustering algorithms. The optimal number of clusters can be determined using silhouette coefficients, thereby identifying typical behavioral pattern categories for patients, such as "regular medication taker," "occasionally missed dose," "frequently missed dose," and "well-coordinated diet and exercise." The clustering results can help the system determine whether a patient's current behavior deviates from their usual pattern, rather than just from general standards.

[0052] Simultaneously, association rule mining algorithms (such as Apriori or FP-Growth) were employed to mine association rules between behavioral patterns and medication adherence from patients' historical behavioral data. Daily behavioral events (such as "taking medication after meals," "taking medication on an empty stomach," "exercising after meals," and "missing a medication dose") were treated as item sets. By calculating support, confidence, and lift, strong association rules such as "taking medication after meals and exercising after meals → blood glucose reaches target levels" or "taking medication on an empty stomach → stomach discomfort" were discovered. These rules not only reveal the relationship between behavior and medication efficacy but also provide actionable causal clues for subsequent interventions.

[0053] Specifically, by combining the results of the aforementioned cross-correlation analysis, Granger causality test, cluster analysis, and association rule mining, we can comprehensively identify historical behavioral deviations in patients' actual medication use compared to the prescribed dosage. These deviations may include: medication timing deviations (e.g., taking medication before meals when prescribed after meals), dosage deviations (e.g., consistently taking half a tablet when prescribed a full tablet), medication frequency deviations (e.g., taking medication three times a day when only twice a day is recommended), and deviations from specific dosage requirements (e.g., chewing tablets when prescribed to be swallowed whole). For each identified deviation, its health risk level is assessed according to pre-defined risk assessment rules (e.g., degree of deviation, duration, intensity of impact on physiological indicators, etc.), classifying it as low-risk, medium-risk, or high-risk. For example, an occasional 15-minute delay in medication use might be considered low-risk, while missing one dose daily for a week and resulting in a significant increase in blood glucose levels would be considered high-risk.

[0054] Finally, based on the identified types of historical behavioral deviations and their corresponding health risk levels, personalized behavioral benchmarks are established for patients. These benchmarks are no longer fixed, universally applicable thresholds, but are adjusted to the specific circumstances and habits of each patient. For example, if cross-correlation analysis reveals that a patient habitually takes medication 45 minutes after lunch, while the drug instructions only require "taken after meals," the benchmark medication time can be set at 45 minutes after lunch, with a tolerance range of ±15 minutes. For patients with long-term dose deviations, if the assessed risk is manageable and physiological indicators are stable, the benchmark dose can be appropriately adjusted after confirmation by a physician. For medication intervals, a personalized benchmark interval can be set between the average interval in the patient's historical medication records and the minimum safe interval specified by the drug. Simultaneously, a tolerance range for each benchmark item (medication time, dosage, frequency, and interval) is set, for example, a tolerance of ±20 minutes for time deviation and ±10% for dosage deviation. This personalized behavioral benchmark respects the pharmacological requirements of the drug while fully adapting to the patient's lifestyle and individual differences, providing a reasonable and flexible reference system for subsequent dynamic adjustments and anomaly detection.

[0055] S130 employs a federated learning architecture, enabling collaborative training of a multimodal temporal anomaly detection model while keeping local patient data from multiple participants within their respective local storage environments.

[0056] In some embodiments, the federated learning architecture, which collaboratively trains a multimodal temporal anomaly detection model without the local patient data of multiple participants leaving their respective local storage environments, includes: A three-layer federated learning architecture is constructed, consisting of an edge node layer, a regional aggregator layer, and a central server layer. Each medical institution acts as an edge node, a medical consortium or medical group acts as a regional aggregator, and the central server is responsible for the initialization, distribution, and aggregation of global model parameters. The central server initializes the model parameters of the global multimodal temporal anomaly detection model and encrypts the global model parameters before sending them to each edge node. After receiving the global model parameters, each edge node uses the locally stored patient multimodal health behavior data to train a copy of the locally held multimodal temporal anomaly detection model. During the training process, a differential privacy protection mechanism is applied to generate local model gradients, and the local model gradients are encrypted and then uploaded to the corresponding regional aggregator. Each regional aggregator collects the encrypted model gradients uploaded by the edge nodes under its jurisdiction, performs weighted average aggregation using a federated averaging algorithm, generates regional model parameters, and uploads the regional model parameters to the central server. The central server collects the regional model parameters uploaded by each regional aggregator, performs global aggregation calculations to update the global model parameters, and then encrypts and sends the updated global model parameters to each edge node again. The above steps are executed iteratively until the global model converges or the preset number of communication rounds is reached; Each edge node uses local data to fine-tune the converged global model to generate a personalized multimodal temporal anomaly detection model that adapts to the distribution characteristics of local patient data.

[0057] Specifically, a federated learning architecture is used for the collaborative training of a multimodal temporal anomaly detection model. Its core objective is to fully utilize local data distributed across multiple medical institutions (such as tertiary hospitals, community health service centers, and family doctor teams) without disclosing the original patient data, and to jointly train a high-quality model with strong generalization ability that can adapt to the data distribution of various institutions. Specifically, a three-tiered federated learning architecture needs to be constructed first: the bottom layer is the edge node layer, typically comprised of various medical institutions (such as hospital information departments or clinical data centers), with each edge node possessing locally stored multimodal health behavior data of patients; the middle layer is the regional aggregator layer, usually undertaken by medical consortia or medical groups, responsible for aggregating model updates from multiple edge nodes within their jurisdiction; the top layer is the central server layer, generally deployed in national or industry-level trusted computing centers, responsible for global model initialization, parameter distribution, global aggregation, and iteration management. This layered design effectively reduces the communication pressure on the central server while complying with my country's policy requirements for hierarchical management of medical data.

[0058] Before training begins, the central server needs to initialize a global multimodal temporal anomaly detection model. This model's network structure typically includes a multimodal encoder (e.g., a medication record LSTM encoder, motion data CNN encoder, dietary information BERT encoder, physiological indicator TCN encoder), a cross-modal attention fusion layer, and a reconstruction-based anomaly detection network (e.g., a variational autoencoder or Transformer decoder). Initialization can employ Xavier or Kaiming uniform initialization methods to ensure the model parameters are within a reasonable range. After initialization, the central server encrypts the global model parameters. Common encryption methods include transport layer encryption using TLS / SSL protocols, or more advanced cryptographic techniques such as homomorphic encryption and secret sharing to ensure the confidentiality of the parameters during transmission. The encrypted global model parameters are then distributed to regional aggregators via the network, and further distributed by the regional aggregators to their respective edge nodes.

[0059] After receiving the global model parameters, each edge node loads them into its locally held model copy. Then, each edge node begins local training using locally stored patient multimodal health behavior data. To protect patient privacy, a differential privacy protection mechanism is applied during training. Specifically, during each mini-batch training, the calculated gradients are pruned, limiting the upper limit of each sample's contribution to the gradient (e.g., setting the gradient L2 norm pruning threshold to 1.0). Random noise conforming to a Laplace or Gaussian distribution is then added to the pruned gradient. The scale of the noise is determined based on the privacy budget ε and δ, typically using the Differential Privacy Stochastic Gradient Descent (DP-SGD) algorithm. The noisy gradient prevents attackers from deriving any specific information about a single patient from the gradient, thus providing mathematically quantifiable privacy protection. After gradient noisification, the edge node encrypts the local model gradient (i.e., the partial derivatives of the loss function with respect to the model parameters), using Paillier homomorphic encryption or elliptic curve cryptography algorithms to ensure the security of the gradient during upload. The encrypted gradient is uploaded to the corresponding regional aggregator via a secure channel.

[0060] Specifically, the regional aggregator collects the encrypted model gradients uploaded by all edge nodes within its jurisdiction and then performs regional-level model aggregation. The aggregation algorithm typically employs a variant of the FedAvg algorithm, which weights the gradients of each edge node. The weights can be determined based on the amount of data, data quality, or number of training epochs for each node. For example, nodes with larger amounts of data are assigned higher weights, as shown in the formula: ,in Let be the number of samples for the i-th edge node. This provides the model parameters or gradients. Since the gradients are encrypted, the region aggregator can perform addition and scalar multiplication on the encrypted domain to obtain encrypted region model parameters. Subsequently, the region aggregator uploads the encrypted region model parameters to the central server. This region aggregation step effectively reduces the communication burden on the central server and simultaneously achieves the initial fusion of models within the region.

[0061] Specifically, after collecting the regional model parameters uploaded by all regional aggregators, the central server performs a global aggregation operation. Global aggregation also employs a weighted average strategy, with weights set based on the total number of patients or data distribution characteristics in each region. To defend against potential malicious attacks (such as poisoning attacks), the central server can also introduce a secure aggregation protocol to remove abnormal regional model updates. After aggregation, the central server obtains a new round of global model parameters and redistributes them to each regional aggregator and edge node via an encrypted channel. This process requires multiple iterations, typically referred to as a communication round. In actual training, the maximum number of communication rounds can be set to 50 to 100, or an early stopping condition can be set: training stops when the anomaly detection accuracy of the global model on the validation set no longer improves for five consecutive rounds. After each round, the central server can monitor convergence by evaluating model performance on a public test set.

[0062] Specifically, after multiple rounds of iterative training, the global model gradually converges. At this point, the central server saves the final global model parameters and distributes them to all participants. However, due to differences in the distribution of local data among various medical institutions (for example, tertiary hospitals treat more critically ill patients, while community health service centers mainly manage chronic diseases), directly using a unified global model may not achieve optimal local performance. Therefore, each edge node can use its own local data to fine-tune the converged global model. Fine-tuning typically uses a small learning rate (e.g., one-tenth of the global training learning rate) and fewer training epochs (e.g., 1 to 3 cycles), updating only the parameters of the last few layers of the model or making lightweight adjustments to all parameters. The resulting personalized multimodal temporal anomaly detection model absorbs general behavioral pattern knowledge from cross-institutional data while fully adapting to the specific behavioral characteristics of patients within the institution, thus achieving higher anomaly detection accuracy and lower false positive rate in actual deployment. Through this federal collaborative training process, multiple healthcare institutions can jointly build a robust and geographically adaptable patient behavior analysis model without sharing original patient data and while complying with data privacy regulations. This provides a solid technical foundation for subsequent anomaly detection and personalized intervention.

[0063] S140, obtain the patient's most recent medication information, and dynamically adjust the benchmark for subsequent behavior judgment based on the deviation between the most recent medication information and the personalized behavior benchmark, and generate a dynamically adjusted judgment benchmark.

[0064] In some embodiments, obtaining the patient's most recent medication information and dynamically adjusting the benchmark for subsequent behavior judgment based on the deviation between the most recent medication information and the personalized behavioral benchmark, to generate a dynamically adjusted judgment benchmark, includes: Real-time monitoring of the patient's most recent medication time, dosage, and administration method, generating a recent medication record; Determine the baseline dosing interval based on the drug dosing interval requirements retrieved from the drug knowledge base; When it is detected that the actual time of a patient's medication is delayed compared to the baseline medication time in the personalized behavioral baseline, the judgment baseline time for the next medication is postponed by the same amount of time as the delay. When a patient's actual medication time is detected to be earlier than the baseline medication time, the safety risks of early medication are assessed, and it is checked whether the minimum safe dosing interval requirement for the drug is violated as obtained from the drug knowledge base. A time interval tolerance range is set for the dynamically adjusted judgment benchmark time. When the actual medication interval exceeds the tolerance range, it is determined to be an abnormal medication time and the abnormal event is recorded. The system continuously monitors the time interval deviation of multiple medication administrations and calculates the cumulative deviation value. When the cumulative deviation value exceeds a preset warning threshold, a warning signal is triggered, and the process of re-establishing the personalized behavioral benchmark is initiated. Acquire the patient's current meal time characteristics, exercise characteristics, and physiological indicators; adaptively adjust the baseline time for postprandial medication based on changes in meal time characteristics; assess the impact of changes in exercise intensity and duration on drug metabolism rate and adjust the baseline time for medication accordingly; assess the impact on medication safety based on significant changes in physiological indicators relative to the normal range and adjust the baseline time for medication accordingly.

[0065] Specifically, through active data entry via smart pillboxes, electronic medication record devices, or patient mobile terminals, the system monitors in real time the patient's most recent medication information, including the time of administration (accurate to the minute), dosage (e.g., one tablet, half a tablet, or milligrams), and administration method (e.g., swallowed whole, chewed, or taken with food), generating a structured recent medication record. This record will serve as the time reference point for subsequent dynamic adjustments. Simultaneously, the system retrieves the corresponding dosing interval requirements for the drug from the drug knowledge base, such as specifying "every 8 hours" or "every 12 hours," and uses this as the baseline dosing interval. It is important to note that the baseline dosing interval may differ from the baseline dosing interval in the personalized behavioral benchmark: the interval in the personalized benchmark already considers minor adjustments based on the patient's historical habits, while the interval retrieved here from the knowledge base is the standard interval specified in the drug's instructions. Both will be used in combination in subsequent calculations.

[0066] Specifically, when a delay is detected in the patient's actual medication time compared to the baseline medication time set in the personalized behavioral baseline (e.g., the baseline time is 8:00 AM, the actual medication time is 10:30 AM, a delay of 2.5 hours), the system will automatically postpone the benchmark time for the next medication by the same amount of time as the delay. The specific calculation logic is: new benchmark time for the next medication time = actual medication time + baseline medication interval. For example, if the second medication was originally scheduled for 4:00 PM, after the delay, the new benchmark time becomes 10:30 + 8:00 = 6:30 PM, and the corresponding reminder time will be adjusted from around 3:30 PM to around 6:00 PM. This postponement mechanism ensures that the blood drug concentration in the body can be maintained at a relatively stable level, avoiding subsequent medication intervals that are too short or too long due to a single delay. If a patient's actual medication time is detected to be earlier than the baseline medication time (e.g., the baseline time is 8:00 AM, but the patient actually took the medication at 7:00 AM), the safety risk of early medication needs to be assessed. In this case, it's necessary to calculate not only the length of the early dose but also to check whether the interval between this dose and the previous actual dose violates the minimum safe dosing interval requirement specified for the medication (e.g., some medications require an interval of at least 6 hours). If the actual interval is less than the minimum safe interval, a safety warning should be issued immediately, indicating the patient may face the risk of medication overdose, and consultation with a doctor should be recommended. If the actual interval is still within the safe range, the early dosing can be recorded without adjusting the subsequent baseline, or the next judgment baseline can be slightly adjusted to avoid cumulative bias.

[0067] Specifically, to enhance the robustness of the judgment, a tolerance range for the time interval needs to be set for the dynamically adjusted judgment benchmark time. This tolerance range can be set based on the drug's half-life and therapeutic window width. For example, a tolerance of ±1 hour can be set for drugs with a longer half-life, while a tolerance of ±15 minutes can be set for drugs with a shorter half-life. When the actual dosing interval (the difference between the current actual dosing time and the previous actual dosing time) exceeds this tolerance range, it is judged as an abnormal dosing time, and the abnormal event (including the abnormality type, degree of deviation, and time of occurrence) is recorded. Simultaneously, the time interval deviation of multiple dosing sessions is continuously monitored, and the cumulative deviation value is calculated. The cumulative deviation can be obtained by summing up multiple consecutive interval deviations; for example, if the first delay is 1 hour, the second is 0.5 hours, and the third is 0.2 hours, the cumulative deviation is 1.3 hours. When the cumulative deviation value exceeds a preset warning threshold (e.g., exceeding 3 hours cumulatively or causing a diurnal rhythm shift of more than 2 days), a warning signal is triggered, indicating that the patient's medication habits have undergone a significant systematic shift, and the original personalized behavioral benchmark may no longer be applicable. At this point, the process of re-establishing personalized behavioral benchmarks will be automatically initiated, which involves re-executing the correlation analysis and cluster analysis in step S120 to generate updated benchmark parameters.

[0068] In addition to dynamic adjustments based on the medication event itself, more refined adaptive adjustments are needed by incorporating the patient's current meal timing characteristics, exercise patterns, and physiological indicators. Specifically, by acquiring the patient's current meal timing characteristics in real time (e.g., extracting the start time of the most recent meal from dietary records), a new baseline time for post-meal medication can be automatically calculated for medications that require "post-meal administration." For example, if a patient's dinner time is delayed from the usual 18:30 to 20:00, the baseline time for post-meal medication administration will be adjusted accordingly from 19:00 to 20:30, and a reminder will be issued around 20:00 stating, "A change in your dinner time has been detected; please take metformin 30 minutes after dinner." Simultaneously, the impact on drug metabolism rate can be assessed based on changes in exercise intensity and duration. For example, if a patient engages in high-intensity, prolonged exercise after taking medication, it may accelerate drug metabolism and lower blood drug concentration. In this case, the baseline interval for the next dose can be appropriately shortened (e.g., adjusted from 8 hours to 7.5 hours) or the patient may be advised to increase the dosage (subject to physician confirmation). Conversely, if a patient remains sedentary for extended periods, their metabolism slows down, and the benchmark interval can be appropriately extended. Furthermore, it is necessary to assess the impact on medication safety based on significant changes in physiological indicators relative to the normal range. For example, if a patient's blood glucose level is consistently monitored as significantly below the normal range (hypoglycemia risk), the timing or dosage of the current hypoglycemic medication should be adjusted, and if necessary, an alert should be issued and immediate glucose supplementation or contact with a doctor recommended. Through these multi-dimensional adaptive adjustments, the dynamic benchmark adjustment process not only responds to time deviations in individual medication administration but also fully integrates changes in life circumstances and physiological states, ensuring that the behavioral judgment benchmark remains in a dynamic balance that matches the patient's actual condition, thereby greatly improving the accuracy and safety of medication management.

[0069] S150, the patient's current multimodal behavioral characteristics are input into the trained multimodal temporal anomaly detection model, and the real-time multimodal temporal behavioral deviation is identified in combination with the dynamically adjusted judgment benchmark. Based on the preset medical knowledge graph, reasoning is performed according to the identification results to generate intervention basis.

[0070] In some embodiments, inputting the patient's current multimodal behavioral characteristics into the trained multimodal temporal anomaly detection model, and combining the dynamically adjusted judgment benchmark to identify real-time multimodal temporal behavioral deviations includes: Load a multimodal temporal anomaly detection model obtained through federated collaborative training; wherein, the multimodal temporal anomaly detection model includes a multimodal encoder, a cross-modal attention fusion layer, and a reconstruction-based anomaly detection network; The patient's current multimodal behavioral features, collected and extracted in real time, are input into the multimodal temporal anomaly detection model, and the following processing is performed: The medication record encoder, motion data encoder, diet information encoder, and physiological indicator encoder in the multimodal encoder are used to perform temporal encoding on the medication time feature, medication dosage feature, medication frequency feature, exercise intensity feature, exercise duration feature, meal time feature, and physiological indicator feature in the multimodal behavioral features, respectively, to generate temporal feature vectors corresponding to each modality. The cross-modal attention fusion layer performs multi-head attention calculation on the temporal feature vectors of each modality to determine the correlation weight matrix between different modal features, and then performs weighted fusion on the temporal feature vectors of each modality according to the correlation weight matrix to generate a unified multimodal fusion feature representation. The multimodal fusion feature representation is input into the reconstruction-based anomaly detection network, and anomaly scores are generated by calculating the reconstruction error of the input features. Based on the abnormality score and the dynamic abnormality threshold set in the dynamically adjusted judgment criteria, it is determined whether the patient's current behavior is abnormal and the corresponding risk level is classified. A temporal pattern clustering algorithm is used to compare the similarity between the multimodal fusion feature representation corresponding to the current behavior and the feature representation of the patient's historical normal behavior pattern to identify specific real-time multimodal temporal behavior deviation categories; wherein, the behavior deviation categories include medication time deviation, medication dosage deviation, medication frequency deviation, and special medication requirement deviation.

[0071] Specifically, the multimodal temporal anomaly detection model obtained through federated collaborative training typically comprises three main structural parts: a multimodal encoder, a cross-modal attention fusion layer, and a reconstruction-based anomaly detection network. The multimodal encoder consists of multiple parallel sub-networks. The medication record encoder can employ a Long Short-Term Memory (LSTM) network or a Gated Recurrent Unit (GRU) to capture the temporal dependencies of medication time, dosage, and frequency features. The motion data encoder can use a one-dimensional convolutional neural network (1D-CNN) combined with LSTM to extract local patterns and long-term trends from motion intensity and duration sequences. The diet information encoder can use a pre-trained language model (such as BioBERT or ClinicalBERT) to semantically encode descriptions of meal times and dietary content. The physiological indicator encoder can use a temporal convolutional network (TCN) or a Transformer encoder to handle long-term dependencies in continuous monitoring data such as blood glucose and blood pressure. Each encoder outputs a corresponding temporal feature vector, which has transformed the original behavioral data into a high-dimensional semantic representation.

[0072] The temporal feature vectors of each modality are input into the cross-modal attention fusion layer. This fusion layer is based on a multi-head attention mechanism. First, it calculates the correlation weight matrix between any two modal feature vectors. The weights reflect which modality is more important for judging behavioral abnormalities at the current time (for example, when judging medication time deviation, medication record features should be highly correlated with meal time features). Through weighted summation, a unified multimodal fusion feature representation is generated. This fusion method is more interpretable and robust than simple feature concatenation, and can dynamically highlight key modalities and suppress noisy modalities.

[0073] Subsequently, the multimodal fusion feature representation is input into a reconstruction-based anomaly detection network. This network can employ a variational autoencoder (VAE) or a Transformer-based autoencoder structure. During training, the model uses only normal behavioral data (i.e., historical data not marked as anomalous) for reconstruction learning, enabling the model to encode and decode normal behavioral patterns with high accuracy. During inference, the current feature representation is input, and the reconstruction error (e.g., mean squared error or negative log-likelihood) is calculated. A larger reconstruction error indicates that the current behavior deviates more from the normal pattern. Additionally, an anomaly scoring network (e.g., a binary classifier) ​​can be combined to output an anomaly score between 0 and 1. To integrate the dynamically adjusted judgment benchmark, the dynamic anomaly threshold generated in the dynamic benchmark adjustment step (e.g., a dynamic threshold calculated based on cumulative bias and tolerance range) is compared with the model's output anomaly score. If the anomaly score exceeds the dynamic threshold, the behavior is judged as abnormal, and a risk level (low, medium, high) is assigned based on the score. This approach, combining model scores with dynamic rule thresholds, leverages the strong fitting capabilities of deep learning while preserving the flexibility and safety of clinical rules.

[0074] Specifically, after identifying abnormal behavior, it is necessary to further clarify the specific deviation category. For this purpose, temporal pattern clustering algorithms can be used, such as HDBSCAN (density-based hierarchical clustering) or K-means clustering based on dynamic time warping (DTW). The multimodal fusion feature representation corresponding to the current behavior is compared with a feature representation library of the patient's historical normal behavior patterns. The historical normal pattern library is obtained by clustering behavioral features that were not marked as abnormal within a past period (e.g., 30 days), with each cluster center representing a typical behavioral pattern (e.g., "taking medication on time after breakfast"). The shortest distance between the current feature representation and each cluster center is calculated. If the distance exceeds a preset threshold, it is identified as a new deviation type; otherwise, the deviation category is determined based on the nearest cluster label, specifically including medication time deviation (e.g., delayed or early beyond the allowable range), medication dosage deviation (e.g., taking half a tablet when prescribed a full tablet), medication frequency deviation (e.g., once daily but actually taken twice), and deviations from special dosage requirements (e.g., not taking after meals or not swallowing the whole tablet). This identification result will serve as input for subsequent knowledge graph reasoning.

[0075] In some embodiments, the reasoning based on the recognition results using a preset medical knowledge graph to generate intervention criteria includes: Construct a medical knowledge graph; wherein the medical knowledge graph includes disease entity nodes, behavior entity nodes, intervention entity nodes, drug entity nodes, medication requirement entity nodes, and directed relation edges between the above nodes; the knowledge sources of the medical knowledge graph include at least one of clinical diagnosis and treatment guidelines, medical literature, expert experience knowledge, historical case data, and drug instructions; Based on the collected multimodal health behavior data and patients' electronic medical records, a patient profile is constructed; wherein, the patient profile includes the patient's basic demographic information, disease diagnosis information, health status information, historical behavioral characteristics, psychological status assessment information, social support information, and medication history information; Using the key features in the patient profile and the real-time multimodal temporal behavioral deviation category as query conditions, a graph traversal search is performed in the medical knowledge graph to query the disease-behavior-intervention path related to the current patient status and current behavioral deviation; Based on the retrieved graph path and graph traversal results, a reasoning operation is performed using rule-based reasoning or path-sorting-based reasoning algorithms to generate a health risk assessment report that may be caused by the patient's current behavior, output preliminary intervention suggestions, and combine the health risk assessment report and the preliminary intervention suggestions as the basis for intervention.

[0076] Specifically, the medical knowledge graph contains various entity nodes, such as disease entities (e.g., diabetes, hypertension), behavioral entities (e.g., medication, exercise, diet), intervention entities (e.g., reminders, education, consultation), drug entities (e.g., metformin), and dosage requirement entities (e.g., after meals, once daily). Directed edges between nodes define their logical or causal relationships, such as "disease-need-behavior," "behavior-leads-health risk," "drug-has-dosage requirement," and "intervention-applies to-bias." Knowledge sources include clinical guidelines, medical literature, expert experience, historical case data, and drug instructions. In practical construction, graph databases (e.g., Neo4j) can be used to store these triples, and natural language processing techniques can be used to automatically extract entities and relationships from unstructured text.

[0077] Simultaneously, dynamic patient profiles are constructed based on collected multimodal health behavior data and patients' electronic medical records. These profiles include not only basic demographic information (age, gender), disease diagnosis information (diabetes type, complications), health status information (recent average blood glucose and blood pressure), and historical behavioral characteristics (compliance trends, frequency of deviations), but also psychological assessment information (such as depression scale scores), social support information (family contact information, whether the patient lives alone), and medication history information (past adverse drug reactions, allergy history). The patient profiles are linked to patient nodes in the medical knowledge graph in the form of an attribute graph.

[0078] Specifically, using key features in the patient profile (e.g., "type 2 diabetes," "glimepiride") and real-time identified behavioral deviation categories (e.g., "take medication 30 minutes after breakfast, but actually take it 60 minutes after breakfast") as query conditions, a graph traversal search is performed in the medical knowledge graph. Graph traversal can employ breadth-first search (BFS) or depth-first search (DFS), combined with path constraints (e.g., relation type, step limit) to query disease-behavior-intervention paths related to the current patient status and current behavioral deviation. For example, starting from the "glimepiride" entity, the search proceeds along the "have a requirement to take it" edge to reach the "take it at breakfast" requirement, then connects to the "delayed medication" deviation entity through the "violation" relation, then connects to the "poor blood sugar control" risk entity through the "cause" relation, and finally connects to intervention entities such as "adjust reminder time" or "contact doctor" through the "suggestion" relation.

[0079] Based on the retrieved graph paths, inference operations are performed using rule-based reasoning or path-ranking-based reasoning algorithms. Rule-based reasoning can predefine SWRL rules (e.g., "If a patient takes glimepiride and the actual medication time deviates from the prescribed time by more than 30 minutes, the health risk level is medium risk, and the recommended intervention is an advance reminder"). Path-ranking-based algorithms (such as the Path Ranking Algorithm, PRA) predict the most appropriate intervention by learning the weights of different paths. The results of the inference include: generating a health risk assessment report that may result from the patient's current behavior (e.g., "Delaying medication after breakfast by more than 45 minutes for three consecutive days may lead to elevated blood glucose after lunch, increasing the risk of hypoglycemia"), and outputting preliminary intervention recommendations (e.g., "Adjust the medication reminder time after breakfast from 30 minutes to 45 minutes after breakfast, and increase the frequency of blood glucose monitoring"). Combining the health risk assessment report and the preliminary intervention recommendations forms a structured intervention basis. This basis not only tells patients and healthcare professionals "what to do," but also explains "why to do it," thus significantly enhancing the interpretability and clinical credibility of the system. Ultimately, the intervention criteria will be passed to the reinforcement learning optimization step to generate personalized intervention strategies.

[0080] S160: Using a reinforcement learning algorithm, a personalized intervention strategy is generated based on the intervention criteria and the patient's feedback data on historical interventions, and personalized reminder content is generated and pushed to the patient accordingly.

[0081] In some embodiments, the step of using a reinforcement learning algorithm to generate a personalized intervention strategy based on the intervention criteria and patient feedback data on historical interventions includes: Define the state space and action space of reinforcement learning; wherein, the state space includes the patient's physiological state parameters, behavioral state parameters, environmental state parameters, psychological state parameters, drug state parameters, and dynamic state parameters; the action space includes intervention type, reminder time parameter, reminder frequency parameter, reminder content template, and reminder tone parameter; A multi-objective reward function is designed based on the state space and the action space, and a hierarchical reinforcement learning architecture is constructed. The hierarchical reinforcement learning architecture includes a high-level policy network and a low-level policy network. The high-level policy network is used to output the intervention type in each decision cycle, and the low-level policy network is used to optimize the specific action parameters corresponding to the selected intervention type according to the output of the high-level policy network in the decision cycle. The strategy gradient algorithm is used to calculate the cumulative reward value based on the patient's real-time feedback data on historical interventions, and the parameters of the high-level strategy network and the low-level strategy network are updated based on the cumulative reward value. The intervention criteria are used as constraints or prior knowledge of the initial strategy to guide the strategy exploration and optimization direction of the strategy gradient algorithm, thereby generating an optimized personalized intervention strategy.

[0082] Specifically, the state space is a high-dimensional vector encompassing all observable factors affecting the effectiveness of patient behavioral interventions, including: physiological state parameters (such as current blood glucose, blood pressure, heart rate and their trends), behavioral state parameters (such as deviation from the baseline for the most recent medication administration, regularity of meal times, and achievement of exercise goals), environmental state parameters (such as current time, whether it is a holiday, geographical location, and weather), psychological state parameters (such as self-rated mood scores collected through mobile applications or stress levels inferred from wearable devices), medication state parameters (such as the half-life, therapeutic window, and dosage requirements of the currently administered medication), and dynamic state parameters (such as the cumulative deviation generated in the dynamic baseline adjustment steps and the response to the three most recent interventions). The action space defines all possible intervention actions, including intervention type (medication reminders, dietary advice, exercise encouragement, blood glucose monitoring, doctor consultation, etc.), reminder time parameters (the amount of advance or delay relative to the dynamic baseline), reminder frequency parameters (number of reminders per day, repetition interval), reminder content templates (selected from a preset library), and reminder tone parameters (encouraging, warning, neutral, empathetic, etc.).

[0083] Based on the aforementioned state and action spaces, a multi-objective reward function needs to be designed. The reward function aims not only to maximize patient adherence to a single reminder, but also to balance multiple, sometimes conflicting, optimization objectives, such as: improving medication adherence (rewards proportional to whether the patient takes medication on time and in the correct dosage), increasing patient satisfaction with interventions (measured through active ratings or indirect behaviors like click-through rates and closure rates), and improving patient health monitoring indicators (such as blood glucose levels returning to normal and blood pressure decreasing). Each objective is assigned a weight coefficient (e.g., adherence weight 0.4, satisfaction weight 0.3, health outcome weight 0.3), and the total reward is calculated using a weighted summation or Pareto optimization. To avoid reward sparsity, intermediate rewards can be introduced, such as a small positive reward when the patient opens the reminder message and a moderate positive reward when the patient actively reports completion of medication.

[0084] Specifically, to address the challenges of large decision spaces and long timescales in patient behavior interventions, a hierarchical reinforcement learning architecture can be employed. The high-level policy network (such as a deep Q-network or a policy gradient network) outputs a macro-level intervention type at each decision cycle (e.g., every morning), such as whether a medication reminder is needed today, whether dietary advice should be sent, or whether a follow-up phone call should be scheduled. The low-level policy network (such as a Gaussian-based continuous action space policy) then further optimizes specific action parameters within the same decision cycle based on the intervention type output by the high-level network. For example, if the high-level network selects a medication reminder, the low-level network needs to decide how many minutes before the baseline time to send the reminder (reminder timing), the interval between repetitions (frequency), which template to use, and whether to use an encouraging or cautionary tone. The high-level and low-level policy networks can be jointly trained or pre-trained with fine-tuning. Commonly used policy gradient algorithms include Proximal Policy Optimization (PPO) and Flexible Actor-Critic (SAC), which can stably handle continuous action spaces and improve sample efficiency through experience replay and advantage estimation.

[0085] During training, the intervention criteria generated in step S150 are used as constraints or prior knowledge for the initial policy to guide the exploration and optimization direction of the policy gradient algorithm. Specifically, this can be achieved by adding a bias term based on the intervention criteria to the output layer of the policy network, causing the policy to tend to select intervention types and parameters consistent with the knowledge graph inference results; or by using reward shaping to transform the risk assessment level in the intervention criteria into additional reward signals (e.g., higher rewards for correct interventions corresponding to high-risk biases). This leverages the causal logic of the knowledge graph to avoid extensive ineffective exploration in the initial stage while preserving the ability of reinforcement learning to learn personalized preferences from individual feedback. After multiple rounds of interaction (e.g., accumulating intervention-feedback data for each patient over several weeks to months), the policy network parameters gradually converge through gradient updates, generating a personalized intervention strategy for that patient.

[0086] In some embodiments, generating personalized reminder content and pushing it to patients includes: The type of intervention to be performed and the corresponding reminder parameters are determined based on the personalized intervention strategy. Based on the patient's educational background, health knowledge level, and psychological state assessment information, determine the complexity level and expression method of the reminder content; Select the corresponding reminder template from the preset reminder template library according to the intervention type; wherein, the reminder template library includes medication reminder templates, dietary advice templates, exercise encouragement templates, and health monitoring reminder templates; Based on the patient's specific behavioral deviations and the dynamically adjusted judgment criteria, the variable fields in the reminder template are filled in to generate personalized reminder text. Based on the preset urgency level of the reminder content, the patient's preset push preferences, and the dynamically adjusted judgment benchmark time, a push channel is selected and the optimal push timing is determined; wherein, the push channel includes at least one of SMS push, APP message push, voice call push, and smart wearable device push.

[0087] Specifically, after generating a personalized intervention strategy, it needs to be translated into specific reminder content and pushed to the patient. First, based on the intervention type and reminder parameters determined by the strategy, the action to be performed is identified. Then, based on the patient's educational background (e.g., junior high, high school, university), health knowledge level (assessed through a prior questionnaire), and psychological state assessment information (e.g., anxiety or depression scores), the complexity level and expression of the reminder content are dynamically determined. For example, for patients with lower levels of education or the elderly, simple and straightforward language, large font, and clear instructions are used; for younger patients with rich health knowledge, appropriate medical terminology can be used, along with explanations of underlying principles. For patients with poor psychological states (e.g., high anxiety), the tone of the reminder should be gentler and more encouraging, avoiding commanding or threatening tones. Finally, a corresponding template is selected from a pre-set reminder template library based on the intervention type. The template library includes medication reminder templates (e.g., "Please take [drug name] [dosage], [precautions] at [time]"), dietary advice templates (e.g., "We recommend [food type], avoid [food type], because [reason]"), exercise encouragement templates (e.g., "We recommend exercising for [duration] minutes today, [exercise type] contributes to [health benefits]"), and health monitoring reminder templates (e.g., "Please measure [physiological indicators] at [time] and record the results"). Then, based on the patient's specific behavioral deviations and dynamically adjusted judgment criteria, the variable fields in the templates are populated. For example, if a patient is detected to have delayed taking medication after breakfast for three consecutive days, and the dynamic benchmark has postponed the next judgment time by 45 minutes, the populated reminder would be: "We have detected a recent delay in your medication time after breakfast. We recommend that you take glimepiride within 45 minutes after breakfast (approximately 8:15 AM), which helps stabilize your postprandial blood glucose." Such reminders not only inform the patient of the specific action but also explain the reason, enhancing the patient's understanding and acceptance.

[0088] Finally, based on the preset urgency level of the reminder content (high-risk deviation corresponds to emergency reminders, medium-risk to routine reminders, and low-risk to general prompts), the patient's preset push preferences (e.g., the patient actively chose APP push as the preferred channel), and the dynamically adjusted judgment benchmark time (e.g., for medications taken after meals, the reminder should be set after the expected meal time plus the benchmark post-meal interval), select the push channel and determine the optimal push time. Push channels include SMS (suitable for emergencies, unstable network conditions, or elderly patients), APP message push (suitable for daily reminders, supporting rich text and interaction), voice call push (suitable for high-risk situations or visually impaired patients), and smart wearable device push (suitable for exercise supervision and instant vibration reminders). When determining the timing, for medication reminders, it is usually set 5-15 minutes before the dynamic benchmark time; for dietary suggestions, it can be set 30 minutes before the meal time; for exercise supervision, it can be set 15 minutes before the patient's usual exercise time. After the push is completed, a push log needs to be recorded, including feedback information such as whether it was successfully delivered, whether the patient opened it, and whether the reminder was followed.

[0089] In some embodiments, the method further includes: Collect patient feedback data on push notifications; wherein, the feedback data includes medication adherence feedback data, patient satisfaction feedback data, health indicator change feedback data, and behavior pattern change feedback data; The collected feedback data is cleaned, labeled, and integrated to generate a standardized feedback dataset. Based on the standardized feedback dataset, the anomaly detection accuracy and recall of the multimodal temporal anomaly detection model, as well as the effectiveness metrics of the intervention strategy generated by the reinforcement learning algorithm, are evaluated; the effectiveness metrics include changes in medication adherence, behavioral deviation correction rate, and patient satisfaction score. When the evaluation results meet the preset model update trigger conditions, at least one of the following will be updated: the model parameters of the multimodal temporal anomaly detection model, the policy network parameters of the reinforcement learning algorithm, the benchmark parameters in the personalized behavior benchmark, and the entities and relationships in the medical knowledge graph. Through the federated learning architecture, the updated model parameters of each participant are encrypted, uploaded, and aggregated, and the aggregated global model parameters are then distributed back to each participant.

[0090] Specifically, to form a closed-loop optimization, this disclosure also includes feedback collection and model update steps. Specifically, it collects patient feedback data on push notifications, including medication adherence feedback data (obtained through smart pillboxes, electronic medication records, or manual confirmation by the patient), patient satisfaction feedback data (through pop-up ratings after the push notification or indirect behaviors such as click-through rate and repeated closure rate), health indicator change feedback data (subsequent changes in physiological indicators such as blood glucose and blood pressure obtained through wearable devices or home monitoring devices), and behavioral pattern change feedback data (such as whether the patient adjusted their meal times or exercise intensity after receiving the reminder). The collected raw feedback data is cleaned (removing obviously erroneous or duplicate records), labeled (marking positive and negative feedback), and integrated (aligning feedback from different data sources by time and patient ID) to generate a standardized feedback dataset.

[0091] Based on the standardized feedback dataset, the performance of each core module is evaluated: the anomaly detection accuracy and recall of the multimodal temporal anomaly detection model are assessed (by comparing the model's predicted anomalies with those actually confirmed by patients), and the effectiveness metrics of the intervention strategies generated by the reinforcement learning algorithm are evaluated (changes in medication adherence, behavioral deviation correction rate, and patient satisfaction scores). When the evaluation results meet preset model update trigger conditions, such as anomaly detection accuracy being below 85% for three consecutive days, or medication adherence improvement stagnating for more than a week, a model update is triggered. Updates may include one or more of the following: updating the model parameters of the multimodal temporal anomaly detection model (e.g., incrementally training the reconstructed network using the latest normal behavior data), updating the policy network parameters of the reinforcement learning algorithm (continuing to train the policy network using newly accumulated interaction data), updating the baseline parameters in the personalized behavior baseline (e.g., recalculating the baseline interval based on the patient's recent average medication interval), and updating entities and relationships in the medical knowledge graph (e.g., extracting new knowledge from new clinical guidelines or medical literature). Finally, through a federated learning architecture, the locally updated model parameters of each participant (edge ​​nodes of each medical institution) are encrypted, uploaded, and aggregated (following the hierarchical aggregation process described in S130). The aggregated global model parameters are then redistributed to each participant, enabling the entire system to continuously learn and evolve from the collective experience of all participants without leaking original patient data. Through this continuous feedback-driven closed-loop optimization, the patient's reminder strategy becomes increasingly personalized and precise, ultimately achieving a behavioral shift from "passively receiving reminders" to "actively managing health."

[0092] The following section uses the medication management of hypertension patients as an example to provide a detailed description of the patient behavior reminder method 100 based on federated learning and multimodal temporal anomaly detection provided in this embodiment of the present disclosure, as detailed below: Suppose that patient Wang, 65 years old, suffers from essential hypertension and has been taking amlodipine besylate tablets (5 mg once daily, prescribed to be taken in the morning) for a long time, while also taking irbesartan (150 mg once daily, also recommended to be taken in the morning). Wang lives in the jurisdiction of a community health service center, and his medication data, blood pressure monitoring data, exercise and diet records are collected in real time through a smart pillbox, wearable wristband and mobile APP.

[0093] First, a knowledge base was constructed by parsing the drug's instruction manual. Taking amlodipine as an example, after obtaining its electronic instruction manual, a named entity recognition model was used to extract "amlodipine besylate tablets" as the drug name, "5 mg" as the single dose, "once daily" as the frequency of administration, and "morning" as the recommended time of administration. Specific requirements such as "swallow whole tablet" and "avoid taking with grapefruit" were also identified. The relation extraction model further associated these entities into triplets such as (amlodipine, with dose, 5 mg), (amlodipine, with frequency, once daily), and (amlodipine, with time, morning), and stored in the drug knowledge base. For irbesartan, similar information such as "once daily," "morning," and "can be taken on an empty stomach or after a meal" was extracted. The knowledge base also incorporates the latest hypertension treatment guidelines through a dynamic update mechanism, supplementing rules such as "avoid strenuous exercise for 1-2 hours after taking the medication."

[0094] Meanwhile, Wang's multimodal health behavior data was continuously collected through multi-source sensors. The smart pillbox recorded the time and dosage of each medication dispensing, the wearable bracelet monitored daily steps, exercise duration, heart rate, and sleep quality, and the mobile app allowed Wang to manually enter breakfast time and systolic / diastolic blood pressure values ​​measured by the blood pressure monitor. After cleaning, time alignment, and standardization, these raw data were used to extract multidimensional behavioral characteristics: medication time characteristics (actual time of medication, such as 7:55 a.m. on a certain day), medication dosage characteristics (whether one tablet was completely taken), medication frequency characteristics (once a day, whether a dose was missed), meal time characteristics (breakfast time is usually between 7:00 and 7:30 a.m.), exercise intensity characteristics (average daily steps of about 6,000), sleep duration characteristics (about 7 hours), and blood pressure index characteristics (morning blood pressure of about 135 / 85 mmHg).

[0095] Next, the matching degree and correlation between medication administration requirements and behavioral characteristics were analyzed to identify historical behavioral deviations and establish personalized behavioral benchmarks. The administration requirements for amlodipine and irbesartan were retrieved from the drug knowledge base: the prescribed administration time is in the morning (no strict postprandial requirement), the prescribed dose is 5mg + 150mg, and the frequency is once daily. Wang's actual medication use over the past 30 days was matched daily with these requirements, and the average medication time deviation was calculated to be +15 minutes (i.e., usually 15 minutes later than the prescribed time), the dose deviation was almost zero, and the frequency deviation was occasional missed doses (approximately twice per month). Further cross-correlation analysis of the medication time data series and the breakfast time data series revealed that the largest correlation coefficient between Wang's actual medication time and breakfast time occurred 30 minutes after breakfast, indicating that he habitually takes the medication half an hour after breakfast, while the drug instructions only require "morning" and do not mandate postprandial administration. Granger causality tests showed that when Wang's medication was delayed by more than 30 minutes for two consecutive days, his morning diastolic blood pressure increased by an average of 3-4 mmHg the following day, which was statistically significant. Cluster analysis categorized Wang as exhibiting a pattern of "regular medication use, but slightly delayed," while association rule mining revealed a strong rule: "taking medication 30 minutes after breakfast and exceeding 8000 steps that day → stable blood pressure upon waking." Based on these results, the identified historical behavioral deviation was primarily delayed medication use (an average delay of 15 minutes, occasionally exceeding 1 hour), assessed as low to medium risk. Accordingly, a personalized behavioral baseline was established for Wang: the baseline medication time was set at 30 minutes after breakfast (i.e., if breakfast was at 7:30, medication was taken at 8:00), the baseline dose remained 5mg + 150mg, the baseline frequency was once daily, and the allowable deviation range for time was set at ±20 minutes, with an allowable deviation range for interval (±30 minutes) for missed doses.

[0096] Because Wang's data is stored at the community health service center, while similar hypertension patients' data is scattered across multiple tertiary hospitals and family doctor teams, a federated learning architecture is used to collaboratively train a multimodal temporal anomaly detection model in order to fully utilize cross-institutional data to improve model performance while protecting privacy. The central server initializes a global model containing an LSTM medication encoder, a CNN motion encoder, a BERT diet encoder, and a TCN physiological index encoder, and distributes encrypted parameters to each edge node. The community health service center, as one of the edge nodes, uses local anonymized data of Wang and other patients for local training under differential privacy protection: after calculating the gradient in each mini-batch, the L2 norm is clipped to 1.0, and Gaussian noise (ε=0.5, δ=1e-5) is added to generate encrypted gradients, which are then uploaded to the regional aggregator (the city's medical consortium platform). The regional aggregator collects gradients from multiple institutions within its jurisdiction, weights them using a federated averaging algorithm, and uploads them to the central server. After 50 rounds of communication iterations, the global model converges. The community health service center then uses local data to fine-tune the global model, obtaining a personalized anomaly detection model adapted to the characteristics of patients in its jurisdiction.

[0097] In daily management, the system obtains Wang's most recent medication information in real time and dynamically adjusts subsequent judgment benchmarks based on deviations from his personalized behavioral benchmarks. One day, Wang took his medication at 6:30 AM to catch an early flight, 1.5 hours earlier than the benchmark time (8:00 AM). Upon detecting this early medication event, the system first assesses the safety risk: it checks whether the interval between the last medication and this dose is less than the minimum safe interval (amlodipine has a long half-life, and the minimum safe interval is usually 12 hours; this time, the interval between the last dose and the current dose exceeded 20 hours, so it is safe). Therefore, it only records the early behavior without forcibly adjusting the next benchmark. On another occasion, Wang took his medication at 12:00 PM due to a meeting delay (the benchmark was 8:00 AM, a delay of 4 hours). The system automatically postpones the next day's medication judgment benchmark time by 4 hours; that is, the original suggestion to take the medication at 8:00 AM is changed to 12:00 PM, and the reminder time is adjusted accordingly. Meanwhile, monitoring showed that the cumulative deviation reached 6 hours over three consecutive days, exceeding the preset warning threshold (3 hours). This triggered an alert and restarted the personalized baseline establishment process, adjusting the new baseline time to 45 minutes after breakfast (based on the average breakfast time of 7:15 over the past three days, medication should be taken at 7:45). In addition, the system also acquired Wang's current mealtime characteristics (dinner was delayed to 20:00 on a certain day), automatically adjusting the "bedtime medication" reminder to 20:30 (if applicable), and assessing the potential for accelerated drug metabolism based on the day's exercise intensity (1 hour of high-intensity exercise), suggesting appropriately shortening the next medication interval or increasing the monitoring frequency.

[0098] For real-time anomaly detection, Wang's current multimodal behavioral characteristics are input into a pre-trained multimodal temporal anomaly detection model. The model's medication encoder encodes the medication time deviation sequence of the past 7 days (e.g., [+15min, +10min, +4h, +3.5h, +20min, +18min, +22min]), the motion encoder encodes the step count sequence, and the physiological index encoder encodes the daily morning blood pressure sequence. A cross-modal attention fusion layer calculates the high correlation weight between the current time period's medication characteristics and blood pressure characteristics (because blood pressure tends to rise after delayed medication), generating a fused feature representation. The reconstruction network calculates the reconstruction error, obtaining an anomaly score of 0.82 (threshold 0.7), classifying it as an anomaly with a medium-to-high risk level. Temporal pattern clustering compares the current behavioral characteristics with historical normal patterns (delay <20 minutes), identifying the specific deviation category as "medication time deviation (delay exceeding 3 hours)."

[0099] Subsequently, intervention criteria were generated based on a pre-set medical knowledge graph. This knowledge graph includes nodes related to hypertension, amlodipine medication, delayed medication behavior, and the risk of elevated blood pressure, as well as intervention nodes such as "recommend measuring blood pressure" and "adjusting medication reminders." The graph stores the clinical rule that "delaying medication by more than 2 hours may lead to an increase of 5-10 mmHg in morning blood pressure the next day." Using Mr. Wang's patient profile (stage 2 hypertension, no complications, living alone) and real-time deviation category (delay of 4 hours) as query conditions, a graph traversal was performed to find the path "delayed amlodipine medication → fluctuation in blood drug concentration → risk of systolic blood pressure fluctuation → recommendation to increase the frequency of home blood pressure monitoring and adjust the reminder time the next day." The inference algorithm outputs a health risk assessment report: "Delaying medication by 4 hours today may increase diastolic blood pressure by 3-5 mmHg overnight and the following morning, increasing the risk of cardiovascular events." It also outputs preliminary intervention recommendations: "Please take your blood pressure again before bed tonight; it is recommended to take your medication at 7:00 AM tomorrow morning to restore your circadian rhythm; tomorrow's reminder will be sent at 6:45 AM." The combination of these two recommendations serves as the basis for intervention.

[0100] Next, a reinforcement learning algorithm was used to generate a personalized intervention strategy based on the intervention criteria and Wang's feedback on historical interventions. The state space included Wang's current blood pressure (135 / 85), behavioral state (delayed by 4 hours), environment (on a business trip), psychological state (moderate anxiety score), medication status (amlodipine half-life 30-50 hours), and dynamic cumulative bias (6 hours). The action space included intervention type (medication reminder, blood pressure monitoring reminder, telephone follow-up), reminder time (earlier or later), and tone (encouraging or warning). The multi-objective reward function weights were set to compliance 0.4, satisfaction 0.3, and health effect 0.3. The high-level policy network of the hierarchical reinforcement learning output a combination of "medication reminder + blood pressure monitoring reminder," while the low-level policy network output: the reminder time is set at 6:45 am the next day (15 minutes earlier than the new baseline), the tone is "encouraging," and the reminder template is a composite template of "medication + monitoring." Since the intervention criteria indicated high risk, the policy gradient algorithm used this criteria as prior knowledge, giving higher initial probabilities to actions consistent with the knowledge graph suggestions during exploration. After training on Wang's feedback data from the past three months (such as his higher acceptance of voice reminders than text reminders), the policy network updated its parameters and generated a personalized strategy: Wang was given dual reminders via APP push and vibration bracelet, with a gentle tone and blood pressure monitoring prompts.

[0101] Personalized reminders are generated based on this strategy. A medication reminder template is selected, and variables are filled in: "Mr. Wang, because you took your medication late this morning, to maintain stable blood pressure, it is recommended that you take one tablet each of amlodipine and irbesartan tomorrow morning around 6:45. Please measure and record your blood pressure before taking the medication. Have a pleasant business trip!" Based on the patient profile of "high school education, good understanding of medical terminology," the content includes a reasonable medical explanation; based on the urgency level (medium to high risk), APP push notifications and wristband vibration are selected, with the push timing set 10 minutes before 6:45, i.e., 6:35.

[0102] Finally, after receiving the reminder, Mr. Wang did take his medication and measure his blood pressure (128 / 82) at 6:50 the next day, thus recording positive feedback on compliance. He rated the reminder content in the app as 4.5 out of 5, indicating positive satisfaction. This feedback data, after cleaning and integration, was used to evaluate the model: the anomaly detection model's recall for this delayed event was correct, and the compliance rate brought about by the reinforcement learning strategy increased from 82% to 89%. The model update trigger condition was met (compliance rate improvement stalled), thus triggering an update: the anomaly detection model was incrementally trained with new normal behavior data, the policy network continued training with interaction data from the last two weeks, and the tolerance range in the personalized behavior baseline narrowed from ±20 minutes to ±15 minutes (due to the patient's good response). Simultaneously, through the federated learning architecture, the community health service center uploaded the encrypted model update to the regional aggregator for global aggregation, and then distributed the aggregated global model, achieving cross-institutional co-evolution.

[0103] Through the complete process described above, this method provides hypertensive patient Wang with a full-chain intelligent medication management system, encompassing knowledge construction, personalized benchmarks, federated collaboration, dynamic adjustment, anomaly detection, knowledge reasoning, and reinforcement learning optimization. This significantly reduces blood pressure fluctuations caused by medication time deviations and improves patient compliance and satisfaction.

[0104] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this disclosure is not limited to the described order of actions, because according to this disclosure, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to this disclosure.

[0105] The above is an introduction to the method embodiments. The following describes the solution described in this disclosure further through device embodiments.

[0106] Figure 2 A structural diagram of a patient behavior alert device based on federated learning and multimodal temporal anomaly detection provided in an embodiment of this disclosure is shown, as follows: Figure 2 As shown, the patient behavior alerting device 200 based on federated learning and multimodal temporal anomaly detection may include: The knowledge base construction and feature extraction module 210 is used to construct a drug knowledge base and collect patients' multimodal health behavior data to extract multimodal behavioral features.

[0107] The behavior-drug correlation analysis module 220 is used to analyze the matching degree and correlation between the drug administration requirements information in the drug knowledge base and the multimodal behavioral characteristics, so as to identify the patient's historical behavioral deviations and establish personalized behavioral benchmarks for the patient based on the historical behavioral deviations.

[0108] The federated collaborative training module 230 is used to collaboratively train a multimodal temporal anomaly detection model using a federated learning architecture, provided that the local patient data of multiple participants does not leave their respective local storage environments.

[0109] The dynamic benchmark adjustment module 240 is used to obtain the patient's most recent medication information and dynamically adjust the benchmark for subsequent behavior judgment based on the deviation between the most recent medication information and the personalized behavior benchmark, thereby generating a dynamically adjusted judgment benchmark.

[0110] The anomaly detection and reasoning module 250 is used to input the patient's current multimodal behavioral characteristics into the trained multimodal temporal anomaly detection model, identify real-time multimodal temporal behavioral deviations in combination with the dynamically adjusted judgment criteria, and generate intervention basis based on the identification results according to the preset medical knowledge graph.

[0111] The strategy generation and push module 260 is used to generate personalized intervention strategies based on the intervention criteria and patient feedback data on historical interventions using reinforcement learning algorithms, and to generate personalized reminder content to push to the patient accordingly.

[0112] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the described module can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0113] The acquisition, storage, and application of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0114] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0115] Figure 3A schematic block diagram of an electronic device 300 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0116] Electronic device 300 includes a computing unit 301, which can perform various appropriate actions and processes according to a computer program stored in ROM 302 or a computer program loaded into RAM 303 from storage unit 308. RAM 303 can also store various programs and data required for the operation of electronic device 300. The computing unit 301, ROM 302, and RAM 303 are interconnected via bus 304. I / O interface 305 is also connected to bus 304.

[0117] Multiple components in electronic device 300 are connected to I / O interface 305, including: input unit 306, such as keyboard, mouse, etc.; output unit 307, such as various types of displays, speakers, etc.; storage unit 308, such as disk, optical disk, etc.; and communication unit 309, such as network card, modem, wireless transceiver, etc. Communication unit 309 allows electronic device 300 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0118] The computing unit 301 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 301 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 301 performs the various methods and processes described above, such as method 100. For example, in some embodiments, method 100 may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 308. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 300 via ROM 302 and / or communication unit 309. When the computer program is loaded into RAM 303 and executed by the computing unit 301, one or more steps of method 100 described above may be performed. Alternatively, in other embodiments, the computing unit 301 may be configured to perform method 100 by any other suitable means (e.g., by means of firmware).

[0119] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0120] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0121] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0122] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including voice input, speech input, or tactile input).

[0123] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0124] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.

[0125] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this disclosure can be achieved, and this is not limited herein.

[0126] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A patient behavior reminder method based on federated learning and multimodal temporal anomaly detection, characterized in that, include: A drug knowledge base was constructed, and patients' multimodal health behavior data was collected to extract multimodal behavioral features; The matching degree and correlation between the drug administration requirements information in the drug knowledge base and the multimodal behavioral characteristics are analyzed to identify the patient's historical behavioral deviations and establish personalized behavioral benchmarks for the patient based on the historical behavioral deviations. A federated learning architecture is adopted to collaboratively train a multimodal temporal anomaly detection model without leaving the local storage environment of the local patient data of multiple participants. Obtain the patient's most recent medication information, and dynamically adjust the benchmark for subsequent behavior judgment based on the deviation between the most recent medication information and the personalized behavior benchmark, thereby generating a dynamically adjusted judgment benchmark. The patient's current multimodal behavioral characteristics are input into the trained multimodal temporal anomaly detection model, and the real-time multimodal temporal behavioral deviations are identified in combination with the dynamically adjusted judgment criteria. Based on the preset medical knowledge graph, the identification results are used to generate intervention basis. Using reinforcement learning algorithms, personalized intervention strategies are generated based on the intervention criteria and patient feedback data on historical interventions, and personalized reminders are then pushed to patients accordingly.

2. The method according to claim 1, characterized in that, The construction of the drug knowledge base and the collection of patients' multimodal health behavior data to extract multimodal behavioral features include: Obtain the raw text data of the drug instruction manual and preprocess the raw text data; Using a named entity recognition model trained based on deep learning, key entities are extracted from preprocessed text data. The key entities include drug name, active ingredient, indications, usage and dosage description, and precautions description. The relationship between the key entities is identified using a relation extraction model. The relationship includes: the dosage relationship between the drug and the single dose, the frequency relationship between the drug and the number of times it is taken per day, the time relationship between the drug and the time point of administration, the interval relationship between the drug and the duration of the administration interval, and the special requirement relationship between the drug and the instructions to take it before or after meals. The extracted key entities and their relationships are structured and stored in the form of triples to construct the drug knowledge base; The system continuously collects patients' health behavior data through multiple sensors, and preprocesses and extracts features from the health behavior data to generate multimodal behavioral features. The multimodal behavioral features include: medication time features, medication dosage features, medication frequency features, meal time features, exercise intensity features, exercise duration features, sleep duration features, and physiological indicator features. The health behavior data includes at least medication record data, dietary record data, exercise monitoring data, sleep monitoring data, and physiological indicator monitoring data.

3. The method according to claim 2, characterized in that, The analysis of the matching degree and correlation between the drug administration requirements information in the drug knowledge base and the multimodal behavioral characteristics, in order to identify the patient's historical behavioral deviations, and to establish personalized behavioral benchmarks for the patient based on the historical behavioral deviations, includes: Based on the name of the medication the patient is currently taking, the corresponding medication usage requirements are retrieved from the drug knowledge base. These requirements include the prescribed time of administration, the prescribed dosage, the prescribed frequency of administration, the prescribed interval between administrations, and any special requirements. The extracted medication time feature, medication dosage feature, and medication frequency feature from the multimodal behavioral features are matched with the prescribed medication time, prescribed medication dosage, and prescribed medication frequency in the medication administration requirement information, and the matching deviation values ​​for each feature are calculated. Cross-correlation analysis was performed on the patient's medication time data series and meal time data series to determine the lag time relationship between the two. Granger causality tests were performed on the time series of patients' medication behavior and the time series of physiological indicators to analyze whether changes in medication behavior constituted a statistical predictive cause of changes in physiological indicators. Clustering algorithms were used to perform cluster analysis on the historical multimodal behavioral characteristics of patients to identify the typical behavioral pattern categories of patients; An association rule mining algorithm was used to extract association rules between behavioral patterns and medication requirements from patients' historical behavioral data. Based on the results of the above cross-correlation analysis, Granger causality test, cluster analysis, and association rule mining, historical behavioral deviations of patients' actual medication behavior relative to the prescribed medication requirements are identified; wherein, the historical behavioral deviations include medication time deviations, medication dosage deviations, medication frequency deviations, and deviations from special medication requirements; Based on the preset risk assessment rules, assess the health risk level corresponding to each historical behavioral deviation; Based on the identified types of historical behavioral deviations and their corresponding health risk levels, personalized behavioral benchmarks are established for patients; wherein, the personalized behavioral benchmarks include benchmark medication time, benchmark medication dosage, benchmark medication frequency, benchmark medication interval, and the allowable deviation range for each benchmark item.

4. The method according to claim 3, characterized in that, The federated learning architecture, which collaboratively trains a multimodal temporal anomaly detection model without requiring local patient data from multiple participants to leave their respective local storage environments, includes: A three-layer federated learning architecture is constructed, consisting of an edge node layer, a regional aggregator layer, and a central server layer. Each medical institution acts as an edge node, a medical consortium or medical group acts as a regional aggregator, and the central server is responsible for the initialization, distribution, and aggregation of global model parameters. The central server initializes the model parameters of the global multimodal temporal anomaly detection model and encrypts the global model parameters before sending them to each edge node. After receiving the global model parameters, each edge node uses the locally stored patient multimodal health behavior data to train a copy of the locally held multimodal temporal anomaly detection model. During the training process, a differential privacy protection mechanism is applied to generate local model gradients, and the local model gradients are encrypted and then uploaded to the corresponding regional aggregator. Each regional aggregator collects the encrypted model gradients uploaded by the edge nodes under its jurisdiction, performs weighted average aggregation using a federated averaging algorithm, generates regional model parameters, and uploads the regional model parameters to the central server. The central server collects the regional model parameters uploaded by each regional aggregator, performs global aggregation calculations to update the global model parameters, and then encrypts and sends the updated global model parameters to each edge node again. The above steps are executed iteratively until the global model converges or the preset number of communication rounds is reached; Each edge node uses local data to fine-tune the converged global model to generate a personalized multimodal temporal anomaly detection model that adapts to the distribution characteristics of local patient data.

5. The method according to claim 3, characterized in that, The process of obtaining the patient's most recent medication information and dynamically adjusting the benchmark for subsequent behavioral judgments based on the deviation between the most recent medication information and the personalized behavioral benchmark, generating dynamically adjusted judgment benchmarks, includes: Real-time monitoring of the patient's most recent medication time, dosage, and administration method, generating a recent medication record; Determine the baseline dosing interval based on the drug dosing interval requirements retrieved from the drug knowledge base; When it is detected that the actual time of a patient's medication is delayed compared to the baseline medication time in the personalized behavioral baseline, the judgment baseline time for the next medication is postponed by the same amount of time as the delay. When a patient's actual medication time is detected to be earlier than the baseline medication time, the safety risks of early medication are assessed, and it is checked whether the minimum safe dosing interval requirement for the drug is violated as obtained from the drug knowledge base. A time interval tolerance range is set for the dynamically adjusted judgment benchmark time. When the actual medication interval exceeds the tolerance range, it is determined to be an abnormal medication time and the abnormal event is recorded. The system continuously monitors the time interval deviation of multiple medication administrations and calculates the cumulative deviation value. When the cumulative deviation value exceeds a preset warning threshold, a warning signal is triggered, and the process of re-establishing the personalized behavioral benchmark is initiated. Acquire the patient's current meal time characteristics, exercise characteristics, and physiological indicators; adaptively adjust the baseline time for postprandial medication based on changes in meal time characteristics; assess the impact of changes in exercise intensity and duration on drug metabolism rate and adjust the baseline time for medication accordingly; assess the impact on medication safety based on significant changes in physiological indicators relative to the normal range and adjust the baseline time for medication accordingly.

6. The method according to claim 4, characterized in that, The step of inputting the patient's current multimodal behavioral characteristics into the trained multimodal temporal anomaly detection model, and combining it with the dynamically adjusted judgment benchmark to identify real-time multimodal temporal behavioral deviations includes: Load a multimodal temporal anomaly detection model obtained through federated collaborative training; wherein, the multimodal temporal anomaly detection model includes a multimodal encoder, a cross-modal attention fusion layer, and a reconstruction-based anomaly detection network; The patient's current multimodal behavioral features, collected and extracted in real time, are input into the multimodal temporal anomaly detection model, and the following processing is performed: The medication record encoder, motion data encoder, diet information encoder, and physiological indicator encoder in the multimodal encoder are used to perform temporal encoding on the medication time feature, medication dosage feature, medication frequency feature, exercise intensity feature, exercise duration feature, meal time feature, and physiological indicator feature in the multimodal behavioral features, respectively, to generate temporal feature vectors corresponding to each modality. The cross-modal attention fusion layer performs multi-head attention calculation on the temporal feature vectors of each modality to determine the correlation weight matrix between different modal features, and then performs weighted fusion on the temporal feature vectors of each modality according to the correlation weight matrix to generate a unified multimodal fusion feature representation. The multimodal fusion feature representation is input into the reconstruction-based anomaly detection network, and anomaly scores are generated by calculating the reconstruction error of the input features. Based on the abnormality score and the dynamic abnormality threshold set in the dynamically adjusted judgment criteria, it is determined whether the patient's current behavior is abnormal and the corresponding risk level is classified. A temporal pattern clustering algorithm is used to compare the similarity between the multimodal fusion feature representation corresponding to the current behavior and the feature representation of the patient's historical normal behavior pattern to identify specific real-time multimodal temporal behavior deviation categories; wherein, the behavior deviation categories include medication time deviation, medication dosage deviation, medication frequency deviation, and special medication requirement deviation.

7. The method according to claim 6, characterized in that, The method of generating intervention criteria based on the recognition results using a pre-set medical knowledge graph includes: Construct a medical knowledge graph; wherein the medical knowledge graph includes disease entity nodes, behavior entity nodes, intervention entity nodes, drug entity nodes, medication requirement entity nodes, and directed relation edges between the above nodes; the knowledge sources of the medical knowledge graph include at least one of clinical diagnosis and treatment guidelines, medical literature, expert experience knowledge, historical case data, and drug instructions; Based on the collected multimodal health behavior data and patients' electronic medical records, a patient profile is constructed; wherein, the patient profile includes the patient's basic demographic information, disease diagnosis information, health status information, historical behavioral characteristics, psychological status assessment information, social support information, and medication history information; Using the key features in the patient profile and the real-time multimodal temporal behavioral deviation category as query conditions, a graph traversal search is performed in the medical knowledge graph to query the disease-behavior-intervention path related to the current patient status and current behavioral deviation; Based on the retrieved graph path and graph traversal results, a reasoning operation is performed using rule-based reasoning or path-sorting-based reasoning algorithms to generate a health risk assessment report that may be caused by the patient's current behavior, output preliminary intervention suggestions, and combine the health risk assessment report and the preliminary intervention suggestions as the basis for intervention.

8. The method according to claim 1, characterized in that, The method of using reinforcement learning algorithms to generate personalized intervention strategies based on the intervention criteria and patient feedback data on historical interventions includes: Define the state space and action space of reinforcement learning; wherein, the state space includes the patient's physiological state parameters, behavioral state parameters, environmental state parameters, psychological state parameters, drug state parameters, and dynamic state parameters; the action space includes intervention type, reminder time parameter, reminder frequency parameter, reminder content template, and reminder tone parameter; A multi-objective reward function is designed based on the state space and the action space, and a hierarchical reinforcement learning architecture is constructed. The hierarchical reinforcement learning architecture includes a high-level policy network and a low-level policy network. The high-level policy network is used to output the intervention type in each decision cycle, and the low-level policy network is used to optimize the specific action parameters corresponding to the selected intervention type according to the output of the high-level policy network in the decision cycle. The strategy gradient algorithm is used to calculate the cumulative reward value based on the patient's real-time feedback data on historical interventions, and the parameters of the high-level strategy network and the low-level strategy network are updated based on the cumulative reward value. The intervention criteria are used as constraints or prior knowledge of the initial strategy to guide the strategy exploration and optimization direction of the strategy gradient algorithm, thereby generating an optimized personalized intervention strategy.

9. The method according to claim 8, characterized in that, The generation of personalized reminders to be pushed to patients includes: The type of intervention to be performed and the corresponding reminder parameters are determined based on the personalized intervention strategy. Based on the patient's educational background, health knowledge level, and psychological state assessment information, determine the complexity level and expression method of the reminder content; Select the corresponding reminder template from the preset reminder template library according to the intervention type; wherein, the reminder template library includes medication reminder templates, dietary advice templates, exercise encouragement templates, and health monitoring reminder templates; Based on the patient's specific behavioral deviations and the dynamically adjusted judgment criteria, the variable fields in the reminder template are filled in to generate personalized reminder text. Based on the preset urgency level of the reminder content, the patient's preset push preferences, and the dynamically adjusted judgment benchmark time, a push channel is selected and the optimal push timing is determined; wherein, the push channel includes at least one of SMS push, APP message push, voice call push, and smart wearable device push.

10. The method according to any one of claims 1-9, characterized in that, The method further includes: Collect patient feedback data on push notifications; wherein, the feedback data includes medication adherence feedback data, patient satisfaction feedback data, health indicator change feedback data, and behavior pattern change feedback data; The collected feedback data is cleaned, labeled, and integrated to generate a standardized feedback dataset. Based on the standardized feedback dataset, the anomaly detection accuracy and recall of the multimodal temporal anomaly detection model, as well as the effectiveness metrics of the intervention strategy generated by the reinforcement learning algorithm, are evaluated; the effectiveness metrics include changes in medication adherence, behavioral deviation correction rate, and patient satisfaction score. When the evaluation results meet the preset model update trigger conditions, at least one of the following will be updated: the model parameters of the multimodal temporal anomaly detection model, the policy network parameters of the reinforcement learning algorithm, the benchmark parameters in the personalized behavior benchmark, and the entities and relationships in the medical knowledge graph. Through the federated learning architecture, the updated model parameters of each participant are encrypted, uploaded, and aggregated, and the aggregated global model parameters are then distributed back to each participant.