Reinforcement learning-based health care robot knowledge graph dynamic updating method

By constructing a dual value network and a counterfactual reasoning mechanism, virtual experience samples are generated. Combined with natural language interaction, the security and efficiency issues of the knowledge graph of the health and wellness robot are solved, and the precise updating of personalized health and wellness services is realized.

CN122050790BActive Publication Date: 2026-07-03SHANDONG PROVINCIAL HOSPITAL AFFILIATED TO SHANDONG FIRST MEDICAL UNIVERSITY (SHANDONG PROVINCIAL HOSPITAL)
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHANDONG PROVINCIAL HOSPITAL AFFILIATED TO SHANDONG FIRST MEDICAL UNIVERSITY (SHANDONG PROVINCIAL HOSPITAL)
Filing Date
2026-04-20
Publication Date
2026-07-03

AI Technical Summary

Technical Problem

Existing knowledge graphs for health and wellness robots are ill-suited to adapting to individual differences and changes in physiological state among users. This leads to risks of trial-and-error exploration, a lack of perception of knowledge ambiguity, and blind spots in subjective knowledge updates, resulting in safety hazards and low learning efficiency.

Method used

By employing a reinforcement learning-based approach, a dual-value network is constructed. Virtual experience samples are generated through counterfactual reasoning, and combined with natural language interaction and real experience samples, the knowledge graph is dynamically updated, ensuring the safety and accuracy of the exploration.

Benefits of technology

It enables zero-risk virtual exploration, accurately identifies areas of knowledge ambiguity, improves the efficiency and accuracy of knowledge updates, and ensures the personalization and safety of health and wellness services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122050790B_ABST
    Figure CN122050790B_ABST
Patent Text Reader

Abstract

The application discloses a health-care robot knowledge graph dynamic updating method based on reinforcement learning, and relates to the technical field of artificial intelligence and rehabilitation medical treatment. The application realizes zero-risk virtual exploration through a counterfactual reasoning mechanism, avoids the security risks caused by the trial-and-error process to users from the root, and uses the current knowledge graph as a world model. When a high-fuzziness knowledge node is identified, instead of directly executing unknown actions in a real environment, a virtual result generated by executing a virtual action is deduced based on similar user group records and association rules in the knowledge graph, a virtual reward value is calculated, virtual experience samples are stored in an experience replay pool to participate in subsequent network training, the reinforcement learning intelligent agent can learn the potential value of the action without ever executing the action in reality, and thus a virtual learning environment is constructed in the whole knowledge exploration process, and the security threat of the trial-and-error process to users is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of interdisciplinary technology of artificial intelligence and rehabilitation medicine, specifically a method for dynamically updating the knowledge graph of rehabilitation robots based on reinforcement learning. Background Technology

[0002] With the accelerating aging of the population, healthcare robots are playing an increasingly important role in personalized rehabilitation training and daily health monitoring. Knowledge graphs, as structured knowledge bases, can provide healthcare robots with semantic representations of domain concepts, entities, and their relationships, supporting the robots in understanding healthcare knowledge and performing rehabilitation tasks. However, traditional knowledge graphs are usually statically constructed, making it difficult to adapt to individual differences among users and dynamic changes in the physiological state of the same user, thus hindering the robots from providing truly personalized healthcare services.

[0003] Existing technologies attempt to introduce reinforcement learning to achieve dynamic updates of knowledge graphs, but they have the following drawbacks:

[0004] The health and wellness sector involves personal safety. The trial-and-error exploration mechanism in the early stages of reinforcement learning may cause the robot to try the wrong massage intensity or rehabilitation movements, causing secondary harm to the user. The contradiction between exploration risk and learning efficiency has not been effectively resolved.

[0005] Existing methods lack a quantitative perception mechanism for knowledge ambiguity. Robots cannot identify blind spots or uncertain areas in their own knowledge and usually use random strategies for exploration, resulting in wasted computing resources and low learning efficiency.

[0006] For knowledge involving users' subjective feelings, existing technologies struggle to accurately quantify it using sensor data and lack proactive interaction mechanisms with users, resulting in blind spots in the updating of this crucial knowledge.

[0007] Therefore, there is an urgent need for a dynamic update method for knowledge graphs that can both safely test and learn to reduce exploration risks and accurately locate and efficiently update ambiguous knowledge areas, so as to achieve the adaptive evolution of the knowledge system of health care robots. Summary of the Invention

[0008] The purpose of this invention is to overcome the shortcomings of existing technologies and provide a method for dynamically updating the knowledge graph of a health and wellness robot based on reinforcement learning. This method enables zero-risk virtual exploration through counterfactual reasoning, fundamentally avoiding the safety hazards to users caused by trial and error. The mechanism uses the current knowledge graph as a world model. When highly fuzzy knowledge nodes are identified, it does not directly execute unknown actions in the real environment. Instead, based on similar user group records and association rules in the knowledge graph, it infers the virtual results generated by executing virtual actions and calculates virtual reward values. By storing virtual experience samples in an experience replay pool for subsequent network training, the reinforcement learning agent can learn the potential value of an action without ever actually executing it, ensuring that the actions and results of counterfactual reasoning are within safe boundaries. This constructs a virtual learning environment throughout the entire knowledge exploration process, solving the safety threats to users during trial and error.

[0009] To solve the above-mentioned technical problems, this invention provides the following technical solution: a method for dynamically updating the knowledge graph of a rehabilitation and elderly care robot based on reinforcement learning, the specific steps of which are as follows:

[0010] S100. Constructing the Initial Knowledge Graph and Dual Value Network: Constructing a knowledge graph that includes the types of health and wellness entities and the relationships between them. and establish a first value network. Second Value Network A dual-value network architecture;

[0011] S200, Knowledge Uncertainty Measurement Based on Dual Time-Series Differential Error: This involves measuring the current user state... Input the first value network Second Value Network ,calculate and For actions to be performed The output difference is used to obtain the dual time-series difference error. ,Will With the first threshold Second threshold A comparison is made, and a knowledge exploration path is determined based on the comparison results. The knowledge exploration path includes counterfactual virtual exploration, active inquiry-based exploration, and direct decision-making.

[0012] S300, Secure Generation of Virtual Experiences Based on Counterfactual Reasoning: In response to the counterfactual virtual exploration in S200, the counterfactual reasoning module is initiated, based on the current knowledge graph. As a world model, it utilizes similar entity relationships and similar user group records in the graph to generate virtual actions that are not actually executed. The corresponding virtual experience samples are then stored in the experience replay pool. ;

[0013] S400, Proactive Query-Based Exploration and Revision: In response to the proactive query-based exploration in S200, semantic feedback from the user is obtained through natural language interaction, the semantic feedback is transformed into structured knowledge triples, and the knowledge graph is revised. The confidence level of the corresponding knowledge node;

[0014] S500, Joint Experience Replay and Dynamic Update of Knowledge Graph: From the aforementioned experience replay pool The first value network is trained by dynamically mixing real and virtual experience samples in a mixed manner. Second Value Network The knowledge graph is dynamically updated based on the reassessment of knowledge value by the trained value network. The topological structure of knowledge triples.

[0015] Furthermore, in S100, the types of health and wellness entities include user entities, health and wellness action entities, state entities, and result entities, and the relationships between entities include causal relationships and similarity relationships.

[0016] Furthermore, the specific types of health and wellness entities mentioned are:

[0017] The attributes of a user entity include the user's basic health information, medical history, medication history, rehabilitation indications, and tolerance attributes;

[0018] The attributes of a health and wellness exercise entity include the type of exercise, operating parameters, applicable conditions, and contraindications.

[0019] The attributes of a state entity include quantitative indicators of the user's physiological state, emotional state, and environmental state.

[0020] The attributes of the resulting entity include health and wellness effects, adverse reactions, and the level and probability of risk events;

[0021] The relationships between the entities are specifically as follows:

[0022] Similarity relationships include similarity relationships between user entities, similarity relationships between state entities, and similarity relationships between action entities, which are used for similarity matching in counterfactual reasoning;

[0023] Causality is the relationship between health and wellness action entities and outcome entities, with confidence weights, used for counterfactual inference and graph updating.

[0024] Furthermore, in S200, the dual timing difference error ,in, For the first value network in parameters The current user state in the state entity below Take action Expected return assessment value, For the second value network in parameters Next, for the same state and actions The expected return assessment value, the The larger the value, the more likely it is to be a first-value network. With the second value network The greater the disagreement on the value assessment of the same action, the higher the knowledge graph. Neutral State and actions The higher the uncertainty of the relevant knowledge nodes.

[0025] Furthermore, in S200, Compared with the first threshold and the second threshold, when When that happens, it switches to S400 for proactive inquiry-based exploration. At that time, it will switch to S300 for counterfactual virtual exploration. In such cases, decisions are made directly based on current knowledge.

[0026] Furthermore, in S300, the specific steps of the counterfactual virtual exploration are as follows:

[0027] The core entities anchored to real-world experience samples include: the current user entity U, the current state entity s, and the executed action. Result Status B, Actual Reward ;

[0028] In knowledge graph In the process, retrieve a set of similar users whose similarity to user entity U is higher than a preset user similarity threshold. retrieval and status A set of similar states whose similarity exceeds a preset state similarity threshold. Retrieve the set of candidate counterfactual actions that belong to the same category as action A. ;

[0029] For each candidate counterfactual action C, based on similar users in the knowledge graph In similar states Record the causal relationship of the next action C, and deduce the state. The expected result state D and expected reward corresponding to the next action C. Generate counterfactual samples , The next state after performing action C;

[0030] The counterfactual samples are subjected to risk verification and confidence screening. Counterfactual samples with an inference confidence level higher than a preset confidence threshold are retained and stored in the reinforcement learning experience replay pool. It is used for joint training of the value network in S500.

[0031] Furthermore, in S400, the specific steps of the active inquiry-based exploration are as follows:

[0032] Generate targeted query statements based on the current context information, and ask users questions that are directly related to the current knowledge node;

[0033] It receives and parses users' natural language feedback, and uses a knowledge graph-based semantic understanding model to transform the user's feedback content into structured knowledge triples. At the same time, the actual reward value for this interaction is generated. ;

[0034] Update the knowledge graph using the generated structured knowledge triples. The confidence weights of the corresponding knowledge nodes and the strength parameters of the relation edges are used, along with the real-world interaction experience. Stored as a sample in the experience playback pool These are marked as real verification samples to distinguish them from virtual samples. For the current state entity, These are rehabilitation movements that have actually been performed. To perform realistic actions The real user state in the next moment after being perceived.

[0035] Furthermore, in S500, the first value network is trained. Second Value Network The specific steps are as follows:

[0036] From the experience replay pool According to the preset dynamic ratio A mixture of real-world and virtual experience samples is used to construct training samples for minimizing the first-value network. Second Value Network The joint loss function;

[0037] The first value network is updated based on empirical samples using a temporal difference learning algorithm. parameters Second Value Network parameters Update the target to make and Differences in the evaluation of the same state-action convergence;

[0038] Monitoring the updated first value network Second Value Network Knowledge Triples The change in value assessment, when the change in value assessment exceeds a preset knowledge update threshold. If so, then the knowledge triplet needs to be updated, where, The head entities in the knowledge graph include user entities, health and wellness action entities, state entities, and result entities. For connector entity With tail entity Semantic relationships, including causal relationships and similarity relationships, The tail entity is the target entity in the knowledge graph;

[0039] When knowledge triples need to be updated, the knowledge graph... Perform atomic update operations to add new knowledge triples and delete those with confidence levels below the lower bound threshold. Knowledge triples.

[0040] Furthermore, the process of making direct decisions based on current knowledge is as follows: the current state is input into the trained policy network to generate an action to be executed. The health care robot executes the action and interacts with the user through sensors. Implicit feedback from the user is collected to calculate immediate rewards. The real experience tuple formed by this interaction is stored in the experience replay pool.

[0041] Compared with existing technologies, this reinforcement learning-based method for dynamically updating the knowledge graph of elderly care robots has the following advantages:

[0042] I. This invention achieves zero-risk virtual exploration through a counterfactual reasoning mechanism, fundamentally avoiding the security risks to users caused by the trial-and-error process. This mechanism uses the current knowledge graph as a world model. When highly fuzzy knowledge nodes are identified, it does not directly execute unknown actions in the real environment. Instead, based on similar user group records and association rules in the knowledge graph, it infers the virtual results generated by executing virtual actions and calculates virtual reward values. By storing virtual experience samples in an experience replay pool to participate in subsequent network training, the reinforcement learning agent can learn the potential value of an action without ever actually executing it. This ensures that the actions and results of counterfactual reasoning are within safety boundaries, thereby constructing a virtual learning environment throughout the entire knowledge exploration process and solving the security threats to users caused by the trial-and-error process.

[0043] Second, this invention utilizes a knowledge uncertainty quantification mechanism based on dual temporal difference errors to achieve precise allocation of computing resources and user interaction costs. By calculating the evaluation differences of two networks for the same state-action pair, the divergence degree of the value network is used as a quantitative indicator of the fuzziness of the corresponding node in the knowledge graph. This ensures that computing resources and user interaction costs are only allocated to highly fuzzy knowledge nodes that truly need optimization, avoiding ineffective investment in mature knowledge and fundamentally improving the efficiency and accuracy of knowledge updates.

[0044] Other advantages, objectives and features of the invention will be set forth in part in the description which follows, and in part will be apparent to those skilled in the art from the following examination or study, or may be learned from the practice of the invention. Attached Figure Description

[0045] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are merely some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without any creative effort.

[0046] Figure 1 A flowchart illustrating the steps of a reinforcement learning-based method for dynamically updating the knowledge graph of elderly care robots.

[0047] Figure 2 This is a flowchart illustrating the steps of counterfactual virtual exploration in an embodiment of the present invention;

[0048] Figure 3 This is a flowchart illustrating the steps of the active inquiry-based exploration in an embodiment of the present invention. Detailed Implementation

[0049] To better understand the above technical solutions, a detailed description of the solutions will be provided below in conjunction with the accompanying drawings and specific embodiments. Obviously, the described embodiments are merely some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0050] To address the shortcomings of existing static knowledge graph construction methods for elderly care robots, which struggle to adapt to individual user differences and dynamic changes in physiological states, and which suffer from technical deficiencies such as exploration safety risks, lack of awareness of knowledge ambiguity, and subjective knowledge update blind spots in reinforcement learning update models, this invention provides a dynamic knowledge graph update method for elderly care robots based on reinforcement learning. This method aims to accurately measure knowledge uncertainty by constructing a dual value network, achieve zero-risk virtual exploration by combining counterfactual reasoning, conduct proactive inquiry-based exploration through natural language interaction, and optimize the value network and dynamically iterate the knowledge graph through joint training of real and virtual experiences. This constructs an intelligent update system encompassing knowledge ambiguity identification, safe exploration learning, and adaptive evolution of the knowledge graph, enabling the knowledge system of elderly care robots to continuously evolve to meet personalized user needs, thereby improving the accuracy and safety of elderly care services.

[0051] This invention is mainly applied to scenarios such as home-based rehabilitation robots and intelligent rehabilitation equipment in rehabilitation institutions, addressing the needs of personalized rehabilitation training, daily health monitoring, and adaptation of rehabilitation movements for rehabilitation services. Traditional methods rely on static knowledge graphs and undifferentiated reinforcement learning exploration, which are prone to problems such as poor service adaptability, trial and error that harms users, and low knowledge update efficiency. This invention achieves safe, efficient, and accurate dynamic updates of the knowledge graph of rehabilitation robots by quantifying the knowledge uncertainty of dual value networks, securely generating virtual experience through counterfactual reasoning, fusing semantic feedback from proactively asked users, and jointly training real and virtual experiences, thereby significantly improving the personalization level and intelligent decision-making capabilities of rehabilitation services.

[0052] Specifically, such as Figure 1 As shown, a method for dynamically updating the knowledge graph of a healthcare robot based on reinforcement learning is proposed. This method includes the following steps:

[0053] S100. Constructing the Initial Knowledge Graph and Dual Value Network: Constructing a knowledge graph that includes the types of health and wellness entities and the relationships between them. and establish a first value network. Second Value Network A dual-value network architecture;

[0054] S200, Knowledge Uncertainty Measurement Based on Dual Time-Series Differential Error: This involves measuring the current user state... Input First Value Network Second Value Network ,calculate and For actions to be performed The output difference is used to obtain the dual time-series difference error. ,Will With the first threshold Second threshold Comparisons are made, and knowledge exploration paths are determined based on the comparison results. Knowledge exploration paths include counterfactual virtual exploration, proactive inquiry-based exploration, and direct decision-making.

[0055] S300, Secure Generation of Virtual Experiences Based on Counterfactual Reasoning: In response to the counterfactual virtual exploration in S200, the counterfactual reasoning module is initiated, based on the current knowledge graph. As a world model, it utilizes similar entity relationships and similar user group records in the graph to generate virtual actions that are not actually executed. The corresponding virtual experience samples are then stored in the experience replay pool. ;

[0056] S400, Proactive Inquiry-Based Exploration and Revision: Responding to the proactive inquiry-based exploration in S200, semantic feedback from users is obtained through natural language interaction, transformed into structured knowledge triples, and the knowledge graph is revised in real time. The confidence level of the corresponding knowledge node;

[0057] S500, Joint Experience Replay and Dynamic Updates of Knowledge Graph: From the Experience Replay Pool The first value network is trained by dynamically mixing real and virtual experience samples in a mixed sampling process. Second Value Network The knowledge graph is dynamically updated based on the reassessment of knowledge value by the trained value network. The topological structure of knowledge triples.

[0058] In the specific implementation process, the health and wellness robot initiates the health and wellness service process. First, it completes the construction of an initial knowledge graph and a dual value network. In this embodiment, it first sorts out the core knowledge system in the health and wellness field, determines the types of health and wellness entities, namely user entities, health and wellness action entities, state entities, and result entities, and defines the core relationships between entities as causal relationships and similarity relationships. Based on this, an initial knowledge graph G is constructed. The attributes of the user entity cover quantitative and qualitative information such as basic health information, medical history, medication history, rehabilitation indications, and tolerance, such as the hypertension history, joint rehabilitation indications, and massage intensity tolerance level of elderly users. The attributes of the health and wellness action entity include action type, operation parameters, applicable conditions, and contraindications. The attributes of the state entity are quantitative indicators of the user's physiological state, emotional state, and environmental state. Physiological and environmental indicators are collected through robot sensors, and emotional states are converted into quantitative values ​​of 0-10 through facial expression recognition and user feedback. The attributes of the result entity include the level and probability of health and wellness effects, adverse reactions, and risk events, all represented by normalized values ​​of 0-1.

[0059] The similarity relationships between entities include the similarity between user entities, state entities, and action entities. This is achieved through cosine similarity calculation of entity attributes and is used for similarity matching in counterfactual reasoning. For example, users with similar age, medical history, and rehabilitation indicators are judged as similar user entities. The causal relationship is the association between health care action entities and result entities. It is used for counterfactual inference and graph update with confidence weights of 0-1.

[0060] Initial knowledge graph Use a graph database for storage, in triplets All knowledge is recorded in the form of [the document], in which For the head entity, For semantic relations, This is the tail entity. Simultaneously, a first value network is constructed. Second Value Network The dual value network architecture uses deep neural network structures for both value networks. The input is the feature vector of the current user state s, and the output is the health and wellness action to be performed. The expected return evaluation value. In this embodiment, the hidden layers of the network are all set to 3 layers, with the number of neurons being 128, 64, and 32 respectively. The activation function is RelU, and the output layer uses a linear activation function. Initial parameters of the two value networks. and Different random seeds are used for initialization to ensure that the initial assessments are different, laying the foundation for subsequent knowledge uncertainty measurement.

[0061] After completing the initial knowledge graph and dual-value network construction, the health and wellness robot collects the user's real-time physiological and environmental status through sensors, determines the current user state s by combining user commands, performs a knowledge uncertainty measurement based on dual temporal difference error, and inputs the feature vector of the current user state s into the first value network. Second Value Network For the health and wellness activities to be performed The expected return evaluation values ​​of the two networks were obtained respectively. and Calculate the dual time-series difference error , The larger the value, the greater the discrepancy in the value assessments of the same state-action pair between the two value networks, i.e., the greater the difference in the knowledge graph's value assessment. In relation to the current user state and actions to be performed The higher the uncertainty of the relevant knowledge nodes.

[0062] In this embodiment, a first threshold is preset. Second threshold , calculate The decision-making knowledge exploration path is compared with two thresholds: when When this indicates extremely high uncertainty in knowledge, the system switches to S400 for proactive inquiry-based exploration, seeking direct feedback from the user; when When this indicates a certain degree of uncertainty in knowledge, the system transitions to S300 for counterfactual virtual exploration, generating virtual experience learning; when This indicates a high degree of certainty in knowledge, allowing for direct decision-making based on current knowledge and the execution of corresponding health and wellness actions.

[0063] If determined to be a counterfactual virtual exploration, execute step S300, such as... Figure 2 As shown, in this embodiment, the specific process of counterfactual virtual exploration is as follows:

[0064] The core entities that define real-world experience samples include the current user entity U, the current state entity s, and the action to be performed. Preset Result Status B, Actual Reward ;

[0065] In knowledge graph In the process, cosine similarity retrieval is used to obtain the user entity. A set of similar users with a similarity score higher than a similarity threshold , and state A set of similar states with a similarity higher than a threshold , and action Set of candidate counterfactual actions in the same category ;

[0066] For each action C in the candidate counterfactual action set, based on similar users in the knowledge graph In similar states The causal relationship record of the next action C is used to deduce the current state using the weighted average method. The expected result state D and expected reward corresponding to the next action C. This leads to the generation of counterfactual samples. ,in The next state after performing action C;

[0067] The generated counterfactual samples undergo risk verification and confidence screening. Based on a preset inference confidence threshold, the inference confidence of each counterfactual sample is obtained by calculating the product of the similar entity matching degree and the causal relationship confidence during the inference process. Samples with confidence scores lower than the inference confidence threshold are removed. The qualified virtual experience samples are stored in the reinforcement learning experience replay pool B for subsequent joint training of the value network in S500.

[0068] If the inquiry is determined to be an active inquiry-based exploration, proceed with step S400, such as... Figure 3As shown, in this embodiment, the specific process of active inquiry-based exploration is as follows:

[0069] Generate targeted natural language queries based on the current context information and send them to the user;

[0070] The voice interaction module of the health and wellness robot receives natural language feedback from users and uses the BERT semantic understanding model based on knowledge graph enhancement to transform the user's feedback into structured knowledge triples.

[0071] Update the knowledge graph using the generated structured knowledge triples. The confidence weight of the corresponding knowledge node and the strength parameters of the relation edge;

[0072] And this real interactive experience Stored as a sample in the experience playback pool These are labeled as real verification samples to distinguish them from virtual samples. These are the revised health and wellness exercises. This is the preset user state after execution.

[0073] If the decision is determined to be based on current knowledge, the current state is input into the trained policy network to generate a definite action. The health and wellness robot then performs health and wellness services according to the action parameters. Simultaneously, it interacts with the user through sensors, collecting implicit feedback such as changes in the user's physiological indicators and limb movements. Combined with a preset reward function, it calculates an immediate reward and stores the real experience tuple formed from this interaction into an experience replay pool. In this context, it serves as a sample of real-world interactive experiences.

[0074] After completing the above exploration or decision-making steps, the S500 joint experience replay and knowledge graph dynamic update steps are executed to optimize the value network and iterate the knowledge graph. In this embodiment, the experience replay pool is first used... According to the preset dynamic ratio Hybrid sampling is used to construct a training set to minimize the first value network. Second Value Network The joint loss function is a weighted sum of the temporal difference losses of the two networks. Based on the training set, the first value network is updated using the temporal difference learning algorithm (Q-Learning). parameters Second Value Network parameters Update the target bit to make and Differences in the evaluation of the same state-action Convergence, in essence, involves gradually reducing the evaluation discrepancy between the two networks, thereby increasing the certainty of knowledge. After training, the updated first value network is monitored. Second Value Network Knowledge graph All knowledge triples The change in value assessment, with a preset knowledge update threshold. When the change in the value assessment of a certain knowledge triple exceeds When this happens, it is determined that the knowledge triplet needs to be updated.

[0075] When it is determined that a knowledge triple needs to be updated, the knowledge graph is... Perform atomic update operations. For triples whose value assessment is improved, increase the confidence weight of their causal or similarity relationships. For newly added user feedback knowledge, add new knowledge triples. For triples whose value assessment is significantly reduced and whose confidence is below the lower threshold, perform atomic update operations. The knowledge triples are directly deleted from the knowledge graph, thus completing the dynamic update of the knowledge graph topology.

[0076] In summary, this invention addresses the challenges of static knowledge graphs, high security risks in reinforcement learning updates, lack of awareness of knowledge ambiguity, and lack of access to subjective knowledge updates in healthcare robots through a comprehensive design encompassing the entire chain from initial knowledge graph and dual value network construction, knowledge uncertainty quantification, secure exploratory learning, joint experience training, and dynamic graph updates. Specifically, this method achieves precise quantification of knowledge ambiguity through temporal difference errors in the dual value network, ensuring accurate allocation of exploration resources; it uses counterfactual reasoning to treat the knowledge graph as a world model, enabling zero-risk virtual experience generation and fundamentally avoiding the harm to users caused by trial and error; it integrates subjective user feedback through proactive inquiry-based exploration via natural language interaction, filling the blind spots in subjective knowledge updates; and it achieves efficient optimization of the value network and adaptive evolution of the knowledge graph through joint training of real and virtual experiences. The entire method is highly intelligent, enabling the healthcare robot's knowledge system to continuously adapt to users' personalized needs and status changes, transforming from static knowledge support to dynamic intelligent evolution, significantly improving the service accuracy and user experience of healthcare robots.

[0077] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any simple modifications, equivalent changes and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention.

Claims

1. A method for dynamically updating a health-care robot knowledge graph based on reinforcement learning, characterized in that, The specific steps of this method are as follows: S100. Construct a knowledge graph that includes the types of health and wellness entities and the relationships between them. and establish a first value network. Second Value Network A dual-value network architecture; S200, Change the current user status Input the first value network Second Value Network ,calculate and For actions to be performed The output difference is used to obtain the dual time-series difference error. ,Will With the first threshold Second threshold A comparison is made, and a knowledge exploration path is determined based on the comparison results. The knowledge exploration path includes counterfactual virtual exploration, active inquiry-based exploration, and direct decision-making. S300, in response to the counterfactual virtual exploration in S200, initiates the counterfactual deduction module, based on the current knowledge graph. As a world model, it utilizes similar entity relationships and similar user group records in the graph to generate virtual actions that are not actually executed. The corresponding virtual experience samples are then stored in the experience replay pool. ; S400, responding to the proactive inquiry-based exploration in S200, obtains semantic feedback from the user through natural language interaction, transforms the semantic feedback into structured knowledge triples, and corrects the knowledge graph. The confidence level of the corresponding knowledge node; S500, the experience playback pool The first value network is trained by dynamically mixing real and virtual experience samples in a mixed manner. Second Value Network The knowledge graph is dynamically updated based on the reassessment of knowledge value by the trained value network. The topological structure of knowledge triples in Chinese; In S100, the types of health and wellness entities include user entities, health and wellness action entities, state entities, and result entities, and the relationships between entities include causal relationships and similarity relationships. In S200, the dual timing difference error ,in, For the first value network in parameters The current user state in the state entity below Take action Expected return assessment value, For the second value network in parameters Next, for the same state and actions Expected return assessment value; In S200, Compared with the first threshold and the second threshold, when When that happens, it switches to S400 for proactive inquiry-based exploration. At that time, it will switch to S300 for counterfactual virtual exploration. In such cases, decisions are made directly based on current knowledge. The specific steps of the counterfactual virtual exploration in S300 are as follows: The core entities anchored to real-world experience samples include: the current user entity U, the current state entity s, and the executed action. Result Status B, Actual Reward ; In knowledge graph In the process, retrieve a set of similar users whose similarity to user entity U is higher than a preset user similarity threshold. retrieval and status A set of similar states whose similarity exceeds a preset state similarity threshold. Retrieve the set of candidate counterfactual actions that belong to the same category as action A. ; For each candidate counterfactual action C, based on similar users in the knowledge graph In similar states Record the causal relationship of the next action C, and deduce the state. The expected result state D and expected reward corresponding to the next action C. Generate counterfactual samples , The next state after performing action C; The counterfactual samples are subjected to risk verification and confidence screening. Counterfactual samples with an inference confidence level higher than a preset confidence threshold are retained and stored in the reinforcement learning experience replay pool. It is used for joint training of the value network in S500.

2. The method for dynamically updating the knowledge graph of a rehabilitation and elderly care robot based on reinforcement learning according to claim 1, characterized in that, The specific types of health and wellness entities mentioned are: The attributes of a user entity include the user's basic health information, medical history, medication history, rehabilitation indications, and tolerance attributes; The attributes of a health and wellness exercise entity include the type of exercise, operating parameters, applicable conditions, and contraindications. The attributes of a state entity include quantitative indicators of the user's physiological state, emotional state, and environmental state. The attributes of the resulting entity include health and wellness effects, adverse reactions, and the level and probability of risk events; The relationships between the entities are specifically as follows: Similarity relationships include similarity relationships between user entities, similarity relationships between state entities, and similarity relationships between action entities, which are used for similarity matching in counterfactual reasoning; Causality is the relationship between health and wellness action entities and outcome entities, with confidence weights, used for counterfactual inference and graph updating.

3. The method for dynamically updating the knowledge graph of a rehabilitation and elderly care robot based on reinforcement learning according to claim 1, characterized in that, In S400, the specific steps of the active inquiry-based exploration are as follows: Generate targeted query statements based on the current context information, and ask users questions that are directly related to the current knowledge node; It receives and parses users' natural language feedback, and uses a knowledge graph-based semantic understanding model to transform the user's feedback content into structured knowledge triples. At the same time, the actual reward value for this interaction is generated. ; Update the knowledge graph using the generated structured knowledge triples. The confidence weights of the corresponding knowledge nodes and the strength parameters of the relation edges are used, along with the real-world interaction experience. Stored as a sample in the experience playback pool These are marked as real verification samples to distinguish them from virtual samples. For the current state entity, These are rehabilitation movements that have actually been performed. To perform realistic actions The real user state in the next moment after being perceived.

4. The method for dynamically updating the knowledge graph of a rehabilitation and elderly care robot based on reinforcement learning according to claim 1, characterized in that, In step S500, the first value network is trained. Second Value Network The specific steps are as follows: From the experience replay pool According to the preset dynamic ratio A mixture of real-world and virtual experience samples is used to construct training samples for minimizing the first-value network. Second Value Network The joint loss function; The first value network is updated based on empirical samples using a temporal difference learning algorithm. parameters Second Value Network parameters Update the target to make and Differences in the evaluation of the same state-action convergence; Monitoring the updated first value network Second Value Network Knowledge Triples The change in value assessment, when the change in value assessment exceeds a preset knowledge update threshold. If so, then the knowledge triplet needs to be updated, where, The head entities in the knowledge graph include user entities, health and wellness action entities, state entities, and result entities. For connector entity With tail entity Semantic relationships, including causal relationships and similarity relationships, The tail entity is the target entity in the knowledge graph; When knowledge triples need to be updated, the knowledge graph... Perform atomic update operations to add new knowledge triples and delete those with confidence levels below the lower bound threshold. Knowledge triples.

5. The method for dynamically updating the knowledge graph of a rehabilitation and elderly care robot based on reinforcement learning according to claim 1, characterized in that, The process of making direct decisions based on current knowledge is as follows: the current state is input into the trained policy network to generate an action to be executed. The health care robot executes the action and interacts with the user through sensors. Implicit feedback from the user is collected to calculate the immediate reward. The real experience tuple formed by this interaction is stored in the experience replay pool.

Citation Information

Patent Citations

  • Interactive clinical decision support system and method based on large model knowledge enhancement

    CN118335360A

  • Decision optimization method fusing enhanced multi-modal learning and knowledge graph

    CN120409642A