Surgical robot control method and system based on world model and reinforcement learning and robot system

By combining the world model and reinforcement learning algorithm, a robot control strategy optimization model was established, which solved the problem of insufficient flexibility of traditional surgical robots in complex environments, achieved efficient adaptation to the surgical environment and accurate operation, and improved the success rate and safety of the surgery.

CN120791749APending Publication Date: 2025-10-17HANGLOK-TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510901037.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Traditional surgical robot control algorithms lack flexibility and adaptability in complex surgical environments and are unable to cope with sudden changes. Learning-based control algorithms require large amounts of data and time to adapt and may lead to operational errors when faced with emergencies.

Method used

Combining the world model and reinforcement learning algorithm, a robot control strategy optimization model is established through the Actor-Critic algorithm. The state parameters are collected using the sensor device, the world model predicts environmental changes, and the policy parameters and value function parameters are optimized through the policy gradient and TD error method to realize action selection.

Benefits of technology

It improves the predictive ability and adaptability of surgical robots in complex environments, ensures the accuracy and safety of operations, reduces dependence on large amounts of training data, shortens the adaptation cycle, and improves the success rate and safety of operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120791749A_ABST
    Figure CN120791749A_ABST
Patent Text Reader

Abstract

The invention discloses a surgical robot control method and system based on a world model and reinforcement learning and a robot system.The method comprises the steps that a robot control strategy optimization model is established based on an Actor-Critic algorithm of reinforcement learning, and an action selection strategy of a surgical robot is output according to strategy parameters and value function parameters; a sensing device is used for collecting state parameters in the current operation environment; using a world model to learn state parameters in the current operation environment so as to predict and output state change data in the operation process; state parameters in the current surgical environment and environment state change data are input into a robot control strategy optimization model; the strategy optimization model analyzes the state parameters and the environment state change data in the current operation environment and outputs a corresponding action selection strategy. According to the method, the world model is integrated into the reinforcement learning algorithm, so that the adaptability and predictive ability of the surgical robot in a complex surgical environment are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of artificial intelligence, and in particular to a surgical robot control method, control system and robot system based on a world model and reinforcement learning. BACKGROUND

[0002] In a complex surgical environment, the control system of a surgical robot faces many challenges. Traditional surgical robot control algorithms, such as Model Predictive Control (MPC), mainly rely on pre-set rules and parameters. These algorithms can perform well in stable and predictable environments. However, in a variable surgical environment, these traditional algorithms often lack sufficient flexibility and adaptability. For example, when unexpected situations occur during surgery, such as sudden changes in patient physiological parameters or unexpected movement of surgical instruments, traditional algorithms may not be able to respond appropriately in a timely manner, affecting the accuracy and safety of the surgery.

[0003] On the other hand, learning-based control algorithms, especially reinforcement learning algorithms, learn optimal operation strategies through continuous interaction with the environment, enabling robots to gradually master skills to complete specific tasks. These algorithms optimize their decisions through a trial-and-error process in the hope of achieving higher performance. Although this approach can improve the adaptability and flexibility of surgical robots in some cases, learning-based control algorithms still face challenges in the surgical environment due to the complexity of the environment and the extremely high requirements for safety. These algorithms often require a large amount of data and a large amount of time to complete learning. Moreover, when faced with sudden changes, learning-based control algorithms may not be able to make quick and accurate predictions and adjustments, which can lead to unexpected operational errors during surgery.

[0004] The disclosure of the above background art is only used to assist in understanding the concept and technical solutions of the present application, and does not necessarily belong to the prior art of the present application, nor does it necessarily provide technical teaching. In the absence of explicit evidence that the above background art has been disclosed before the filing date of the present application, the above background art should not be used to evaluate the novelty and inventiveness of the present application. SUMMARY

[0005] The purpose of the present application is to provide an innovative surgical robot control method that integrates a world model into a reinforcement learning algorithm to improve the adaptability and predictive ability of surgical robots in complex surgical environments.

[0006] To achieve the above purpose, the technical solutions adopted by the present application are as follows:

[0007] A surgical robot control method based on a world model and reinforcement learning, comprising the following steps:

[0008] An Actor-Critic algorithm based on reinforcement learning is used to establish a robot control strategy optimization model configured to output an action selection strategy of a surgical robot according to policy parameters and value function parameters;

[0009] State parameters in a current surgical environment are collected by a sensing device;

[0010] The state parameters in the current surgical environment are learned by a world model to predict state change results in a surgical process and output predicted environmental state change data;

[0011] The state parameters in the current surgical environment collected by the sensing device and the predicted environmental state change data predicted by the world model are input into the robot control strategy optimization model;

[0012] The robot control strategy optimization model analyzes the state parameters in the current surgical environment and the environmental state change data and outputs a corresponding action selection strategy.

[0013] Further, any of the above technical solutions or a combination of multiple technical solutions, the robot control strategy optimization model analyzes the state parameters in the current surgical environment and the environmental state change data and outputs a corresponding action selection strategy, comprising:

[0014] The robot control strategy optimization model pre-initializes policy parameters and value function parameters;

[0015] According to the policy parameters and value function parameters, an advantage function is calculated to determine the degree of advantage of taking an alternative action compared to taking an expected action under the state parameters in the current surgical environment;

[0016] The policy parameters are updated according to the advantage function using a policy gradient method, wherein the policy gradient method guides the policy parameters to increase the probability of actions with higher advantages;

[0017] The value of the updated policy parameters to the environmental state change data is estimated, and the value function parameters are updated using a minimum TD error method;

[0018] The advantage function is iteratively calculated and the policy parameters and value function parameters are iteratively updated using the updated policy parameters and value function parameters until a convergence condition is reached;

[0019] According to the policy parameters and value function parameters after the convergence condition is reached, a corresponding action selection strategy is determined.

[0020] Further, any of the above technical solutions or a combination of multiple technical solutions, the robot control strategy optimization model updates the policy parameters by:

[0021] The robot control strategy optimization model estimates the expected return obtained after taking an alternative action a and following a strategy Str in the current state s in the surgical environment, denoted as Q Str (s, a);

[0022] The expected return obtained after following a strategy Str in the current state s in the surgical environment is estimated, denoted as V Str (s), and the expected return obtained after following a strategy Str in the next state s' following the environmental state change data is estimated, denoted as V Str (s');

[0023] The advantage function A is calculated by the following formula Str (s, a) = Q Str (s, a) - V Str (s) + β1 x V Str (s'), where β1 is a preset discount factor.

[0024] The calculation results of the advantage function corresponding to each alternative action a are compared to determine the optimal action a*.

[0025] The policy parameter is defined as θ, and the policy function is defined as Str θ (a|s), the log gradient of the policy parameter to improve the probability of the optimal action a* is calculated, denoted as

[0026] The policy gradient is calculated by multiplying the log gradient and the advantage function.

[0027] The policy parameter is updated according to the following formula: where θ t+1 is the updated policy parameter, θ t is the policy parameter before updating in the current state s in the surgical environment, and α1 is a preset learning rate.

[0028] Further, any one of the technical solutions or a combination of the technical solutions, the log gradient is calculated in the policy network through a backpropagation algorithm.

[0029] The policy gradient is calculated by the following formula: where J(θ) represents the performance function of the policy Str θ , and E Strθ represents the expectation under the policy Str θ .

[0030] Furthermore, based on any one of the above technical solutions or a combination of multiple technical solutions, the robot control strategy optimization model updates the value function parameters in the following manner:

[0031] The robot control strategy optimization model estimates the immediate value of taking action a under the state parameter s in the current surgical environment, denoted as R t ;

[0032] Define the state value function as Among them, S is the state parameter, is the value function parameter; calculate the state value of the state parameter s in the current surgical environment, recorded as And calculate the state value of the next state s' following the environmental state change data, recorded as

[0033] The TD error is calculated using the following formula: Among them, β2 is the preset discount factor;

[0034] Calculate the gradient of the state value function under the state parameter s in the current surgical environment

[0035] Update the value function parameters according to the following formula: in, is the updated value function parameter, is the value function parameter under the state parameter s in the current surgical environment before updating, and α2 is the preset learning rate.

[0036] Furthermore, according to any one of the above technical solutions or a combination of multiple technical solutions, the policy function related to the policy parameters is implemented by a neural network, and the training goal of the network is to maximize the expectation of cumulative return;

[0037] The value function associated with the value function parameters is implemented by a neural network, the training goal of which is to estimate the value of each state under the current policy;

[0038] The calculation of the state value is completed in the value network through forward propagation.

[0039] Furthermore, based on any one of the above technical solutions or a combination of multiple technical solutions, the world model is constructed in the following manner:

[0040] Utilizing a variety of sensing devices to detect visual and / or audio and / or temperature information in the surgical environment, and extracting time domain features and / or frequency domain features therefrom to obtain a variety of surgical environment feature data;

[0041] Evaluate the importance of various operating environment characteristics to the operating environment monitoring, screen out key operating environment characteristics, and fuse to obtain comprehensive environment characteristic data;

[0042] A recurrent neural network model is used to construct a world model, the world model is trained using the collected comprehensive environment characteristic data, and the world model learns the dynamic changes of the environment based on a self-supervised learning method to predict future environmental state changes.

[0043] Further, any of the technical solutions or combinations of multiple technical solutions described above, the world model evaluates the prediction performance of the model by a cross-validation method, and the indicators of the prediction performance include the area under the receiver operating characteristic curve, sensitivity and specificity.

[0044] Further, any of the technical solutions or combinations of multiple technical solutions described above, the key operating environment characteristics include intraoperative patient vital signs, and the corresponding future environmental state changes are intraoperative massive hemorrhage risk states;

[0045] The key operating environment characteristics also include patient preoperative signs, and the corresponding future environmental state changes of the patient preoperative signs and intraoperative patient vital signs are postoperative adverse event risk states.

[0046] According to another aspect of the present application, the present application provides a surgical robot control system based on a world model and reinforcement learning, comprising:

[0047] A sensing device configured to collect state parameters in a current operating environment;

[0048] A world model configured to learn the state parameters in the current operating environment to predict state change results during the operation and output predicted environmental state change data;

[0049] A robot control strategy optimization model based on an Actor-Critic algorithm of reinforcement learning, which outputs an action selection strategy of the surgical robot according to policy parameters and value function parameters;

[0050] The sensing device and the world model are both electrically connected to the input end of the robot control strategy optimization model, and the robot control strategy optimization model analyzes the state parameters in the current operating environment and the environmental state change data, and outputs the corresponding action selection strategy.

[0051] Further, any of the technical solutions or combinations of multiple technical solutions described above, the robot control strategy optimization model includes a policy parameter optimization module and a value function parameter optimization module, wherein:

[0052] The policy parameter optimization module initializes policy parameters, and the value function parameter optimization module initializes value function parameters;

[0053] The robot control policy optimization model calculates an advantage function according to the policy parameters and the value function parameters, to determine the degree of advantage of taking an alternative action compared to taking an expected action under state parameters in the current surgical environment;

[0054] The policy parameter optimization module updates the policy parameters according to the advantage function by using a policy gradient method, wherein the policy gradient method guides the policy parameters to increase the probability of actions that bring higher advantage;

[0055] The value function parameter optimization module estimates the value of the updated policy parameters to the environment state change data, and updates the value function parameters by using a minimum TD error method;

[0056] The robot control policy optimization model iteratively calculates the advantage function and iteratively updates the policy parameters and the value function parameters by using the updated policy parameters and the value function parameters, until a convergence condition is reached; and determines a corresponding action selection policy according to the policy parameters and the value function parameters after the convergence condition is reached.

[0057] Further, any of the above technical solutions or a combination of the above technical solutions, the policy parameter optimization module is configured to perform:

[0058] Estimate the expected return obtained after taking an alternative action a and following a policy Str under state parameters s in the current surgical environment, denoted as Q Str (s,a);

[0059] Estimate the expected return obtained after following a policy Str under state parameters s in the current surgical environment, denoted as V Str (s), and estimate the expected return obtained after following a policy Str under the next state s' following the environment state change data, denoted as V Str (s');

[0060] Calculate the advantage function A Str (s,a)=Q Str (s,a)-V Str (s)+β1×V Str (s'),wherein β1 is a preset discount factor;

[0061] Compare the calculation results of the advantage functions corresponding to each alternative action a to determine the optimal action a*;

[0062] Define the policy parameters as θ, and define the policy function as Str θ(a|s), the log gradient of the policy parameter on the probability of promoting the optimal action a* is calculated, denoted as

[0063] The policy gradient is calculated by calculating the product of the log gradient and the advantage function

[0064] The policy parameter is updated according to the following formula: Where θ t+1 is the updated policy parameter, θ t is the policy parameter before updating under the state parameter s in the current surgical environment, and α1 is a preset learning rate.

[0065] Further, any of the technical solutions or combinations of the technical solutions described above, the value function parameter optimization module is configured to perform:

[0066] The immediate value obtained by taking action a under state parameter s in the current surgical environment is estimated, denoted as R t ;

[0067] The state value function is defined as Where S is the state parameter, is the value function parameter; the state value of the state parameter s in the current surgical environment is calculated, denoted as And the state value of the next state s' following the environmental state change data is calculated, denoted as

[0068] The TD error is calculated by the following formula: Where β2 is a preset discount factor;

[0069] The gradient of the state value function under the state parameter s in the current surgical environment is calculated

[0070] The value function parameter is updated according to the following formula: Where is the updated value function parameter, is the value function parameter before updating under the state parameter s in the current surgical environment, and α2 is a preset learning rate.

[0071] According to another aspect of the present application, a surgical robot system is provided, comprising a surgical robot and a surgical robot control system based on a world model and reinforcement learning as described above, the surgical robot being configured to execute the action selection policy output by the surgical robot control system.

[0072] The technical solutions provided by the present application have the following beneficial effects:

[0073] a. Simulate the dynamic changes of the environment with the world model, and provide effective prediction of future states, so that the robot can adjust the operation strategy more accurately;

[0074] b. By integrating the world model into the reinforcement learning algorithm, the robot not only can better foresee and adapt to sudden changes during the operation process, but also can present higher accuracy and safety when performing tasks;

[0075] c. The reinforcement learning system integrated with the world model can continuously learn and optimize, and through iterative improvement based on feedback of actual operation data, this adaptive ability enables the surgical robot to maintain efficient operation performance in a changing surgical environment. BRIEF DESCRIPTION OF DRAWINGS

[0076] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or prior art description will be briefly introduced as follows. Obviously, the drawings in the following description are only some embodiments described in the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0077] Figure 1 The flow chart of the surgical robot control method based on the world model and the reinforcement learning provided for an exemplary embodiment of the present application;

[0078] Figure 2 The flow chart of the action selection strategy determined by the control strategy optimization model provided for an exemplary embodiment of the present application;

[0079] Figure 3 The flow chart of the model update strategy parameter provided for an exemplary embodiment of the present application;

[0080] Figure 4 The flow chart of the model update value function parameter provided for an exemplary embodiment of the present application;

[0081] Figure 5 The schematic block diagram of the surgical robot system provided for an exemplary embodiment of the present application. DETAILED DESCRIPTION

[0082] In order to make the person skilled in the art better understand the present application, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor should be within the scope of protection of the present application.

[0083] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and in the above drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the application described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, device, product or apparatus including a series of steps or units does not have to be limited to only those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or apparatuses.

[0084] As surgical robots are increasingly widely used in minimally invasive surgery, the intelligent level of their control systems becomes particularly critical. The core task of a surgical robot control system is to accurately perform various complex surgical operations, thereby improving surgical outcomes and ensuring patient safety.

[0085] Current traditional MPC control algorithms or reinforcement learning control algorithms have limited predictive ability and limited adaptability when faced with complex surgical environments, that is, traditional control algorithms are based on system dynamics models for predictive control, but are difficult to adapt to changing environments; reinforcement learning algorithms mainly learn decision-making based on data, lack effective predictive ability for future states, and rely on a large amount of training data, resulting in a long training period to adapt to new environments, making it difficult to meet real-time needs during surgery, especially when dealing with unexpected situations in the surgical environment, which can lead to decision-making errors, affecting the safety and effectiveness of surgery.

[0086] The present application aims to provide a strategy for controlling surgical robots based on a world model combined with a reinforcement learning algorithm to enhance the predictive ability and decision-making intelligence of surgical robots, thereby improving their operation precision and safety in dynamic surgical environments.

[0087] In one embodiment of the present application, a surgical robot control method based on a world model and reinforcement learning is provided, as shown in Figure 1 The control method comprises the following steps:

[0088] S1: based on the Actor-Critic algorithm of reinforcement learning, a robot control strategy optimization model is established, which is configured to output an action selection strategy for the surgical robot according to policy parameters and value function parameters;

[0089] S2: using a sensing device to collect state parameters in the current surgical environment;

[0090] S3: learning state parameters in the current surgical environment using a world model to predict state change results during surgery and output predicted environmental state change data;

[0091] S4: inputting state parameters in the current surgical environment collected by the sensing device and environmental state change data predicted by the world model into the robot control strategy optimization model;

[0092] S5: the robot control strategy optimization model analyzes the state parameters in the current surgical environment and the environmental state change data, and outputs the corresponding action selection strategy.

[0093] The world model can predict possible state changes in the surgical process by learning data in the surgical environment. Through this module, the surgical robot not only relies on current state information for decision-making, but also can foresee possible situations in advance, so as to consider more extensive factors in the decision-making process and improve the accuracy of decision-making, thereby achieving more accurate control operation in the dynamic surgical process.

[0094] Based on the traditional reinforcement learning algorithm, the prediction results of the world model are introduced into the strategy optimization process, so that the surgical robot can better adapt to dynamic changes in the complex surgical environment and optimize its operation strategy.

[0095] By combining the accurate prediction of the world model with advanced reinforcement learning algorithms, intelligent optimization of surgical robot operation decisions is achieved. This optimization strategy enables the robot to quickly adapt and make intelligent operation decisions in a complex and variable surgical environment, improving the success rate and safety of surgery.

[0096] As shown in Figure 2 The robot control strategy optimization model analyzes the state parameters in the current surgical environment and the environmental state change data, and outputs the corresponding action selection strategy, including:

[0097] S510: the robot control strategy optimization model pre-initializes strategy parameters θ0 and value function parameters

[0098] S520: according to the strategy parameters and value function parameters, calculate the advantage function to determine the pros and cons of taking alternative actions compared to taking expected actions under the state parameters in the current surgical environment;

[0099] S530: update the strategy parameters according to the advantage function using the policy gradient method, wherein the policy gradient method guides the strategy parameters to increase the probability of actions with higher advantages;

[0100] S540: estimate the value of the updated policy parameter to the environment state change data, and update the value function parameter by using a minimum TD error method;

[0101] S550: iteratively calculate the advantage function and iteratively update the policy parameter and the value function parameter by using the updated policy parameter and the value function parameter, until a convergence condition is reached;

[0102] S560: determine the corresponding action selection policy according to the policy parameter and the value function parameter after the convergence condition is reached.

[0103] As shown in the robot control policy optimization model updates the policy parameter in the following way: Figure 3

[0104] S531: the robot control policy optimization model estimates the expected return obtained by taking an alternative action a and following a policy Str in the state parameter s in the current surgical environment, denoted as Q Str (s,a);

[0105] S532: estimate the expected return obtained by following the policy Str in the state parameter s in the current surgical environment, denoted as V Str (s), and estimate the expected return obtained by following the policy Str in the next state s' following the environment state change data, denoted as V Str (s');

[0106] S533: calculate the advantage function A Str (s,a) = Q Str (s,a) - V Str (s) + β1 x V Str (s'), where β1 is a preset discount factor;

[0107] S534: compare the calculation results of the advantage function corresponding to each alternative action a to determine the optimal action a*;

[0108] S535: define the policy parameter as θ, and define the policy function as Str θ (a|s), calculate the log gradient of the policy parameter to the probability of promoting the optimal action a*, denoted as The policy function is realized by a neural network, and the training target of the network is to maximize the expected cumulative return;

[0109] S536: calculate the policy gradient by calculating the product of the log gradient and the advantage function

[0110] S537: update the policy parameter according to the following formula: ​wherein θ t+1 is the updated policy parameter, θ t is the policy parameter before updating under the state parameter s in the current surgical environment, and a1 is a preset learning rate.

[0111] Specifically, the log gradient is calculated in the policy network through a back propagation algorithm.

[0112] The policy gradient is calculated by the following formula: wherein J(θ) represents a performance function of the policy Str θ , and E Strθ represents an expectation under the policy Str θ .

[0113] As shown in FIG. 5, the robot control policy optimization model updates the value function parameter in the following manner: Figure 4

[0114] S541: The robot control policy optimization model estimates an immediate value obtained by taking an action a under a state parameter s in the current surgical environment, denoted as R t .

[0115] S542: Define a state value function as V(s, s') = R + γV(s', s') wherein s is a state parameter, and V is a value function parameter; calculate a state value of the state parameter s in the current surgical environment, denoted as V(s, s) = R + γV(s, s') wherein the state value function is realized by a neural network, and a training target of the network is to estimate the value of each state under the current policy, and the calculation of the state value is completed in the value network through forward propagation;

[0116] S543: Calculate a TD error by the following formula: wherein β2 is a preset discount factor;

[0117] S544: Calculate a gradient of the state value function under the state parameter s in the current surgical environment, denoted as

[0118] S545: Update the value function parameter according to the following formula: wherein V is the updated value function parameter, V is the value function parameter before updating under the state parameter s in the current surgical environment, and a2 is a preset learning rate.

[0119] World Models is a model in the field of artificial intelligence, aiming to simulate the way humans and animals naturally learn about the operation of the world through observation and interaction. The core of World Models lies in its ability to estimate the state information of the world that perception does not provide and predict the possible changes in the future state. World Models usually consists of two main components: state representation and transition model. The state representation is used to capture the current state of the environment, while the transition model is used to predict the next state of the environment. In specific implementation, World Models can exist in various forms such as probabilistic models, physical models, and generative models, each with different structures and characteristics. However, the core goal of World Models is to form predictions of future events and states through learning and understanding of historical data. In this embodiment, World Models are constructed in the following ways:

[0120] Various sensing devices are used to detect visual and / or audio and / or temperature information in the surgical environment, and time domain features and / or frequency domain features are extracted therefrom to obtain various surgical environment feature data;

[0121] The importance of various surgical environment features for surgical environment monitoring is evaluated to screen out key surgical environment features, and comprehensive environment feature data is fused to better describe the state of the surgical environment;

[0122] A recurrent neural network model (RNN) is used to construct World Models, and the collected comprehensive environment feature data is used to train the World Models, so that the World Models learn the dynamic changes of the environment based on a self-supervised learning method to predict future environmental state changes.

[0123] Further, the World Models evaluate the prediction performance of the model through cross-validation method to ensure that the model has good generalization ability on different data sets. The indicators of prediction performance include the area under the receiver operating characteristic curve, sensitivity and specificity.

[0124] In a specific application example, the key surgical environment features include intraoperative patient vital signs, and the corresponding future environmental state changes are intraoperative massive hemorrhage risk states;

[0125] The key surgical environment features also include preoperative patient signs, and the corresponding future environmental state changes of the preoperative patient signs and intraoperative patient vital signs are postoperative adverse event risk states.

[0126] The steps of constructing a model for predicting the risk state of intraoperative massive bleeding and the risk state of postoperative adverse events typically include the following: 1. Data collection, including various clinical indicators of patients before, during and after surgery, such as gender, age, preoperative lactate level, whether there are comorbidities, surgical method, operation time, postoperative ICU stay time, total hospitalization time, etc., combined with various data sources, such as electronic medical records, laboratory test results, imaging test results, etc. which can fully reflect the patient's health status; 2. Data preprocessing, such as extracting meaningful features, handling missing values and outliers, and selecting features significantly related to the risk state of intraoperative massive bleeding or the risk state of postoperative adverse events; 3. Traditional statistical models such as Logistic regression model, or machine learning models such as XGBoost, random forest, etc. can be used to train and validate the predictive performance of the statistical model or machine learning model using the collected and preprocessed data.

[0127] The world model that has completed training and passed validation can make real-time predictions: during the operation, the state parameters of the surgical environment are collected in real time and input into the trained world model. The model predicts future state changes based on the state parameters in the current surgical environment, and the output of the world model is used as one of the inputs of the robot control strategy optimization model to realize forward-looking strategy planning.

[0128] In an embodiment of the present application, a surgical robot control system based on a world model and reinforcement learning is provided, as shown in Figure 5 The control system comprises:

[0129] A sensing device configured to collect state parameters in the current surgical environment;

[0130] A world model configured to learn the state parameters in the current surgical environment to predict the state change results during the operation and output predicted environmental state change data;

[0131] A robot control strategy optimization model based on the Actor-Critic algorithm of reinforcement learning, which outputs the action selection strategy of the surgical robot according to the policy parameters and the value function parameters;

[0132] The sensing device and the world model are electrically connected to the input end of the robot control strategy optimization model, and the robot control strategy optimization model analyzes the state parameters in the current surgical environment and the environmental state change data to output the corresponding action selection strategy.

[0133] Specifically, the robot control strategy optimization model comprises a policy parameter optimization module and a value function parameter optimization module, wherein:

[0134] The policy parameter optimization module initializes policy parameters, and the value function parameter optimization module initializes value function parameters;

[0135] The robot control strategy optimization model calculates an advantage function according to the policy parameters and the value function parameters, to determine the pros and cons of taking an alternative action compared to taking an expected action under state parameters in the current surgical environment;

[0136] The policy parameter optimization module updates the policy parameters according to the advantage function using a policy gradient method, wherein the policy gradient method guides the policy parameters to increase the probability of actions that bring higher advantages;

[0137] The value function parameter optimization module estimates the value of the updated policy parameters to the environment state change data, and updates the value function parameters using a minimum TD error method;

[0138] The robot control strategy optimization model iteratively calculates the advantage function and iteratively updates the policy parameters and the value function parameters using the updated policy parameters and the value function parameters until a convergence condition is reached; and determines a corresponding action selection strategy according to the policy parameters and the value function parameters after the convergence condition is reached.

[0139] Further, any of the above technical solutions or a combination of the above technical solutions, the policy parameter optimization module is configured to perform as Figure 3 The flowchart shown:

[0140] Estimate the expected return obtained after taking an alternative action a and following a policy Str under state parameters s in the current surgical environment, denoted as Q Str (s,a);

[0141] Estimate the expected return obtained after following a policy Str under state parameters s in the current surgical environment, denoted as V Str (s), and estimate the expected return obtained after following a policy Str under the next state s' following the environment state change data, denoted as V Str (s');

[0142] The advantage function A Str (s,a) is calculated by the following formula: Str (s,a) - V Str (s) + β1×V Str (s') where β1 is a preset discount factor;

[0143] Compare the calculation results of the advantage functions corresponding to each alternative action a to determine the optimal action a*;

[0144] Define a policy parameter as θ, and define a policy function as Str θ (a|s), calculate the log gradient of the policy parameter on the probability of promoting the optimal action a*, denoted as

[0145] Calculate the policy gradient by calculating the product of the log gradient and the advantage function

[0146] Update the policy parameter according to the following formula: Where θ t+1 is the updated policy parameter, θ t is the policy parameter before updating under the state parameter s in the current surgical environment, and α1 is the preset learning rate.

[0147] Further, any of the technical solutions or combinations of the above technical solutions, the value function parameter optimization module is configured to perform the flow as shown in Figure 4

[0148] Estimate the immediate value obtained by taking action a under state parameter s in the current surgical environment, denoted as R t ;

[0149] Define a state value function as Where S is the state parameter, is the value function parameter; calculate the state value of the state parameter s in the current surgical environment, denoted as And calculate the state value of the next state s' following the environmental state change data, denoted as

[0150] Calculate the TD error by the following formula: Where β2 is a preset discount factor;

[0151] Calculate the gradient of the state value function under the state parameter s in the current surgical environment

[0152] Update the value function parameter according to the following formula: Where, is the updated value function parameter, is the value function parameter before updating under the state parameter s in the current surgical environment, and α2 is the preset learning rate.

[0153] ​The surgical robot control system based on the world model and reinforcement learning provided by the embodiment belongs to the same inventive concept as the surgical robot control method based on the world model and reinforcement learning provided by the above-mentioned embodiment. The content of the surgical robot control method based on the world model and reinforcement learning is incorporated herein by reference in its entirety, and will not be repeated here.

[0154] As shown in Figure 5 The embodiment of the application also provides a surgical robot system, which comprises a surgical robot and a surgical robot control system based on a world model and reinforcement learning as described above, and the surgical robot is configured to execute the action selection strategy output by the surgical robot control system.

[0155] Traditional reinforcement learning algorithms mainly make decisions based on the current state and historical experience, lacking effective prediction ability for future states. This limitation may lead to decision-making errors when dealing with complex surgical environments. By introducing a world model, the application can predict possible future states during surgery, providing more accurate decision-making basis. This prediction ability enables the surgical robot to better cope with complex and unexpected surgical situations, improving the precision and safety of operations.

[0156] Existing reinforcement learning algorithms often require a long training time to adapt to new environments when facing dynamic surgical environments. By combining the prediction results of the world model with reinforcement learning strategy optimization, the surgical robot can quickly adjust strategies in complex environments and optimize operation results. This optimization capability not only improves the adaptability of the system, but also reduces the dependence on a large amount of training data and shortens the adaptation period.

[0157] By combining the prediction results of the world model and the strategy optimization of the reinforcement learning algorithm, the application can provide more accurate operation decisions. In complex surgical environments, this improvement in accuracy can significantly improve the success rate and safety of surgery, thereby meeting the requirements of modern surgery for high-precision operations.

[0158] The present application aims to provide an improved surgical robot control system that integrates a world model and reinforcement learning algorithm, improving the predictive ability and adaptability of the surgical robot to the surgical environment, greatly improving the success rate and safety of surgery, reducing the risks that may be caused by improper operation, and providing more reliable technical support for complex surgery, providing more detailed operation recommendations and adjustment strategies for the robot. For example, through comprehensive modeling of the surgical environment, the world model can identify potential risks or changes in advance, helping the robot make reasonable predictions and adjustments before operation. This integrated solution enables the surgical robot to have higher flexibility and accuracy when facing complex and dynamic surgical environments. The robot can adjust the operation strategy in real time according to the prediction results of the world model, thereby improving the response capability to unexpected events and the execution accuracy of complex tasks. Ultimately, this improvement not only improves the success rate of surgery, but also significantly enhances the safety and effectiveness of the surgical process, providing more reliable technical support for complex surgery.

[0159] It should be noted that, in this document, relational terms such as first and second and the like can be used solely to distinguish one entity or action from another entity or action without necessarily requiring or implying any actual such relationship or order between such entities or actions. Moreover, the terms "comprises", "comprising", or any other variation thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can include other elements not expressly listed or inherent to such process, method, article, or apparatus. Without limitation, an element preceded by "comprises... a" does not, without more constraints, foreclose the existence of additional identical elements in the process, method, article, or apparatus that comprises the element.

[0160] The above description is merely one specific implementation of the application. It should be noted that, for those of ordinary skill in the art, without departing from the principles of the application, a number of improvements and refinements can be made, and these improvements and refinements should also be considered within the scope of protection of the application.

Claims

1. A surgical robot control method based on world model and reinforcement learning, characterized in that: The following steps are involved: Based on the Actor-Critic algorithm of reinforcement learning, a robot control strategy optimization model is established, which is configured to output the action selection strategy of the surgical robot according to the policy parameters and the value function parameters; Using sensor devices to collect status parameters in the current surgical environment; Using the world model to learn state parameters in the current surgical environment to predict state change results during the surgical process, and output predicted environmental state change data; Inputting the state parameters of the current surgical environment collected by the sensing device and the environmental state change data predicted by the world model into the robot control strategy optimization model; The robot control strategy optimization model analyzes state parameters and environmental state change data in the current surgical environment, and outputs a corresponding action selection strategy.

2. The surgical robot control method based on world model and reinforcement learning according to claim 1, characterized in that: The robot control strategy optimization model analyzes the state parameters and environmental state change data of the current surgical environment and outputs a corresponding action selection strategy, including: The robot control strategy optimization model pre-initializes strategy parameters and value function parameters; Calculating an advantage function based on the strategy parameters and the value function parameters to determine the superiority of taking an alternative action compared to taking a desired action under the state parameters in the current surgical environment; Using a policy gradient method to update policy parameters based on the advantage function, wherein the policy gradient method guides the policy parameters to increase the probability of actions that bring higher advantages; estimating the value of the updated policy parameters to the environmental state change data, and updating the value function parameters using a TD error minimization method; Using the updated policy parameters and value function parameters, iteratively calculating the advantage function and iteratively updating the policy parameters and value function parameters until a convergence condition is reached; Determine the corresponding action selection strategy based on the strategy parameters and value function parameters after reaching the convergence conditions.

3. The surgical robot control method based on world model and reinforcement learning according to claim 2, characterized in that: The robot control strategy optimization model updates the strategy parameters in the following way: The robot control strategy optimization model estimates the expected reward obtained by taking the alternative action a and following the strategy Str under the state parameter s in the current surgical environment, denoted as Q Str (s,a); Estimate the expected return after following the strategy Str under the state parameters s in the current surgical environment, denoted as V Str (s), and estimate the expected return obtained by following the strategy Str in the next state s' following the environmental state change data, denoted as V Str (s'); The advantage function A is calculated by the following formula Str (s,a)=Q Str (s,a)-V Str (s)+β1×V Str (s'), where β1 is the preset discount factor; Compare the calculation results of the advantage function corresponding to each alternative action a and determine the optimal action a*; Define the policy parameter as θ and the policy function as Str θ (a|s), calculate the logarithmic gradient of the policy parameters with respect to the probability of improving the optimal action a*, denoted as The policy gradient is calculated by multiplying the logarithmic gradient and the advantage function The policy parameters are updated according to the following formula: Among them, θ t+1 is the updated policy parameter, θ t is the strategy parameter under the state parameter s in the current surgical environment before updating, and α1 is the preset learning rate.

4. The surgical robot control method based on world model and reinforcement learning according to claim 3, characterized in that: The logarithmic gradient Calculated in the policy network through the back-propagation algorithm; The policy gradient is calculated using the following formula: Among them, J(θ) represents the strategy Str θ The performance function, E Strθ Indicates that in the strategy Str θ The expectations below.

5. The surgical robot control method based on world model and reinforcement learning according to claim 2, characterized in that: The robot control strategy optimization model updates the value function parameters in the following way: The robot control strategy optimization model estimates the immediate value of taking action a under the state parameter s in the current surgical environment, denoted as R t ; Define the state value function as Among them, S is the state parameter, is the value function parameter; calculate the state value of the state parameter s in the current surgical environment, recorded as And calculate the state value of the next state s' following the environmental state change data, recorded as The TD error is calculated using the following formula: Among them, β2 is the preset discount factor; Calculate the gradient of the state value function under the state parameter s in the current surgical environment Update the value function parameters according to the following formula: in, is the updated value function parameter, is the value function parameter under the state parameter s in the current surgical environment before updating, and α2 is the preset learning rate.

6. The surgical robot control method based on world model and reinforcement learning according to claim 5, characterized in that: The policy function associated with the policy parameters is implemented by a neural network, and the training objective of the network is to maximize the expected cumulative return; The value function associated with the value function parameters is implemented by a neural network, the training goal of which is to estimate the value of each state under the current policy; The calculation of the state value is completed in the value network through forward propagation.

7. The surgical robot control method based on world model and reinforcement learning according to any one of claims 1 to 6, characterized in that: The world model is constructed in the following way: Utilizing a variety of sensing devices to detect visual and / or audio and / or temperature information in the surgical environment, and extracting time domain features and / or frequency domain features therefrom to obtain a variety of surgical environment feature data; Assess the importance of various surgical environment characteristics to surgical environment monitoring, screen out key surgical environment characteristics, and integrate them to obtain comprehensive environmental characteristic data; A recurrent neural network model is used to construct a world model, and the collected comprehensive environmental feature data is used to train the world model, so that the world model learns the dynamic changes of the environment based on a self-supervised learning method to predict future changes in environmental states.

8. The surgical robot control method based on world model and reinforcement learning according to claim 7, characterized in that: The world model is evaluated for its predictive performance by a cross-validation method, and the predictive performance indicators include the area under the receiver operating characteristic curve, sensitivity, and specificity.

9. The surgical robot control method based on world model and reinforcement learning according to claim 7, characterized in that: The key surgical environment characteristics include intraoperative patient vital signs, and the corresponding future environmental state changes are intraoperative massive bleeding risk states; The key surgical environment characteristics also include the patient's preoperative physical signs, and the future environmental state changes corresponding to the patient's preoperative physical signs and intraoperative patient vital signs are the risk states of postoperative adverse events.

10. A surgical robot control system based on world model and reinforcement learning, characterized in that: include: A sensing device configured to collect status parameters in a current surgical environment; a world model configured to learn state parameters in the current surgical environment to predict state change results during the surgical procedure and output predicted environment state change data; A robot control strategy optimization model, based on the Actor-Critic algorithm of reinforcement learning, outputs the action selection strategy of the surgical robot according to the policy parameters and value function parameters; The sensing device and the world model are both electrically connected to the input end of the robot control strategy optimization model. The robot control strategy optimization model analyzes the state parameters in the current surgical environment and the environmental state change data, and outputs a corresponding action selection strategy.

11. The surgical robot control system based on world model and reinforcement learning according to claim 10, characterized in that: The robot control strategy optimization model includes a strategy parameter optimization module and a value function parameter optimization module, wherein: The strategy parameter optimization module initializes the strategy parameters, and the value function parameter optimization module initializes the value function parameters; The robot control strategy optimization model calculates an advantage function based on the strategy parameters and the value function parameters to determine the superiority of taking an alternative action compared to taking a desired action under the state parameters in the current surgical environment; The policy parameter optimization module updates the policy parameters according to the advantage function using a policy gradient method, wherein the policy gradient method guides the policy parameters to increase the probability of actions that bring higher advantages; The value function parameter optimization module estimates the value of the updated strategy parameters to the environmental state change data, and updates the value function parameters using a TD error minimization method; The robot control strategy optimization model uses the updated strategy parameters and value function parameters to iteratively calculate the advantage function and iteratively update the strategy parameters and value function parameters until the convergence conditions are reached; and determines the corresponding action selection strategy based on the strategy parameters and value function parameters after the convergence conditions are reached.

12. The surgical robot control system based on world model and reinforcement learning according to claim 11, characterized in that: The strategy parameter optimization module is configured to perform: Estimate the expected reward after taking alternative action a and following strategy Str under the state parameter s in the current surgical environment, denoted as Q Str (s,a); Estimate the expected return after following the strategy Str under the state parameters s in the current surgical environment, denoted as V Str (s), and estimate the expected return obtained by following the strategy Str in the next state s' following the environmental state change data, denoted as V Str (s'); The advantage function A is calculated by the following formula Str (s,a)=Q Str (s,a)-V Str (s)+β1×V Str (s'), where β1 is the preset discount factor; Compare the calculation results of the advantage function corresponding to each alternative action a and determine the optimal action a*; Define the policy parameter as θ and the policy function as Str θ (a|s), calculate the logarithmic gradient of the policy parameters with respect to the probability of improving the optimal action a*, denoted as The policy gradient is calculated by multiplying the logarithmic gradient and the advantage function The policy parameters are updated according to the following formula: Among them, θ t+1 is the updated policy parameter, θ t is the strategy parameter under the state parameter s in the current surgical environment before updating, and α1 is the preset learning rate.

13. The surgical robot control system based on world model and reinforcement learning according to claim 11, characterized in that: The cost function parameter optimization module is configured to perform: Estimate the immediate value of taking action a under the state parameter s in the current surgical environment, denoted as R t ; Define the state value function as Among them, S is the state parameter, is the value function parameter; calculate the state value of the state parameter s in the current surgical environment, recorded as And calculate the state value of the next state s' following the environmental state change data, recorded as The TD error is calculated using the following formula: Among them, β2 is the preset discount factor; Calculate the gradient of the state value function under the state parameter s in the current surgical environment Update the value function parameters according to the following formula: in, is the updated value function parameter, is the value function parameter under the state parameter s in the current surgical environment before updating, and α2 is the preset learning rate.

14. A surgical robot system, characterized in that: The invention comprises a surgical robot and a surgical robot control system based on a world model and reinforcement learning as claimed in any one of claims 10 to 13, wherein the surgical robot is configured to execute an action selection strategy output by the surgical robot control system.

Citation Information

Patent Citations

  • Robot motion control method and system and electronic equipment

    CN117506889A

  • Mechanical arm control method based on selective state space and model reinforcement learning

    CN118721208A

  • Automatic driving decision-making method and system based on generative world large model and multi-step reinforcement learning

    CN118790287A

  • AGV path planning method and device based on world model hidden variables and reinforcement learning

    CN118839831A

  • Robot control policy

    US20250100135A1