Intelligent surgical robot control method based on reinforcement learning
By combining medical image processing and dynamic graph convolutional neural networks with the lightning search algorithm to optimize the control method of surgical robots, the problem of autonomous decision-making and strategy optimization of surgical robots in complex environments was solved, and high-precision and stable surgical operations were achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- LITTLE BUTLER (SUZHOU) HEALTH TECHNOLOGY CO LTD
- Filing Date
- 2026-03-17
- Publication Date
- 2026-05-12
AI Technical Summary
Existing surgical robot systems lack deep understanding and autonomous decision-making capabilities in complex surgical environments, making it difficult to adapt to individual patient differences and changes in tissue physiological state. Traditional reinforcement learning methods suffer from local optima problems in high-dimensional continuous control spaces, and the simulation environment lacks realism and dynamic response mechanisms, resulting in poor policy transfer performance.
By combining medical image processing, dynamic graph convolutional neural networks, and lightning search algorithm, a dynamic graph structure is constructed to represent the intraoperative state. A continuous control strategy is generated through a reinforcement learning framework, and the strategy is trained in a high-fidelity simulation environment. Multidimensional performance evaluation and iterative optimization are introduced.
It enhances the perception and decision-making capabilities of surgical robots, enabling autonomous and high-precision operation in complex scenarios, strengthening the adaptability and transferability of strategies, and ensuring the stability and efficiency of control.
Smart Images

Figure CN122005101A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent surgical control technology, and in particular to an intelligent surgical robot control method based on reinforcement learning. Background Technology
[0002] With the rapid development of modern medical and artificial intelligence technologies, surgical robot systems have gradually become important tools for high-precision surgical procedures. Compared to traditional surgical methods, surgical robots offer advantages such as strong operational stability, high control precision, and good repeatability, and are widely used in many high-risk fields such as urology, cardiothoracic surgery, and neurosurgery. Existing surgical robots primarily rely on remote operation by surgeons, using robotic arms to perform tasks such as anatomical structure recognition, instrument guidance, tissue cutting, and suturing, greatly improving surgical safety and success rates. However, in practical applications, robot operation is still mainly based on human-computer interaction control, with limited intelligence and a lack of deep understanding of the surgical scenario and autonomous decision-making capabilities, thus restricting their adaptability and operational flexibility in complex environments.
[0003] In existing technologies, medical image processing has become one of the main means for surgical robot systems to perceive the external environment. Modeling and analyzing preoperative and intraoperative tissue structures using MRI, CT, and ultrasound medical imaging technologies can provide crucial information for surgical path planning and operational decisions. However, most existing methods rely solely on static images or regular templates for 3D reconstruction and anatomical structure recognition, making it difficult to fully reflect tissue deformation, instrument interference, and real-time changes during surgery. Furthermore, traditional image processing methods primarily depend on convolutional neural networks for feature extraction, failing to effectively capture the spatial topological relationships and temporal series features between tissues, thus limiting their performance in dynamic and complex surgical scenarios.
[0004] Meanwhile, existing robot control strategies mostly rely on preset path planning or motion prediction based on supervised learning training. While these methods are effective to some extent in fixed procedures and standard scenarios, in actual surgical operations, factors such as individual patient differences, changes in tissue physiological states, and sudden disturbances lead to a high degree of uncertainty in the operating environment. Due to the lack of real-time modeling of the environmental state and adaptive update mechanisms for the strategy, existing control methods struggle to meet the dual demands of precise operation and flexible response in complex surgical tasks, easily leading to decreased control accuracy, unstable movements, or even operational failure.
[0005] In recent years, reinforcement learning, as an important technology for solving decision-making and action selection problems of intelligent agents, has achieved extensive results in fields such as autonomous driving, robot navigation, and intelligent game theory, and is gradually being introduced into the field of medical robot control. Through the reinforcement learning framework, intelligent agents can continuously adjust their strategies to obtain higher rewards during interaction with the environment, thereby achieving autonomous optimization of complex tasks. However, the application of traditional reinforcement learning methods in medical surgical scenarios still faces several challenges, such as high-dimensional state spaces, sparse feedback signals, difficulties in obtaining training samples, and weak policy transfer capabilities. Especially in high-risk, high-precision surgical environments, how to construct stable, interpretable, and real-time feedback-capable reinforcement learning control models remains a bottleneck in current research.
[0006] Furthermore, existing reinforcement learning methods mostly model discrete states and discrete action spaces, making them difficult to adapt to the continuous control requirements of surgical procedures. For example, surgical robots need to control actuators with multiple degrees of freedom simultaneously, including position adjustment, clamping force, and angle control. The control parameters of each dimension need to be finely adjusted in a continuous space. Traditional methods are prone to getting trapped in local optima in high-dimensional continuous action spaces, making it difficult to achieve global policy optimization. At the same time, the stability and efficiency of the training process are also difficult to guarantee.
[0007] In policy optimization, existing methods largely rely on traditional optimization algorithms such as gradient descent, including stochastic gradient descent and Adam. These methods are widely used in deep reinforcement learning training, but they often suffer from slow convergence, getting trapped in local optima, and sensitivity to initial values when dealing with complex conditions such as high-dimensional parameter spaces, non-convex objective functions, or dynamic environmental perturbations. Therefore, introducing more efficient optimization algorithms with greater global search capabilities to improve the global convergence and generalization ability of policy parameter training has become a critical technical problem that urgently needs to be solved.
[0008] Furthermore, the construction of the simulation environment is a significant factor limiting the development of existing technologies when training surgical control strategy models. Currently used surgical simulation environments often lack sufficient realism and dynamic response mechanisms, failing to effectively simulate intraoperative instrument-tissue interactions, tissue deformation, physiological feedback, and sudden disturbances. This results in poor transfer performance of strategy models trained in simulations in real-world environments, leading to a high risk of strategy failure. Therefore, there is an urgent need to construct a high-fidelity surgical simulation environment that supports multi-source data stream input and possesses physiological feedback and disturbance injection capabilities, providing a more realistic training platform for reinforcement learning strategies.
[0009] Therefore, how to provide a reinforcement learning-based intelligent surgical robot control method is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0010] One objective of this invention is to propose a reinforcement learning-based intelligent surgical robot control method. This invention combines medical image processing, dynamic graph convolutional neural network models, and lightning search algorithm techniques. By constructing a dynamic graph structure to represent the intraoperative state, it generates continuous control strategies, thereby improving the surgical robot's perception and decision-making capabilities. The system trains the strategy in a high-fidelity simulation environment and iteratively optimizes the control model based on multi-dimensional performance evaluation results, achieving autonomous and high-precision surgical operations in complex scenarios.
[0011] A method for controlling an intelligent surgical robot based on reinforcement learning according to an embodiment of the present invention includes the following steps: S1. Acquire medical image data of the area to be operated on and preprocess it to generate a standardized medical image; S2. Construct a dynamic graph convolutional neural network model based on standardized medical images to model the intraoperative target region as a dynamic graph structure; S3. Construct a control framework based on reinforcement learning, using dynamic graph structure features as state input, establish a surgical neural network model and output continuous control actions; S4. The lightning search algorithm is used to optimize the control parameters of the surgical neural network model. An electric field induction unit and a guiding movement path are configured for each individual lightning body. Each lightning body represents a set of control strategy parameter configurations. The performance score of each set of parameter configurations during the strategy training process is calculated through a centralized evaluation mechanism to generate the optimal strategy parameter set. S5. Perform multiple rounds of strategy training in the constructed surgical simulation environment, optimize the parameter configuration of each group of the surgical neural network model using the optimal strategy parameter set, and generate the optimal neural network model. S6. Deploy the optimal surgical neural network model to the intelligent surgical robot system, continuously collect intraoperative sensor data, and update the dynamic graph structure. S7. After the surgery is completed, record the robot's execution trajectory, tissue deformation data and operation parameters, and feed the performance evaluation results back to the reinforcement learning-based control framework for iteration.
[0012] Optionally, the medical image data specifically includes preoperative MRI images, CT scan images, ultrasound images, and endoscopic images.
[0013] Optionally, the control strategy parameter configuration specifically includes the initial weights of the strategy network, the learning rate, and the discount factor.
[0014] Optionally, S2 specifically includes: S21. Input standardized medical images into an image segmentation network to automatically identify key anatomical structures during surgery and label the structure categories and boundary contours. S22. Based on the positional relationships and spatial relative distances of anatomical structures in standardized medical images, construct an initial graph structure containing nodes and edges, where nodes represent anatomical structural units and edges represent the connection relationships between structures. S23. Obtain the predefined task phase sequence of the intraoperative operation process, encode the current surgical phase as a phase label, and embed the phase label into the additional attribute vector of each node. S24. Set the temporal sliding window length and update step size of the graph structure. Extract a snapshot of the graph structure at a fixed time step during the operation. Generate a time frame sequence by changing the node state and edge weight between the snapshots to construct an initial dynamic graph structure with temporal dependence. S25. Apply the initial dynamic graph structure to the attention-enhanced convolutional processing unit, dynamically adjust the edge weights and perform graph convolution operations based on node feature similarity, stage label weights and structural positional relationships. S26. Output the dynamic graph structure after attention-enhanced convolution; Optionally, S3 specifically includes: S31. Extract dynamic graph structure information, including node features, edge weight information, stage attributes, and time frame sequence; S32. The extracted dynamic graph structure information is used as the state input and input to the state encoding unit to uniformly encode the topological relationship and temporal characteristics between nodes, and generate a state representation for decision-making. Let the state representation be a state vector. S33. Construct an action generation unit based on a neural network. The action generation unit receives a state vector as input and establishes a surgical neural network model consisting of an input layer, multiple nonlinear computation layers and an output layer. S34. In the output layer, a continuous control action command vector is constructed based on the state vector and the internal weight matrix: ; in, Represents the continuous control action vector. This represents the candidate vector of actions to be optimized. Represents the action candidate space, Indicates the total number of time steps. Indicates the first The state vector at each time step Indicates the first The output layer weight matrix corresponding to each time step Indicates the first Time step attention weight tensor The linear transformation matrix that maps action vectors to the state-target space. Represents element-wise tensor product. This represents the square of the Euclidean distance. Regularization coefficient, Let n denote the third norm of the j-th action vector, n denote the dimension of the action vector, and j denote the control action vector. The index of the j-th component in the data. Indicates a time step. This represents a small positive constant used to prevent the denominator from being zero; S35. Construct a value evaluation unit to evaluate the expected reward value of the control action of the action generation unit under the state vector, and predict the reward trend in future states. S36. The action generation unit and the value evaluation unit are combined to form a reinforcement learning control framework, which receives performance feedback indicators and updates and optimizes the neural network structure and parameters.
[0015] Optionally, S4 specifically includes: S41. Construct an initial electric potential field mapping structure, randomly distribute several individual lightning bodies in the parameter space, and each individual lightning body corresponds to a set of control parameters of the surgical neural network model; S42. Configure an electric field induction structure unit for each lightning body, calculate the local electric field intensity vector based on the parameter distance and potential difference between adjacent lightning bodies, construct a voltage gradient guiding tensor, and record the directional migration tendency of each lightning body. S43. Introduce a perturbation-reconstruction mechanism, apply a small-range random perturbation to the current position of each lightning element, and construct a perturbation feedback map by combining it with the historical potential trajectory, and dynamically adjust the next movement path; S44. Use the control parameter set corresponding to each lightning body to initialize the surgical neural network model, and run a fixed number of policy training sessions in the same simulation training environment. Record the performance indicators of each set of parameters in four dimensions: control accuracy, motion smoothness, convergence rate, and tissue feedback response. S45. Construct a centralized evaluation mechanism, jointly inputting performance indicators and control action vectors into the evaluation convergence unit to calculate the comprehensive performance score of each group of individual lightning bodies: ; in, This represents the overall performance score corresponding to the control parameters. This represents the i-th component in the continuous control action vector. This represents the dimension of the control action vector. Indicates the total number of time steps. This represents the control stability weight of the i-th control variable at time step t. This represents the target action reference value of the i-th control variable of the lightning body at the t-th time step, where i represents the i-th dimension in the control action vector, j represents the j-th component in the control action vector, and t represents the scoring time step index. S46. Mark the set of control parameters corresponding to the lightning body with the lowest overall performance score as the optimal strategy parameter set.
[0016] Optionally, S5 specifically includes: S51. Construct a surgical simulation environment with 3D modeling based on MRI data, and integrate a real-time physiological feedback simulation unit and an interference injection unit in the environment; S52. Load the optimal strategy parameter set into the surgical neural network model, initialize the weight tensor, bias tensor and time step decision weight mapping table, and set up an environmental disturbance recording unit to track the impact of environmental fluctuations on the control strategy during training. S53. Perform a multi-batch reinforcement training process in a surgical simulation environment, collect the continuous state transition sequence, environmental disturbance response trajectory and control action output trajectory generated by the robot's actions, and construct a sequence batch processing tensor in the form of timestamps; S54. Based on the comprehensive performance score, dynamically relabel the weights of the training sample batches and perform cross-time step residual alignment update operations on the sequence batch processing tensors: ; in, This represents the trainable parameter tensor of the surgical neural network model in the current round. This represents the surgical neural network parameter tensor obtained after parameter updates. This represents the learning rate coefficient. Indicates the total number of time steps. Let represent the dimension of the control action vector, where i represents the i-th dimension in the control action vector. This represents the overall performance score corresponding to the control parameters. Indicates the first The actual output value of the i-th control action in each time step. Indicates a time step. Indicates the first The target reference value for the i-th control action in -1 time steps. Let i represent the i-th component of the dynamic graph structure state vector at time step t. No. The i-th component of the dynamic graph structure state vector in time step -1. This represents the overall performance score corresponding to the control parameters. represents a small positive constant used to prevent the denominator from being zero, and i represents the i-th dimension in the control action vector; S55. During the training process, a parameter sparsity self-adjustment mechanism is introduced to perform round-by-round gating compression on the low-amplitude units in the updated surgical neural network parameter tensor. S56. Based on three dimensions—action performance stability, model structure sparsity, and disturbance response sensitivity—a final evaluation index system is constructed. The models are then ranked according to the evaluation index to output the optimal neural network model.
[0017] Optionally, S6 specifically includes: S61. Deploy the optimal neural network model into the control and decision-making unit of the intelligent surgical robot; S62. Load a dynamic graph convolutional neural network structure and a state perception unit into the surgical robot, and establish a data input-output connection with the control decision unit. S63. Initialize the robot end effector position, joint state, graph structure buffer, and environmental sensor interface before the operation begins; S64. Collect real-time visual images, force feedback, instrument touch position and tissue deformation data generated during the operation, and input them to the status sensing unit at a fixed frequency to update the dynamic graph structure.
[0018] Optionally, S64 specifically includes: S641. The state-aware unit is used to convert the collected data into a graph structure input, including constructing nodes, edges, spatial connection relationships and stage attributes, and generating a dynamic graph state representation at each time step; S642. Input the dynamic graph structure of the current time step into the control decision unit as the state vector input of the surgical neural network model; S643. The surgical neural network model outputs the continuous control action command vector for the current time step, including multi-degree-of-freedom action values and corresponding control amplitudes. S644: Transmit the control motion command vector to the robot's underlying actuator interface to drive the end effector to perform operations, including position movement, attitude adjustment and clamping force control; S645: Receive actuator feedback signals and sensor monitoring values, and record the current status and action result information; S646. Repeat steps S641 to S642 throughout the entire surgical procedure to form a real-time state-action loop, thereby realizing an intelligent control execution process driven by a dynamic graph structure.
[0019] Optionally, S7 specifically includes: S71. After the surgery is completed, extract the complete execution trajectory from the robot, including the displacement path of the end effector, the change of attitude angle and the time stamp sequence; S72. Extract tissue deformation data collected during surgery, including the degree of tissue deformation, instrument contact point trajectory, and force feedback records; S73. Record the operation parameters during the surgical procedure, including the control action amplitude sequence, response delay, interference error value and state transition frequency, and construct a multi-dimensional operation dataset; S74. Perform unified format processing on all recorded data in steps S71 to S73 to construct a performance evaluation data matrix. S75. Set up a multi-index evaluation system that includes control accuracy, tissue protection, energy consumption level and motion stability, and perform weighted aggregation on the performance evaluation data matrix to generate performance evaluation results. S76. Feed the performance evaluation results back to the reinforcement learning-based control framework for iteration.
[0020] The beneficial effects of this invention are: This invention constructs a highly intelligent and dynamically adaptive surgical robot control method by integrating medical image processing, dynamic graph convolutional neural network modeling, reinforcement learning control strategies, and lightning search optimization algorithms. This significantly improves the robot's perception, decision-making, and operational accuracy in complex surgical scenarios. Through standardized processing and anatomical structure segmentation of preoperative multi-source medical image data such as MRI, CT, and ultrasound, it effectively achieves automatic identification and graph structure modeling of key target areas during surgery. Combined with the dynamic graph modeling mechanism, it can dynamically capture tissue structure, operational stages, and temporal changes, enabling continuous updating and accurate perception of the surgical environment state, providing stable and expressive state input for the control strategy.
[0021] The reinforcement learning-based control framework proposed in this invention uses dynamic graph structure features as state input. It not only possesses end-to-end autonomous decision-making capabilities but also continuously optimizes control strategies during training, thereby adapting to complex situations such as tissue deformation, mechanical interference, and visual blur. Compared to traditional preset path or rule-based control strategies, this invention achieves more flexible action generation through the self-learning process of the policy network, supporting continuous control of multi-degree-of-freedom actuators and effectively improving the smoothness and execution accuracy of robot movements.
[0022] To address the challenge of optimizing high-dimensional policy parameter spaces, this invention introduces the Lightning Search algorithm. By constructing a simulated electric potential field and guiding paths, it guides the collective evolution optimization of policy parameters, achieving faster convergence speed and higher parameter performance during policy training. This avoids getting trapped in local optima and improves the policy's generalization ability and robustness. Simultaneously, the integration of real-time physiological feedback simulation and interference injection mechanisms within a high-fidelity surgical simulation environment makes the training process more closely resemble the real intraoperative environment, enhancing the model's adaptability and transferability to real-world scenarios.
[0023] This invention realizes a complete closed-loop process from preoperative image modeling, intraoperative dynamic state perception, strategy generation and optimization, to postoperative behavior evaluation and iteration, comprehensively improving the autonomous control intelligence level of surgical robots. During task execution, the robot can flexibly adjust its control actions according to the current state, and after surgery, it performs performance evaluation and strategy iteration based on historical execution data, providing continuous optimization intelligent support for subsequent surgeries. It has broad clinical application prospects and promotional value. Attached Figure Description
[0024] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings: Figure 1 This is a flowchart of a reinforcement learning-based intelligent surgical robot control method proposed in this invention; Figure 2 This is a schematic diagram of a reinforcement learning-based intelligent surgical robot control method proposed in this invention; Figure 3 This is a data flow diagram of a reinforcement learning-based intelligent surgical robot control method proposed in this invention. Detailed Implementation
[0025] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.
[0026] refer to Figure 1-3 A method for controlling an intelligent surgical robot based on reinforcement learning includes the following steps: S1. Acquire medical image data of the area to be operated on and preprocess it to generate a standardized medical image; S2. Construct a dynamic graph convolutional neural network model based on standardized medical images to model the intraoperative target region as a dynamic graph structure; S3. Construct a control framework based on reinforcement learning, using dynamic graph structure features as state input, establish a surgical neural network model and output continuous control actions; S4. The lightning search algorithm is used to optimize the control parameters of the surgical neural network model. An electric field induction unit and a guiding movement path are configured for each individual lightning body. Each lightning body represents a set of control strategy parameter configurations. The performance score of each set of parameter configurations during the strategy training process is calculated through a centralized evaluation mechanism to generate the optimal strategy parameter set. S5. Perform multiple rounds of strategy training in the constructed surgical simulation environment, optimize the parameter configuration of each group of the surgical neural network model using the optimal strategy parameter set, and generate the optimal neural network model. S6. Deploy the optimal surgical neural network model to the intelligent surgical robot system, continuously collect intraoperative sensor data, and update the dynamic graph structure. S7. After the surgery is completed, record the robot's execution trajectory, tissue deformation data and operation parameters, and feed the performance evaluation results back to the reinforcement learning-based control framework for iteration.
[0027] This invention employs a lightning search algorithm to construct an electric potential field guidance mechanism, performing a global search within the policy parameter space to optimize the key parameter configuration of the reinforcement learning control network. By constructing a dynamic graph structure to represent the surgical state and combining it with attention map convolution to extract high-dimensional features, continuous control action generation and real-time policy updates are achieved. The system completes multiple rounds of training iterations in a simulation environment, effectively improving the control accuracy, response speed, and interference robustness of the surgical robot, ensuring high-precision, autonomous intelligent control in complex intraoperative environments.
[0028] In this embodiment, the medical image data specifically includes preoperative MRI images, CT scan images, ultrasound images, and endoscopic images.
[0029] This invention integrates preoperative MRI, CT, ultrasound, and endoscopic images to construct a multimodal medical image input system, improving the completeness and accuracy of anatomical structure recognition. Through cross-modal image fusion and standardized processing, the system can comprehensively capture the spatial relationships and tissue features of the surgical area, achieving accurate modeling and dynamic graph construction of key structures. This provides high-quality state input for subsequent control strategy generation, significantly enhancing the system's adaptability and robustness to complex intraoperative scenarios.
[0030] In this embodiment, the control strategy parameter configuration specifically includes the initial weights of the strategy network, the learning rate, and the discount factor.
[0031] This invention introduces three core parameters—initial weights, learning rate, and discount factor—into the control strategy parameter configuration. By constructing an optimization space through multi-dimensional combinations, it achieves refined control over the reinforcement learning training process. The initial weights influence the starting point of the policy search, the learning rate controls the magnitude of parameter updates, and the discount factor balances short-term and long-term returns. By dynamically adjusting the parameter configuration, the system improves the convergence speed and stability of policy training, ensuring that the control model possesses stronger generalization ability and response robustness in complex surgical environments.
[0032] In this embodiment, S2 specifically includes: S21. Input standardized medical images into an image segmentation network to automatically identify key anatomical structures during surgery and label the structure categories and boundary contours. S22. Based on the positional relationships and spatial relative distances of anatomical structures in standardized medical images, construct an initial graph structure containing nodes and edges, where nodes represent anatomical structural units and edges represent the connection relationships between structures. S23. Obtain the predefined task phase sequence of the intraoperative operation process, encode the current surgical phase as a phase label, and embed the phase label into the additional attribute vector of each node. S24. Set the temporal sliding window length and update step size of the graph structure. Extract a snapshot of the graph structure at a fixed time step during the operation. Generate a time frame sequence by changing the node state and edge weight between the snapshots to construct an initial dynamic graph structure with temporal dependence. S25. Apply the initial dynamic graph structure to the attention-enhanced convolutional processing unit, dynamically adjust the edge weights and perform graph convolution operations based on node feature similarity, stage label weights and structural positional relationships. S26. Output the dynamic graph structure after attention-enhanced convolution; This invention utilizes a standardized medical image-driven dynamic graph modeling mechanism, combining image segmentation, structural topology construction, and surgical stage embedding to achieve accurate modeling of key intraoperative anatomical structures. A sliding window is employed to generate a temporal graph structure, combined with attention-enhanced convolution to dynamically adjust the edge weights between nodes, improving the ability to model interactions between structures. This method effectively captures the temporal changes and state evolution of anatomical structures, providing high-dimensional, semantically rich input features for control strategies, and enhancing the system's understanding and response to complex intraoperative situations.
[0033] In this embodiment, S3 specifically includes: S31. Extract dynamic graph structure information, including node features, edge weight information, stage attributes, and time frame sequence; S32. The extracted dynamic graph structure information is used as the state input and input to the state encoding unit to uniformly encode the topological relationship and temporal characteristics between nodes, and generate a state representation for decision-making. Let the state representation be a state vector. S33. Construct an action generation unit based on a neural network. The action generation unit receives a state vector as input and establishes a surgical neural network model consisting of an input layer, multiple nonlinear computation layers and an output layer. S34. In the output layer, a continuous control action command vector is constructed based on the state vector and the internal weight matrix: ; in, Represents the continuous control action vector. This represents the candidate vector of actions to be optimized. Represents the action candidate space, Indicates the total number of time steps. Indicates the first The state vector at each time step Indicates the first The output layer weight matrix corresponding to each time step Indicates the first Time step attention weight tensor The linear transformation matrix that maps action vectors to the state-target space. Represents element-wise tensor product. This represents the square of the Euclidean distance. Regularization coefficient, Let n denote the third norm of the j-th action vector, n denote the dimension of the action vector, and j denote the control action vector. The index of the j-th component in the data. Indicates a time step. This represents a small positive constant used to prevent the denominator from being zero; S35. Construct a value evaluation unit to evaluate the expected reward value of the control action of the action generation unit under the state vector, and predict the reward trend in future states. S36. The action generation unit and the value evaluation unit are combined to form a reinforcement learning control framework, which receives performance feedback indicators and updates and optimizes the neural network structure and parameters.
[0034] This invention extracts node features, edge weights, and time frame sequences from a dynamic graph structure, uniformly encodes topological and temporal information, constructs a reinforcement learning control framework, and generates continuous action commands. The action generation unit and value evaluation module are co-optimized, combining attention mechanisms and high-dimensional state representation to achieve accurate perception and action prediction of complex states in surgical scenarios. This method improves the expressiveness and adaptability of control strategies in a multivariable continuous space, significantly enhancing the system's responsiveness to dynamic environmental changes and operational stability.
[0035] In this embodiment, S4 specifically includes: S41. Construct an initial electric potential field mapping structure, randomly distribute several individual lightning bodies in the parameter space, and each individual lightning body corresponds to a set of control parameters of the surgical neural network model; S42. Configure an electric field induction structure unit for each lightning body, calculate the local electric field intensity vector based on the parameter distance and potential difference between adjacent lightning bodies, construct a voltage gradient guiding tensor, and record the directional migration tendency of each lightning body. S43. Introduce a perturbation-reconstruction mechanism, apply a small-range random perturbation to the current position of each lightning element, and construct a perturbation feedback map by combining it with the historical potential trajectory, and dynamically adjust the next movement path; S44. Use the control parameter set corresponding to each lightning body to initialize the surgical neural network model, and run a fixed number of policy training sessions in the same simulation training environment. Record the performance indicators of each set of parameters in four dimensions: control accuracy, motion smoothness, convergence rate, and tissue feedback response. S45. Construct a centralized evaluation mechanism, jointly inputting performance indicators and control action vectors into the evaluation convergence unit to calculate the comprehensive performance score of each group of individual lightning bodies: ; in, This represents the overall performance score corresponding to the control parameters. This represents the i-th component in the continuous control action vector. This represents the dimension of the control action vector. Indicates the total number of time steps. This represents the control stability weight of the i-th control variable at time step t. This represents the target action reference value of the i-th control variable of the lightning body at the t-th time step, where i represents the i-th dimension in the control action vector, j represents the j-th component in the control action vector, and t represents the scoring time step index. S46. Mark the set of control parameters corresponding to the lightning body with the lowest overall performance score as the optimal strategy parameter set.
[0036] This invention introduces a lightning search algorithm to construct an electric potential field mapping structure. In the parameter space, it achieves global exploration and targeted optimization of control strategy parameters through simulated electric field induction and voltage gradient guidance mechanisms. Combining a perturbation-reconstruction mechanism with historical potential trajectories, it dynamically adjusts individual migration paths, improving strategy diversity and global convergence. The system integrates multi-dimensional performance indicators such as control accuracy and motion smoothness through a centralized evaluation mechanism to select the optimal set of strategy parameters, significantly improving the training efficiency and control performance of the surgical neural network model in high-dimensional parameter spaces.
[0037] In this embodiment, S5 specifically includes: S51. Construct a surgical simulation environment with 3D modeling based on MRI data, and integrate a real-time physiological feedback simulation unit and an interference injection unit in the environment; S52. Load the optimal strategy parameter set into the surgical neural network model, initialize the weight tensor, bias tensor and time step decision weight mapping table, and set up an environmental disturbance recording unit to track the impact of environmental fluctuations on the control strategy during training. S53. Perform a multi-batch reinforcement training process in a surgical simulation environment, collect the continuous state transition sequence, environmental disturbance response trajectory and control action output trajectory generated by the robot's actions, and construct a sequence batch processing tensor in the form of timestamps; S54. Based on the comprehensive performance score, dynamically relabel the weights of the training sample batches and perform cross-time step residual alignment update operations on the sequence batch processing tensors: ; in, This represents the trainable parameter tensor of the surgical neural network model in the current round. This represents the surgical neural network parameter tensor obtained after parameter updates. This represents the learning rate coefficient. Indicates the total number of time steps. Let represent the dimension of the control action vector, where i represents the i-th dimension in the control action vector. This represents the overall performance score corresponding to the control parameters. Indicates the first The actual output value of the i-th control action in each time step. Indicates a time step. Indicates the first The target reference value for the i-th control action in -1 time steps. Let i represent the i-th component of the dynamic graph structure state vector at time step t. No. The i-th component of the dynamic graph structure state vector in time step -1. This represents the overall performance score corresponding to the control parameters. represents a small positive constant used to prevent the denominator from being zero, and i represents the i-th dimension in the control action vector; S55. During the training process, a parameter sparsity self-adjustment mechanism is introduced to perform round-by-round gating compression on the low-amplitude units in the updated surgical neural network parameter tensor. S56. Based on three dimensions—action performance stability, model structure sparsity, and disturbance response sensitivity—a final evaluation index system is constructed. The models are then ranked according to the evaluation index to output the optimal neural network model.
[0038] This invention constructs a surgical simulation environment integrating MRI 3D modeling, real-time physiological feedback, and interference injection mechanisms to simulate real intraoperative conditions and conduct multi-batch reinforcement training. Through a dynamic relabeling weight mechanism and residual alignment update strategy, the model's sensitivity to state transitions and training stability are improved. A parameter sparsity self-adjustment mechanism is introduced to compress invalid connections, improving model inference efficiency. The system selects the optimal model based on three-dimensional indicators: action stability, structural sparsity, and interference response, achieving high-performance transfer and execution of strategies in complex environments.
[0039] In this embodiment, S6 specifically includes: S61. Deploy the optimal neural network model into the control and decision-making unit of the intelligent surgical robot; S62. Load a dynamic graph convolutional neural network structure and a state perception unit into the surgical robot, and establish a data input-output connection with the control decision unit. S63. Initialize the robot end effector position, joint state, graph structure buffer, and environmental sensor interface before the operation begins; S64. Collect real-time visual images, force feedback, instrument touch position and tissue deformation data generated during the operation, and input them to the status sensing unit at a fixed frequency to update the dynamic graph structure.
[0040] This invention deploys an optimal neural network model to the control unit of an intelligent surgical robot, loading a dynamic graph convolutional structure and a state perception module to achieve real-time perception and decision-making linkage during surgery. The system ensures consistency of the robot's state before operation by initializing the actuator state and sensor interfaces; during surgery, it continuously collects visual, force, and tissue deformation data, constructs a dynamic graph structure, and updates it in real time, providing high-precision state input to the control model, thereby improving the robot's operational stability, response speed, and adaptability.
[0041] In this embodiment, S64 specifically includes: S641. The state-aware unit is used to convert the collected data into a graph structure input, including constructing nodes, edges, spatial connection relationships and stage attributes, and generating a dynamic graph state representation at each time step; S642. Input the dynamic graph structure of the current time step into the control decision unit as the state vector input of the surgical neural network model; S643. The surgical neural network model outputs the continuous control action command vector for the current time step, including multi-degree-of-freedom action values and corresponding control amplitudes. S644: Transmit the control motion command vector to the robot's underlying actuator interface to drive the end effector to perform operations, including position movement, attitude adjustment and clamping force control; S645: Receive actuator feedback signals and sensor monitoring values, and record the current status and action result information; S646. Repeat steps S641 to S642 throughout the entire surgical procedure to form a real-time state-action loop, thereby realizing an intelligent control execution process driven by a dynamic graph structure.
[0042] This invention converts multi-source intraoperative data into a dynamic graph structure input through a state-aware unit, constructing node, edge, and stage attributes in real time and generating a high-dimensional state representation. The control model receives the dynamic graph state vector and outputs continuous control actions to drive the robot actuators to complete multi-degree-of-freedom operations. During execution, the system continuously senses feedback, cyclically updating the state and actions to achieve real-time state-action linkage control. This method significantly improves the accuracy and response speed of robot operations, ensuring efficient, autonomous, and dynamically adjustable capabilities in complex intraoperative environments.
[0043] In this embodiment, S7 specifically includes: S71. After the surgery is completed, extract the complete execution trajectory from the robot, including the displacement path of the end effector, the change of attitude angle and the time stamp sequence; S72. Extract tissue deformation data collected during surgery, including the degree of tissue deformation, instrument contact point trajectory, and force feedback records; S73. Record the operation parameters during the surgical procedure, including the control action amplitude sequence, response delay, interference error value and state transition frequency, and construct a multi-dimensional operation dataset; S74. Perform unified format processing on all recorded data in steps S71 to S73 to construct a performance evaluation data matrix. S75. Set up a multi-index evaluation system that includes control accuracy, tissue protection, energy consumption level and motion stability, and perform weighted aggregation on the performance evaluation data matrix to generate performance evaluation results. S76. Feed the performance evaluation results back to the reinforcement learning-based control framework for iteration.
[0044] This invention constructs a multi-dimensional operational dataset by extracting the execution trajectory, tissue deformation data, and operational parameters after surgery, and generates a performance evaluation data matrix using a unified format. The system sets comprehensive evaluation indicators including control accuracy, tissue protection, energy consumption, and motion stability to quantitatively analyze the surgical process. The evaluation results are fed back to a reinforcement learning control framework, driving continuous iterative optimization of the policy model. This method effectively improves the closed-loop efficiency of model training and enhances the robot's long-term adaptability to complex surgical scenarios and its control strategy evolution capabilities.
[0045] Example 1: To verify the feasibility of this invention in practice, it was applied to a robot-assisted surgical procedure for a brain tumor resection performed in the neurosurgery department of a tertiary general hospital. This type of surgery is complex, involving precise localization of deep intracranial tumors, path planning, soft tissue puncture, and lesion resection. It places extremely high demands on the operator's spatial awareness, operational precision, and real-time feedback response capabilities. Traditional surgical methods rely on preoperative static images and the surgeon's subjective experience, and require real-time manual judgment of the operational path when operating the surgical robot. This makes them prone to errors, such as puncture trajectory deviation, insufficient control of resection boundaries, and accidental contact with critical nerve areas, especially in cases of complex tissue structures or unexpected situations.
[0046] In this scenario, the implementation team first collected the patient's preoperative MRI and CT images, and combined them with intraoperative real-time ultrasound and endoscopic images. Using the image standardization processing and graph structure modeling technology of this invention, they performed three-dimensional modeling of key anatomical structures surrounding the tumor, and constructed a dynamic graph representation structure based on their spatial relationships. By embedding surgical procedure stage labels, nerve protection priorities, and anatomical connection edge weights, a multi-timestep dynamic graph sequence was constructed and input into a dynamic graph convolutional neural network for feature extraction to obtain a state representation vector.
[0047] A surgical neural network model is employed to continuously generate control actions for the state at each time step, covering multi-degree-of-freedom control parameters such as the gripping force, propulsion rate, resection path, and attitude angle of the robot actuator. Simultaneously, a lightning search algorithm is introduced to optimize key hyperparameters of the policy network, such as the learning rate, discount factor, and initial weights, and an electric potential field transfer mechanism is constructed to achieve faster convergence speed and better control performance.
[0048] Prior to the surgery, the system underwent 72 hours of intensive training in a high-fidelity simulation environment. The simulation system integrated near-realistic physiological feedback and noise interference. During training, metrics such as the control accuracy of the strategy actions, convergence speed, stability of the disturbance response, and tissue damage rate were collected to ultimately generate the optimal strategy network model for real-world execution.
[0049] During the actual surgery, the system acquires visual images, force data, and instrument contact positions every 50 milliseconds, converting them into dynamic graphical states and inputting them into the control model. The surgical robot then adjusts the path and controls the end effector according to the model's output. The surgeon only needs to perform overall monitoring and confirm key points, significantly reducing the frequency of human-machine interaction during the surgery. The entire resection process took 95 minutes, compared to the conventional 115-130 minutes for this type of surgery, representing an efficiency improvement of approximately 20%. Furthermore, the volume of tissue damage during the operation was controlled to within 0.8 cm³, significantly reduced compared to the average 1.5 cm³ of traditional procedures. Postoperative patient recovery was faster, with good consciousness recovery and no significant neurological dysfunction within 48 hours post-surgery.
[0050] Table 1. Comparison of Optimization Effects of Reinforcement Learning-Based Intelligent Surgical Robot Control Methods
[0051] Table 1 shows that during the actual surgery, visual images, force data, and instrument contact positions were acquired every 50 milliseconds and converted into a dynamic graph structure, which was then input into the surgical neural network model. The surgical robot, based on the control actions output by the surgical neural network model, completed the path adjustment and operation control of the end effector. The surgeon only needed to perform overall monitoring and key point confirmation, significantly reducing the frequency of human-machine interaction intervention during the surgery. The entire resection process took a total of 95 minutes, compared to the conventional time of 115-130 minutes for this type of surgery, representing an improvement in operational efficiency of approximately 20%. Furthermore, the volume of tissue damage during the operation was controlled within 0.8 cm³, significantly reduced compared to the average 1.5 cm³ of traditional procedures. Postoperative patient recovery was faster, with good recovery of consciousness and no significant neurological dysfunction within 48 hours post-surgery.
[0052] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
Claims
1. A control method for an intelligent surgical robot based on reinforcement learning, characterized in that, Includes the following steps: S1. Acquire medical image data of the area to be operated on and preprocess it to generate a standardized medical image; S2. Construct a dynamic graph convolutional neural network model based on standardized medical images to model the intraoperative target region as a dynamic graph structure; S3. Construct a control framework based on reinforcement learning, using dynamic graph structure features as state input, establish a surgical neural network model and output continuous control actions; S4. The lightning search algorithm is used to optimize the control parameters of the surgical neural network model. An electric field induction unit and a guiding movement path are configured for each individual lightning body. Each lightning body represents a set of control strategy parameter configurations. The performance score of each set of parameter configurations during the strategy training process is calculated through a centralized evaluation mechanism to generate the optimal strategy parameter set. S5. Perform multiple rounds of strategy training in the constructed surgical simulation environment, optimize the parameter configuration of each group of the surgical neural network model using the optimal strategy parameter set, and generate the optimal neural network model. S6. Deploy the optimal surgical neural network model to the intelligent surgical robot system, continuously collect intraoperative sensor data, and update the dynamic graph structure. S7. After the surgery is completed, record the robot's execution trajectory, tissue deformation data and operation parameters, and feed the performance evaluation results back to the reinforcement learning-based control framework for iteration.
2. The intelligent surgical robot control method based on reinforcement learning according to claim 1, characterized in that, The medical image data specifically includes preoperative MRI images, CT scan images, ultrasound images, and endoscopic images.
3. The intelligent surgical robot control method based on reinforcement learning according to claim 1, characterized in that, The control strategy parameter configuration specifically includes the initial weights of the policy network, the learning rate, and the discount factor.
4. The intelligent surgical robot control method based on reinforcement learning according to claim 1, characterized in that, S2 specifically includes: S21. Input standardized medical images into an image segmentation network to automatically identify key anatomical structures during surgery and label the structure categories and boundary contours. S22. Based on the positional relationships and spatial relative distances of anatomical structures in standardized medical images, construct an initial graph structure containing nodes and edges, where nodes represent anatomical structural units and edges represent the connection relationships between structures. S23. Obtain the predefined task phase sequence of the intraoperative operation process, encode the current surgical phase as a phase label, and embed the phase label into the additional attribute vector of each node. S24. Set the temporal sliding window length and update step size of the graph structure. Extract a snapshot of the graph structure at a fixed time step during the operation. Generate a time frame sequence by changing the node state and edge weight between the snapshots to construct an initial dynamic graph structure with temporal dependence. S25. Apply the initial dynamic graph structure to the attention-enhanced convolutional processing unit, dynamically adjust the edge weights and perform graph convolution operations based on node feature similarity, stage label weights and structural positional relationships. S26. Output the dynamic graph structure after attention-enhanced convolution.
5. The intelligent surgical robot control method based on reinforcement learning according to claim 1, characterized in that, S3 specifically includes: S31. Extract dynamic graph structure information, including node features, edge weight information, stage attributes, and time frame sequence; S32. The extracted dynamic graph structure information is used as the state input and input to the state encoding unit to uniformly encode the topological relationship and temporal characteristics between nodes, and generate a state representation for decision-making. Let the state representation be a state vector. S33. Construct an action generation unit based on a neural network. The action generation unit receives a state vector as input and establishes a surgical neural network model consisting of an input layer, multiple nonlinear computation layers and an output layer. S34. In the output layer, a continuous control action command vector is constructed based on the state vector and the internal weight matrix: ; in, Represents the continuous control action vector. This represents the candidate vector of actions to be optimized. Represents the action candidate space, Indicates the total number of time steps. Indicates the first The state vector at each time step Indicates the first The output layer weight matrix corresponding to each time step Indicates the first Time step attention weight tensor The linear transformation matrix that maps action vectors to the state-target space. Represents element-wise tensor product. Represents the squared Euclidean distance. Regularization coefficient, Let n denote the third norm of the j-th action vector, n denote the dimension of the action vector, and j denote the control action vector. The index of the j-th component in the data. Indicates a time step. This represents a small positive constant used to prevent the denominator from being zero; S35. Construct a value evaluation unit to evaluate the expected reward value of the control action of the action generation unit under the state vector, and predict the reward trend in future states. S36. The action generation unit and the value evaluation unit are combined to form a reinforcement learning control framework, which receives performance feedback indicators and updates and optimizes the neural network structure and parameters.
6. The intelligent surgical robot control method based on reinforcement learning according to claim 1, characterized in that, S4 specifically includes: S41. Construct an initial electric potential field mapping structure, randomly distribute several individual lightning bodies in the parameter space, and each individual lightning body corresponds to a set of control parameters of the surgical neural network model; S42. Configure an electric field induction structure unit for each lightning body, calculate the local electric field intensity vector based on the parameter distance and potential difference between adjacent lightning bodies, construct a voltage gradient guiding tensor, and record the directional migration tendency of each lightning body. S43. Introduce a perturbation-reconstruction mechanism, apply a small-range random perturbation to the current position of each lightning element, and construct a perturbation feedback map by combining it with the historical potential trajectory, and dynamically adjust the next movement path; S44. Use the control parameter set corresponding to each lightning body to initialize the surgical neural network model, and run a fixed number of policy training sessions in the same simulation training environment. Record the performance indicators of each set of parameters in four dimensions: control accuracy, motion smoothness, convergence rate, and tissue feedback response. S45. Construct a centralized evaluation mechanism, jointly inputting performance indicators and control action vectors into the evaluation convergence unit to calculate the comprehensive performance score of each group of individual lightning bodies: ; in, This represents the overall performance score corresponding to the control parameters. This represents the i-th component in the continuous control action vector. This represents the dimension of the control action vector. Indicates the total number of time steps. This represents the control stability weight of the i-th control variable at time step t. This represents the target action reference value of the i-th control variable of the lightning body at the t-th time step, where i represents the i-th dimension in the control action vector, j represents the j-th component in the control action vector, and t represents the scoring time step index. S46. Mark the set of control parameters corresponding to the lightning body with the lowest overall performance score as the optimal strategy parameter set.
7. The intelligent surgical robot control method based on reinforcement learning according to claim 1, characterized in that, S5 specifically includes: S51. Construct a surgical simulation environment with 3D modeling based on MRI data, and integrate a real-time physiological feedback simulation unit and an interference injection unit in the environment; S52. Load the optimal strategy parameter set into the surgical neural network model, initialize the weight tensor, bias tensor and time step decision weight mapping table, and set up an environmental disturbance recording unit to track the impact of environmental fluctuations on the control strategy during training. S53. Perform a multi-batch reinforcement training process in a surgical simulation environment, collect the continuous state transition sequence, environmental disturbance response trajectory and control action output trajectory generated by the robot's actions, and construct a sequence batch processing tensor in the form of timestamps; S54. Based on the comprehensive performance score, dynamically relabel the weights of the training sample batches and perform cross-time step residual alignment update operations on the sequence batch processing tensors: ; in, This represents the trainable parameter tensor of the surgical neural network model in the current round. This represents the surgical neural network parameter tensor obtained after parameter updates. This represents the learning rate coefficient. Indicates the total number of time steps. Let represent the dimension of the control action vector, where i represents the i-th dimension in the control action vector. This represents the overall performance score corresponding to the control parameters. Indicates the first The actual output value of the i-th control action in each time step. Indicates a time step. Indicates the first The target reference value for the i-th control action in -1 time steps. Let i represent the i-th component of the dynamic graph structure state vector at time step t. No. The i-th component of the dynamic graph structure state vector in time step -1. This represents the overall performance score corresponding to the control parameters. represents a small positive constant used to prevent the denominator from being zero, and i represents the i-th dimension in the control action vector; S55. During the training process, a parameter sparsity self-adjustment mechanism is introduced to perform round-by-round gating compression on the low-amplitude units in the updated surgical neural network parameter tensor. S56. Based on three dimensions—action performance stability, model structure sparsity, and disturbance response sensitivity—construct a final evaluation index system, rank the models based on the evaluation index, and output the optimal neural network model.
8. The intelligent surgical robot control method based on reinforcement learning according to claim 1, characterized in that, S6 specifically includes: S61. Deploy the optimal neural network model into the control and decision-making unit of the intelligent surgical robot; S62. Load a dynamic graph convolutional neural network structure and a state perception unit into the surgical robot, and establish a data input-output connection with the control decision unit. S63. Initialize the robot end effector position, joint state, graph structure buffer, and environmental sensor interface before the operation begins; S64. Collect real-time visual images, force feedback, instrument touch position and tissue deformation data generated during the operation, and input them to the status sensing unit at a fixed frequency to update the dynamic graph structure.
9. The intelligent surgical robot control method based on reinforcement learning according to claim 8, characterized in that, S64 specifically includes: S641. The state-aware unit is used to convert the collected data into a graph structure input, including constructing nodes, edges, spatial connection relationships and stage attributes, and generating a dynamic graph state representation at each time step; S642. Input the dynamic graph structure of the current time step into the control decision unit as the state vector input of the surgical neural network model; S643. The surgical neural network model outputs the continuous control action command vector for the current time step, including multi-degree-of-freedom action values and corresponding control amplitudes. S644: Transmit the control motion command vector to the robot's underlying actuator interface to drive the end effector to perform operations, including position movement, attitude adjustment and clamping force control; S645: Receive actuator feedback signals and sensor monitoring values, and record the current status and action result information; S646. Repeat steps S641 to S642 throughout the entire surgical procedure to form a real-time state-action loop, thereby realizing an intelligent control execution process driven by a dynamic graph structure.
10. The intelligent surgical robot control method based on reinforcement learning according to claim 1, characterized in that, Specifically, S7 includes: S71. After the surgery is completed, extract the complete execution trajectory from the robot, including the displacement path of the end effector, the change of attitude angle and the time stamp sequence; S72. Extract tissue deformation data collected during the operation, including the degree of tissue deformation, instrument contact point trajectory, and force feedback records; S73. Record the operation parameters during the surgical procedure, including the control action amplitude sequence, response delay, interference error value and state transition frequency, and construct a multi-dimensional operation dataset; S74. Perform unified format processing on all recorded data in steps S71 to S73 to construct a performance evaluation data matrix. S75. Set up a multi-index evaluation system that includes control accuracy, tissue protection, energy consumption level and motion stability, and perform weighted aggregation on the performance evaluation data matrix to generate performance evaluation results. S76. Feed the performance evaluation results back to the reinforcement learning-based control framework for iteration.