Neurosurgical robot obstacle avoidance method based on reinforcement learning

By applying the obstacle avoidance control model based on reinforcement learning in neurosurgery robots, the problems of insufficient obstacle avoidance accuracy, real-timeness and adaptability in the prior art are solved, and more efficient and safer robot surgical operations are achieved.

CN120053078AActive Publication Date: 2025-05-30BEIHANG UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510199603.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-05-30
Estimated Expiration
2045-02-24

AI Technical Summary

Technical Problem

The existing intelligent navigation obstacle avoidance technology has problems with insufficient accuracy, real-time and adaptability in neurosurgery robots.

Method used

Using a reinforcement learning-based method, the obstacle avoidance control model is trained through a deep deterministic strategy gradient algorithm, combined with three-dimensional environmental images and real-time sensor data, the robot motion path is dynamically adjusted to achieve intelligent obstacle avoidance.

Benefits of technology

It improves the accuracy and safety of surgical robots, enhances the real-time and adaptive ability of obstacle avoidance, and ensures that obstacles can be avoided in a timely and accurate manner in complex surgical environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120053078A_ABST
    Figure CN120053078A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of robot control, in particular to a neurosurgery robot obstacle avoidance method based on reinforcement learning, which comprises the following steps: determining a target space based on a three-dimensional environment image, determining state space information of a neuroendoscope, and determining action space information of the neuroendoscope; determining a reward model and a penalty model of the motion action of the surgical robot; a depth deterministic strategy gradient algorithm is adopted for training, and an obstacle avoidance control model is obtained; operating the neurosurgical robot, acquiring a real-time state by the position sensor and the force sensor, and inputting the real-time state into the obstacle avoidance control model to obtain a control action; the surgical robot is controlled by the control action, movement obstacle avoidance of the surgical robot is completed, and the obstacle avoidance precision, real-time performance and adaptability of the neurosurgery surgical robot can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of robot control, and particularly to an obstacle avoidance method for a neurosurgical robot based on reinforcement learning. Background Art

[0002] The neuroendoscopic robot is an advanced medical assistance technology. It uses a high-precision tracking locator and a precision robotic arm as the "eyes" and "hands" of the operator, and can achieve precise positioning and flexible operation in a narrow physiological space. The most significant advantage of the neuroendoscopic robot is that it can determine the three-dimensional anatomical structure with the help of image information and achieve precise control of the robotic arm through force feedback technology. It can clearly identify the spatial positions of key tissues and automatically adjust the operation force and operation direction through an intelligent control strategy, thereby protecting the surrounding tissues to the greatest extent. Compared with traditional methods, this intelligent assistance system has the characteristics of clear vision, small trauma, high precision, etc., and has become an important technological innovation in the medical field.

[0003] The existing intelligent navigation and obstacle avoidance technologies are mainly divided into two categories: spatial navigation technology based on pre-set images and robot-assisted technology based on real-time force feedback. Among them, the navigation technology based on pre-operative images pre-establishes a spatial positioning model through pre-collected three-dimensional image data (such as CT or MRI scans). The system uses image registration and segmentation methods to construct a spatial relationship map of each target object for real-time guiding the movement of the device. However, due to the deformation of soft tissues, the possible change of the position of the target object, and the deviation generated during the movement of the device, these factors will all cause a deviation between the pre-set spatial model and the actual situation, affecting the positioning accuracy. The robot-assisted technology based on force feedback is equipped with a force sensor at the end of the robotic arm and performs obstacle avoidance control by detecting the contact force, but there are also various problems: the force signal needs to be transmitted through multiple mechanical structures to the sensor, resulting in insufficient real-time performance; the forces from all directions in a complex structure will interfere with each other, causing a relatively large measurement error; at the same time, the system is sensitive to the control force applied by the operator, making the adaptive ability limited. In a rapidly changing working environment, the existing systems are difficult to make timely and accurate obstacle avoidance decisions. In short, the existing intelligent navigation and obstacle avoidance technologies still need to be improved in terms of accuracy, real-time performance, and adaptive ability. Summary of the Invention

[0004] In view of the above problems, the present invention provides an obstacle avoidance method for a neurosurgical robot based on reinforcement learning, which solves the problems of poor accuracy, real-time performance, and adaptability in obstacle avoidance of surgical robots in the prior art.

[0005] The present invention provides an obstacle avoidance method for a neurosurgical robot based on reinforcement learning. The neurosurgical robot includes a neuroendoscope, a robotic arm, a position sensor, and a force sensor. The robotic arm has multiple robotic joints, and one end of the robotic arm is connected to the neuroendoscope. The method is characterized by the following steps:

[0006] Step S1: Determine the target space based on the three-dimensional environmental image, determine the state space information of the neuroendoscope, and determine the action space information of the neuroendoscope;

[0007] Step S2: Determine the reward model of the surgical robot's motion action based on the distance between the position of the neuroendoscope and the motion end point in the target space, and determine the penalty model of the surgical robot's motion action based on the distance between the neuroendoscope and the key environmental structures in the target space and the contact force between the neuroendoscope and the contactable structures in the target space;

[0008] Step S3: According to the state space information and the action space information, and based on the reward model and penalty model of the surgical robot's motion action, use the deep deterministic policy gradient algorithm for training to obtain an obstacle avoidance control model;

[0009] Step S4: Run the neurosurgical robot, obtain the real-time state by the position sensor and the force sensor, input it into the obstacle avoidance control model, and obtain the control action;

[0010] Step S5: Control the surgical robot by the control action to complete the motion obstacle avoidance of the surgical robot.

[0011] Preferably, step S1 specifically includes:

[0012] Step S1-1: Obtain the tomographic scan image of the area to be analyzed through a medical imaging device; perform three-dimensional reconstruction on the tomographic scan image to obtain a three-dimensional model; determine the target space based on the three-dimensional model, and the target space is the motion space of the surgical robot;

[0013] Step S1-2: Determine the position state and force state of the neuroendoscope as the state space information;

[0014] Step S1-3: Determine the motion state of the robotic arm as the action space information.

[0015] Preferably, the state space information in step S1-2 specifically includes:

[0016] Neuroendoscope position P e 、Key environmental structure position P a 、Contact force F between the neuroendoscope and the contactable structure axis And motion end point position P t ;

[0017] The action space information described in step S1-3 specifically includes: the rotational angle increments of the 6 joints of the robotic arm [Δθ 1 , Δθ 2, , Δθ 3, , Δθ 4, , Δθ 5 , Δθ 6 .

[0018] Preferably, step S2 specifically includes:

[0019] Step S2-1, determining the in-place reward, where the closer the distance between the neuroendoscope and the movement end point, the greater the in-place reward;

[0020] Step S2-2, determining the collision penalty based on the size relationship between the distance between the neuroendoscope and the key environmental structure and the safety distance and warning distance;

[0021] Step S2-3, determining the contact force penalty based on the size relationship between the contact force between the neuroendoscope and the accessible structure in the target space and the first and second contact force thresholds.

[0022] Preferably, in step S2-1, the calculation expression of the in-place reward is:

[0023]

[0024] where R1 represents the in-place reward, Dis(P e , P t ) represents the distance between the neuroendoscope position P e and the movement end point position P t , and α represents the in-place reward normalization coefficient;

[0025] In step S2-2, the calculation expression of the collision penalty is:

[0026]

[0027] where R2 represents the collision penalty, Dis(P e , P a ) represents the distance between the neuroendoscope position P e and the key environmental structure position P a , L1 represents the warning distance, L2 represents the safety distance, and β represents the collision penalty normalization coefficient;

[0028] In step S2-3, the calculation expression of the contact force penalty is:

[0029]

[0030] where R3 represents the contact force penalty, Faxis Indicates the contact force between the neuroendoscope and the contactable structure. T1 and T2 respectively represent the first and second contact force thresholds, and γ is the contact force penalty normalization coefficient. Is the first constant.

[0031] Preferably, the step S3 specifically includes:

[0032] Taking the state space of the neuroendoscope and the action space of the neuroendoscope as the state space and action space of the DDPG algorithm, and taking the rewards and penalties of the surgical robot's motion actions as the rewards and penalties of the DDPG algorithm, performing reinforcement learning training, and finally obtaining an obstacle avoidance control model;

[0033] During the reinforcement learning training of the DDPG algorithm, applying the current moment action a t to the current moment state s t , obtaining the next moment state s t+1 , and determining the in-place reward, collision penalty, and contact force penalty at the current moment according to the neuroendoscope position, key environmental structure position, contact force between the neuroendoscope and the contactable structure, and motion end position in the next moment state s t+1 , and determining the current moment reward from the in-place reward, collision penalty, and contact force penalty at the current moment.

[0034] Preferably, the specific calculation method for determining the current moment reward from the in-place reward, collision penalty, and contact force penalty at the current moment is:

[0035] For any moment, the reward r is calculated using the following expression:

[0036] r = R1 + R2 + R3

[0037] where R1, R2, and R3 are the in-place reward, collision penalty, and contact force penalty respectively.

[0038] Preferably, the step S4 specifically includes:

[0039] Real-time collecting the neuroendoscope position P e , key environmental structure position P a , and motion end position P t through the position sensor, and simultaneously real-time collecting the contact force F axis between the neuroendoscope and the contactable structure through the force sensor, as the real-time state;

[0040] Inputting the real-time state into the obstacle avoidance control model, and outputting corresponding control actions, where the control actions include the rotation angle increments of each joint of the surgical robot's manipulator.

[0041] Compared with the prior art, the present invention has at least the following beneficial effects:

[0042] (1) Through the obstacle avoidance method of the neurosurgical robot based on reinforcement learning, the present invention improves the accuracy and safety of the surgery. By determining the target space and the state and action space information of the neuroendoscope from the three-dimensional environmental image, it provides an accurate environmental perception basis for subsequent obstacle avoidance control.

[0043] (2) By designing a reward model and a penalty model, and comprehensively considering the distance between the neuroendoscope and the target space as well as the contact force, the present invention ensures that the robot can intelligently avoid key environmental structures during movement. The robot can continuously optimize its obstacle avoidance strategy during the surgery, automatically learn the best obstacle avoidance path, thereby improving the obstacle avoidance efficiency and accuracy of the operation of the surgical robot.

[0044] (3) Through training with the deep deterministic policy gradient algorithm to form an obstacle avoidance control model, the robot can adjust the movement path in real time. Combining the real-time data provided by the position sensor and the force sensor, the robot can dynamically respond to the changes in the surgical environment and timely adjust the obstacle avoidance strategy. It can maintain the obstacle avoidance accuracy and adaptability of the surgical robot. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] The drawings are only for the purpose of showing specific embodiments and are not considered as a limitation of the present invention.

[0046] Figure 1 It is a schematic diagram of the obstacle avoidance method of the neurosurgical robot based on reinforcement learning provided by the present invention;

[0047] Figure 2 It is a schematic diagram of the principle of the deep deterministic policy gradient method provided by the present invention.

[0048] Figure 3 It is a schematic diagram of the intracranial structure constructed using medical images provided by the present invention.

[0049] Figure 4 It is a schematic diagram of the neuroendoscope manipulator model provided by the present invention.

[0050] Figure 5 It is a schematic diagram of the neuroendoscope collision penalty provided by the present invention.

[0051] Figure 6 It is a schematic diagram of the neuroendoscope contact force penalty provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0052] To more clearly understand the above objects, features, and advantages of the present invention, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be noted that, without conflict, the embodiments of the present invention and the features in the embodiments can be combined with each other. In addition, the present invention can also be implemented in other ways different from those described herein. Therefore, the protection scope of the present invention is not limited by the specific embodiments disclosed below.

[0053] To illustrate the effectiveness of the method proposed by the present invention, the above technical solutions of the present invention will be described in detail below through a specific embodiment. A specific embodiment of the present invention discloses an obstacle avoidance method for a neurosurgical robot based on reinforcement learning, as Figure 1 , Figure 4 shown. The neurosurgical robot includes a robotic arm, a neuroendoscope is connected to the end of the robotic arm, and there are multiple robot joints on the robotic arm, and the robotic arm controls the movement of the neuroendoscope. The neurosurgical robot also includes a position sensor, a force sensor, etc. The specific implementation steps are as follows:

[0054] Step S1: Determine the target space based on the three-dimensional environmental image, determine the state space information of the neuroendoscope, and determine the action space information of the neuroendoscope.

[0055] In this step, tomographic images of the area to be analyzed are obtained through medical imaging devices (such as MRI, CT, etc.), that is, three-dimensional environmental images. The tomographic images are spatially registered and fused through a three-dimensional reconstruction algorithm to generate a three-dimensional model including tubular structures, abnormal regions, and key environmental structures, as Figure 3 shown.

[0056] Based on the three-dimensional model, the target space is determined. The target space is the space where the surgical robot moves. In some embodiments, a three-dimensional coordinate system of the target space can be established, including the spatial coordinate positions of tubular structures, abnormal regions, and key structures, etc.

[0057] The present invention defines the state space and action space of the neuroendoscope during the subsequent reinforcement learning training process.

[0058] The state space represents all possible states of the neuroendoscope in the current environment. A certain state in the state space includes the following data:

[0059] (1) The position P of the neuroendoscope e : Represents the current spatial position of the neuroendoscope, which can be obtained through the position sensor.

[0060] (2) The position P of the key environmental structure a: The positions in the target space, such as a tubular structure, an abnormal area, and a key structure, can be obtained by the position sensor.

[0061] (3) Contact force P between the neuroendoscope and the contactable structure axis : The feedback data indicating whether the neuroendoscope is in contact with the contactable structure and the contact force between the neuroendoscope and the contactable structure can be obtained by the force sensor.

[0062] (4) Movement end position P t : Represents the target position of the current neuroendoscope movement.

[0063] The action space represents the movement information of each joint of the robotic arm of the neurosurgical robot. In the neurosurgical robot-assisted endoscopic surgery, the neuroendoscope is mounted at the end of the robot, and all movements of the neuroendoscope are controlled by the robotic arm. In some embodiments, the robotic arm has 6 joints, and the rotation angle increment a = [Δθ 1 , Δθ 2, , Δθ 3, , Δθ 4, , Δθ 5 , Δθ 6 of each joint of the robotic arm is determined as a certain action in the action space.

[0064] Through the above steps, the present invention establishes a target space model, a state space, and an action space of the neuroendoscope based on three-dimensional imaging for reinforcement learning training.

[0065] Step S2: Determine a reward model for the movement actions of the surgical robot based on the distance between the position of the neuroendoscope and the movement end point in the target space, and determine a penalty model for the movement actions of the surgical robot based on the distance between the neuroendoscope and the key environmental structures in the target space and the contact force between the neuroendoscope and the contactable structures in the target space;

[0066] In reinforcement learning, the reward is a key factor determining how the model learns, and a good reward design can guide the system to learn to make appropriate decisions in a complex environment.

[0067] In the scenario of controlling the obstacle avoidance movement of the neurosurgical robot, the present invention designs the calculation methods of rewards and penalties according to the complex situations in the scenario, including:

[0068] (1) In-place reward: The training goal is to send the neuroendoscope to the movement end position. Therefore, the closer to the movement end position, the greater the reward. After reaching the position, the reward is 1, indicating that the task is completed.

[0069] The expression of the in-place reward is:

[0070]

[0071] Among them, R1 represents the in-place reward, and Dis(P e , P t ) represents the distance between the neuroendoscope position P e and the movement end point position P t , and α represents the in-place reward normalization coefficient.

[0072] (2) Collision penalty: During the movement of the neuroendoscope, it is necessary to ensure that it does not touch the key environmental structures. To further improve safety, a warning distance L1 is set. When the distance is less than L1, it is considered to enter a dangerous state and the task fails; a safety distance L2 is set. When the distance is greater than L2, it is considered that there is no collision risk, and at this time the penalty is set to 0, which means not participating in the feedback, and the training efficiency can be further improved; between L1 and L2, it is considered that a penalty is required, which is expressed as a function of the distance, as shown in Figure 5 .

[0073] The expression of the collision penalty is:

[0074]

[0075] Among them, R2 represents the collision penalty, and Dis(P e , P a ) represents the distance between the neuroendoscope position P e and the position of the key environmental structure P a , L1 represents the warning distance, L2 represents the safety distance, and β represents the collision penalty normalization coefficient.

[0076] (3) Contact force penalty: There will be contact between the endoscope and the contactable structure during movement. To ensure safety. When the contact force is greater than T2, it is considered dangerous, and at this time the task fails; if the contact force is less than T1, it is considered safe and does not participate in the feedback; when the force is between T1 and T2, a penalty is required, and the penalty base is increased, and it is considered that the contact force is more worthy of attention during movement, as shown in Figure 6 .

[0077] The expression of the contact force penalty is:

[0078]

[0079] Among them, R3 represents the contact force penalty, and F axis represents the contact force between the neuroendoscope and the contactable structure, T1 and T2 respectively represent the first and second contact force thresholds, γ is the contact force penalty normalization coefficient, is the first constant.

[0080] Step S3: According to the state space information and the action space information, and based on the reward model and penalty model of the surgical robot's motion actions, use the Deep Deterministic Policy Gradient algorithm for training to obtain an obstacle avoidance control model;

[0081] The present invention performs reinforcement learning based on the state space information, the action space information, and the rewards and penalties of the surgical robot's motion actions.

[0082] The present invention uses the Deep Deterministic Policy Gradient (DDPG) method. The action space output by the method is a continuous deterministic action. DDPG is implemented based on the Actor-Critic framework. The Actor network includes a main network μ(s|θ μ ) and a target network μ′(s|θ μ′ ), where θ μ and θ μ′ represent the parameters of the main network and the target network respectively. The main network and the target network of the Actor network can take the current state s as input and obtain the corresponding actions. The Critic network also includes a main network Q(s,a|θ Q ) and a target network Q′(s,a|θ Q′ ), where θ Q and θ Q′ represent the parameters of the main network and the target network respectively. The main network and the target network of the Critic network can take the current state s and the current action a as input and obtain an evaluation value of the quality of the action a in the state s. The specific training process based on DDPG is as Figure 2 shown.

[0083] Specifically in the present invention, the state space of the neuroendoscope and the action space of the neuroendoscope are used as the state space and action space of the DDPG algorithm, and the rewards and penalties of the surgical robot's motion actions are used as the rewards and penalties of the DDPG algorithm for reinforcement learning training, and finally an obstacle avoidance control model is obtained.

[0084] The pseudo-code of the specific training process based on DDPG is as follows:

[0085]

[0086] In the above DDPG training process, the action a t at the current moment includes the rotation angle increment of each joint of the robotic arm at the current moment, and the state s t at the current moment includes the position of the neuroendoscope at the current moment, the position of the key environmental structures, the contact force between the neuroendoscope and the contactable structures, and the motion end position.

[0087] In the DDPG training process, the current moment action a is utilized t to obtain the current moment reward r t The specific implementation method is as follows:

[0088] Apply the current moment action a t to the current moment state s t to obtain the next moment state s t+1 Based on the neuroendoscope position, key environmental structure position, contact force between the neuroendoscope and the contactable structure, and movement end position in the next moment state s t+1 determine the current moment in-place reward, collision penalty, and contact force penalty, and determine the current moment reward r from the current moment in-place reward, collision penalty, and contact force penalty t .

[0089] For any moment, the expression for calculating the reward r in the present invention is:

[0090] r = R1 + R2 + R3

[0091] Through the training of the above steps, the present invention obtains an obstacle avoidance control model for a neurosurgical robot. This model can take the current state s of the robot as input and obtain the corresponding action

[0092] Step S4: During the operation of the neurosurgical robot, obtain the real-time state by the position sensor and the force sensor, input it into the obstacle avoidance control model, and obtain the control action

[0093] In this step, during the actual operation of the neurosurgical robot, the neuroendoscope position P e , key environmental structure position P a and movement end position P t are collected in real time by the position sensor, and at the same time, the contact force F between the neuroendoscope and the contactable structure is collected in real time by the force sensor axis .

[0094] Input the collected position information and force information as state space information into the pre-trained obstacle avoidance control model. Based on the current state space information, this obstacle avoidance control model calculates and outputs the corresponding control action to guide the surgical robot to adjust the movement trajectory and realize the intelligent avoidance of obstacles during the movement of the surgical robot

[0095] The control action may include the rotation angle increment of each joint of the robotic arm of the surgical robot: [Δθ 1 , Δθ 2, , Δθ 3, , Δθ 4, , Δθ 5 , Δθ 6.

[0096] Step S5: Control the surgical robot by the control action to complete the motion obstacle avoidance of the surgical robot.

[0097] In this step, the instruction corresponding to the control action is transmitted to the drive modules of each robotic arm of the neurosurgical robot. The control action includes the rotational angle increments of the joints of the robotic arm of the surgical robot, specifically indicating the rotation direction and angle of each joint, so as to ensure that the robot can move along a predetermined path when performing surgical tasks and avoid obstacles in the environment in real time. Through precise joint angle adjustment, the surgical robot can smoothly and flexibly avoid possible collisions or interferences, ensuring the smooth progress of the surgical procedure. In addition, the control system also dynamically adjusts the control strategy according to the real-time sensor feedback data to cope with different surgical environment changes, ensuring that the robot always maintains the best motion trajectory and safety.

[0098] Although the specific implementation manners of the present invention depict each action or step in a specific order, it should be understood that such actions or steps are required to be executed in the specific order shown or in a sequential order, or all the illustrated actions or steps should be executed to achieve the desired result. In certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the above description, these should not be construed as limiting the scope of the present disclosure. Certain features described in the context of separate embodiments can also be implemented in combination in a single implementation. On the contrary, various features described in the context of a single implementation can also be implemented separately or in any suitable sub-combination in multiple implementations. The above is only the preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered within the protection scope of the present invention.

[0099] The above is only the preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered within the protection scope of the present invention.

Claims

1. A neurosurgery robot obstacle avoidance method based on reinforcement learning, wherein the neurosurgery robot comprises a neuroendoscope, a robotic arm, a position sensor and a force sensor, wherein the robotic arm has a plurality of robot joints, and one end of the robotic arm is connected to the neuroendoscope, wherein: The following steps are involved: Step S1, determining the target space based on the three-dimensional environment image, determining the state space information of the neuroendoscope, and determining the action space information of the neuroendoscope; Step S2, determining a reward model for the surgical robot's motion based on the distance between the position of the neuroendoscope and the motion endpoint in the target space, and determining a penalty model for the surgical robot's motion based on the distance between the neuroendoscope and the key environmental structure in the target space and the contact force between the neuroendoscope and the contactable structure in the target space; Step S3, according to the state space information and the action space information, and based on the reward model and penalty model of the surgical robot's motion action, a deep deterministic policy gradient algorithm is used for training to obtain an obstacle avoidance control model; Step S4, running the neurosurgery robot, obtaining the real-time status from the position sensor and the force sensor, inputting the obstacle avoidance control model, and obtaining the control action; Step S5: Control the surgical robot by the control action to complete the movement and obstacle avoidance of the surgical robot.

2. The neurosurgical robot obstacle avoidance method based on reinforcement learning according to claim 1 is characterized in that: The step S1 specifically includes: Step S1-1, obtaining a tomographic image of the area to be analyzed by a medical imaging device; performing three-dimensional reconstruction on the tomographic image to obtain a three-dimensional model; determining a target space based on the three-dimensional model, wherein the target space is a motion space of the surgical robot; Step S1-2, determining the position state and force state of the neuroendoscope as the state space information; Step S1-3: determining the motion state of the robotic arm as the action space information.

3. The neurosurgical robot obstacle avoidance method based on reinforcement learning according to claim 2 is characterized in that: The state space information in step S1-2 specifically includes: Neuroendoscopic position P e 、Key environmental structure location P a , contact force F between neuroendoscope and contactable structure axis And the end position P t ; The action space information in step S1-3 specifically includes: the rotation angle increments of the six joints of the robot arm [Δθ1, Δθ 2, ,Δθ 3, ,Δθ 4, ,Δθ5,Δθ6].

4. The neurosurgical robot obstacle avoidance method based on reinforcement learning according to claim 3 is characterized in that: The step S2 specifically includes: Step S2-1, determining a reward for reaching the target, wherein the closer the distance between the neuroendoscope and the end point of the movement is, the greater the reward for reaching the target; Step S2-2, determining a collision penalty based on the relationship between the distance between the neuroendoscope and the key environmental structure and the safety distance and the warning distance; Step S2-3: determining a contact force penalty based on a relationship between the contact force between the neuroendoscope and the contactable structure in the target space and the first and second contact force thresholds.

5. The obstacle avoidance method for a neurosurgery robot based on reinforcement learning according to claim 4 is characterized in that: In step S2-1, the calculation expression of the arrival reward is: Among them, R1 represents the reward in place, Dis(P e ,P t ) represents the position of the neuroendoscope P e and the end position P t , α represents the normalization coefficient of the arrival reward; In step S2-2, the calculation expression of the collision penalty is: Among them, R2 represents the collision penalty, Dis(P e ,P a ) represents the position of the neuroendoscope P e and key environmental structure location P a distance, L1 represents the warning distance, L2 represents the safety distance, and β represents the collision penalty normalization coefficient; In step S2-3, the calculation expression of the contact force penalty is: Among them, R3 represents the contact force penalty, F axis represents the contact force between the neuroendoscope and the contactable structure, T1 and T2 represent the first and second contact force thresholds, respectively, γ is the contact force penalty normalization coefficient, is the first constant.

6. The neurosurgical robot obstacle avoidance method based on reinforcement learning according to claim 5 is characterized in that: The step S3 specifically includes: The state space and action space of the neuroendoscope are used as the state space and action space of the DDPG algorithm, and the rewards and penalties of the surgical robot's motion actions are used as the rewards and penalties of the DDPG algorithm, and reinforcement learning training is performed to finally obtain an obstacle avoidance control model; In the DDPG algorithm reinforcement learning training process, the current action a t Acting on the current state s t , get the next moment state s t+1 , according to the next moment state s t+1 The position of the neuroendoscope, the position of the key environmental structure, the contact force between the neuroendoscope and the contactable structure, and the end point position of the movement determine the current moment's arrival reward, collision penalty, and contact force penalty. The current moment's reward is determined by the current moment's arrival reward, collision penalty, and contact force penalty.

7. The neurosurgery robot obstacle avoidance method based on reinforcement learning according to claim 6 is characterized in that: The specific calculation method of determining the current moment reward from the current moment's arrival reward, collision penalty, and contact force penalty is as follows: For any time, the reward r is calculated using the following expression: r=R1+R2+R3 Among them, R1, R2, and R3 are the arrival reward, collision penalty, and contact force penalty, respectively.

8. The neurosurgery robot obstacle avoidance method based on reinforcement learning according to claim 7, characterized in that: The step S4 specifically includes: The position sensor is used to collect the position P of the neuroendoscopy in real time. e 、Key environmental structure location P a and the end position P t At the same time, the contact force F between the neuroendoscope and the contactable structure is collected in real time by a force sensor. axis , as the real-time status; The real-time state is input into the obstacle avoidance control model, and a corresponding control action is output, wherein the control action includes a rotation angle increment of each joint of the surgical robot mechanical arm.

Citation Information

Patent Citations

  • Artificial intelligence capsule endoscopy examination method and system based on deep reinforcement learning

    CN108784636A

  • Mechanical arm control method and system based on deep reinforcement learning algorithm

    CN117140527A

  • Track planning method and track planning device for collision avoidance of equipotential operation mechanical arm

    CN119388424A

  • Mechanical arm dynamic obstacle avoidance method based on intelligent task supervision

    CN119458386A

  • Method and apparatus for intelligently controlling mechanical arm

    US20240351199A1