A surgery robot control method and device based on embodied intelligence

By obtaining reward functions and differential constraints through inverse reinforcement learning, and using multi-objective optimization to train the operating strategy of the surgical robot arm, the problem of adapting the surgical robot's operating experience to the differences in new surgical robot data is solved, thereby improving the accuracy and safety of surgical robot control.

CN119302746BActive Publication Date: 2025-12-19LONGWOOD VALLEY MEDICAL TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411211823.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-30
Publication Date
2025-12-19
Estimated Expiration
2044-08-30

AI Technical Summary

Technical Problem

Existing technologies struggle to generate new control centers that can adapt to the differences between surgical robot operating experience and new surgical robot data, leading to adaptation difficulties.

Method used

By obtaining the reward function through inverse reinforcement learning, determining the differences in constraints between past experience scenarios and application scenarios, and using multi-objective optimization to train the operating strategy of the surgical robot arm through reinforcement learning, the operating trajectory with the highest cumulative reward is selected.

Benefits of technology

By transforming differences into constraints and using multi-objective optimization for policy training, the adaptation problem caused by the differences between surgical robot operating experience and new surgical robot data is solved, thereby improving the accuracy and safety of surgical robot control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119302746B_ABST
    Figure CN119302746B_ABST
Patent Text Reader

Abstract

The application provides a surgical robot control method and device based on embodied intelligence, the method comprising: acquiring a reward function processed based on inverse reinforcement learning; determining a constraint condition based on differences between past experience scenarios and application scenarios; training an operation strategy of a surgical robot mechanical arm through reinforcement learning based on multi-objective optimization; after the training is completed, selecting an operation trajectory with the highest cumulative reward to control the surgical robot mechanical arm to execute the operation trajectory; the multi-objective optimization comprises multiple objective functions, and the objective functions at least include the reward function and the constraint condition. In the application, the differences are converted into constraint conditions, and the strategy training is performed through the multi-objective optimization, thereby solving the problem that it is difficult to adapt due to the differences between the operation experience and the data of the new surgical robot.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of image processing, in particular to a surgery robot control method and device based on embodied intelligence. BACKGROUND

[0002] Embodied intelligence refers to an intelligent agent with a body and supporting interaction with the physical world, such as robots, unmanned vehicles, etc. The intelligent agent is driven by motion instructions generated by a control center, such as a large model, by processing multiple sensor data inputs, replacing the traditional rule-based or mathematical formula-based motion driving mode, and realizing the deep integration of virtual and reality.

[0003] The surgery robot can be a perfect carrier of embodied intelligence, and the function of the surgery robot is realized based on embodied intelligence. In the specific implementation process, the operation experience of the past surgery robot can be summarized first, and then a new control center is generated according to the summarized experience.

[0004] However, there will be certain differences between the operation experience of the past surgery robot and the data of the new surgery robot. How to generate a new control center based on the summarized experience such as the reward function under the current situation is a current difficult problem to solve. SUMMARY

[0005] The problem solved by the present application is that it is currently very difficult to generate a new control center through a reward function.

[0006] To solve the above problems, the first aspect of the present application provides a surgery robot control method based on embodied intelligence, comprising:

[0007] obtaining a reward function processed based on inverse reinforcement learning;

[0008] determining a constraint condition based on the difference between the past experience scene and the application scene;

[0009] training the running strategy of the surgery robot mechanical arm based on multi-objective optimization through reinforcement learning;

[0010] After the training is completed, the running trajectory with the highest cumulative reward is selected to control the surgery robot mechanical arm to execute the running trajectory; the multi-objective optimization includes multiple objective functions, and the objective functions at least include the reward function and the constraint condition.

[0011] The second aspect of the present application provides a surgery robot control device based on embodied intelligence, comprising:

[0012] a reward acquisition module for obtaining a reward function processed based on inverse reinforcement learning;

[0013] a difference determination module configured to determine constraint conditions based on differences between past experience scenarios and application scenarios;

[0014] a strategy training module configured to train a running strategy of the surgical robot manipulator by reinforcement learning based on multi-objective optimization;

[0015] a manipulator control module configured to, after the training, select a running trajectory with the highest cumulative reward, and control the surgical robot manipulator to execute the running trajectory; the multi-objective optimization includes multiple objective functions, and the objective functions at least include the reward function and the constraint conditions.

[0016] The third aspect of the present application provides an electronic device, comprising a memory and a processor;

[0017] The memory is configured to store a program;

[0018] The processor is coupled to the memory and configured to execute the program, so as to:

[0019] obtain a reward function processed based on inverse reinforcement learning;

[0020] determine constraint conditions based on differences between past experience scenarios and application scenarios;

[0021] train a running strategy of the surgical robot manipulator by reinforcement learning based on multi-objective optimization;

[0022] After the training, a running trajectory with the highest cumulative reward is selected, and the surgical robot manipulator is controlled to execute the running trajectory; the multi-objective optimization includes multiple objective functions, and the objective functions at least include the reward function and the constraint conditions.

[0023] The fourth aspect of the present application provides a computer readable storage medium, which stores a computer program, and the program is executed by a processor to implement the above-mentioned surgical robot control method based on embodied intelligence.

[0024] In the present application, the differences are converted into constraint conditions, and the strategy training is performed by the way of multi-objective optimization, so as to solve the problem that it is difficult to adapt due to the differences between the operation experience and the data of the new surgical robot. BRIEF DESCRIPTION OF DRAWINGS

[0025] Figure 1 A flowchart of the surgical robot control method based on embodied intelligence according to the embodiments of the present application;

[0026] Figure 2 A flowchart of the strategy training of the surgical robot control method based on embodied intelligence according to the embodiments of the present application;

[0027] Figure 3 A structural block diagram of a surgery robot control device based on embodied intelligence according to an embodiment of the present application is shown in FIG. 1.

[0028] Figure 4 A structural block diagram of an electronic device according to an embodiment of the present application is shown in FIG. 2. DETAILED DESCRIPTION

[0029] In order to make the above objectives, features and advantages of the present application more apparent, specific embodiments of the present application will be described in detail below with reference to the accompanying drawings. Although exemplary embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments set forth herein. Rather, these embodiments are provided so that the present application can be more thoroughly understood and so that the scope of the present application can be accurately conveyed to those skilled in the art.

[0030] It should be noted that, unless otherwise specified, technical terms or scientific terms used in the present application should be understood as their common meanings to those skilled in the art to which the present application pertains.

[0031] The surgery robot can be a perfect carrier of embodied intelligence, and the function of the surgery robot can be realized based on embodied intelligence. In the specific implementation process, the operation experience of the past surgery robot can be summarized first, and then a new control center can be generated according to the summarized experience.

[0032] However, for the surgery robot, the past experience learned is the experience of manual control of the surgery robot, and the application scenario based on the experience is the scenario of automatic control of the surgery robot. There is a difference in the use environment between the past experience and the application scenario. How to reduce the influence of this difference in the process of generating a new control center is a current difficult problem to solve.

[0033] In view of the above problems, the present application provides a new bone registration scheme, which performs coarse registration in a digital twinning manner to solve the problem of low accuracy of current feature point registration.

[0034] An embodiment of the present application provides a surgery robot control method based on embodied intelligence. The specific scheme of the method is shown in FIG. 3, and the method can be executed by a surgery robot control device based on embodied intelligence. The surgery robot control device based on embodied intelligence can be integrated in a computer, a server, a computer cluster, a data center, or other electronic devices. As shown in FIG. 4, it is a flowchart of a surgery robot control method based on embodied intelligence according to an embodiment of the present application; wherein the surgery robot control method based on embodied intelligence comprises the following steps. Figures 1-2 Figure 1

[0035] ​​S101, acquire a reward function based on inverse reinforcement learning processing;

[0036] S102, determine a constraint condition based on a difference between a past experience scene and an application scene;

[0037] For example, in the experience scene, the historical trajectory data is obtained by expert doctor operating a robotic arm with soft control. In the operation scene, the expert doctor holds the robotic arm and applies force to the robotic arm to control the operation of the robotic arm. However, in the active control application scene of the robotic arm, obstacle avoidance control is required. In this case, the robotic arm needs to be controlled to move away from the expert doctor to avoid interference, which is contrary to the past experience scene.

[0038] Therefore, it is necessary to determine the corresponding constraint condition based on the difference between the past experience scene and the application scene, so as to avoid the adverse consequences caused by the difference.

[0039] S103, training the operation strategy of the surgical robot arm through reinforcement learning based on multi-objective optimization;

[0040] S104, after the training is completed, selecting an operation trajectory with the highest cumulative reward to control the surgical robot arm to execute the operation trajectory; the multi-objective optimization includes multiple objective functions, and the objective functions at least include the reward function and the constraint condition.

[0041] In the present application, the difference is converted into a constraint condition, and the strategy training is performed through multi-objective optimization, thereby solving the problem that it is difficult to adapt due to the difference between the operation experience and the data of the new surgical robot.

[0042] In an embodiment, the determination of the constraint condition based on the difference between the past experience scene and the application scene comprises:

[0043] Acquiring past experience scene information and application scene information;

[0044] Determining the constraint direction of the past scene based on the past experience scene information;

[0045] Determining the constraint direction of the application scene based on the application scene information;

[0046] Determining a difference direction according to the constraint direction of the past scene and the application scene, the difference direction belonging to the application scene and not belonging to the past scene;

[0047] Determining the constraint condition according to the difference direction.

[0048] In the present application, the past experience scene information can be video information and preoperative planning information of the past scene.

[0049] In the present application, the application scenario information can be video information of the current application scenario and preoperative planning information.

[0050] In the present application, the past experience scenario information determines the constraint direction of the past scenario, which can be obtained by qualitative analysis of the past experience scenario information, or can be summarized by expert doctors.

[0051] In the present application, the constraint direction of the application scenario is determined based on the application scenario information, which can be obtained by qualitative analysis of the application scenario information, or can be summarized by expert doctors.

[0052] In the present application, the difference direction is not possessed in the past scenario application scenario, but is possessed in the application scenario information. If it is possessed in the past scenario application scenario, but is not possessed in the application scenario information, it is generally not a difference direction, unless the difference direction conflicts with the application scenario information or its constraints.

[0053] In an embodiment, the constraint condition includes an obstacle constraint, and / or a safety boundary constraint, and / or an end tool orientation constraint, and / or a mechanical constraint.

[0054] In the present application, when there are multiple constraint conditions, multiple constraint conditions can be combined by weight addition, or the case of satisfying all constraint conditions is considered to satisfy the constraint.

[0055] The mechanical constraint can be: tool-bone contact force: during cutting, the robot must accurately control the force exerted by the tool on the bone to prevent bone fragmentation or unnecessary damage; vibration control: minimize tool vibration during cutting to improve surgical precision and patient comfort.

[0056] The safety boundary constraint can be: to avoid deviation of the tool from the path during actual execution.

[0057] The end tool orientation constraint can be: to avoid the tool from facing the staff during actual execution.

[0058] In the present application, the specific settings of the safety boundary constraint, and / or the end tool orientation constraint, and / or the mechanical constraint can be determined according to the actual situation, and the present application does not limit and constrain.

[0059] In an embodiment, the constraint condition includes an obstacle constraint, and / or a safety boundary constraint, and / or an end tool orientation constraint, and / or a mechanical constraint. Figure 2 As shown, the S103 trains the operation strategy of the surgical robot manipulator by reinforcement learning based on multi-objective optimization, including:

[0060] S301, acquiring operation environment data and planning data, and constructing a simulation space;

[0061] S302, constructing a state space and an action space based on the inverse reinforcement learning process and the running environment data and the planning data; the action space comprises a plurality of joint axes, each joint axis having a forward rotation action and a reverse rotation action;

[0062] S303, constructing a deep Q network model;

[0063] S304, performing first training on the deep Q network model based on the deep Q network learning with the obtained reward function;

[0064] S305, performing second training on the deep Q network model based on the deep Q network learning with the differential constraint condition as the reward function;

[0065] S306, alternately performing the first training and the second training until a preset condition is met.

[0066] In this way, by alternately training with the obtained reward function and the differential constraint condition, the reward function is first optimized, and then the strategy is gradually adjusted to approach the optimized reward under the premise of meeting the constraint condition, while trying to meet the constraint as much as possible. This is alternately performed until a solution that meets both the reward function and the constraint condition is found.

[0067] In an embodiment, S304, performing first training on the deep Q network model based on the deep Q network learning with the obtained reward function;

[0068] The obtained reward function is used as the reward function;

[0069] An experience replay pool is constructed, the experience replay pool comprising actions, states, rewards, and next states of the agent;

[0070] Sampling is performed from the experience replay pool, and a Q value is calculated based on the deep Q network model;

[0071] The weights of the deep Q network model are updated by minimizing a loss function until the training is completed.

[0072] In an embodiment, S305, performing second training on the deep Q network model based on the deep Q network learning with the differential constraint condition as the reward function, comprises:

[0073] A reward function is constructed according to the differential constraint condition;

[0074] An experience replay pool is constructed, the experience replay pool comprising actions, states, rewards, and next states of the agent;

[0075] Sampling is performed from the experience replay pool, and a Q value is calculated based on the deep Q network model;

[0076] The weights of the deep Q network model are updated by minimizing the loss function until the training is completed.

[0077] In the present application, the specific process of deep Q reinforcement learning and the corresponding setting of the deep Q network model can be carried out with reference to the prior art, which will not be described herein.

[0078] In the present application, the specific process of the first training and the second training is similar, and only the reward function is different.

[0079] In an embodiment, the calculation process of the obstacle constraint is:

[0080] Obtain the pose data of the robot arm in the current state;

[0081] Obtain the environment data in the current state;

[0082] Determine the minimum distance between the robot arm and the environmental obstacles according to the robot arm pose data and the environment data;

[0083] In the case where the minimum distance between the robot arm and the environmental obstacles is greater than the preset distance, it is determined that the obstacle constraint is satisfied.

[0084] The preset distance can be determined according to the actual situation.

[0085] It should be noted that the shapes of the robot arm and the environmental obstacles are irregular, and it is difficult to accurately calculate and estimate them.

[0086] In an embodiment, the determination of the minimum distance between the robot arm and the environmental obstacles according to the robot arm pose data and the environment data comprises:

[0087] Equivalent processing is performed on the robot arm pose data, and the robot arm is equivalent to a plurality of line segments and a preset thickness connected in sequence;

[0088] Equivalent processing is performed on the environment data, and the obstacles in the environment are equivalent to a sphere and a cylinder wrapping the obstacles;

[0089] The distance between each line segment of the robot arm and the sphere center and the cylinder axis segment is calculated;

[0090] The minimum distance is determined according to the calculated distance, the preset thickness, the radius of the sphere, and the radius of the cylinder.

[0091] In the present application, the robot arm and the obstacles are equivalent to simple geometric bodies, thereby greatly reducing the workload of the minimum distance calculation.

[0092] Thus, the robot arm is equivalent to a line segment, and the obstacle is equivalent to a point or a line segment. Only the distance between the line segment and the line segment and the distance between the line segment and the point need to be calculated, and the distance between the robot arm itself and the obstacle can be obtained.

[0093] The distance between the line segment and the point is to calculate the projection point of the point to the line segment. If the projection point is on the line segment, the distance between the point and the projection point is the distance to be calculated. If the projection point is not on the line segment, the smaller value of the distance between the point and the two end points of the line segment is the distance to be calculated.

[0094] The distance between the line segment and the line segment is to parameterize the line segment and use the least square method to solve two parameters to determine the distance between the nearest points on the line segment, or to check the distance between the end points of the line segment.

[0095] In S101, a reward function based on inverse reinforcement learning processing is obtained, including:

[0096] Record data of any object operation surgical robot arm is obtained, and the record data includes historical trajectory data;

[0097] The historical trajectory data is preprocessed;

[0098] Based on the inverse reinforcement learning algorithm, the corresponding reward function is inferred from the preprocessed historical trajectory data;

[0099] The reward function is verified based on the reinforcement learning algorithm, and the final reward function is determined after verification.

[0100] In an embodiment, the robot arm is a six-axis robot arm, and the historical trajectory data is the angle change data of the corresponding six motors of the six-axis robot arm over time.

[0101] In an embodiment, the record data further includes operating environment data; and the historical trajectory data is the motion data of the robot arm under compliant control.

[0102] In an embodiment, the reward function corresponding to the preprocessed historical trajectory data is inferred based on the inverse reinforcement learning algorithm, including:

[0103] Based on the operating environment data, a simulation space is constructed;

[0104] Based on the operating environment data and the historical trajectory data, a state space is constructed;

[0105] An action space is set, and the action space includes a plurality of joint axes, each joint axis having a forward rotation action and a reverse rotation action;

[0106] Based on the state space and the action space, a complete operating trajectory is constructed;

[0107] Construct a reward function model;

[0108] By modeling the maximum entropy of the trajectory probability, a maximum entropy model is obtained;

[0109] By maximizing the log-likelihood of the expert trajectory, the parameters of the reward function model are iteratively optimized to obtain the optimal reward function parameters.

[0110] In one implementation, the reward function model is:

[0111] R(s t ,a t )=α·exp(-‖p t -g‖ 2 )-β·‖Δq t || 2 -γ·exp(‖u t || 2 )

[0112] Where R is the reward value, a t For the t-th action, s t Let p be the t-th state, where α, β, and γ are weighting coefficients. t Let g be the position of the robotic arm's end effector, g be the target position, and Δq be the position of the robotic arm's end effector. t u is the t-th comprehensive change value of the joint angle. t This is the joint angle vector.

[0113] In one implementation, the maximum entropy model is:

[0114]

[0115] Where τ is the trajectory. For possible trajectories, R(s) t ,a t ) to perform action a t Get state s t The reward value is a normalized function. To perform the action Get the state The reward value, Let P(τ) be the possible state and possible action corresponding to the t-th state and action, and let P(τ) be the probability of trajectory τ.

[0116] In one implementation, the preprocessing of the historical trajectory data includes:

[0117] The unit time point is determined based on the maximum value of the angle change data;

[0118] Segmenting angle change data based on unit time points;

[0119] Based on the segmented angle change data, the corresponding motor angle value at each unit time point is determined.

[0120] The historical trajectory data includes the angle change data of the motor over time. Based on this angle change data, a unit of measurement is first selected, and the angle change data of all joint motors over time are statistically analyzed. The shortest time required for the joint motor to change by one unit (this unit can be selected according to the actual situation or set by the user to facilitate accuracy constraints) is calculated. This shortest time is taken as the unit time point. In this way, the angle change of all motors within the unit time point does not exceed 1.

[0121] After determining the unit time point, the angle change data is segmented with the unit time point as the duration; each unit time point corresponds to an angle change data and a current angle value (determined by the angle value at the start or end of the time point); the current angle value corresponding to each time point is processed by approximation or rounding down to obtain an approximate angle value. The approximate angle value is an integer, and the angle values ​​of adjacent time points differ by ±1 or 0.

[0122] In this way, after preprocessing, it is easy to obtain the corresponding action and state of each joint motor.

[0123] In this application, by preprocessing historical trajectory data, the continuous joint motor movements are converted into discrete movements that are easy to identify and use, thereby greatly reducing the number of movements.

[0124] It should be noted that in this application, if the motor motion of each joint of the six-axis robotic arm is set to three actions (±1 or 0), then the total number of possible combinations of motions from the six motors is 3 to the power of 6. This results in an excessively large motion space in both reinforcement learning and inverse reinforcement learning, requiring a longer training time.

[0125] In one implementation, the unit time point is divided into six virtual time points, and the motor actions from the first joint to the end joint are executed respectively.

[0126] In other words, at the first virtual time point, the motor action of the first joint is executed (this motor action corresponds to the action at the unit time point), but the motors of other joints remain unchanged; other virtual time points remain unchanged. In this way, each virtual time point corresponds to three actions, and six virtual time points correspond to 18 actions, thereby greatly reducing the combination of motion space.

[0127] Preferably, after dividing each unit time point into six virtual time points, the action corresponding to each virtual time point is determined, and the state corresponding to each virtual time point is generated through the action and state.

[0128] In the present application, the state can be divided into an operating environment state and a robot arm state, and after each unit time point is divided into six virtual time points, the operating environment state remains unchanged (is set to be unchanged), and the robot arm state changes correspondingly with the motor action, so that the changed robot arm state and the unchanged operating environment state are combined to form a new state corresponding to the virtual time point.

[0129] Preferably, the reward function is determined based on the pre-processed inverse reinforcement learning method, and after the new action sequence is obtained by training the new reinforcement learning method based on the reward function, the actions corresponding to the six consecutive virtual time points are combined into an action corresponding to a unit time point, so that the actions of the robot arm are coherent.

[0130] The embodiment of the present application provides a surgical robot control device based on embodied intelligence, which is used to execute the surgical robot control method based on embodied intelligence described in the foregoing content of the present application. The surgical robot control device based on embodied intelligence is described in detail as follows.

[0131] As shown in Figure 3 The surgical robot control device based on embodied intelligence comprises:

[0132] The reward acquisition module 101 is configured to acquire a reward function based on inverse reinforcement learning processing.

[0133] The difference determination module 102 is configured to determine a constraint condition based on the difference between the past experience scene and the application scene.

[0134] The strategy training module 103 is configured to train the operation strategy of the surgical robot arm based on multi-objective optimization through reinforcement learning.

[0135] The robot arm control module 104 is configured to select an operation trajectory with the highest cumulative reward after the training is completed, and control the surgical robot arm to execute the operation trajectory. The multi-objective optimization comprises a plurality of objective functions, and the objective functions at least include the reward function and the constraint condition.

[0136] In an implementation manner, the difference determination module 102 is further configured to:

[0137] acquire past experience scene information and application scene information; determine a constraint direction of the past scene based on the past experience scene information; determine a constraint direction of the application scene based on the application scene information; determine a difference direction according to the constraint directions of the past scene and the application scene, the difference direction belongs to the application scene and does not belong to the past scene; and determine the constraint condition according to the difference direction.

[0138] In an implementation, the constraints include obstacle constraints, and / or, safety boundary constraints, and / or, end tool orientation constraints, and / or, mechanical constraints.

[0139] In an implementation, the policy training module 103 is further configured to:

[0140] obtain running environment data and planning data, construct a simulation space; based on an inverse reinforcement learning process and the running environment data, planning data, construct a state space and an action space; the action space contains a plurality of joint axes, each joint axis has a forward rotation action and a reverse rotation action; construct a deep Q network model; based on the obtained reward function, the deep Q network learning is used to perform first training on the deep Q network model; the differential constraint condition is used as the reward function, and the deep Q network learning is used to perform second training on the deep Q network model; the first training and the second training are alternately performed until a preset condition is met.

[0141] In an implementation, the policy training module 103 is further configured to:

[0142] construct a reward function according to the differential constraint condition; construct an experience replay pool, the experience replay pool includes actions, states, rewards, next states of the agent; sample from the experience replay pool, and calculate Q values based on the deep Q network model; update the weights of the deep Q network model by minimizing the loss function until the training is completed.

[0143] In an implementation, the policy training module 103 is further configured to:

[0144] obtain the current state of the robot arm pose data; obtain the current state of the environment data; determine the minimum distance between the robot arm and the environmental obstacles according to the robot arm pose data and the environment data; in the case that the minimum distance between the robot arm and the environmental obstacles is greater than the preset distance, it is determined that the obstacle constraint is satisfied.

[0145] In an implementation, the policy training module 103 is further configured to:

[0146] equivalent processing of the robot arm pose data, equivalent to the robot arm as a plurality of line segments and a preset thickness connected in turn; equivalent processing of the environment data, equivalent to the obstacles in the environment as a sphere and a cylinder wrapping the obstacles; calculate the distance between each line segment of the robot arm and the sphere center, the cylinder axis segment; according to the calculated distance, the preset thickness, the radius of the sphere, the radius of the cylinder, determine the minimum distance.

[0147] The above-mentioned embodiments of the present application provide a surgical robot control device based on embodied intelligence, which has a corresponding relationship with the surgical robot control method based on embodied intelligence provided by the embodiments of the present application. Therefore, the specific contents in the device have a corresponding relationship with the surgical robot control method based on embodied intelligence, and the specific contents can be referred to the records in the surgical robot control method based on embodied intelligence. Here, the specific contents will not be described again.

[0148] The above-mentioned embodiments of the present application provide a surgical robot control device based on embodied intelligence, which has the same beneficial effects as the method adopted, run or implemented by the application program stored therein, based on the same inventive concept as the surgical robot control method based on embodied intelligence provided by the embodiments of the present application.

[0149] The above describes the internal functions and structures of the surgical robot control device based on embodied intelligence. As shown in Figure 4 In practice, the surgical robot control device based on embodied intelligence can be implemented as an electronic device, which includes a memory 301 and a processor 303.

[0150] The memory 301 can be configured to store programs.

[0151] In addition, the memory 301 can also be configured to store other various data to support the operation on the electronic device. Examples of these data include instructions for any application or method operating on the electronic device, contact data, phonebook data, messages, pictures, videos, etc.

[0152] The memory 301 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk or optical disk.

[0153] The processor 303 is coupled to the memory 301 and is configured to execute the programs in the memory 301 for:

[0154] Obtaining a reward function based on inverse reinforcement learning processing;

[0155] Determining a constraint condition based on the difference between the past experience scene and the application scene;

[0156] Training the operation strategy of the surgical robot mechanical arm through reinforcement learning based on multi-objective optimization;

[0157] After the training is completed, a running track with the highest cumulative reward is selected, and the surgical robot manipulator is controlled to execute the running track; the multi-objective optimization includes multiple objective functions, and the objective functions at least include the reward function and the constraint condition.

[0158] In an embodiment, the processor 303 is further configured to:

[0159] acquire past experience scene information and application scene information; determine a constraint direction of a past scene based on the past experience scene information; determine a constraint direction of an application scene based on the application scene information; determine a difference direction according to the constraint directions of the past scene and the application scene, the difference direction belonging to the application scene and not belonging to the past scene; and determine a constraint condition according to the difference direction.

[0160] In an embodiment, the constraint condition includes an obstacle constraint, and / or a safety boundary constraint, and / or an end tool orientation constraint, and / or a mechanical constraint.

[0161] In an embodiment, the processor 303 is further configured to:

[0162] acquire running environment data and planning data, and construct a simulation space; construct a state space and an action space based on an inverse reinforcement learning process and the running environment data and the planning data; the action space includes multiple joint axes, each joint axis having a forward rotation action and a reverse rotation action; construct a deep Q network model; perform a first training on the deep Q network model based on deep Q network learning with a reward function acquired; perform a second training on the deep Q network model based on deep Q network learning with a difference constraint condition as the reward function; alternately perform the first training and the second training until a preset condition is met.

[0163] In an embodiment, the processor 303 is further configured to:

[0164] construct a reward function according to the difference constraint condition; construct an experience replay pool, the experience replay pool including actions, states, rewards, and next states of an agent; sample from the experience replay pool and calculate Q values based on the deep Q network model; and update weights of the deep Q network model by minimizing a loss function until the training is completed.

[0165] In an embodiment, the processor 303 is further configured to:

[0166] acquire manipulator pose data in a current state; acquire environment data in the current state; determine a minimum distance between the manipulator and an environmental obstacle according to the manipulator pose data and the environment data; and determine that an obstacle constraint is met when the minimum distance between the manipulator and the environmental obstacle is greater than a preset distance.

[0167] In an embodiment, the processor 303 is further configured to:

[0168] The mechanical arm pose data is equivalently processed to equivalently connect a plurality of line segments and a preset thickness; the environment data is equivalently processed to equivalently connect obstacles in the environment to a sphere and a cylinder; the distance between each line segment of the mechanical arm and the sphere center and the cylinder axis segment is calculated; and the minimum distance is determined according to the calculated distance, the preset thickness, the radius of the sphere, and the radius of the cylinder.

[0169] In the present application, the processor is further configured to perform all the processes and steps of the above-mentioned surgical robot control method based on embodied intelligence, and the specific content can be referred to the record in the surgical robot control method based on embodied intelligence. In the present application, this will not be described again.

[0170] In the present application, Figure 4 The electronic device shown in FIG. 1 only shows some components, and does not mean that the electronic device only includes Figure 4 The components shown in FIG. 1.

[0171] The electronic device provided in the embodiment has the same beneficial effects as the method adopted, run or implemented by the application program stored therein, based on the same inventive concept as the surgical robot control method based on embodied intelligence provided in the embodiment of the present application.

[0172] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer usable program code.

[0173] The present application is described with reference to flowcharts and / or block diagrams according to the method, device (system), and computer program product of the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of the flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device produce a device that implements the functions specified in the flowcharts and / or block diagrams. Figure 1 The functions specified in one flow or a plurality of flows and / or blocks Figure 1 The device that implements the functions specified in one flow or a plurality of flows and / or blocks

[0174] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0175] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0176] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0177] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0178] This application also provides a computer-readable storage medium corresponding to the embodied intelligence-based surgical robot control method provided in the foregoing embodiments, wherein a computer program (i.e., a program product) is stored thereon, and the computer program, when run by a processor, executes the embodied intelligence-based surgical robot control method provided in any of the foregoing embodiments.

[0179] Computer-readable media includes permanent and non-permanent, movable and non-movable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device. According to the definition herein, computer-readable media does not include transitory media such as modulated data signals and carriers.

[0180] The computer-readable storage medium provided by the above embodiments of the present application has the same beneficial effects as the method adopted, run or implemented by the application program stored therein.

[0181] It should be noted that in the specification provided herein, a large number of specific details are explained. However, it can be understood that the embodiments of the present application can be practiced without these specific details. In some examples, well-known structures and techniques are not shown in detail in order not to obscure the understanding of the present specification.

[0182] It should also be noted that the term "include", "contain" or any other variant thereof is intended to cover non-exclusive inclusion, so that the process, method, product or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, product or device. Without more limitations, the element defined by the statement "including a" does not exclude the presence of another identical element in the process, method, product or device including the element.

[0183] The above only describes the embodiments of the present application and is not intended to limit the present application. The present application can have various modifications and changes for those skilled in the art. Any modification, equivalent replacement, improvement, etc. within the spirit and principle of the present application shall be included in the scope of claims of the present application.

Claims

1. A body-aware surgical robot control method, characterized by, The method comprises the following steps: obtaining a reward function based on inverse reinforcement learning processing; determining a constraint condition based on the difference between past experience scenarios and application scenarios; training the operation strategy of the surgical robot manipulator through reinforcement learning based on multi-objective optimization; after training, selecting the operation trajectory with the highest cumulative reward to control the surgical robot manipulator to execute the operation trajectory; the multi-objective optimization includes multiple objective functions, and the objective functions at least include the reward function and the constraint condition; the training of the operation strategy of the surgical robot manipulator through reinforcement learning based on multi-objective optimization comprises the following steps: obtaining operation environment data and planning data to construct a simulation space; constructing a state space and an action space based on the operation environment data, the planning data, and an inverse reinforcement learning process; the action space includes multiple joint axes, and each joint axis has a forward rotation action and a reverse rotation action; constructing a deep Q network model; first training the deep Q network model based on deep Q network learning with the obtained reward function; second training the deep Q network model based on deep Q network learning with the constraint condition as the reward function; alternately performing the first training and the second training until a preset condition is met.

2. The embodiment according to claim 1, wherein, The determination of the constraint condition based on the difference between past experience scenarios and application scenarios comprises the following steps: obtaining past experience scenario information and application scenario information; determining the constraint direction of the past scenarios based on the past experience scenario information; determining the constraint direction of the application scenarios based on the application scenario information; determining a difference direction according to the constraint directions of the past scenarios and the application scenarios, wherein the difference direction belongs to the application scenarios and does not belong to the past scenarios; determining the constraint condition according to the difference direction. 3.The embodiment of the present application according to claim 2, wherein, The constraint condition includes an obstacle constraint, and / or a safety boundary constraint, and / or an end tool orientation constraint, and / or a mechanical constraint.

4. The embodiment 3, wherein The second training of the deep Q network model based on deep Q network learning with the constraint condition as the reward function comprises the following steps: constructing a reward function according to the constraint condition; constructing an experience replay pool, wherein the experience replay pool includes the actions, states, rewards, and next states of an agent; sampling from the experience replay pool and calculating Q values based on the deep Q network model; updating the weights of the deep Q network model by minimizing a loss function until the training is completed.

5. The embodiment according to claim 4, characterized in that, The calculation process of the obstacle constraint is as follows: obtaining manipulator pose data in a current state; obtaining environment data in the current state; determining the minimum distance between the manipulator and environmental obstacles according to the manipulator pose data and the environment data; determining that the obstacle constraint is met when the minimum distance between the manipulator and the environmental obstacles is greater than a preset distance. 6.The embodiment of the present application according to claim 5, wherein, The determination of the minimum distance between the manipulator and the environmental obstacles according to the manipulator pose data and the environment data comprises the following steps: equivalently processing the manipulator pose data to equivalently connect multiple line segments and a preset thickness; equivalently processing the environment data to equivalently connect the obstacles in the environment into spheres and cylinders; calculating the distance between each line segment of the manipulator and the sphere center and the cylinder axis segment; Determine the minimum distance according to the calculated distance, the preset thickness, the radius of the sphere, and the radius of the cylinder.

7. A body-aware surgical robot control apparatus, comprising: Comprise: The reward acquisition module is used for acquiring a reward function based on inverse reinforcement learning processing; The difference determination module is used for determining a constraint condition based on the difference between the past experience scene and the application scene; The strategy training module is used for training the operation strategy of the surgical robot mechanical arm through reinforcement learning based on multi-objective optimization; The mechanical arm control module is used for selecting the operation trajectory with the highest cumulative reward after the training is completed, and controlling the surgical robot mechanical arm to execute the operation trajectory; the multi-objective optimization includes multiple objective functions, and the objective functions at least include the reward function and the constraint condition; The strategy training module is also used for: acquiring operation environment data and planning data, and constructing a simulation space; Based on the inverse reinforcement learning process and the operation environment data and the planning data, a state space and an action space are constructed; the action space includes multiple joint axes, each joint axis has a forward rotation action and a reverse rotation action; a deep Q network model is constructed; the deep Q network learning is used to perform first training on the deep Q network model based on the acquired reward function; the deep Q network learning is used to perform second training on the deep Q network model based on the constraint condition of the difference as the reward function; the first training and the second training are alternately executed until a preset condition is met.

8. An electronic device, comprising: Comprise: Memory and processor; The memory is used for storing programs; The processor is coupled to the memory and is used for executing the programs, so as to: Acquire a reward function based on inverse reinforcement learning processing; Determine a constraint condition based on the difference between the past experience scene and the application scene; Train the operation strategy of the surgical robot mechanical arm through reinforcement learning based on multi-objective optimization; After the training is completed, select the operation trajectory with the highest cumulative reward, and control the surgical robot mechanical arm to execute the operation trajectory; the multi-objective optimization includes multiple objective functions, and the objective functions at least include the reward function and the constraint condition; The training of the operation strategy of the surgical robot mechanical arm through reinforcement learning based on multi-objective optimization comprises: Acquire operation environment data and planning data, and construct a simulation space; Based on the inverse reinforcement learning process and the operation environment data and the planning data, a state space and an action space are constructed; the action space includes multiple joint axes, each joint axis has a forward rotation action and a reverse rotation action; Construct a deep Q network model; The deep Q network learning is used to perform first training on the deep Q network model based on the acquired reward function; The deep Q network learning is used to perform second training on the deep Q network model based on the constraint condition of the difference as the reward function; The first training and the second training are alternately executed until a preset condition is met.

9. A computer readable storage medium having stored thereon a computer program, characterized in that, The programs are executed by the processor to realize the surgical robot control method based on embodied intelligence according to any one of claims 1-6.

Citation Information

Patent Citations

  • Artificial intelligence system for efficiently learning robotic control policies

    US10926408B1