Surgical robot control inverse reinforcement learning method and device based on embodied intelligence

By summarizing the historical trajectory data of the surgical robot through inverse reinforcement learning, the reward function is inferred and verified, which solves the problem of difficulty in summarizing past experience in existing technologies and realizes efficient control and smooth movement of the robotic arm.

CN119388413BActive Publication Date: 2025-12-26LONGWOOD VALLEY MEDICAL TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411210950.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-30
Publication Date
2025-12-26
Estimated Expiration
2044-08-30

AI Technical Summary

Technical Problem

Existing technologies cannot effectively summarize past experience in operating surgical robots, making it difficult to improve the control accuracy and efficiency of robotic arms.

Method used

By employing inverse reinforcement learning, historical trajectory data of the surgical robot arm is acquired, preprocessed, and then a reward function is inferred based on the inverse reinforcement learning algorithm. This function is then verified through reinforcement learning, and the optimal reward function is finally determined to train the robot arm's action execution strategy.

Benefits of technology

This represents a perfect summary of past experience, improving the control accuracy and efficiency of the robotic arm and ensuring efficient iteration and smooth movement of the robotic arm in application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119388413B_ABST
    Figure CN119388413B_ABST
Patent Text Reader

Abstract

The application provides a surgery robot control inverse reinforcement learning method and device based on embodied intelligence, the method comprises the following steps: acquiring record data of any object operating a surgery robot mechanical arm, the record data comprises historical trajectory data; preprocessing the historical trajectory data; based on an inverse reinforcement learning algorithm, inferring a corresponding reward function from the preprocessed historical trajectory data; verifying the reward function based on a reinforcement learning algorithm, and determining a final reward function after verification. In the application, the historical record of the surgery robot is preprocessed, and the corresponding reward function is inferred from the processed historical record based on the inverse reinforcement learning method.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of image processing, in particular to a surgical robot control inverse reinforcement learning method and device based on embodied intelligence. BACKGROUND

[0002] Embodied intelligence refers to an intelligent agent with a body and supporting interaction with the physical world, such as robots, unmanned vehicles, etc. By processing multiple sensor data inputs, a motion instruction is generated by a control center, such as a large model, to drive the intelligent agent, replacing the traditional rule-based or mathematical formula-based motion driving method, and realizing the deep integration of virtual and reality.

[0003] Surgical robots can be a perfect carrier of embodied intelligence, and the functions of surgical robots can be realized based on embodied intelligence. In the specific implementation process, the operation experience of the past surgical robots can be summarized first, and then a new control center can be generated according to the summarized experience.

[0004] However, how to summarize the past surgical robots is a current difficult problem to solve. SUMMARY

[0005] The problem solved by the present application is that it is currently very difficult to summarize the past surgical robots.

[0006] To solve the above problems, the first aspect of the present application provides a surgical robot control inverse reinforcement learning method based on embodied intelligence, comprising:

[0007] Obtaining recorded data of any object operating a surgical robot manipulator, the recorded data comprising historical trajectory data;

[0008] Preprocessing the historical trajectory data;

[0009] Based on an inverse reinforcement learning algorithm, an corresponding reward function is inferred from the preprocessed historical trajectory data;

[0010] Verifying the reward function based on a reinforcement learning algorithm, and determining a final reward function after verification.

[0011] The second aspect of the present application provides a surgical robot control inverse reinforcement learning device based on embodied intelligence, comprising:

[0012] A data acquisition module for acquiring recorded data of any object operating a surgical robot manipulator, the recorded data comprising historical trajectory data;

[0013] A data preprocessing module for preprocessing the historical trajectory data;

[0014] a reward estimation module configured to estimate a corresponding reward function from the preprocessed historical trajectory data based on a inverse reinforcement learning algorithm;

[0015] a reward verification module configured to verify the reward function based on a reinforcement learning algorithm, and determine a final reward function after verification.

[0016] A third aspect of the present application provides an electronic device, comprising a memory and a processor;

[0017] the memory is configured to store a program;

[0018] the processor is coupled to the memory and configured to execute the program, so as to:

[0019] obtain recorded data of any object operation surgical robot manipulator, the recorded data comprising historical trajectory data;

[0020] preprocess the historical trajectory data;

[0021] estimate a corresponding reward function from the preprocessed historical trajectory data based on a inverse reinforcement learning algorithm;

[0022] verify the reward function based on a reinforcement learning algorithm, and determine a final reward function after verification.

[0023] A fourth aspect of the present application provides a computer readable storage medium having a computer program stored thereon, the program being executed by a processor to implement the above-mentioned surgical robot control inverse reinforcement learning method based on embodied intelligence.

[0024] In the present application, the historical records of the surgical robot are preprocessed, and then the corresponding reward function is estimated from the preprocessed historical records based on the inverse reinforcement learning method. The reward function is a summary of past experience. Thus, data generation from data to reward function is completed, and perfect summary in data aspect is realized. BRIEF DESCRIPTION OF DRAWINGS

[0025] Figure 1 FIG. 1 is a schematic diagram of a surgical robot manipulator joint motion according to the surgical robot control inverse reinforcement learning method based on embodied intelligence of embodiments of the present application;

[0026] Figure 2 FIG. 2 is a flowchart of the surgical robot control inverse reinforcement learning method based on embodied intelligence according to embodiments of the present application;

[0027] Figure 3 FIG. 3 is a flowchart of reward estimation of the surgical robot control inverse reinforcement learning method based on embodied intelligence according to embodiments of the present application;

[0028] Figure 4 This is a structural block diagram of an embodied intelligence-based surgical robot control inverse reinforcement learning device according to an embodiment of this application;

[0029] Figure 5 This is a structural block diagram of an electronic device according to an embodiment of this application. Detailed Implementation

[0030] To make the above-mentioned objects, features, and advantages of this application more apparent and understandable, specific embodiments of this application will be described in detail below with reference to the accompanying drawings. Although exemplary embodiments of this application are shown in the drawings, it should be understood that this application can be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of this application and to fully convey the scope of this application to those skilled in the art.

[0031] It should be noted that, unless otherwise stated, the technical or scientific terms used in this application shall have the ordinary meaning as understood by one of ordinary skill in the art to which this application pertains.

[0032] To address the aforementioned issues, this application provides a novel inverse reinforcement learning scheme for surgical robot control, which solves the problem of the difficulty in summarizing past surgical robots through inverse reinforcement learning.

[0033] This application provides an embodiment of an inverse reinforcement learning method for controlling surgical robots based on embodied intelligence. The specific scheme of this method is as follows: Figures 1-3 As shown, this method can be executed by an embodied intelligence-based surgical robot-controlled inverse reinforcement learning device, which can be integrated into electronic devices such as computers, servers, computer clusters, and data centers. Combined with... Figure 1 , Figure 2 The diagram shows a flowchart of an inverse reinforcement learning method for surgical robot control based on embodied intelligence, according to an embodiment of this application; wherein the inverse reinforcement learning method for surgical robot control based on embodied intelligence includes:

[0034] S101, acquire the recorded data of any object operating the surgical robot arm, the recorded data including historical trajectory data;

[0035] In one embodiment, the robotic arm is a six-axis robotic arm, and the historical trajectory data is the angle change data of the six motors corresponding to the six-axis robotic arm over time.

[0036] In one implementation, the recorded data further includes operating environment data; the historical trajectory data is robotic arm motion data under compliant control.

[0037] Among them, compliant control is a control strategy mainly used in the operation of robots and mechanical arms to achieve soft interaction with the environment. It allows the robot to exhibit "compliance" when contacting uncertain or dynamically changing environments, thereby avoiding the application of excessive force or causing damage. In this application, the compliant control is the active control of the surgical robot arm under the operation of the expert doctor.

[0038] In this application, the operating environment data can include application scene data captured by a camera during the operation of the mechanical arm.

[0039] Preferably, the operating environment data further includes preoperative planning data of the mechanical arm.

[0040] In this way, based on the operating environment data, the experience of processing preoperative planning of expert data is learned, high consistency with experts is achieved, and the accuracy of mechanical arm control is improved.

[0041] S102, preprocessing the historical trajectory data;

[0042] S103, based on the inverse reinforcement learning algorithm, inferring the corresponding reward function from the preprocessed historical trajectory data;

[0043] S104, verifying the reward function based on the reinforcement learning algorithm, and determining the final reward function after verification.

[0044] In this application, the planning corresponding to the historical trajectory data is trained by reinforcement learning based on the reward function. If the trajectory data obtained by training has high similarity with the historical trajectory data, the verification is passed; or the trajectory data obtained by training is judged by the expert doctor. If it is feasible, the verification is passed. Otherwise, reiterate the inverse reinforcement learning.

[0045] In this application, the historical records of the surgical robot are preprocessed, and then the corresponding reward function is inferred from the preprocessed historical records based on the inverse reinforcement learning method. The reward function is a summary of past experience. Thus, data generation from data to reward function is completed, and perfect summary of data is achieved.

[0046] In one embodiment, after determining the final reward function, it further includes training the action execution strategy in the application scene based on the final reward function through reinforcement learning.

[0047] In this application, the final reward function is obtained by inverse reinforcement learning, and perfect summary of data is achieved. Then, reinforcement learning training is performed based on the final reward function, so as to apply the perfect summary of data to the application scene of the mechanical arm, and realize efficient iteration of the mechanical arm.

[0048] Preferably, in the present application, the motion sequence of the joint motor is determined according to the action execution strategy, and the corresponding mechanical arm end path is obtained through the motion sequence; the end path is the mechanical arm path to which the summarized experience is applied.

[0049] It should be noted that the application of the final reward function can be performed according to actual conditions, which will not be described in detail in the present application.

[0050] In an embodiment, the pre-processing of the historical trajectory data comprises:

[0051] determining a unit time point according to the maximum value of the angle change data;

[0052] segmenting the angle change data based on the unit time point;

[0053] determining the angle value of the corresponding motor at each unit time point based on the segmented angle change data.

[0054] In the historical trajectory data, the angle change data of the motor over time is included; based on the angle change data, the measurement unit is first selected, and the angle change data of all joint motors over time is counted, the shortest time required for the joint motor to change one unit (the unit can be selected according to actual conditions, or can be set by oneself to facilitate precision constraint) is determined as the unit time point, so that the angle change of all motors within the unit time point is not more than 1.

[0055] After determining the unit time point, the angle change data is segmented with the unit time point as the time length; each unit time corresponds to an angle change data and a current angle value (determined by the angle value at the starting point or the ending point of the time point); the current angle value corresponding to each time point is processed by approximation or downward rounding to obtain an approximate angle value, which is an integer, and the angle values of adjacent time points differ by ±1 or 0.

[0056] In this way, after pre-processing, the corresponding motion and state of each joint motor can be easily obtained.

[0057] In the present application, the continuous joint motor motion is converted into discrete motion which is easy to identify and use by pre-processing the historical trajectory data, greatly reducing the number of motions.

[0058] It should be noted that in the present application, if the motor motion of each joint of the six-axis mechanical arm is set to three motions of ±1 or 0, all motion combinations of the six motors are three six times. This will result in a too large action space in the entire reinforcement learning and inverse reinforcement learning, and a longer training time is required.

[0059] In an embodiment, each unit time point is divided into six virtual time points, and motor actions of the first joint to the end joint are executed respectively.

[0060] That is, the first virtual time point executes the motor action of the first joint (the motor action is the action of the corresponding unit time point), but the motors of other joints remain unchanged; other virtual time points remain unchanged. In this way, each virtual time point corresponds to three actions, and six virtual time points correspond to 18 actions, thereby greatly reducing the combination of action space.

[0061] Preferably, after each unit time point is divided into six virtual time points, the action corresponding to each virtual time point is determined, and the state corresponding to each virtual time point is generated through the action and the state.

[0062] In this application, the state can be divided into running environment state and robot arm state. After each unit time point is divided into six virtual time points, the running environment state remains unchanged (it is set to be unchanged), and the robot arm state changes correspondingly with the motor action, so that the changed robot arm state and the unchanged running environment state are combined into a new state corresponding to the virtual time point.

[0063] Preferably, the reward function is determined based on the pre-processed inverse reinforcement learning method, and after the new action sequence is trained by the new reinforcement learning method based on the reward function, the actions corresponding to the six consecutive virtual time points need to be combined into an action corresponding to a unit time point, so as to make the actions of the robot arm coherent.

[0064] In an embodiment, the method further comprises: Figure 3 as shown in the figure,

[0065] The S103 comprises:

[0066] S301, constructing a simulation space based on the running environment data;

[0067] In this application, the simulation space can be a virtual three-dimensional space, which is used for inverse reinforcement learning in the simulation space.

[0068] S302, constructing a state space based on the running environment data and the historical trajectory data;

[0069] In this application, the state of the robot arm can be represented by the angles of each joint. If the robot arm has n joints, the state space is n-dimensional. It can also include end effector position data and the like.

[0070] Preferably, the overall state of the robot arm includes the running environment state determined by the running environment data and the robot arm state determined by the historical trajectory data, so that the state of the robot arm is best simulated and the correlation between the state and the environment is maximized.

[0071] In the state space, the state of the robot arm can be represented as a vector.

[0072] S303, setting an action space, the action space containing a plurality of joint axes, each joint axis having a forward rotation action and a reverse rotation action;

[0073] In the action space, the robot arm action represents the change of the joint angle of the robot arm at each time step, which can also be a vector.

[0074] S304, based on the state space and the action space, constructing a complete running trajectory;

[0075] Wherein, the running trajectory is composed of a series of state and action sequences.

[0076] S305, constructing a reward function model;

[0077] Wherein, the reward function model is a nonlinear function, so that the model is iterated by adjusting the weight parameter in the subsequent iteration.

[0078] S306, modeling the maximum entropy of the trajectory probability to obtain a maximum entropy model;

[0079] In an embodiment, the maximum entropy model is:

[0080]

[0081] Wherein, τ is the trajectory, is the possible trajectory, R(s t ,a t ) is the reward value of the state s t obtained by executing the action a t , and is a normalization function, is the reward value of the state obtained by executing the action , and is the possible state and possible action corresponding to the tth state and action pair, and P(τ) is the probability of the trajectory τ.

[0082] In this application, the calculation of the maximum entropy involves the summation of all possible trajectories.

[0083] S307, iteratively optimizing the parameters of the reward function model by maximizing the log-likelihood of the expert trajectory to obtain the optimal reward function parameters.

[0084] In the present application, the calculation of the maximum entropy involves the summation of all possible trajectories. Therefore, the optimization is performed using the maximization of the log-likelihood of the expert trajectory.

[0085] where the log-likelihood of the expert trajectory is determined and optimized in the direction of the maximum value by gradient descent or gradient ascent until the optimization is completed.

[0086] In the present application, the specific optimization process of gradient descent or ascent is not described again.

[0087] In the present application, the expert trajectory is the complete running trajectory constructed.

[0088] In the present application, after learning the optimal reward function parameters, the reinforcement learning algorithm (such as value iteration, policy iteration) can be used to learn the optimal policy under the reward function.

[0089] Specifically, the reward function is used to optimize the motion trajectory of the robot arm to maximize the cumulative reward, thereby generating new trajectories. These new trajectories should be similar to the demonstration trajectory of the expert and can perform similar tasks.

[0090] In one embodiment, the reward function model is:

[0091] R(s t ,a t )=α·exp(-‖p t -g‖ 2 )-β·‖Δq t ‖ 2 -γ·exp(‖u t ‖ 2 )

[0092] where R is the reward value, a t is the tth action, s t is the tth state, α, β, γ are weight coefficients, p t is the position of the end effector of the robot arm, g is the target position, Δq t is the tth integrated change value of the joint angle, u t is the joint angle vector.

[0093] In the present application, through the reward function model, the target proximity, motion smoothness and energy efficiency in the motion of the robot arm are considered, and the robot arm is encouraged to approach the target object as much as possible while maintaining smooth motion and low energy consumption.

[0094] Target proximity: Encourage the end effector to approach the target object.

[0095] Smoothness of movement: Punishes drastic changes in joint angles to maintain the smoothness of the robotic arm's movements.

[0096] Energy efficiency: penalizes high energy consumption.

[0097] By optimizing the reward function, a robotic arm motion trajectory can be learned based on expert trajectories that can effectively approach the target object while maintaining smooth movements and conserving energy.

[0098] This application provides an embodiment of a surgical robot control inverse reinforcement learning device based on embodied intelligence, used to execute the surgical robot control inverse reinforcement learning method based on embodied intelligence described above. The following is a detailed description of the surgical robot control inverse reinforcement learning device based on embodied intelligence.

[0099] like Figure 4 As shown, the embodied intelligence-based surgical robot control inverse reinforcement learning device includes:

[0100] Data acquisition module 101 is used to acquire recorded data of any object operating the surgical robot arm, the recorded data including historical trajectory data;

[0101] Data preprocessing module 102 is used to preprocess the historical trajectory data;

[0102] The reward prediction module 103 is used to predict the corresponding reward function from the preprocessed historical trajectory data based on the inverse reinforcement learning algorithm;

[0103] The reward verification module 104 is used to verify the reward function based on the reinforcement learning algorithm, and determine the final reward function after the verification is successful.

[0104] In one embodiment, the robotic arm is a six-axis robotic arm, and the historical trajectory data is the angle change data of the six motors corresponding to the six-axis robotic arm over time.

[0105] In one embodiment, the data preprocessing module 102 is further configured to:

[0106] The unit time point is determined based on the maximum value of the angle change data; the angle change data is segmented based on the unit time point; and the angle value of the corresponding motor at each unit time point is determined based on the segmented angle change data.

[0107] In one implementation, the recorded data further includes operating environment data; the historical trajectory data is robotic arm motion data under compliant control.

[0108] In one implementation, the reward prediction module 103 is further configured to:

[0109] Based on the running environment data, an analog space is constructed; based on the running environment data and the historical trajectory data, a state space is constructed; an action space is set, the action space contains a plurality of joint shafts, each joint shaft has a forward rotation action and a reverse rotation action; based on the state space and the action space, a complete running trajectory is constructed; a reward function model is constructed; the maximum entropy of the trajectory probability is modeled to obtain a maximum entropy model; by maximizing the logarithmic likelihood of the expert trajectory, the parameters of the reward function model are iteratively optimized to obtain the optimal reward function parameters.

[0110] In an embodiment, the reward function model is:

[0111] R(s t ,a t )=α·exp(-‖p t -g‖ 2 )-β·‖Δq t ‖ 2 -γ·exp(‖u t ‖ 2 )

[0112] Wherein, R is a reward value, a t is the tth action, s t is the tth state, alpha, beta, gamma are weight coefficients, p t is the position of the end effector of the robot arm, g is the target position, Delta q t is the tth comprehensive change value of the joint angle, u t is the joint angle vector.

[0113] In an embodiment, the maximum entropy model is:

[0114]

[0115] Wherein, tau is a trajectory, is a possible trajectory, R(s t ,a t ) is a reward value obtained by executing action a t to state s t , is a normalization function, is a reward value obtained by executing action to state , is a possible state and a possible action corresponding to the tth state and action, P(tau) is the probability of trajectory tau.

[0116] The embodiment of the above-mentioned application provides the surgical robot control inverse reinforcement learning device based on the embodiment of the above-mentioned application and the surgical robot control inverse reinforcement learning method based on the embodiment of the above-mentioned application, and the specific content in the device has a corresponding relationship with the surgical robot control inverse reinforcement learning method based on the embodiment of the above-mentioned application, and the specific content can be referred to the record in the surgical robot control inverse reinforcement learning method based on the embodiment of the above-mentioned application, and the above-mentioned application will not be described here.

[0117] The embodiment of the above-mentioned application provides the surgical robot control inverse reinforcement learning device based on the embodiment of the above-mentioned application and the surgical robot control inverse reinforcement learning method based on the embodiment of the above-mentioned application, and the specific content in the device has a corresponding relationship with the surgical robot control inverse reinforcement learning method based on the embodiment of the above-mentioned application, and the specific content can be referred to the record in the surgical robot control inverse reinforcement learning method based on the embodiment of the above-mentioned application, and the above-mentioned application will not be described here.

[0118] The above describes the internal function and structure of the surgical robot control inverse reinforcement learning device based on the embodiment of the above-mentioned application, as shown in Figure 5 The surgical robot control inverse reinforcement learning device based on the embodiment of the above-mentioned application can be realized as an electronic device, including a memory 301 and a processor 303.

[0119] The memory 301 can be configured to store programs.

[0120] In addition, the memory 301 can also be configured to store other various data to support the operation on the electronic device. Examples of these data include instructions for any application or method operating on the electronic device, contact data, phonebook data, messages, pictures, videos, etc.

[0121] The memory 301 can be realized by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk or optical disk.

[0122] The processor 303 is coupled to the memory 301 and is used to execute the program in the memory 301, so as to:

[0123] Obtain the record data of any object operating the surgical robot mechanical arm, and the record data includes historical trajectory data;

[0124] The historical trajectory data is preprocessed;

[0125] Based on the inverse reinforcement learning algorithm, the corresponding reward function is inferred from the preprocessed historical trajectory data;

[0126] The reward function is verified based on a reinforcement learning algorithm, and after verification, a final reward function is determined.

[0127] In an embodiment, the mechanical arm is a six-axis mechanical arm, and the historical trajectory data is angle change data of six motors corresponding to the six-axis mechanical arm over time.

[0128] In an embodiment, the processor 303 is further configured to:

[0129] A unit time point is determined according to a maximum value of the angle change data, the angle change data is segmented based on the unit time point, and an angle value of the corresponding motor at each unit time point is determined based on the segmented angle change data.

[0130] In an embodiment, the record data further includes operating environment data, and the historical trajectory data is motion data of the mechanical arm under compliant control.

[0131] In an embodiment, the processor 303 is further configured to:

[0132] An analog space is constructed based on the operating environment data, a state space is constructed based on the operating environment data and the historical trajectory data, an action space is set, the action space includes a plurality of joint axes, each joint axis has a forward rotation action and a reverse rotation action, a complete operating trajectory is constructed based on the state space and the action space, a reward function model is constructed, a maximum entropy of the trajectory probability is modeled to obtain a maximum entropy model, and parameters of the reward function model are iteratively optimized by maximizing a logarithmic likelihood of the expert trajectory to obtain optimal reward function parameters.

[0133] In an embodiment, the reward function model is:

[0134] R(s t ,a t )=α·exp(-‖p t -g‖ 2 )-β·‖Δq t ‖ 2 -γ·exp(‖u t ‖ 2 )

[0135] wherein R is a reward value, a t is the tth action, s t is the tth state, α, β, γ are weight coefficients, p t is a position of an end effector of the mechanical arm, g is a target position, Δq t is a tth integrated change value of joint angles, and u t is a joint angle vector.

[0136] In one implementation, the maximum entropy model is:

[0137]

[0138]

[0139] Where τ is the trajectory. For possible trajectories, R(s) t ,a t ) to perform action a t Get state s t The reward value is a normalized function. To perform the action Get the state The reward value, Let P(τ) be the possible state and possible action corresponding to the t-th state and action, and let P(τ) be the probability of trajectory τ.

[0140] In this application, the processor is also specifically used to execute all the processes and steps of the above-mentioned inverse reinforcement learning method for surgical robot control based on embodied intelligence. For details, please refer to the records in the inverse reinforcement learning method for surgical robot control based on embodied intelligence. This application will not repeat them here.

[0141] In this application, Figure 5 The diagram only shows some components and does not mean that the electronic device includes only these components. Figure 5 The components shown.

[0142] The electronic device provided in this embodiment is based on the same inventive concept as the embodied intelligence-based surgical robot control inverse reinforcement learning method provided in this application embodiment, and has the same beneficial effects as the methods adopted, run or implemented by the application stored therein.

[0143] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-readable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0144] The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. Figure 1 The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. Figure 1 The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks.

[0145] The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. Figure 1 The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. Figure 1 The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks.

[0146] The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. Figure 1 The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. Figure 1 The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks.

[0147] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.

[0148] The memory can include non-persistent memory, random access memory (RAM), and / or non-volatile memory, etc. in the form of a computer-readable medium, such as read only memory (ROM) or flash memory (Flash RAM). The memory is an example of computer-readable media.

[0149] The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks.

[0150] Computer-readable media includes permanent and non-permanent, movable and non-movable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device. According to the definition herein, computer-readable media does not include transitory media such as modulated data signals and carriers.

[0151] The computer-readable storage medium provided by the above embodiments of the present application has the same beneficial effects as the method adopted, run or implemented by the application program stored therein based on the somatic intelligence of the surgical robot control inverse reinforcement learning method.

[0152] It should be noted that in the specification provided herein, a large number of specific details are explained. However, it can be understood that the embodiments of the present application can be practiced without these specific details. In some examples, well-known structures and techniques are not shown in detail in order not to obscure the understanding of the present specification.

[0153] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or apparatus including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such a process, method, article or apparatus. Without more limitations, the element defined by the statement "comprising a" does not exclude the presence of additional identical elements in the process, method, article or apparatus including the element.

[0154] The above only describes the embodiments of the present application and is not intended to limit the present application. The present application can have various modifications and changes for those skilled in the art. Any modification, equivalent replacement, improvement, etc. within the spirit and principle of the present application shall be included in the scope of claims of the present application.

Claims

1. A control inverse reinforcement learning method for a surgical robot based on embodied intelligence, characterized in that, The method comprises the following steps: acquiring record data of any object operating surgical robot manipulator, the record data comprising historical trajectory data; preprocessing the historical trajectory data; speculating a corresponding reward function from the preprocessed historical trajectory data based on an inverse reinforcement learning algorithm; verifying the reward function based on a reinforcement learning algorithm, and determining a final reward function after verification; the manipulator is a six-axis manipulator, and the historical trajectory data is angle change data of six motors corresponding to the six-axis manipulator over time; the preprocessing of the historical trajectory data comprises the following steps: determining a unit time point according to the maximum value of the angle change data; segmenting the angle change data based on the unit time point; determining the angle value of the corresponding motor at each unit time point based on the segmented angle change data; secondly dividing the unit time point, and dividing each unit time point into six virtual time points to respectively execute motor actions of the first joint to the end joint.

2. The surgical robotic control inverse reinforcement learning method of claim 1, wherein, The record data further comprises operating environment data, and the historical trajectory data is manipulator motion data under compliant control.

3. The surgical robotic control inverse reinforcement learning method of claim 2, wherein, The speculation of the corresponding reward function from the preprocessed historical trajectory data based on the inverse reinforcement learning algorithm comprises the following steps: constructing a simulation space based on the operating environment data; constructing a state space based on the operating environment data and the historical trajectory data; setting an action space, the action space comprising a plurality of joint axes, each joint axis having a forward rotation action and a reverse rotation action; constructing a complete operating trajectory based on the state space and the action space; constructing a reward function model; modeling the maximum entropy of the trajectory probability to obtain a maximum entropy model; iteratively optimizing the parameters of the reward function model by maximizing the logarithmic likelihood of the expert trajectory to obtain optimal reward function parameters.

4. The surgical robotic control inverse reinforcement learning method of claim 3, wherein, The reward function model is: wherein R is a reward value, is a first action, is a first state, is a weight coefficient, is a position of an end effector of a robot arm, is a target position, is a first overall change value of a joint angle, is a joint angle vector.

5. The surgical robotic control inverse reinforcement learning method of claim 3, wherein, The maximum entropy model is: wherein, is a trajectory, is a possible trajectory, is performing an action results in a state is a reward value, is a normalization function, is performing an action results in a state is a reward value, is a possible state and possible action corresponding to the th state, action, is a trajectory probability.

6. A surgery robot control inverse reinforcement learning apparatus based on embodied intelligence, characterized by, The method comprises the following steps: a data acquisition module is configured to acquire record data of any object operating surgical robot manipulator, the record data comprising historical trajectory data; a data preprocessing module is configured to preprocess the historical trajectory data; a reward speculation module is configured to speculate a corresponding reward function from the preprocessed historical trajectory data based on an inverse reinforcement learning algorithm; a reward verification module is configured to verify the reward function based on a reinforcement learning algorithm, and determine a final reward function after verification; the manipulator is a six-axis manipulator, and the historical trajectory data is angle change data of six motors corresponding to the six-axis manipulator over time; the data preprocessing module is further configured to: determine a unit time point according to the maximum value of the angle change data; segment the angle change data based on the unit time point; determine the angle value of the corresponding motor at each unit time point based on the segmented angle change data; secondly divide the unit time point, and divide each unit time point into six virtual time points to respectively execute motor actions of the first joint to the end joint.

7. An electronic device, comprising: The method comprises the following steps: a memory and a processor; the memory is configured to store a program; the processor is coupled to the memory and is configured to execute the program, so as to: Acquire record data of any object operating surgical robot manipulator, the record data includes historical trajectory data; Preprocess the historical trajectory data; Based on the inverse reinforcement learning algorithm, the corresponding reward function is inferred from the preprocessed historical trajectory data; Based on the reinforcement learning algorithm, the reward function is verified, and the final reward function is determined after verification; The manipulator is a six-axis manipulator, and the historical trajectory data is the angle change data of the corresponding six motors of the six-axis manipulator over time; The preprocessing of the historical trajectory data comprises: Determine the unit time point according to the maximum value of the angle change data; Based on the unit time point, the angle change data is cut; Based on the cut angle change data, the angle value of the corresponding motor at each unit time point is determined; The unit time point is divided into six virtual time points, and the motor action of the first joint to the end joint is executed respectively.

8. A computer-readable storage medium having stored thereon a computer program, characterized in that, The program is executed by the processor to realize the inverse reinforcement learning method of the surgical robot control based on the embodiment intelligence according to any one of claims 1-5.

Citation Information

Patent Citations

  • Motion planning method based on inverse reinforcement learning

    CN118504808A