A multi-mode comprehensive control method and device for surgical robot body intelligence

By combining inverse reinforcement learning and reinforcement learning, a multi-reward function is generated to train the operating strategy of the surgical robot arm. This solves the problem of integrating historical operating experience and visual obstacle avoidance control mode, and realizes the comprehensive control and precise operation of the robot arm.

CN119655875BActive Publication Date: 2025-12-26LONGWOOD VALLEY MEDICAL TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411537120.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-31
Publication Date
2025-12-26
Estimated Expiration
2044-10-31

AI Technical Summary

Technical Problem

Existing technologies struggle to achieve comprehensive control of surgical robot arms, especially the integration of historical operational experience and visual obstacle avoidance control modes.

Method used

The first reward function is generated by inverse reinforcement learning. The second and third reward functions are constructed by combining the target position and obstacle information. The operation strategy of the robotic arm is trained by reinforcement learning, and the operation trajectory with the highest cumulative reward is selected for control.

Benefits of technology

The system achieves comprehensive control of the surgical robot arm, improving the accuracy and integration of control, and ensuring stable operation of the robot arm in different environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119655875B_ABST
    Figure CN119655875B_ABST
Patent Text Reader

Abstract

The application provides a multi-mode comprehensive control method and device for surgical robot body intelligence, the method comprising: obtaining a first reward function based on inverse reinforcement learning processing; obtaining target position and obstacle information; constructing a second reward function based on the target position and a third reward function based on the obstacle information; training the operation strategy of the mechanical arm through reinforcement learning according to the first reward function, the second reward function and the third reward function; selecting an operation trajectory with the highest cumulative reward and controlling the mechanical arm to execute the operation trajectory. In the application, the first reward function is generated through historical operation experience, the second reward function and the third reward function are generated through visual obstacle avoidance; the overall operation strategy based on different reward functions is realized through the reinforcement learning mode, thereby realizing the comprehensive control of the mechanical arm based on historical operation experience and visual obstacle avoidance.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of image processing, in particular to a multi-mode comprehensive control method and device for surgical robot embodied intelligence. BACKGROUND

[0002] Embodied intelligence refers to an intelligent agent with a body and supporting interaction with the physical world, such as robots, unmanned vehicles, etc. The intelligent agent is driven by motion instructions generated by a control center, such as a large model, by processing multiple sensor data inputs, replacing the traditional motion driving mode based on rules or mathematical formulas, and realizing the deep integration of virtual and reality.

[0003] Surgical robots can be a perfect carrier of embodied intelligence, and the functions of surgical robots can be realized based on embodied intelligence. In the specific implementation process, the operation experience of the past surgical robots can be summarized first, and then the mechanical arm of the surgical robot can be controlled in combination with the visual obstacle avoidance mode.

[0004] However, historical operation experience and visual obstacle avoidance are different control modes, and it is difficult to realize comprehensive control. SUMMARY

[0005] The problem solved by the present application is that it is currently difficult to realize comprehensive control of the mechanical arm.

[0006] To solve the above problems, the first aspect of the present application provides a multi-mode comprehensive control method for the mechanical arm of a surgical robot with embodied intelligence, comprising:

[0007] obtaining a first reward function based on inverse reinforcement learning processing;

[0008] obtaining target position and obstacle information, the obstacle information including static information and dynamic information;

[0009] constructing a second reward function based on the target position and a third reward function based on the obstacle information;

[0010] training the running strategy of the mechanical arm through reinforcement learning according to the first reward function, the second reward function and the third reward function;

[0011] After training, the running trajectory with the highest cumulative reward is selected, and the mechanical arm is controlled to execute the running trajectory.

[0012] The second aspect of the present application provides a multi-mode comprehensive control device for the mechanical arm of a surgical robot with embodied intelligence, comprising:

[0013] an inverse reinforcement learning module for obtaining a first reward function based on inverse reinforcement learning processing;

[0014] an obstacle obtaining module configured to obtain target position and obstacle information, the obstacle information including static information and dynamic information;

[0015] a reward constructing module configured to construct a second reward function based on the target position and a third reward function based on the obstacle information;

[0016] a reinforcement learning module configured to train a running strategy of the robot arm through reinforcement learning according to the first reward function, the second reward function and the third reward function;

[0017] a robot arm control module configured to select a running track with the highest cumulative reward after the training is completed, and control the robot arm to execute the running track.

[0018] The third aspect of the present application provides an electronic device, comprising a memory and a processor;

[0019] the memory is configured to store a program;

[0020] the processor is coupled to the memory and configured to execute the program, so as to:

[0021] obtain a first reward function based on inverse reinforcement learning processing;

[0022] obtain target position and obstacle information, the obstacle information including static information and dynamic information;

[0023] construct a second reward function based on the target position and a third reward function based on the obstacle information;

[0024] train a running strategy of the robot arm through reinforcement learning according to the first reward function, the second reward function and the third reward function;

[0025] select a running track with the highest cumulative reward after the training is completed, and control the robot arm to execute the running track.

[0026] The fourth aspect of the present application provides a computer readable storage medium, which stores a computer program, the program being executed by a processor to realize the above-mentioned robot arm multi-mode comprehensive control method for surgical robot body intelligence.

[0027] In the present application, the first reward function is generated through historical operation experience, the second reward function and the third reward function are generated through visual obstacle avoidance, and the overall running strategy based on different reward functions is realized through reinforcement learning, so as to realize the comprehensive control of the robot arm based on historical operation experience and visual obstacle avoidance. BRIEF DESCRIPTION OF DRAWINGS

[0028] Figure 1A flow chart of a multi-mode comprehensive control method for a surgical robot body intelligence mechanical arm according to an embodiment of the present application is shown in

[0029] Figure 2 A structural block diagram of a multi-mode comprehensive control device for a surgical robot body intelligence mechanical arm according to an embodiment of the present application is shown in

[0030] Figure 3 A structural block diagram of an electronic device according to an embodiment of the present application is shown in DETAILED DESCRIPTION

[0031] In order to make the above objectives, features and advantages of the present application more obvious and understandable, specific embodiments of the present application will be described in detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments described herein. On the contrary, these embodiments are provided so that the present application can be more thoroughly understood and the scope of the present application can be accurately conveyed to those skilled in the art.

[0032] It should be noted that, unless otherwise specified, the technical terms or scientific terms used in the present application should be understood as their usual meanings understood by those skilled in the art to which the present application belongs.

[0033] An embodiment of the present application provides a multi-mode comprehensive control method for a surgical robot body intelligence mechanical arm, and a specific scheme of the method is shown in Figure 1 The method can be executed by a multi-mode comprehensive control device for a surgical robot body intelligence mechanical arm, which can be integrated in a computer, a server, a computer cluster, a data center, etc. electronic device. As shown in Figure 1 A flow chart of a multi-mode comprehensive control method for a surgical robot body intelligence mechanical arm according to an embodiment of the present application is shown in

[0034] S101, obtaining a first reward function based on inverse reinforcement learning processing;

[0035] S102, obtaining target position and obstacle information, the obstacle information including static information and dynamic information;

[0036] S103, constructing a second reward function based on the target position and a third reward function based on the obstacle information;

[0037] S104, training the operation strategy of the mechanical arm through reinforcement learning according to the first reward function, the second reward function and the third reward function;

[0038] S105, after the training is completed, selecting a running track with the highest cumulative reward, and controlling the robot arm to execute the running track.

[0039] In the present application, after the training is completed, all possible running tracks are evaluated by using the trained running strategy, and the track with the highest cumulative reward is selected as the running track of the robot arm, and the robot arm is controlled to execute the track.

[0040] In the present application, the first reward function is generated by historical operation experience, and the second reward function and the third reward function are generated by visual obstacle avoidance; the overall running strategy based on different reward functions is realized by the way of reinforcement learning, so as to realize the comprehensive control of the robot arm based on historical operation experience and visual obstacle avoidance.

[0041] In an embodiment, the first reward function based on inverse reinforcement learning processing is obtained, comprising:

[0042] Obtaining record data of any object operating the robot arm, the record data comprising historical trajectory data;

[0043] Preprocessing the historical trajectory data;

[0044] Based on the inverse reinforcement learning algorithm, the corresponding reward function is inferred from the preprocessed historical trajectory data;

[0045] Based on the reinforcement learning algorithm, the reward function is verified, and the final reward function is determined after verification.

[0046] In an embodiment, the robot arm is a six-axis robot arm, and the historical trajectory data is the angle change data of the corresponding six motors of the six-axis robot arm over time.

[0047] In an embodiment, the record data further comprises running environment data; and the historical trajectory data is the motion data of the robot arm under compliant control.

[0048] In an embodiment, the inverse reinforcement learning algorithm is used to infer the corresponding reward function from the preprocessed historical trajectory data, comprising:

[0049] Based on the running environment data, an analog space is constructed;

[0050] Based on the running environment data and the historical trajectory data, a state space is constructed;

[0051] An action space is set, the action space comprising a plurality of joint axes, each joint axis having a forward rotation action and a reverse rotation action;

[0052] Based on the state space and the action space, a complete running track is constructed;

[0053] constructing a reward function model;

[0054] modeling maximum entropy of trajectory probability to obtain a maximum entropy model;

[0055] iteratively optimizing parameters of the reward function model by maximizing log-likelihood of expert trajectory to obtain optimal reward function parameters.

[0056] In an embodiment, the reward function model is:

[0057] R(s t ,a t )=α·exp(-‖p t -g‖ 2 )-β·‖Δq t ‖ 2 -γ·exp(‖u t ‖ 2 )

[0058] wherein R is a reward value, a t is the tth action, s t is the tth state, α, β, γ are weight coefficients, p t is the position of the end effector of the robot arm, g is the target position, Δq t is the tth integrated change value of joint angle, u t is the joint angle vector.

[0059] In an embodiment, the maximum entropy model is:

[0060]

[0061]

[0062] wherein τ is a trajectory, is a possible trajectory, R(s t ,a t ) is a reward value obtained by executing action a t to obtain state s t , is a normalization function, is a reward value obtained by executing action to obtain state , is a possible state and a possible action corresponding to the tth state and action, P(τ) is the probability of trajectory τ.

[0063] In an embodiment, the real surgical scene video stream is obtained and parsed to obtain the target position and obstacle information, comprising:

[0064] obtaining a surgical scene video stream and segmenting it into continuous depth image frames;

[0065] input the depth image frame into the obstacle identification model to obtain an obstacle identification result; the obstacle identification result comprises a static obstacle, a dynamic obstacle, and a target position;

[0066] input the continuous depth image frame labeled with the obstacle identification result into the space-time prediction model to obtain a spatial motion trajectory of the dynamic obstacle.

[0067] In an embodiment, the inputting of the continuous depth image frame labeled with the obstacle identification result into the space-time prediction model to obtain an obstacle prediction result comprises:

[0068] based on the obstacle identification result, labeling position key points of the dynamic obstacle in the depth image frame;

[0069] generating a continuous feature map according to the position key points of the continuous depth image frame;

[0070] inputting the continuous feature map into the space-time prediction model to obtain a spatial motion trajectory of the dynamic obstacle.

[0071] In an embodiment, the space-time prediction model is a gated dilated causal convolutional neural network.

[0072] In an embodiment, the generating of the continuous feature map according to the position key points of the continuous depth image frame comprises:

[0073] constructing a domain space graph structure according to the position key points;

[0074] constructing a weighted adjacency matrix of the domain space graph structure based on the correlation relationship between the position key points of the same dynamic obstacle;

[0075] generating a feature representation of the domain space graph structure at different time instants according to the domain space graph structure and the corresponding weighted adjacency matrix of the continuous depth image frame, i.e., the continuous feature map.

[0076] In an embodiment, the operation strategy of the robot arm is trained through reinforcement learning according to the first reward function, the second reward function, and the third reward function;

[0077] generating a first control mode reward, a second control mode reward, and a third control mode reward according to the first reward function, the second reward function, and the third reward function;

[0078] preliminarily training the first control mode reward, the second control mode reward, and the third control mode reward based on the reinforcement learning method;

[0079] constructing a switching threshold between the first control mode reward, the second control mode reward, and the third control mode reward;

[0080] According to the switching threshold, the first control mode reward, the second control mode reward and the third control mode reward are trained as a whole to obtain an operation strategy of the robot arm.

[0081] In the present application, the second reward function is constructed based on the target position, and the specific form can be the Euclidean distance between the coordinates of the end tool of the robot arm and the coordinates of the target position.

[0082] In the present application, the third reward function is constructed based on the obstacle information, and the specific form can be the minimum distance between the current end of the robot arm and the obstacle, or the minimum distance between the robot arm itself (including multiple joints) and the obstacle.

[0083] In the present application, the first reward function, the second reward function and the third reward function are not directly used as the reward functions of the three control modes, but the first control mode reward, the second control mode reward and the third control mode reward are generated by the first reward function, the second reward function and the third reward function. Since the first reward function, the second reward function and the third reward function are target-specific reward functions, directly using them as the reward functions of the control modes will make the differences between the three control modes larger, and it is difficult to integrate them.

[0084] In the present application, the first control mode reward, the second control mode reward and the third control mode reward are generated by the first reward function, the second reward function and the third reward function, so as to integrate the first reward function, the second reward function and the third reward function into the three control modes, greatly increase the similarity between the three control modes, and improve the coherence of the comprehensive control.

[0085] In the present application, the switching threshold between the first control mode reward, the second control mode reward and the third control mode reward is constructed, so as to realize the soft switching between the three control modes through the switching threshold, realize the precise training and reward accumulation, and greatly improve the accuracy of the comprehensive control.

[0086] In an embodiment, the first control mode reward is determined by the second reward function and the third reward function, and the second control mode reward and the third control mode reward are both determined by the first reward function, the second reward function and the third reward function.

[0087] In an embodiment, the generation formula of the first control mode reward is:

[0088] R1=αr1+r2

[0089] Wherein, R1 is the first control mode reward, r1 is the first reward function, r2 is the second reward function, and a is the generation coefficient.

[0090] In the present application, the generation coefficient is set before the first reward function to adjust the relative proportion of the first reward function and the second reward function, and to improve the generation effect.

[0091] In the present application, the generation coefficient a in the generation formula of the first control mode reward is preferably 0.3.

[0092] In an embodiment, the generation formula of the second control mode reward is:

[0093] R2 = a r1 + b r2 + r3

[0094] Wherein, R2 is the second control mode reward, r1 is the first reward function, r2 is the second reward function, r3 is the third reward function, and a and b are generation coefficients.

[0095] In an embodiment, the generation formula of the third control mode reward is:

[0096] R3 = r1 + b r2 + g r3

[0097] Wherein, R3 is the third control mode reward, r1 is the first reward function, r2 is the second reward function, r3 is the third reward function, and b and g are generation coefficients.

[0098] It should be noted that the generation formula of different control mode rewards has a generation coefficient, but their specific values are not the same. For example, the generation coefficient a in the generation formula of the first control mode reward is preferably 0.3, and the generation coefficient a in the generation formula of the second control mode reward is preferably 0.5.

[0099] It should be noted that the values of the generation coefficients a, b and g are in the range of 0-1.

[0100] In the present application, the second control mode reward is generated by maximizing the third reward function, and the third control mode reward is generated by maximizing the first reward function, so as to greatly improve the proportional difference between the three control mode rewards and maintain the similarities and differences of the three controls.

[0101] In an embodiment, the reinforcement learning method is used to preliminarily train the first control mode reward, comprising:

[0102] According to the target position and obstacle information, a simulation space is constructed;

[0103] According to the parameters of the robot arm, a state space and an action space are constructed; the action space contains a plurality of joint axes, and each joint axis has a forward rotation action and a reverse rotation action;

[0104] a deep Q-network model and a first experience pool are constructed; the first control mode reward, the second control mode reward, and the third control mode reward have corresponding first, second, and third experience pools;

[0105] The first control mode reward is used as the reward function to sample from the experience replay pool, and the Q value is calculated based on the deep Q-network model.

[0106] By minimizing the loss function, the weights of the deep Q-network model are updated until a predetermined number of iterations is reached.

[0107] In this application, the specific process of deep Q reinforcement learning and the corresponding setting of deep Q-network model can be referred to the prior art, which will not be repeated here.

[0108] In this application, a simulation space is constructed: according to the target position and obstacle information, a simulation environment is created for the training and testing of the robot arm. This can be a 3D simulation environment reflecting the task scenario in the real world.

[0109] In this application, the state space and action space are constructed:

[0110] State space: the state space is constructed according to the current state of the robot arm (e.g. joint angle, velocity, position) and environmental information (target position, obstacle position).

[0111] Action space: define the action space of the robot arm, which contains multiple joint axes. Each joint axis has two actions: forward rotation and reverse rotation, forming a discrete action set.

[0112] In this application, a deep Q-network model and an experience pool are constructed:

[0113] Deep Q-network model: design a deep Q-network (DQN) to estimate the Q value of taking each action in a given state.

[0114] Experience pool: create experience pools for the first control mode reward, the second control mode reward, and the third control mode reward. These experience pools are used to store the experiences of the agent for subsequent sampling and training.

[0115] In this application, sampling and calculating Q value:

[0116] Use the first control mode reward as the reward function to sample experience (state, action, reward, next state) from the first experience pool.

[0117] Based on the deep Q-network model, calculate the Q value to evaluate the expected return of taking a certain action in the current state.

[0118] In this application, the deep Q-network model is updated:

[0119] A loss function, typically mean squared error (MSE) of Q values, is defined to calculate the gap between the network predicted Q values and the target Q values.

[0120] The loss function is minimized by a backpropagation algorithm to update the weights of the deep Q network model.

[0121] The above steps are repeated until a predetermined number of iterations is reached.

[0122] It should be noted that in this application, the first control mode reward is preliminarily trained based on the reinforcement learning method, and the preliminary training of the second control mode and the third control mode is similar, and the difference is that the corresponding second experience pool and third experience pool are selected for training; and the corresponding loss function is different.

[0123] Preferably, in this application, a task mode is left as a buffer between obstacle avoidance (third reward function, second control mode) and other tasks, and in addition to switching modes by threshold, it can also be switched by reward accumulation threshold (that is, the switching threshold is an accumulation threshold, which is accumulated after each training calculation threshold, and switched after reaching a predetermined value). The reward accumulation threshold can leave sufficient buffer space to prevent obstacles from affecting other tasks to the greatest extent, and can quickly react and quickly switch to obstacle avoidance mode when obstacles arrive quickly.

[0124] In one embodiment, the switching threshold includes a first switching threshold and a second switching threshold; and the overall training of the first control mode reward, the second control mode reward, and the third control mode reward according to the switching threshold to obtain the operation strategy of the robot arm includes:

[0125] Training data is sampled from the first experience pool, the second experience pool and the third experience pool according to a preset proportion;

[0126] The first control mode reward is used as a reward function, and the weights of the deep Q network model are iterated according to the training data;

[0127] The current state of the robot arm is determined in real time;

[0128] When the current state of the robot arm meets the first switching threshold, the second control mode reward is used as a reward function, and the weights of the deep Q network model are iterated according to the training data;

[0129] When the current state of the robot arm meets the second switching threshold, the third control mode reward is used as a reward function, and the weights of the deep Q network model are iterated according to the training data;

[0130] In the case that the number of iterations reaches a preset number or the loss function converges, the overall training is completed, and the operation strategy of the robot arm is obtained.

[0131] In the present application, the first switching threshold and the second switching threshold can be threshold values calculated by a specific calculation formula, or accumulated threshold values obtained by accumulating numerical values calculated by a calculation formula.

[0132] In the present application, the weight of the deep Q network model is iterated according to the training data, taking the first control mode reward as the reward function, which is similar to the preliminary training of the first control mode reward based on the reinforcement learning method as described above, except that the number of iterations is limited by the switching threshold.

[0133] In the present application, the first switching threshold is used to determine when to switch from the first control mode to the second control mode.

[0134] In the present application, the second switching threshold is used to determine when to switch from the second control mode to the third control mode.

[0135] In the present application, each control mode is initially trained independently, and a reward function suitable for independent training is given. Subsequently, each control mode is migrated from a single scenario to a multi-switching threshold control overall scenario. In this way, the multi-mode comprehensive control has certain experience in the early stage, thereby serving as a preprocessing and preparing for the final configuration.

[0136] In the present application, the training data is sampled according to a preset proportion from the first, second, and third experience pools. Such sampling can ensure that the model considers the experience of different control modes during training.

[0137] In the present application, the first control mode reward is used as the reward function, and the weight of the deep Q network model is iterated according to the sampled training data. This step is to establish the model's learning of the initial control mode.

[0138] In the present application, in each iteration, the current state of the robot arm is monitored in real time to make a switching decision based on the current environment.

[0139] In the present application, when the current state of the robot arm meets the first switching threshold, the second control mode reward is switched to. The second control mode reward is used as the reward function, and the weight of the deep Q network model is iterated according to the training data.

[0140] In the present application, when the preset number of iterations or the loss function converges, the overall training is completed, and the optimized operation strategy of the robot arm is finally obtained.

[0141] It should be noted that in the present application, the loss function in the overall training process is obtained by adding the loss function proportion of the preliminary training of the first control mode, the second control mode and the third control mode reward based on the reinforcement learning method, so as to unify the loss function of the whole overall training, so as to realize the judgment of whether the loss converges.

[0142] The embodiment of the present application provides a multi-mode comprehensive control device for a surgical robot body intelligence, which is used to execute the multi-mode comprehensive control method for a surgical robot body intelligence as described above.

[0143] As shown in Figure 2 The multi-mode comprehensive control device for a surgical robot body intelligence comprises:

[0144] The inverse reinforcement learning module 101 is configured to obtain a first reward function based on inverse reinforcement learning processing.

[0145] The obstacle acquisition module 102 is configured to obtain target position and obstacle information, wherein the obstacle information comprises static information and dynamic information.

[0146] The reward construction module 103 is configured to construct a second reward function based on the target position and a third reward function based on the obstacle information.

[0147] The reinforcement learning module 104 is configured to train the operation strategy of the robot arm based on the first reward function, the second reward function and the third reward function through reinforcement learning.

[0148] The robot arm control module 105 is configured to select the operation trajectory with the highest cumulative reward after the training is completed, and control the robot arm to execute the operation trajectory.

[0149] In an embodiment, the inverse reinforcement learning module 101 is further configured to:

[0150] acquire record data of any object operating the robot arm, wherein the record data comprises historical trajectory data; pre-process the historical trajectory data; infer the corresponding reward function from the pre-processed historical trajectory data based on an inverse reinforcement learning algorithm; verify the reward function based on a reinforcement learning algorithm, and determine the final reward function after the verification is passed.

[0151] In an embodiment, the obstacle acquisition module 102 is further configured to:

[0152] The surgical scene video stream is acquired and segmented into continuous depth image frames; the depth image frames are input into an obstacle identification model to obtain an obstacle identification result; the obstacle identification result includes static obstacles, dynamic obstacles and target positions; the continuous depth image frames labeled with the obstacle identification result are input into a space-time prediction model to obtain a spatial motion trajectory of the dynamic obstacles.

[0153] In an implementation, the reinforcement learning module 104 is further configured to:

[0154] generate the first control mode reward, the second control mode reward and the third control mode reward according to the first reward function, the second reward function and the third reward function;

[0155] preliminarily train the first control mode reward, the second control mode reward and the third control mode reward based on the reinforcement learning method;

[0156] construct a switching threshold between the first control mode reward, the second control mode reward and the third control mode reward;

[0157] train the first control mode reward, the second control mode reward and the third control mode reward as a whole according to the switching threshold to obtain an operation strategy of the robot arm.

[0158] In an implementation, the first control mode reward is determined by the second reward function and the third reward function, and the second control mode reward and the third control mode reward are both determined by the first reward function, the second reward function and the third reward function.

[0159] In an implementation, the reinforcement learning module 104 is further configured to:

[0160] construct a simulation space according to the target position and the obstacle information; construct a state space and an action space according to parameters of the robot arm; the action space contains a plurality of joint axes, each joint axis having a forward rotation action and a reverse rotation action; construct a deep Q network model and a first experience pool; the first control mode reward, the second control mode reward and the third control mode reward have corresponding first, second and third experience pools; sample from the experience replay pool with the first control mode reward as a reward function, and calculate a Q value based on the deep Q network model; update the weight of the deep Q network model by minimizing a loss function until a predetermined number of iterations is reached.

[0161] In an implementation, the switching threshold includes a first switching threshold and a second switching threshold; the reinforcement learning module 104 is further configured to:

[0162] Sample training data from the first experience pool, the second experience pool and the third experience pool according to a preset ratio; take the first control mode reward as a reward function, and iteratively update the weight of the deep Q network model according to the training data; the current state of the robot arm is determined in real time; when the current state of the robot arm meets the first switching threshold, the second control mode reward is taken as the reward function, and the weight of the deep Q network model is iteratively updated according to the training data; when the current state of the robot arm meets the second switching threshold, the third control mode reward is taken as the reward function, and the weight of the deep Q network model is iteratively updated according to the training data; in the case that the number of iterations reaches a preset number, or the loss function converges, the overall training is completed, and the operation strategy of the robot arm is obtained.

[0163] The above embodiments of the present application provide a mechanical arm multi-mode comprehensive control device for surgical robot body intelligence, which has a corresponding relationship with the method for surgical robot body intelligence, and the specific content of the device has a corresponding relationship with the method for surgical robot body intelligence. The specific content can be referred to in the method for surgical robot body intelligence, and will not be repeated here.

[0164] The above embodiments of the present application provide a mechanical arm multi-mode comprehensive control device for surgical robot body intelligence, which has the same beneficial effects as the method adopted, run or implemented by the application program stored therein.

[0165] The above describes the internal functions and structures of the mechanical arm multi-mode comprehensive control device for surgical robot body intelligence. As shown in Figure 3 The mechanical arm multi-mode comprehensive control device for surgical robot body intelligence can be implemented as an electronic device, including a memory 301 and a processor 303.

[0166] The memory 301 can be configured to store programs.

[0167] In addition, the memory 301 can also be configured to store other various data to support the operation on the electronic device. Examples of these data include instructions for any application or method operating on the electronic device, contact data, phonebook data, messages, pictures, videos, etc.

[0168] The memory 301 can be implemented by any type of volatile or nonvolatile storage devices or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk or optical disk.

[0169] The processor 303 is coupled to the memory 301 and configured to execute programs in the memory 301, so as to:

[0170] Obtain a first reward function based on inverse reinforcement learning processing;

[0171] Obtain target position information and obstacle information, wherein the obstacle information includes static information and dynamic information;

[0172] Construct a second reward function based on the target position information and a third reward function based on the obstacle information;

[0173] Train the operation strategy of the robot arm through reinforcement learning according to the first reward function, the second reward function and the third reward function;

[0174] After the training is completed, an operation trajectory with the highest cumulative reward is selected, and the robot arm is controlled to execute the operation trajectory.

[0175] In an embodiment, the processor 303 is further configured to:

[0176] Obtain recorded data of any object operating the robot arm, wherein the recorded data includes historical trajectory data; pre-process the historical trajectory data; infer a corresponding reward function from the pre-processed historical trajectory data based on an inverse reinforcement learning algorithm; verify the reward function based on a reinforcement learning algorithm, and determine a final reward function after the verification is passed.

[0177] In an embodiment, the processor 303 is further configured to:

[0178] Obtain a surgical scene video stream and segment the video stream into continuous depth image frames; input the depth image frames into an obstacle recognition model to obtain an obstacle recognition result; the obstacle recognition result includes static obstacles, dynamic obstacles and target position information; input the continuous depth image frames labeled with the obstacle recognition result into a space-time prediction model to obtain a spatial motion trajectory of the dynamic obstacles.

[0179] In an embodiment, the processor 303 is further configured to:

[0180] The first control mode reward, the second control mode reward and the third control mode reward are generated according to the first reward function, the second reward function and the third reward function; the first control mode reward, the second control mode reward and the third control mode reward are preliminarily trained based on a reinforcement learning method; a switching threshold between the first control mode reward, the second control mode reward and the third control mode reward is constructed; and the first control mode reward, the second control mode reward and the third control mode reward are overall trained according to the switching threshold, so as to obtain the operation strategy of the robot arm.

[0181] In an embodiment, the first control mode reward is determined by the second reward function and the third reward function, and the second control mode reward and the third control mode reward are both determined by the first reward function, the second reward function and the third reward function.

[0182] In an embodiment, the processor 303 is further configured to:

[0183] According to the target position and the obstacle information, a simulation space is constructed; according to the robot arm parameters, a state space and an action space are constructed; the action space contains a plurality of joint axes, each joint axis having a forward rotation action and a reverse rotation action; a deep Q network model and a first experience pool are constructed; the first control mode reward, the second control mode reward and the third control mode reward have corresponding first, second and third experience pools; the first control mode reward is taken as a reward function, training data is sampled from the experience replay pool, and a Q value is calculated based on the deep Q network model; the weights of the deep Q network model are updated by minimizing a loss function until a predetermined number of iterations is reached.

[0184] In an embodiment, the switching threshold includes a first switching threshold and a second switching threshold; and the processor 303 is further configured to:

[0185] The training data is sampled from the first, second and third experience pools according to a preset ratio; the weights of the deep Q network model are iterated according to the training data and the first control mode reward is taken as a reward function; the current state of the robot arm is determined in real time; when the current state of the robot arm meets the first switching threshold, the weights of the deep Q network model are iterated according to the training data and the second control mode reward is taken as a reward function; when the current state of the robot arm meets the second switching threshold, the weights of the deep Q network model are iterated according to the training data and the third control mode reward is taken as a reward function; and when the number of iterations reaches a preset number or the loss function converges, the overall training is completed, and the operation strategy of the robot arm is obtained.

[0186] In the present application, the processor is also specifically configured to execute all processes and steps of the above-mentioned multi-mode comprehensive control method for the mechanical arm of the surgical robot body intelligence, and specific contents can be referred to the records in the multi-mode comprehensive control method for the mechanical arm of the surgical robot body intelligence, which will not be described here again in the present application.

[0187] In the present application, Figure 3 Only some components are shown in the present application, and it does not mean that the electronic device only includes Figure 3 The components shown.

[0188] The electronic device provided by the embodiment of the present application has the same beneficial effects as the method adopted, run or implemented by the application program stored therein, based on the same inventive concept as the multi-mode comprehensive control method for the mechanical arm of the surgical robot body intelligence provided by the embodiment of the present application.

[0189] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer usable program code.

[0190] The present application is described with reference to flowcharts and / or block diagrams according to the method, device (system) and computer program product of the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of the flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device produce a device that implements the functions specified in the flowcharts and / or block diagrams. Figure 1 The functions specified in one flow or multiple flows and / or blocks Figure 1 The functions specified in one flow or multiple flows and / or blocks

[0191] These computer program instructions can also be stored in a computer readable memory capable of guiding the computer or other programmable data processing device to work in a specific way, so that the instructions stored in the computer readable memory produce a product including instruction devices, which implement the functions specified in the flowcharts and / or block diagrams. Figure 1 The functions specified in one flow or multiple flows and / or blocks Figure 1 The functions specified in one flow or multiple flows and / or blocks

[0192] These computer program instructions can also be loaded into a computer or other programmable data processing devices, so that a series of operational steps are performed on the computer or other programmable devices to generate a computer-implemented process, so that the instructions executed by the computer or other programmable devices provide a process for implementing the functions specified in the flowchart Figure 1 one flow or multiple flows and / or blocks Figure 1 one flow or multiple flows and / or blocks

[0193] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.

[0194] The memory can include non-persistent memory, random access memory (RAM), and / or non-volatile memory, such as read only memory (ROM) or Flash memory, in a computer readable medium. The memory is an example of computer readable media.

[0195] The application also provides a computer readable storage medium corresponding to the multi-mode comprehensive control method for the body intelligence of a surgical robot provided by the foregoing embodiments, and a computer program (i.e., a program product) is stored on the computer readable storage medium. When the computer program is executed by a processor, the multi-mode comprehensive control method for the body intelligence of a surgical robot provided by any of the foregoing embodiments is executed.

[0196] Computer readable media includes permanent and non-permanent, removable and non-removable media, which can be implemented by any method or technology to store information. The information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read only memory (ROM), electrically erasable programmable read only memory (EEPROM), flash memory or other memory technology, compact disc read only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device. According to the definition herein, computer readable media does not include transitory media, such as modulated data signals and carriers.

[0197] The computer readable storage medium provided by the above embodiments of the application has the same beneficial effects as the multi-mode comprehensive control method for the body intelligence of a surgical robot provided by the embodiments of the application, which are adopted, run or implemented by the application programs stored therein.

[0198] It is to be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting. It is also possible in the present disclosure that steps can be executed in different sequence where is logically possible.

[0199] It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting. It is also possible in the present disclosure that steps can be executed in different sequence where is logically possible.

[0200] The embodiments of method, system and apparatus described above are merely possible solutions to the problems and are not intended to limit the protection scope of the present application. Any modification, equivalent replacement or improvement made within the spirit and principle of the present application should be included in the protection scope of the present application.

Claims

1. A multi-mode integrated control method for surgical robot body intelligence, characterized in that, The method comprises the following steps: obtaining a first reward function based on inverse reinforcement learning processing; obtaining target position information and obstacle information, wherein the obstacle information comprises static information and dynamic information; constructing a second reward function based on the target position information and a third reward function based on the obstacle information; training the operation strategy of the robot arm through reinforcement learning according to the first reward function, the second reward function and the third reward function; after the training is completed, selecting an operation trajectory with the highest cumulative reward and controlling the robot arm to execute the operation trajectory; the training of the operation strategy of the robot arm through reinforcement learning according to the first reward function, the second reward function and the third reward function comprises the following steps: generating a first control mode reward, a second control mode reward and a third control mode reward according to the first reward function, the second reward function and the third reward function; preliminarily training the first control mode reward, the second control mode reward and the third control mode reward based on a reinforcement learning method; constructing a switching threshold between the first control mode reward, the second control mode reward and the third control mode reward; training the first control mode reward, the second control mode reward and the third control mode reward as a whole according to the switching threshold to obtain the operation strategy of the robot arm; the switching threshold comprises a first switching threshold and a second switching threshold; the training of the first control mode reward, the second control mode reward and the third control mode reward as a whole according to the switching threshold to obtain the operation strategy of the robot arm comprises the following steps: sampling training data from a first experience pool, a second experience pool and a third experience pool according to a preset proportion; iterating the weight of a deep Q network model according to the training data with the first control mode reward as a reward function; determining the current state of the robot arm in real time; when the current state of the robot arm meets the first switching threshold, iterating the weight of the deep Q network model according to the training data with the second control mode reward as a reward function; when the current state of the robot arm meets the second switching threshold, iterating the weight of the deep Q network model according to the training data with the third control mode reward as a reward function; when the number of iterations reaches a preset number or the loss function converges, the whole training is completed to obtain the operation strategy of the robot arm.

2. The multi-mode integrated control method for surgical robot body intelligence of claim 1, wherein the method comprises the following steps: obtaining record data of any object operating the robot arm, wherein the record data comprises historical trajectory data; preprocessing the historical trajectory data; speculating a corresponding reward function from the historical trajectory data after preprocessing based on an inverse reinforcement learning algorithm; verifying the reward function based on a reinforcement learning algorithm, and determining a final reward function after verification.

3. The multi-mode integrated control method for surgical robot body intelligence of claim 1, wherein, obtaining a real surgery scene video stream and analyzing the target position information and the obstacle information, comprising: obtaining a surgery scene video stream and segmenting it into continuous depth image frames; inputting the depth image frames into an obstacle recognition model to obtain an obstacle recognition result, wherein the obstacle recognition result comprises static obstacles, dynamic obstacles and target positions; Input the continuous depth image frames marked with obstacle recognition results into the space-time prediction model to obtain the spatial motion trajectory of the dynamic obstacle.

4. The multi-mode integrated control method for surgical robot body intelligence of the mechanical arm according to any one of claims 1-3, characterized in that, The first control mode reward is determined by a second reward function and a third reward function, and the second control mode reward and the third control mode reward are both determined by the first reward function, the second reward function and the third reward function.

5. The multi-mode integrated control method for surgical robot body intelligence of claim 1-3, wherein, The preliminary training of the first control mode reward based on the reinforcement learning method comprises: constructing a simulation space according to the target position and the obstacle information; constructing a state space and an action space according to the parameters of the robot arm, wherein the action space comprises a plurality of joint axes, and each joint axis has a forward rotation action and a reverse rotation action; constructing a deep Q network model and a first experience pool, wherein the first control mode reward, the second control mode reward and the third control mode reward have corresponding first, second and third experience pools; sampling from the experience replay pool by taking the first control mode reward as a reward function, and calculating a Q value based on the deep Q network model; updating the weights of the deep Q network model by minimizing a loss function until a predetermined number of iterations is reached.

6. A multi-mode integrated control device for surgical robot body intelligence, characterized in that, comprise: an inverse reinforcement learning module configured to obtain a first reward function processed based on inverse reinforcement learning; an obstacle acquisition module configured to obtain target position and obstacle information, wherein the obstacle information comprises static information and dynamic information; a reward construction module configured to construct a second reward function based on the target position and a third reward function based on the obstacle information; a reinforcement learning module configured to train the operation strategy of the robot arm by reinforcement learning according to the first reward function, the second reward function and the third reward function; a robot arm control module configured to select an operation trajectory with the highest cumulative reward after the training is completed, and control the robot arm to execute the operation trajectory; the reinforcement learning module is further configured to: generate the first control mode reward, the second control mode reward and the third control mode reward according to the first reward function, the second reward function and the third reward function, preliminarily train the first control mode reward, the second control mode reward and the third control mode reward based on the reinforcement learning method, construct a switching threshold between the first control mode reward, the second control mode reward and the third control mode reward, and train the first control mode reward, the second control mode reward and the third control mode reward as a whole according to the switching threshold to obtain the operation strategy of the robot arm; the switching threshold comprises a first switching threshold and a second switching threshold, and the training of the first control mode reward, the second control mode reward and the third control mode reward as a whole according to the switching threshold to obtain the operation strategy of the robot arm comprises: sampling training data from the first experience pool, the second experience pool and the third experience pool according to a preset proportion. The first control mode reward is taken as a reward function, and weights of a deep Q network model are iterated according to the training data; the current state of the mechanical arm is determined in real time; when the current state of the mechanical arm meets a first switching threshold, the second control mode reward is taken as a reward function, and the weights of the deep Q network model are iterated according to the training data; when the current state of the mechanical arm meets a second switching threshold, the third control mode reward is taken as a reward function, and the weights of the deep Q network model are iterated according to the training data; when the number of iterations reaches a preset number or the loss function converges, overall training is completed, and the operation strategy of the mechanical arm is obtained.

7. An electronic device, comprising: Comprise: a memory and a processor; the memory is used for storing programs; the processor is coupled to the memory and is used for executing the programs, so as to: obtain a first reward function based on inverse reinforcement learning processing; obtain target position and obstacle information, the obstacle information including static information and dynamic information; construct a second reward function based on the target position and a third reward function based on the obstacle information; train the operation strategy of the mechanical arm through reinforcement learning according to the first reward function, the second reward function and the third reward function; after the training is completed, select an operation trajectory with the highest cumulative reward, and control the mechanical arm to execute the operation trajectory; the training of the operation strategy of the mechanical arm through reinforcement learning according to the first reward function, the second reward function and the third reward function comprises: generating a first control mode reward, a second control mode reward and a third control mode reward according to the first reward function, the second reward function and the third reward function; preliminarily training the first control mode reward, the second control mode reward and the third control mode reward based on a reinforcement learning method; constructing switching thresholds among the first control mode reward, the second control mode reward and the third control mode reward; training the first control mode reward, the second control mode reward and the third control mode reward according to the switching thresholds to obtain the operation strategy of the mechanical arm; the switching thresholds comprise a first switching threshold and a second switching threshold; and the training of the first control mode reward, the second control mode reward and the third control mode reward according to the switching thresholds to obtain the operation strategy of the mechanical arm comprises: sampling training data from a first experience pool, a second experience pool and a third experience pool according to a preset ratio; taking the first control mode reward as a reward function, and iteratively training weights of a deep Q network model according to the training data; determining the current state of the mechanical arm in real time; when the current state of the mechanical arm meets the first switching threshold, taking the second control mode reward as a reward function, and iteratively training the weights of the deep Q network model according to the training data; when the current state of the mechanical arm meets the second switching threshold, taking the third control mode reward as a reward function, and iteratively training the weights of the deep Q network model according to the training data; when the number of iterations reaches a preset number or the loss function converges, completing overall training to obtain the operation strategy of the mechanical arm.

8. A computer-readable storage medium having stored thereon a computer program, characterized in that, The program is executed by the processor to realize the multi-mode comprehensive control method for the surgical robot body intelligence mechanical arm according to any one of claims 1-5.

Citation Information

Patent Citations

  • Underwater robot path planning method and device

    CN118331316A

  • Motion planning method based on inverse reinforcement learning

    CN118504808A