Bronchoscope robot autonomous navigation system and method based on deep reinforcement learning
Through deep reinforcement learning, combining electromagnetic navigation and endoscopic image data, the bronchoscopic robot autonomous control method is solved, and the problem of insufficient path planning accuracy and poor environmental adaptability is achieved, efficient, safe and accurate bronchoscopic operation is achieved, reducing the operating burden of doctors.
Patent Information
- Application Number
- CN202510354813.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-11
AI Technical Summary
The existing bronchoscopic robot system has insufficient path planning accuracy and poor environmental adaptability in complex environments, heavy operation burden on doctors, and low training efficiency, making it difficult to achieve autonomous navigation and dynamic control.
The autonomous control method based on deep reinforcement learning is adopted, combined with electromagnetic navigation technology and endoscopic image data, and the robot position and image data are obtained in real time. Through feature extraction and attention mechanism fusion, the agent is trained for autonomous navigation, and the bronchial tree map is used for personalized training to improve the autonomy and accuracy of the robot.
It improves the robot's path planning accuracy and autonomy in complex bronchial paths, reduces the operating burden of doctors, reduces artificial errors during the operation, and achieves more efficient, safe and accurate bronchoscopic operation.
Smart Images

Figure CN120284473A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of robot control and autonomous navigation, and particularly relates to a method for autonomous control of a bronchoscope robot based on deep reinforcement learning. The invention trains an intelligent agent through deep reinforcement learning for autonomous navigation of the bronchoscope robot, uses electromagnetic navigation technology and an endoscope camera to obtain pose data and image data in real time, then performs data feature extraction and multi-modal data fusion and inputs them to a reinforcement learning agent, and then obtains a motion output through a deep network, and finally realizes the control of the three degrees of freedom (rotation, bending, advancing and retreating) of the robot. The present invention has broad application prospects, and is particularly suitable for fields such as medical robot navigation and artificial intelligence surgical assistance. Background Art
[0002] As an advanced medical assistance device, bronchoscope robots have been widely used in the medical field in recent years, especially during bronchoscopy examinations and treatments. Traditional bronchoscope operations mostly rely on manual operations by doctors. Although basic operation tasks can be achieved, the flexibility and precision during the operation are limited by the doctor's experience and operation skills, and doctors are prone to fatigue during long-term operations, resulting in a decline in operation accuracy. These problems make bronchoscope robots still face certain challenges, especially in complex and irregular bronchial paths, where doctors need to continuously adjust the operation, increasing the difficulty of the operation and the discomfort of the patient.
[0003] In recent years, with the development of robot technology and artificial intelligence technology, more and more research has begun to explore how to combine robots and intelligent control systems to improve the operation efficiency and accuracy of bronchoscope robots. In the prior art, some research attempts to develop a robot system with autonomous control functions to reduce the burden on doctors and improve the safety of operations. For example, some bronchoscope robot systems use automated navigation technology and rely on pre-set paths before surgery for automatic motion control. However, these systems still have certain limitations, such as the problem of being unable to effectively cope with real-time path adjustment in a complex dynamic environment.
[0004] In the field of automatic control, some research combines sensor technology, image processing and navigation algorithms to optimize the path planning of the robot through real-time positioning and state feedback, thereby improving its motion accuracy and execution efficiency. However, the deficiencies of these systems are that they often rely on traditional path tracking algorithms and lack the ability of real-time autonomous decision-making for complex environments and emergencies. Specifically, existing path planning algorithms usually rely on environmental data provided by external position sensors. However, in some narrow or complex bronchial paths, the sensors may be interfered with or unable to accurately obtain sufficient information, resulting in deviations in path planning and motion control. On the other hand, image data with rich information has not been well utilized.
[0005] At present, control methods based on reinforcement learning are also gradually being applied in some robot applications, but they are less applied in the field of bronchoscope robots. Moreover, although existing methods widely adopt deep neural network models, most of them take single-modal data (such as image or sensor data) as input, and rarely achieve cross-modal information fusion and collaborative optimization. Therefore, how to utilize the existing sensor data and image information, combined with the reinforcement learning algorithm, to train an agent to achieve efficient, precise, and real-time autonomous control of the bronchoscope robot in a complex bronchial path environment has become an important direction in current research and development.
[0006] Some solutions have been proposed at present. For example, the patent invention of "an AI-based AR navigation bronchoscope system capable of real-time monitoring of blind areas" invented an AI-based AR navigation bronchoscope system. The DCNN image processing module extracts bronchial characteristic data in real time, and combines with the DRL navigation module to dynamically optimize the path planning to adapt to the bronchial structures and lesion conditions of different patients. At the same time, the system monitors the respiratory state and adjusts the navigation strategy to reduce respiratory interference and lower the operation risk. Finally, the navigation path is superimposed on the endoscope image through AR technology to provide intuitive real-time navigation guidance for doctors, improving the accuracy and safety of the surgery.
[0007] Existing bronchoscope robot systems mainly adopt the master-slave control mode, in which doctors operate by manually controlling the equipment. Although this control mode can provide an intuitive operation experience, it also has a series of significant defects. In a complex bronchial path, the robot often has difficulty in precisely executing the doctor's instructions, with a large control error. Especially in the case of a narrow or curved path, the accuracy and flexibility of the robot's movement cannot meet the high requirements of the surgery. Doctors need to frequently adjust the operation, resulting in a more cumbersome surgical process, and the operation accuracy will decrease with the doctor's fatigue. In addition, in the master-slave control mode, doctors rely too much on the operation, requiring doctors to have excellent operation skills and rich experience, and the physical consumption and high-intensity operation burden during the operation will affect the doctor's work efficiency and accuracy. Long-term fine operation may also lead to medical safety problems, especially in high-difficulty surgeries, where human errors are likely to occur during the operation.
[0008] In addition, existing control methods have not fully addressed key challenges such as low sample efficiency and sparse reward functions in the training of complex robotic systems. Due to the traditional master-slave control method relying on manual operations by doctors, the robot has not achieved autonomous path planning and dynamic control. Therefore, a large amount of operation data is often required during training, and the collection of this data often depends on the participation of doctors and is difficult to cover the individual differences of different patients and complex bronchial structures, resulting in poor adaptability of the robotic system and difficulty in effectively operating in a real surgical environment. In addition, most traditional path planning methods rely on data provided by external sensors, but in complex bronchial paths, sensor data is often interfered with or limited, resulting in inaccurate path planning and thus affecting the performance of the robot in actual surgeries. Summary of the Invention
[0009] The present invention aims to solve the problems of insufficient control accuracy, excessive doctor operation burden, low training efficiency, and poor environmental adaptability existing in the prior art. The present invention uses an autonomous control method based on deep reinforcement learning, combines electromagnetic navigation technology and endoscopic image data, and obtains robot pose data and image data in real time to achieve autonomous navigation and dynamic path planning of the robot. In the present invention, the screen data and pose data are preprocessed and feature-extracted, and then data fusion is performed through an attention mechanism to be transformed into a high-dimensional vector containing a large amount of information, thereby forming the observation input required for robot control. The deep neural network takes the observation value as the input and the encoding of the motor controllers of each degree of freedom of the bronchoscope robot as the output. Thereafter, a reinforcement learning algorithm is first used to pre-train the model of the intelligent agent, and then for a specific patient, after three-dimensional reconstruction of the bronchial tree map using CT data, further training is carried out so that the model can better adapt to the patient's personalized bronchial structure, thereby improving the generalization ability of the robotic system. The present invention can not only solve the problems of insufficient path planning accuracy and poor environmental adaptability existing in traditional control methods, but also improve the autonomy of the robot, reduce the operation burden of doctors, and reduce human errors during the operation, thus providing a more efficient, safe, and accurate solution in complex bronchoscope robot operations.
[0010] Advantageous Effects
[0011] 1. The present invention uses electromagnetic navigation to collect position data and an endoscopic camera to collect image data, and uses a processing method of feature extraction and attention mechanism fusion to construct the observation input of deep reinforcement learning. This technology ensures that the robot can efficiently process and fuse multi-source data and provide intelligent decision-making support.
[0012] 2. The present invention uses a deep reinforcement learning algorithm to train an agent to move autonomously on a bronchial tree map. The method covers details such as scenario construction, reward function design, and model optimization to ultimately ensure that the robot can complete tasks autonomously and efficiently in an actual surgical environment.
[0013] 3. The present invention uses deep reinforcement learning to achieve autonomous navigation. Traditional robot control relies on doctor operation, and the accuracy is easily affected by fatigue. The present invention trains an agent through a deep reinforcement learning algorithm, and uses the bronchial tree map and the map reconstructed from the patient's CT for autonomous training, enabling the robot to independently make decisions on the path, reducing the dependence on doctors, and improving the autonomy and accuracy of the surgery. The use of multi-source data fusion improves the perception accuracy. The prior art relies on a single sensor to obtain data, resulting in insufficient path planning accuracy. In contrast, the present invention fuses electromagnetic navigation and endoscopic camera data, extracts features, and then fuses them into a feature vector through an attention mechanism, enhancing the robot's perception ability of complex bronchial paths and ensuring more precise control and navigation. By automatically generating a large number of bronchial tree data sets, the problem of difficult training due to scarce data during the learning process is overcome. By automatically creating data sets, not only the richness and diversity of the data are greatly improved, but also the cost and time of manual data collection are significantly reduced, reducing the economic cost. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 is the overall flowchart of the autonomous navigation method of the bronchoscope robot based on deep reinforcement learning according to the present invention.
[0015] Figure 2 is the flowchart of the part for automatically generating the bronchial tree model data set according to the present invention.
[0016] Figure 3 is the overall architecture diagram of the bronchoscope robot for deep reinforcement learning training according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0017] The present invention will be further described below with reference to the drawings and embodiments.
[0018] Embodiment
[0019] As Figure 1 shown, an autonomous navigation system for a bronchoscope robot based on deep reinforcement learning, the autonomous navigation system for the bronchoscope robot based on deep reinforcement learning includes an endoscope module, an image processing module, an electromagnetic navigation positioning module, an attention fusion module, a bronchial tree data set construction module, a reinforcement learning training module, an observation data processing module, a reinforcement learning agent module, a control signal conversion module, and a motor communication module;
[0020] The described endoscope module is used to receive the image information captured in real time by a medical endoscope, and then output the real-time captured image information to the image processing module;
[0021] The described image processing module is used to receive the real-time captured image information input by the endoscope module, crop and enhance the received image information, and then output the cropped and enhanced image information to the attention fusion module;
[0022] The electromagnetic navigation positioning module is used to receive the real-time pose data captured by the electromagnetic navigation sensor, and output the received real-time pose data to the attention fusion module;
[0023] The described attention fusion module is used to receive the cropped and enhanced image information output by the image processing module, and is also used to receive the real-time pose data output by the electromagnetic navigation positioning module, and perform attention fusion on the received multi-image information and real-time pose data to obtain fusion information, and then output the fusion information to the observation data processing module;
[0024] The bronchial tree dataset construction module is used to automatically generate a large number of artificial bronchial tree models, and then output some of the artificial bronchial tree models to the reinforcement learning training module;
[0025] The described reinforcement learning training module is used to receive the artificial bronchial tree models output by the bronchial tree dataset construction module, train the reinforcement learning agent on the artificial bronchial tree models using the deep reinforcement learning method, and then output the trained reinforcement learning agent to the reinforcement learning agent module;
[0026] The described observation data processing module is used to receive the fusion information output by the attention fusion module, and process the received fusion information into the form of input for the standard reinforcement learning observation space to obtain standard observation data, and then output the standard observation data to the reinforcement learning agent module;
[0027] The described reinforcement learning agent module is used to receive the standard observation data output by the observation data processing module, and is also used to receive the reinforcement learning agent output by the reinforcement learning training module, and use the received reinforcement learning agent to reason about the received standard observation data to obtain standard motion data, and then output the obtained standard motion data to the control signal conversion module;
[0028] The described control signal conversion module is used to receive the standard motion data output by the reinforcement learning agent module, and process the received standard motion data into motor encoded data and output it to the motor communication module;
[0029] The described motor communication module is used to receive the motor coding data output by the control signal conversion module, process the received motor coding data into CAN communication messages, and send the CAN communication messages to the bronchoscope robot through the CAN bus. The bronchoscope robot performs autonomous navigation based on the received CAN communication messages.
[0030] The method for the attention fusion module to perform attention fusion on the received multi-image information and real-time pose data is as follows:
[0031] Step S1, obtain the cropped and enhanced image information and real-time pose data;
[0032] Step S2, use a convolutional neural network to process the cropped and enhanced image information, extract the anatomical structure features in the image, and obtain the feature vector I;
[0033] Step S3, use an autoencoder to process the real-time pose data and obtain the feature vector P;
[0034] Step S4, use the attention mechanism for information fusion. Let the feature Q = I, the feature K = the feature V = P, and use the dot product attention formula to obtain the attention score Attention Score;
[0035] Step S5, use the Softmax function to normalize the attention score Attention Score to obtain the normalized attention weight;
[0036] Step S6, use the normalized attention weight to perform weighted summation on the input feature V to obtain the representation output vector, and this representation output vector is the fused information;
[0037] The method for the bronchial tree dataset construction module to automatically generate a large number of artificial bronchial tree models is as follows:
[0038] Step 1, based on medical data, establish a human bronchial tree dictionary D. The dictionary D contains the average length, inner diameter, outer diameter, direction parameters of each main branch of the human bronchial structure, and the variance corresponding to each parameter;
[0039] Step 2, according to the parameter list provided by the human bronchial tree dictionary D, use the normal distribution model to randomly generate a set of airway branch parameter tables. The branch parameters include the length, inner diameter, outer diameter, and direction determined for each branch at each level;
[0040] Step 3: Create an empty boolean matrix and define that if an element in the matrix is 1, it represents that there is an entity at the position of this element, and if an element in the matrix is 0, it represents that there is no entity at the position of this element. Select a definite starting point in the matrix and set the matrix element corresponding to this position to 1. Next, according to the main airway parameters in the airway branch parameter table, move 1 unit length in the direction of the main airway from the previous point and find the nearest integer point, and set the matrix element corresponding to this position to 1. By analogy, all points on the main airway axis will be set to 1. Then, use the last point of the main airway as the starting point of the secondary (left and right lobe bronchi), and operate in the same way as above. Finally, all points on the bronchial tree axis in the matrix will be set to 1. After that, traverse each point on the axis. For each point, traverse the points in the surrounding outer diameter size area and set them to 1 to obtain a solid bronchial tree model.
[0041] Step 4: Traverse each point on the axis. For each point, traverse the points in the surrounding inner diameter size area and set them to 0, and perform opening processing on the starting and ending parts of the bronchial tree to obtain a complete bronchial tree model.
[0042] Repeat steps 2 - 4 to obtain a large number of random bronchial tree models.
[0043] The specific method for training the reinforcement learning agent using the deep reinforcement learning method in the reinforcement learning training module is as follows:
[0044] Use the PPO algorithm of deep reinforcement learning for training, and set the standard observation data output by the observation data processing module as the observation space of the reinforcement learning algorithm. Set the action output of the reinforcement learning algorithm as a three - dimensional vector with a range of [-1, 1];
[0045] Set the reward function R of the reinforcement learning t to be composed as follows: R t = r t + r end . r t : The immediate reward for each step. r end : The immediate reward for each step;
[0046] Among them, r end is determined according to the reason for the end of the task, specifically including the following situations: successfully reaching the goal, r win is taken, abnormal errors such as severe collisions, r c , too many steps, r over ;
[0047] And r t = r cl + r bending + r step . Among them, r step : Single - step penalty. rcl : Distance from the centerline penalty, r bending : Excessive bending penalty to avoid large oscillation amplitudes;
[0048] Sensor information fusion processing part: The endoscope module is the visual perception unit of the bronchoscope robot, integrating a high-resolution endoscope imaging system to capture the dynamic video stream inside the bronchus in real time. Through fiber optic transmission and optoelectronic conversion technology, the optical image is converted into a digital signal to provide the original visual data for subsequent processing. In the actual operation process, a medical thin endoscope is used to obtain images, and in the deep reinforcement learning training platform, a virtual camera is used to obtain images. The image processing module is responsible for the standardized preprocessing of the original endoscope images. First, the contrast of the image is enhanced and noise is removed. Then it is cropped to a suitable size. Subsequently, a convolutional neural network is used to extract the anatomical structure features (such as bronchial bifurcation points, blood vessel textures) in the image and convert them into a structured feature vector I, which provides standardized visual input for multimodal data fusion. The electromagnetic navigation module is based on electromagnetic tracking technology to real-time calculate the pose information (position and attitude angle) of the bronchoscope end in three-dimensional space. Through the collaborative work of the magnetic field generator and the micro sensor, six-degree-of-freedom pose data is output with low latency, and a dynamic mapping relationship with the patient's thoracic coordinate system is established to provide a spatial positioning reference for path planning. In the actual operation process, an electromagnetic navigator is used to obtain the pose information, and in the deep reinforcement learning training platform, a pose acquisition function is used to obtain the pose information. The obtained pose data is processed into a feature vector P through an autoencoder. The attention fusion module will perform spatio-temporal alignment and feature fusion on the visual feature vector and the pose data. Through adaptive weight allocation, the image semantic information and the spatial motion state are dynamically integrated to generate a high-dimensional joint feature representation. In the attention mechanism, the image information I dominates, Q = I. The pose information is used as an auxiliary, K = V = P. Use the dot product attention formula to calculate the attention score. Finally, a characterization vector is output, which serves as the observation input of the deep reinforcement learning agent to support the generation of autonomous decision-making and control instructions.
[0049] Bronchial tree dataset generation part: This part is responsible for automatically generating a large number of artificial bronchial tree models to provide data for deep reinforcement learning training. The specific method is as follows. Establish a dictionary of the human bronchial tree, which contains the average length, inner diameter, outer diameter, direction parameters of each main branch of the human bronchial structure, and the variance corresponding to each parameter. The method for creating a new bronchial tree model is as follows: First, according to the parameter list provided by the human bronchial tree dictionary, use the normal distribution model to randomly generate a set of airway branch parameter tables. The branch parameters include the determined length, inner diameter, outer diameter, and direction of each branch at each level. Create an empty boolean matrix, and stipulate that if an element in the matrix is 1, it represents that there is an entity at the position of this element, and if an element in the matrix is 0, it represents that there is no entity at the position of this element. Select a determined starting point in the matrix and set the corresponding matrix element at this position to 1. Next, according to the main airway parameters in the airway branch parameter table, move 1 unit length in the direction of the main airway from the previous point and find the nearest integer point, and set the corresponding matrix element at this position to 1. And so on, all points on the main airway axis will be set to 1. Then, use the last point of the main airway as the starting point of the secondary (left and right lobe bronchi), and operate in the same way as above. Finally, all points on the bronchial tree axis in the matrix will be set to 1. After that, traverse each point on the axis. For each point, traverse the points in the area of its outer diameter size around it and set them to 1 to obtain a solid bronchial tree model. After completion, traverse each point on the axis again. For each point, traverse the points in the area of its inner diameter size around it and set them to 0, and perform opening processing on the starting and ending parts of the bronchial tree to obtain the final bronchial tree model. By determining different bronchial tree model parameters, a large number of bronchial tree models are generated as a dataset. The overall implementation process is as Figure 2 .
[0050] Bronchoscope robot deep reinforcement learning part: This part uses the deep reinforcement learning algorithm to construct an intelligent autonomous control system to achieve an end-to-end mapping from multi-modal data perception to motion control. Deep Reinforcement Learning (DRL) is an advanced artificial intelligence technology that combines the perception ability of deep learning and the decision-making ability of reinforcement learning. Its core mechanism is to learn the optimal policy mapping from a high-dimensional state space to a continuous action space through the continuous interaction between the agent and the environment. The high-dimensional joint features (visual semantics + pose data) output by the sensor information processing and fusion module and the real-time motion parameters (such as joint angles, speeds) are fused through the attention mechanism to form a composite observation vector containing anatomical structure features, instrument space state, and environmental dynamic changes, which is used as the state space of reinforcement learning. The control commands of the motors of each degree of freedom of the bronchoscope are encoded into a continuous action space, including the forward movement amount, bending joint angle amount, and rotation angle amount, which are output as the action space. The objective function of PPO combines the policy ratio and the advantage function and is defined as: Reward function Rt It is mainly composed of two parts R t = r t + r end . r t : The immediate reward for each step. r end : The immediate reward for each step. r end It depends on the reason for the task to end, specifically including the following situations: successfully reaching the goal r win Take, abnormal errors such as serious collisions r c , too many steps r over . r t It is obtained by summing three parts r t = r step + r cl + r bending . Among them, r step : Single-step penalty. r cl : Penalty for the distance from the center line. r bending : Penalty for excessive bending to avoid excessive oscillation amplitude. Deploy the deep reinforcement learning algorithm to the simulation platform, use a large number of bronchial tree datasets to generate a partially created bronchial tree model, conduct multi-batch training, and save the final model. The deep reinforcement learning architecture is as Figure 3 .
[0051] Intelligent agent autonomous control part: This part is mainly used to perform inferences using the deployed model and convert the inference output into corresponding motor control signals to control the actual motor movement. The observation data processing module receives multi-modal fusion information, normalizes the input data into a 512-dimensional vector with a range of [0,1] as the standard observation input for reinforcement learning. The reinforcement learning agent module delivers the standard observation input to the trained reinforcement learning agent. The model evaluates the current environment based on the current standard input and outputs a set of optimal standard action policies. The standard action output is a 3-dimensional standard vector with a range of [0,1]. It corresponds to the degrees of freedom of robot bending, rotation, and forward / backward movement respectively. The control signal conversion module first transforms the standard action output into control quantities for each degree of freedom. Such as bending [0,90] degrees, rotating [0,360] degrees, and moving forward [-1,1] cm. Then it converts the control quantity for each degree of freedom into the corresponding motor encoding quantity. The motor communication module is responsible for sending the control signal to the motor through the CAN bus protocol and establishing a real-time communication connection with the motor. This module first initializes the CAN bus interface, configures parameters such as the baud rate to ensure normal communication with the motor. After receiving the motor encoding signal from the control signal conversion module, the module encapsulates it into a data frame format that conforms to the CAN protocol and sends it to the corresponding motor through the CAN bus.
[0052] During the actual deployment process, the system runs in a closed-loop autonomous navigation mode: First, the endoscope and the electromagnetic locator are synchronized to collect the original images and six-degree-of-freedom pose data (X / Y / Z coordinates, Roll / Pitch / Yaw angles). The image processing module performs real-time correction and feature extraction on the video stream. The attention fusion module dynamically weights and fuses the visual features and pose vectors, and outputs a high-dimensional joint feature vector as the observation input of the policy network. Based on the current state, the policy network infers and generates commands for a 3D continuous action space. After being converted into control quantities for each motor, it drives the bronchoscope actuator to complete precise pose adjustment. During the navigation process, the electromagnetic tracking system continuously monitors the actual spatial pose of the robot's end, thus forming an autonomous control closed-loop of "perception - decision - execution - re-perception" to achieve high-precision continuous autonomous navigation.
[0053] A method for autonomous navigation of a bronchoscope robot based on deep reinforcement learning, the steps of which include:
[0054] In the first step, the endoscope module receives the image information captured in real time by the medical endoscope, and then outputs the real-time captured image information to the image processing module;
[0055] In the second step, the image processing module receives the real-time captured image information input by the endoscope module, crops and enhances the received image information, and then outputs the cropped and enhanced image information to the attention fusion module;
[0056] In the third step, the electromagnetic navigation positioning module receives the real-time pose data captured by the electromagnetic navigation sensor, and outputs the received real-time pose data to the attention fusion module;
[0057] In the fourth step, after the attention fusion module receives the image information and the real-time pose data, it performs attention fusion on the received multi-image information and the real-time pose data to obtain fusion information, and then outputs the fusion information to the observation data processing module;
[0058] In the fifth step, the bronchial tree dataset construction module automatically generates a large number of artificial bronchial tree models, and then outputs some of the artificial bronchial tree models to the reinforcement learning training module;
[0059] In the sixth step, the reinforcement learning training module receives the artificial bronchial tree models output by the bronchial tree dataset construction module, trains the reinforcement learning agent on the artificial bronchial tree models using the deep reinforcement learning method, and then outputs the trained reinforcement learning agent to the reinforcement learning agent module;
[0060] In the seventh step, the observation data processing module receives the fusion information output by the attention fusion module, processes the received fusion information into the form of the input of the standard reinforcement learning observation space to obtain standard observation data, and then outputs the standard observation data to the reinforcement learning agent module;
[0061] In the eighth step, the reinforcement learning agent module receives the standard observation data and the reinforcement learning agent, and uses the received reinforcement learning agent to reason about the received standard observation data to obtain standard motion data, and then outputs the obtained standard motion data to the control signal conversion module;
[0062] In the ninth step, the control signal conversion module receives the standard motion data output by the reinforcement learning agent module, and processes the received standard motion data into motor coding data and outputs it to the motor communication module;
[0063] In the tenth step, the motor communication module receives the motor coding data output by the control signal conversion module, processes the received motor coding data into a CAN communication message, and the CAN communication message is sent to the bronchoscope robot through the CAN bus, and the bronchoscope robot performs autonomous navigation according to the received CAN communication message.
[0064] In summary, the above are only the preferred embodiments of the present invention and are not intended to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A bronchoscope robot autonomous navigation system based on deep reinforcement learning, characterized in that: The bronchoscope robot autonomous navigation system based on deep reinforcement learning includes an endoscope module, an image processing module, an electromagnetic navigation positioning module, an attention fusion module, a bronchial tree dataset construction module, a reinforcement learning training module, an observation data processing module, a reinforcement learning agent module, a control signal conversion module, and a motor communication module; The endoscope module is used to receive the images captured by the medical endoscope in real time, and then output the images captured in real time to the image processing module; The image processing module is used to receive the images captured in real time input by the endoscope module, crop and enhance the received images, and then output the cropped and enhanced images to the attention fusion module; The electromagnetic navigation positioning module is used to receive the real-time pose data captured by the electromagnetic navigation sensor, and output the received real-time pose data to the attention fusion module; The attention fusion module is used to receive the cropped and enhanced images output by the image processing module, and is also used to receive the real-time pose data output by the electromagnetic navigation positioning module, perform attention fusion on the received images and real-time pose data to obtain fusion information, and then output the fusion information to the observation data processing module; The bronchial tree dataset construction module is used to automatically generate a large number of artificial bronchial tree models, and then output some of the artificial bronchial tree models to the reinforcement learning training module; The reinforcement learning training module is used to receive the artificial bronchial tree models output by the bronchial tree dataset construction module, train the reinforcement learning agent on the artificial bronchial tree models using the deep reinforcement learning method, and then output the trained reinforcement learning agent to the reinforcement learning agent module; The observation data processing module is used to receive the fusion information output by the attention fusion module, process the received fusion information into the form of input for the standard reinforcement learning observation space to obtain standard observation data, and then output the standard observation data to the reinforcement learning agent module; The reinforcement learning agent module is used to receive the standard observation data output by the observation data processing module, and is also used to receive the reinforcement learning agent output by the reinforcement learning training module, and use the received reinforcement learning agent to reason about the received standard observation data to obtain standard motion data, and then output the obtained standard motion data to the control signal conversion module; The control signal conversion module is used to receive the standard motion data output by the reinforcement learning agent module, and process the received standard motion data into motor encoded data and output it to the motor communication module; The motor communication module is used to receive the motor encoded data output by the control signal conversion module, process the received motor encoded data into CAN communication messages, and the CAN communication messages are sent to the bronchoscope robot through the CAN bus, and the bronchoscope robot performs autonomous navigation according to the received CAN communication messages.
2. The bronchoscope robot autonomous navigation system based on deep reinforcement learning according to claim 1, characterized in that: The method for the attention fusion module to perform attention fusion on the received image and real-time pose data is as follows: Step S1: Use a convolutional neural network to extract the anatomical structure features in the image to obtain a feature vector I. Step S2: Use an autoencoder to process the real-time pose data to obtain a feature vector P. Step S3, use the attention mechanism to fuse the feature vector I obtained in step S1 and the feature vector P obtained in step S2. Specifically: Let feature Q = I, feature K = feature V = feature vector P, and use the dot product attention formula to obtain the attention score Attention Score; Step S4: Use the Softmax function to normalize the attention score Attention Score obtained in Step S3 to obtain a normalized attention weight. Step S5: Use the normalized attention weight obtained in Step S4 to perform weighted summation on the feature V to obtain a representation output vector, which is the fusion information.
3. The autonomous navigation system of a bronchoscope robot based on deep reinforcement learning according to claim 1, wherein: The method for the bronchial tree dataset construction module to automatically generate an artificial bronchial tree model is as follows: Step 1: Establish a human bronchial tree dictionary D. Step 2: According to the parameter list provided by the human bronchial tree dictionary D, use a normal distribution model to randomly generate a set of airway branch parameter tables. Step 3: Create an empty boolean matrix, and stipulate that an element with a value of 1 in the matrix represents that there is an entity at the position of the element, and an element with a value of 0 represents that there is no entity at the position of the element. Select a definite starting point in the matrix and set the corresponding matrix element at this position to 1. Next, according to the main airway parameters in the airway branch parameter table, move 1 unit length in the direction of the main airway from the previous point and find the nearest integer point, and set the corresponding matrix element at this position to 1. By analogy, all points on the main airway axis will be set to 1. Then, use the last point of the main airway as the starting point of the secondary airway and operate in the same way as above. Finally, all points on the bronchial tree axis in the matrix will be set to 1. After that, traverse each point on the axis. For each point, traverse the points in the surrounding outer diameter size area and set them to 1 to obtain a solid bronchial tree model. Step 4: Traverse each point on the axis. For each point, traverse the points in the surrounding inner diameter size area and set them to 0, and perform opening processing on the starting and ending parts of the bronchial tree to obtain a complete bronchial tree model.
4. The autonomous navigation system of a bronchoscope robot based on deep reinforcement learning according to claim 3, wherein: Repeat Steps 2 to 4 to obtain a large number of randomly generated bronchial tree models.
5. The autonomous navigation system of a bronchoscope robot based on deep reinforcement learning according to claim 3, wherein: In Step 1, the dictionary D includes the average length, inner diameter, outer diameter, direction parameters of each main branch of the human bronchial structure, and the variance corresponding to each parameter.
6. The autonomous navigation system of a bronchoscope robot based on deep reinforcement learning according to claim 3, wherein: In Step 2, the airway branch parameters include the length, inner diameter, outer diameter, and direction determined for each branch at each level.
7. The autonomous navigation system of a bronchoscope robot based on deep reinforcement learning according to claim 1, wherein: The method of training the reinforcement learning agent using the deep reinforcement learning method in the reinforcement learning training module is as follows: Use the PPO algorithm of deep reinforcement learning for training. Set the standard observation data output by the observation data processing module as the observation space of the reinforcement learning algorithm. Set the action output of the reinforcement learning algorithm as a three-dimensional vector with a range of [-1, 1].
8. A bronchoscope robot autonomous navigation system based on deep reinforcement learning according to claim 7, characterized in that: Set the reward function R of reinforcement learning t It is composed as follows: R t = r t + r end ; where r t is the immediate reward for each step, and r end is the final reward at the end of the episode.
9. A bronchoscope robot autonomous navigation system based on deep reinforcement learning according to claim 8, characterized in that: r t =r cl +r bending +r step ; Among them, r step is the single-step reward, r cl is the centerline distance reward, r bending is the excessive bending reward; r end Determined according to the state at the end of the round. If the end target position is successfully reached at the end of the round, then r end = r win where r win is the reward obtained when the end target position is successfully reached; At the end of the round, if an exception error occurs, then r end = r c ; r c is the reward obtained when an exception error occurs; When the number of steps in a single round exceeds the set threshold at the end of the round, then r end = r over ; r over is the reward obtained when the number of steps in a single round exceeds the set threshold.
10. A method for autonomous navigation of a bronchoscope robot based on deep reinforcement learning, characterized in that The steps of this method include: First step, the endoscope module receives the image information captured by the medical endoscope in real time, and then outputs the real-time captured image information to the image processing module; Second step, the image processing module receives the real-time captured image information input by the endoscope module, crops and enhances the received image information, and then outputs the cropped and enhanced image information to the attention fusion module; Third step, the electromagnetic navigation positioning module receives the real-time pose data captured by the electromagnetic navigation sensor, and outputs the received real-time pose data to the attention fusion module; Fourth step, after the attention fusion module receives the image information and the real-time pose data, it performs attention fusion on the received multi-image information and real-time pose data to obtain fusion information, and then outputs the fusion information to the observation data processing module; Fifth step, the bronchial tree dataset construction module automatically generates a large number of artificial bronchial tree models, and then outputs some of the artificial bronchial tree models to the reinforcement learning training module; Sixth step, the reinforcement learning training module receives the artificial bronchial tree model output by the bronchial tree dataset construction module, trains the reinforcement learning agent on the artificial bronchial tree model using the deep reinforcement learning method, and then outputs the trained reinforcement learning agent to the reinforcement learning agent module; Seventh step, the observation data processing module receives the fusion information output by the attention fusion module, processes the received fusion information into the form input to the standard reinforcement learning observation space to obtain standard observation data, and then outputs the standard observation data to the reinforcement learning agent module; Eighth step, the reinforcement learning agent module receives the standard observation data and the reinforcement learning agent, and uses the received reinforcement learning agent to reason about the received standard observation data to obtain standard motion data, and then outputs the obtained standard motion data to the control signal conversion module; Ninth step, the control signal conversion module receives the standard motion data output by the reinforcement learning agent module, and processes the received standard motion data into motor encoded data and outputs it to the motor communication module; Tenth step, the motor communication module receives the motor encoded data output by the control signal conversion module, processes the received motor encoded data into a CAN communication message, and the CAN communication message is sent to the bronchoscope robot through the CAN bus. The bronchoscope robot performs autonomous navigation according to the received CAN communication message.
Citation Information
Cited By
Cavity operation robot autonomous navigation decision-making method based on hybrid reinforcement learning and related device
CN120871582A
A cavity operation robot autonomous navigation decision method based on hybrid reinforcement learning and related device
CN120871582B