Intercostal trajectory planning method and system based on AI reinforcement learning robot

Through the robot intercostal trajectory planning method based on AI reinforcement learning, the problem of difficulty in imaging the target area due to the acoustic shadow of ribs and intercostal space in ultrasound imaging is solved, and the probe is accurately moved in the intercostal area, which significantly improves imaging quality and operation efficiency.

CN120125784APending Publication Date: 2025-06-10SHENZHEN BEAUTIFUL RUBIKS CUBE ROBOT CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510185110.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-19
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

In thoracic examination, ultrasound imaging is difficult to image the target area due to the sound and shadow effect of ribs, narrow intercostal space and limited probe posture, which affects the accuracy of diagnosis and operational complexity.

Method used

Using the robot intercostal trajectory planning method based on AI reinforcement learning, a three-dimensional anatomical model is constructed through a deep convolutional image segmentation network and a 3D point cloud detection network, a three-dimensional anatomical model is detected, key feature points are generated, and path point sets are optimized through dual adversarial deep Q learning to ensure that the probe can achieve accurate movement within the intercostal area.

Benefits of technology

Break through the limits of rib sound and shadow, significantly improve the imaging quality of the target area, reduce operator experience dependence, achieve high efficiency and stability, and is suitable for intelligent adaptation in multiple scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120125784A_ABST
    Figure CN120125784A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of human body intercostal model trajectory construction, and discloses an intercostal trajectory planning method and system based on an AI reinforcement learning robot. The inter-costal trajectory planning method is applied to an AI robot, and specifically comprises the following steps: S101, receiving an inter-costal region three-dimensional anatomical model construction instruction sent by a user terminal, obtaining anatomical image data of a thoracic cavity from CT or MRI scanning of a patient, and segmenting the anatomical image data by using a deep convolutional image segmentation network to obtain an anatomical image target region; the CT anatomical model and the reinforcement learning technology are fully combined, the completely observable Markov decision model is constructed, the robot path planning strategy is trained in the virtual environment, the robot can autonomously adjust the posture and position of the ultrasonic probe, it is ensured that the probe achieves accurate movement in the complex intercostal region, and the accuracy of the ultrasonic probe is improved. Therefore, limitation of rib sound shadow is broken through, and imaging quality of a target area is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of constructing the trajectory of the human intercostal model. Specifically, it relates to an intercostal trajectory planning method and system based on an AI reinforcement learning robot. Background Art

[0002] Nowadays, ultrasound technology has been widely applied in clinical practice and has become an important tool for screening visceral organs and guiding interventional therapy.

[0003] However, in thoracic examinations, ultrasound imaging faces significant challenges: the high acoustic impedance characteristics of the rib structure will produce a shadowing effect, making it difficult to image the anatomical area behind the ribs. In addition, the narrow intercostal space, individual differences in patient body types, and limited probe postures further exacerbate the imaging difficulty. These problems not only seriously affect the visualization of the target area and the accuracy of diagnosis, but also significantly increase the complexity of the operation, resulting in a large impact of the experience difference among operators on the results, and restricting the wide application of ultrasound in the diagnosis of thoracic diseases and postoperative monitoring. Summary of the Invention

[0004] The purpose of the present invention is to provide an intercostal trajectory planning method and system based on an AI reinforcement learning robot, which breaks through the limitation of rib shadows, significantly improves the imaging quality of the target area, not only solves the high dependence of traditional ultrasound path planning on the experience of operators, but also has high efficiency and stability, and can achieve intelligent adaptation in multiple scenarios, aiming to solve the problems in the prior art.

[0005] The present invention is implemented as follows. The intercostal trajectory planning method based on an AI reinforcement learning robot is applied to an AI robot and specifically includes the following steps:

[0006] S101: Receive the instruction to construct a three-dimensional anatomical model of the intercostal region sent by the user terminal, obtain the anatomical image data of the chest cavity from the patient's CT or MRI scan, and use a deep convolutional image segmentation network to segment the anatomical image data to obtain the target area of the anatomical image;

[0007] S102: Voxelize the target area of the anatomical image using a fixed resolution to generate a three-channel voxel matrix, then define the surface of the target area of the anatomical image as key feature points, implement key feature point detection through a 3D point cloud detection network, and generate a set of path points by linear interpolation between the key feature points;

[0008] S103: Combine the voxel matrix to represent the position and posture of the current probe of the AI robot in three-dimensional space, determine the next action information of the probe to determine the parameter change value of the reward function, and then learn and optimize the path planning through the parameter change value to complete the update of the objective function;

[0009] S104: Fit the initially generated set of path points with a spline curve to generate a smooth scanning trajectory, evenly sample path points on the trajectory to ensure the continuity and stability of the path, and then calculate the tangent direction and normal vector between adjacent path points for each path point to preliminarily fit the planned path;

[0010] S105: Control the AI robot to move along the planned path, adjust the probe contact depth through force feedback to ensure good acoustic coupling, and then perform frequency-domain analysis on the collected anatomical image data and contact force data of the chest cavity to obtain more accurate target point coordinates and complete the correction and optimization of the planned path.

[0011] Further, in S101, obtain the anatomical image data of the chest cavity from the patient's CT or MRI scan, and use a deep convolutional image segmentation network to segment the anatomical image data to obtain the target region of the anatomical image, including:

[0012] The obtained anatomical image data includes the patient's ribs, skin surface, and target organ;

[0013] Using a deep convolutional image segmentation network to segment the anatomical image data to obtain the target region of the anatomical image is to obtain the three-dimensional structures of the target organ and ribs through a U-Net convolutional neural network. The voxel matrix of the target region extracted by the U-Net convolutional neural network is defined as follows:

[0014]

[0015] Among them, M(x, y, z) is the segmented voxel matrix.

[0016] Further, in S102, voxelize the target region of the anatomical image with a fixed resolution to generate a three-channel voxel matrix, including:

[0017] Voxelize the CT data with a fixed resolution of 1mm 3 to generate a three-channel voxel matrix, and the definition formula of the three-channel voxel matrix is:

[0018] F[x, y, z] = [T(x, y, z), B(x, y, z), P(x, y, z)];

[0019] Among them, T(x, y, z) is the voxel label of the target region; B(x, y, z) is the voxel label of the rib occlusion; P(x, y, z) is the voxel label of the skin surface.

[0020] Further, define the surface of the target region of the anatomical image as key feature points, including:

[0021] Define the surface of the target area of the anatomical image as the key feature points. The position of each key feature point is obtained by network calculation. The position of each key feature point is calculated by the 3D point cloud detection network as follows:

[0022]

[0023] Where: is the i-th monitoring point detected, k is the candidate point set, and p(k) is the probability that the candidate point k is determined to be a key feature point.

[0024] Furthermore, the key feature point detection is realized by the 3D point cloud detection network, and the path point set is generated by using linear interpolation between the key feature points, including:

[0025] Generate a denser path point set by using linear interpolation between the key feature points. The formula for generating the path point set is:

[0026] P(t)=(1 - t)K i +tK i+1 , t ∈ [0, 1];

[0027] Where, K i , K i+1 are adjacent key feature points, P(t) is the interpolation point in the initial path point set. By introducing the 3D key feature point detection neural network, the surface features of the target area of the anatomical image are simplified to a key feature point set to achieve efficient rough positioning. First, each anatomical unit is represented as a key feature point, and the network is trained to detect the positions of these key feature points. The network automatically generates an ordered set of 3D key feature point coordinates by maximizing the key feature point probability of the candidate points, and a denser path point set is generated by using linear interpolation between the key feature points.

[0028] Furthermore, in S103, combine the voxel matrix to represent the position and attitude of the probe of the current AI robot in the three-dimensional space, and judge the next action information of the probe, including:

[0029] Combine the voxel matrix to represent the position and attitude of the probe of the current AI robot in the three-dimensional space is defined as follows:

[0030] s t ={p t , R t , F[x, y, z]};

[0031] Where: p t is the current position of the probe, R t is the current attitude of the probe, and F[x, y, z] is the environmental voxel matrix;

[0032] Judge the next action information of the probe as follows:

[0033] a t =(Δx, Δy, Δz, Δφ, Δψ);

[0034] Among them, Δx, Δy, and Δz are the translation values of the probe, and Δφ and Δψ are the angular adjustment values of the values.

[0035] Furthermore, to determine the parameter change value of the reward function, including:

[0036] The calculation formula for determining the parameter change value of the reward function is as follows:

[0037] r t =r c +α 1 r a +α 2 r s ;

[0038] Coverage reward, n t is the number of voxels covered by the current scan, and N is the total number of target voxels;

[0039] Attenuation minimization reward, d t is the distance between the probe and the target area, and R c is the standard distance;

[0040] r s =1 - p t : Shadow avoidance reward, p t is the proportion of the current occlusion area;

[0041] Among them, α 1 , α 2 are weight coefficients.

[0042] Furthermore, then learn and optimize the path planning through the parameter change value to complete the update of the objective function, including:

[0043] Use double adversarial deep Q learning to optimize the path planning, and the network is updated through the following objective function. The objective function update formula is as follows:

[0044]

[0045] Among them, Q(s t+1 , a t+1 ) is the Q value of the current state-action pair, and γ is the discount factor.

[0046] Further, in S105, the contact depth of the probe is adjusted through force feedback to ensure good acoustic coupling, and then the frequency domain analysis is performed on the collected anatomical image data and contact force data of the chest cavity to obtain more accurate target point coordinates, including:

[0047] Feedback optimization performs frequency domain analysis on the collected anatomical image data and contact force data of the chest cavity, and extracts key feature points through the following filtering equation:

[0048]

[0049] The true value is gradually approximated through the filtering equation to reduce data jumps and interference, and more accurate target point coordinates are obtained after filtering for path correction and optimization.

[0050] Compared with the prior art, the intercostal trajectory planning method and system based on an AI reinforcement learning robot provided by the present invention have the following beneficial effects:

[0051] 1. By fully combining the CT anatomical model and reinforcement learning technology, a fully observable Markov decision model is constructed, and the robot path planning strategy is trained in a virtual environment. Through intelligent control, the robot can autonomously adjust the attitude and position of the ultrasonic probe to ensure that the probe can move precisely in the complex intercostal area, thus breaking through the limitation of rib shadow and significantly improving the imaging quality of the target area. It not only solves the high dependence on the operator's experience in traditional ultrasonic path planning, but also has high efficiency and stability, and can achieve intelligent adaptation in multiple scenarios.

[0052] 2. Coarse positioning of the target area based on key feature point detection introduces the 3D key feature point detection neural network PointNet++ to simplify the surface features of the target area into a set of key feature points, realizing efficient coarse positioning. Each anatomical unit is represented as a key feature point, and the network is trained to detect the positions of these key feature points. The network automatically generates an ordered set of 3D key feature point coordinates by maximizing the key feature point probability of candidate points. To improve the fineness of path planning, linear interpolation is used between key feature points to generate a denser path point set, laying a foundation for subsequent path optimization and attitude calculation. This process effectively reduces the redundant information in the complex anatomical model and provides accurate initial positioning input for the reinforcement learning algorithm at the same time;

[0053] 3. Construct a virtual training environment through CT data, roughly locate the target area by combining the prior information of the anatomical model, realize the initial optimization of path planning, and the introduction of reinforcement learning further refines the path to ensure full coverage scanning of the target area in the limited intercostal acoustic window. The specific implementation of each link includes: constructing a three-dimensional anatomical model of the intercostal area, roughly locating the target area based on 3D key point detection, interpolating to generate an initial set of path points, using double adversarial deep Q-learning to optimize path planning and a reward function that fully considers the influence of shadows to achieve path planning, and refining the path and optimizing the pose by smoothing the path points with a spline curve and calculating the normal, tangent, and lateral directions of the probe to ensure that the probe fits the target surface and achieve accurate and efficient scanning pose adjustment.

[0054] The intercostal trajectory planning system based on an AI reinforcement learning robot executes the above-mentioned intercostal trajectory planning method. The intercostal trajectory planning system includes:

[0055] An acquisition module for receiving the instruction to construct a three-dimensional anatomical model of the intercostal area and the anatomical image data of the chest cavity sent by the user terminal;

[0056] A segmentation module for segmenting the anatomical image data using a deep convolutional image segmentation network to obtain the target area of the anatomical image, and voxelizing the target area of the anatomical image using a fixed resolution to generate a three-channel voxel matrix;

[0057] A path point set module for defining the surface of the target area of the anatomical image as key feature points, implementing key feature point detection through a 3D point cloud detection network, and generating a set of path points by using linear interpolation between the key feature points;

[0058] A calculation module for combining the voxel matrix to represent the position and pose of the probe of the current AI robot in three-dimensional space, judging the next action information of the probe to determine the parameter change value of the reward function, and then learning to optimize the path planning through the parameter change value to complete the update calculation of the objective function;

[0059] A fitting module for fitting the initially generated set of path points with a spline curve to generate a smooth scanning trajectory, and equally spacing path points on the trajectory to initially fit the planned path;

[0060] A path optimization module for adjusting the contact depth of the probe through force feedback, performing frequency domain analysis on the collected anatomical image data and contact force data of the chest cavity to obtain more accurate target point coordinates, and completing the correction and optimization of the planned path. Description of the Drawings

[0061] Figure 1 It is a schematic flow chart of the intercostal trajectory planning method based on an AI reinforcement learning robot proposed by the present invention;

[0062] Figure 2 This is a flowchart showing the process of using a deep convolutional image segmentation network to segment anatomical image data and obtain the target area of the anatomical image in the intercostal trajectory planning method based on an AI reinforcement learning robot proposed by the present invention;

[0063] Figure 3 This is a schematic structural diagram of the intercostal trajectory planning system based on an AI reinforcement learning robot proposed by the present invention. Detailed implementation manners

[0064] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention, but not to limit the present invention.

[0065] The implementation of the present invention will be described in detail below with reference to specific embodiments.

[0066] In the accompanying drawings of this embodiment, the same or similar reference numerals correspond to the same or similar components; in the description of the present invention, it should be understood that if there are terms such as "upper", "lower", "left", "right", etc. indicating the orientation or positional relationship, they are based on the orientation or positional relationship shown in the accompanying drawings. This is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, the terms describing the positional relationship in the accompanying drawings are only for illustrative purposes and cannot be understood as limiting the present invention. For those of ordinary skill in the art, the specific meanings of the above terms can be understood according to specific circumstances.

[0067] Refer to Figure 1-2 As shown, the intercostal trajectory planning method based on an AI reinforcement learning robot is applied to an AI robot and specifically includes the following steps:

[0068] S101: Receive a command for constructing a three-dimensional anatomical model of the intercostal region sent by a user terminal, obtain anatomical image data of the chest cavity from a patient's CT or MRI scan, and use a deep convolutional image segmentation network to segment the anatomical image data to obtain the target area of the anatomical image;

[0069] Among them, obtaining anatomical image data of the chest cavity from a patient's CT or MRI scan and using a deep convolutional image segmentation network to segment the anatomical image data to obtain the target area of the anatomical image includes:

[0070] The obtained anatomical image data includes the patient's ribs, skin surface, and target organs, and the target organs include the liver or heart, etc.;

[0071] Segment the anatomical image data using a deep convolutional image segmentation network. Obtain the target region of the anatomical image by acquiring the three-dimensional structures of the target organ and ribs through a U-Net convolutional neural network. The voxel matrix of the target region extracted by the U-Net convolutional neural network is defined as follows:

[0072]

[0073] M(x, y, z) is the segmented voxel matrix;

[0074] S102: Voxelize the target region of the anatomical image using a fixed resolution to generate a three-channel voxel matrix. Then, define the surface of the target region of the anatomical image as key feature points, implement key feature point detection through a 3D point cloud detection network, and generate a set of path points using linear interpolation between the key feature points;

[0075] Among them, voxelizing the target region of the anatomical image using a fixed resolution to generate a three-channel voxel matrix includes:

[0076] Voxelize the CT data using a fixed resolution of 1mm 3 to generate a three-channel voxel matrix, and the definition formula of the three-channel voxel matrix is:

[0077] F[x, y, z] = T[T(x, y, z), B(x, y, z), P(x, y, z)];

[0078] T(x, y, z) is the voxel label of the target region; B(x, y, z) is the voxel label of the rib occlusion; P(x, y, z) is the voxel label of the skin surface;

[0079] S103: Combine the voxel matrix to represent the position and orientation of the probe of the current AI robot in three-dimensional space, determine the next action information of the probe, so as to determine the parameter change value of the reward function, and then learn and optimize the path planning through the parameter change value to complete the update of the objective function;

[0080] Among them, combining the voxel matrix to represent the position and orientation of the probe of the current AI robot in three-dimensional space and determining the next action information of the probe includes:

[0081] The combination of the voxel matrix to represent the position and orientation of the probe of the current AI robot in three-dimensional space is defined as follows:

[0082] s t = {p t , R t , F[x, y, z]};

[0083] Among them: p t is the current position of the probe, R tis the current pose of the probe, and F[x, y, z] is the environmental voxel matrix;

[0084] The action information of the probe for the next step is judged as follows:

[0085] a t =(Δx, Δy, Δz, Δφ, Δψ);

[0086] Among them, Δx, Δy, and Δz are the translation values of the probe, and Δφ and Δψ are the angular adjustment values of the values.

[0087] To determine the parameter change value of the reward function, including:

[0088] The calculation formula for determining the parameter change value of the reward function is as follows:

[0089] r t =r c +α 1 r a +α 2 r s ;

[0090] Coverage reward, n t is the number of voxels covered by the current scan, and N is the total number of target voxels;

[0091] Attenuation minimization reward, d t is the distance between the probe and the target area, and R c is the standard distance;

[0092] r s =1 - p t : Shadow avoidance reward, p t is the proportion of the current occlusion area;

[0093] Among them, α 1 and α 2 are weight coefficients;

[0094] S104: Fit the initially generated path point set through a spline curve to generate a smooth scanning trajectory, and equally spaced path points are taken on the trajectory to ensure the continuity and stability of the path. Then, for each path point, calculate the tangent direction and normal vector between adjacent path points to initially fit the planned path;

[0095] S105: Control the AI robot to move along the planned path, adjust the contact depth of the probe through force feedback to ensure good acoustic coupling, and then perform frequency domain analysis on the collected anatomical image data and contact force data of the chest cavity to obtain more accurate target point coordinates and complete the correction and optimization of the planned path;

[0096] Among them, the probe contact depth is adjusted through force feedback to ensure good acoustic coupling, and then the frequency-domain analysis is performed on the collected anatomical image data and contact force data of the chest cavity to obtain more accurate target point coordinates, including:

[0097] The feedback optimization performs frequency-domain analysis on the collected anatomical image data and contact force data of the chest cavity, and extracts key feature points through the following filtering equation:

[0098]

[0099] The data jump and interference are reduced by gradually approaching the true value through the filtering equation, and more accurate target point coordinates are obtained after filtering for path correction and optimization. In this technical solution, a virtual training environment is constructed through CT data, and the target area is roughly located by combining the prior information of the anatomical model to achieve the initial optimization of path planning. The introduction of reinforcement learning further refines the path to ensure full coverage scanning of the target area in the limited acoustic window between the ribs. Moreover, the specific implementation of each link includes: constructing a three-dimensional anatomical model of the intercostal area, roughly locating the target area based on 3D key point detection, interpolating to generate an initial path point set, and using double adversarial deep Q-learning to optimize path planning and a reward function that fully considers the shadow effect to achieve path planning.

[0100] In this embodiment, the surface of the anatomical image target area is further defined as key feature points, including:

[0101] Defining the surface of the anatomical image target area as key feature points is to calculate the position of each key feature point through network calculation. The position of each key feature point is calculated by a 3D point cloud detection network as follows:

[0102]

[0103] Where: is the i-th monitoring point detected, K is the candidate point set, and p(k) is the probability that candidate point k is determined as a key feature point.

[0104] In this embodiment, key feature point detection is realized through a 3D point cloud detection network, and a path point set is generated by linear interpolation between key feature points, including:

[0105] A denser path point set is generated by linear interpolation between key feature points, and the formula for generating the path point set is:

[0106] P(t) = (1 - t)K i + tK i+1 , t ∈ [0, 1];

[0107] Where, K i , K i+1Let \(P_1\) and \(P_2\) be adjacent key feature points, and \(P(t)\) be the interpolation points in the initial path point set. By introducing a 3D key feature point detection neural network, the surface features of the target area of the anatomical image are simplified into a set of key feature points to achieve efficient rough positioning. First, each anatomical unit is represented as a key feature point, and the network is trained to detect the positions of these key feature points. The network automatically generates an ordered set of 3D key feature point coordinates by maximizing the key feature point probability of the candidate points, and uses linear interpolation between the key feature points to generate a denser path point set.

[0108] In this embodiment, the path planning is further optimized by learning the parameter change values to complete the update of the objective function, including:

[0109] Use double adversarial deep Q-learning to optimize the path planning. The network is updated through the following objective function, and the update formula of the objective function is as follows:

[0110]

[0111] where \(Q(s t+1 , a t+1 ) is the Q value of the current state-action pair, and \(\gamma\) is the discount factor.

[0112] This technical solution constructs a virtual training environment through CT data, combines the prior information of the anatomical model to perform rough positioning on the target area, realizes the initial optimization of the path planning. The introduction of reinforcement learning further refines the path to ensure full coverage scanning of the target area in the limited intercostal acoustic window. The specific implementation of each link includes: constructing a three-dimensional anatomical model of the intercostal area, rough positioning of the target area based on 3D key point detection, interpolating to generate the initial path point set, using double adversarial deep Q-learning to optimize the path planning and a reward function that fully considers the shadow effect to achieve path planning, and refining the path and optimizing the pose by smoothing the path points with a spline curve and calculating the probe normal, tangent and lateral directions to ensure that the probe fits the target surface and achieve accurate and efficient scanning pose adjustment.

[0113] Refer to Figure 3As shown in the figure, the intercostal trajectory planning system based on the AI reinforcement learning robot executes the above-mentioned intercostal trajectory planning method. The intercostal trajectory planning system includes: an acquisition module, which is used to receive the instruction for constructing a three-dimensional anatomical model of the intercostal region and the anatomical image data of the chest cavity sent by the user terminal; a segmentation module, which is used to segment the anatomical image data by using a deep convolutional image segmentation network to obtain the target region of the anatomical image, and voxelize the target region of the anatomical image with a fixed resolution to generate a three-channel voxel matrix; a path point set module, which is used to define the surface of the target region of the anatomical image as key feature points, implement key feature point detection through a 3D point cloud detection network, and generate a path point set by using linear interpolation between the key feature points; a calculation module, which is used to combine the voxel matrix to represent the position and posture of the probe of the current AI robot in the three-dimensional space, judge the next action information of the probe to determine the parameter change value of the reward function, and then learn and optimize the path planning through the parameter change value to complete the update calculation of the objective function; a fitting module, which is used to fit the initially generated path point set by using a spline curve to generate a smooth scanning trajectory, and equally spaced path points are taken on the trajectory to initially fit the planned path; a path optimization module, which is used to adjust the contact depth of the probe through force feedback, perform frequency domain analysis on the collected anatomical image data and contact force data of the chest cavity to obtain more accurate target point coordinates, and complete the correction and optimization of the planned path. By fully combining the CT anatomical model and the reinforcement learning technology, a fully observable Markov decision model is constructed, and the robot path planning strategy is trained in a virtual environment. Through intelligent control, the robot can autonomously adjust the posture and position of the ultrasonic probe to ensure that the probe can move precisely in the complex intercostal region, thus breaking through the limitation of the rib shadow and significantly improving the imaging quality of the target region. It not only solves the high dependence of traditional ultrasonic path planning on the operator's experience, but also has high efficiency and stability, and can achieve intelligent adaptation in multiple scenarios.

[0114] Specifically, for the rough localization of the target region based on key feature point detection, by introducing the 3D key feature point detection neural network PointNet++, the surface features of the target region are simplified into a set of key feature points to achieve efficient rough localization. Each anatomical unit is represented as a key feature point, and the network is trained to detect the positions of these key feature points. The network automatically generates an ordered set of 3D key feature point coordinates by maximizing the key feature point probability of the candidate points. To improve the fineness of path planning, linear interpolation is used between the key feature points to generate a denser path point set, laying a foundation for subsequent path optimization and pose calculation. This process effectively reduces the redundant information in the complex anatomical model and at the same time provides accurate initial positioning input for the reinforcement learning algorithm.

[0115] Specifically, for the basic principle and structure of PointNet++:

[0116] The core improvement of PointNet++ lies in its multi-level feature extraction structure, which mainly includes the following parts:

[0117] Sampling Layer: Some points are randomly sampled as centroids, which are used for subsequent grouping operations;

[0118] Grouping Layer: Centered on these centroids, the surrounding points are grouped to form local regions;

[0119] PointNet Layer: PointNet is applied to each local region for feature extraction to obtain local features;

[0120] Interpolate: In the segmentation network, through upsampling and downsampling operations, the gradual abstraction and refinement of features are realized.

[0121] In this embodiment, the entire operation process can be controlled by a computer to achieve automated operation control. Moreover, in each operation link, sensors can be set to perform signal feedback to achieve sequential execution of steps. These are all common knowledge of current automated control and will not be elaborated one by one in this embodiment.

[0122] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. An intercostal trajectory planning method based on AI reinforcement learning robot, characterized in that: Applied to AI robots, the specific steps include: S101: receiving a command for constructing a three-dimensional anatomical model of an intercostal region sent by a user terminal, acquiring anatomical image data of the chest cavity from a CT or MRI scan of a patient, and segmenting the anatomical image data using a deep convolutional image segmentation network to obtain an anatomical image target region; S102: voxelizing the target area of ​​the anatomical image using a fixed resolution to generate a three-channel voxel matrix, defining the surface of the target area of ​​the anatomical image as key feature points, detecting the key feature points through a 3D point cloud detection network, and generating a path point set using linear interpolation between the key feature points; S103: combining the voxel matrix to represent the position and posture of the probe of the current AI robot in the three-dimensional space, determining the next action information of the probe to determine the parameter change value of the reward function, and then learning and optimizing the path planning through the parameter change value to complete the objective function update; S104: Fitting the initially generated path point set through a spline curve to generate a smooth scanning trajectory, and taking path points at equal intervals on the trajectory to ensure the continuity and stability of the path, and then calculating the tangent direction and normal vector between adjacent path points for each path point to preliminarily fit the planned path; S105: Control the AI ​​robot to move along the planned path, adjust the probe contact depth through force feedback to ensure good acoustic coupling, and then perform frequency domain analysis on the collected thoracic anatomical image data and contact force data to obtain more accurate target point coordinates and complete the correction and optimization of the planned path.

2. The intercostal trajectory planning method based on AI reinforcement learning robot according to claim 1, characterized in that: In S101, anatomical image data of the chest cavity is obtained from a CT or MRI scan of a patient, and the anatomical image data is segmented using a deep convolutional image segmentation network to obtain an anatomical image target region, including: The acquired anatomical image data includes the patient's ribs, skin surface, and target organs; The anatomical image data is segmented using a deep convolutional image segmentation network, and the target area of ​​the anatomical image is obtained by obtaining the three-dimensional structure of the target organ and ribs through a U-Net convolutional neural network. The voxel matrix of the target area extracted by the U-Net convolutional neural network is defined as follows: Among them, M(x, y, z) is the voxel matrix after segmentation.

3. The intercostal trajectory planning method based on AI reinforcement learning robot according to claim 2, characterized in that: In S102, the target area of ​​the anatomical image is voxelized using a fixed resolution to generate a three-channel voxel matrix, including: Use fixed resolution 1mm 3 The CT data is voxelized to generate a three-channel voxel matrix, and the definition formula of the three-channel voxel matrix is: F[x,y,z]=[T(x,y,z),B(x,y,z),P(x,y,z)]; Among them, T(x, y, z) is the voxel marker of the target area; B(x, y, z) is the voxel marker of the rib occlusion; P(x, y, z) is the voxel marker of the skin surface.

4. The intercostal trajectory planning method based on AI reinforcement learning robot according to claim 3, characterized in that: The surface of the target area of ​​the anatomical image is then defined as key feature points, including: The surface of the target area of ​​the anatomical image is defined as the key feature point. The position of each key feature point is obtained through network calculation. The position of each key feature point is calculated by the 3D point cloud detection network as follows: in: is the i-th monitoring point detected, K is the candidate point set, and p(k) is the probability that candidate point k is determined to be a key feature point.

5. The intercostal trajectory planning method based on AI reinforcement learning robot according to claim 4, characterized in that: Key feature point detection is achieved through the 3D point cloud detection network, and a set of path points is generated using linear interpolation between key feature points, including: Linear interpolation is used between key feature points to generate a denser set of path points. The formula for generating a set of path points is: P(t)=(1-t)K i +tK i+1 ,t∈[0,1]; Among them, K i , K i+1 are adjacent key feature points, P(t) is the interpolation point in the initial point set of the path. By introducing a 3D key feature point detection neural network, the surface features of the target area of ​​the anatomical image are simplified into a set of key feature points to achieve efficient coarse positioning. First, each anatomical unit is represented as a key feature point, and the network is trained to detect the positions of these key feature points. The network automatically generates a set of ordered 3D key feature point coordinates by maximizing the key feature point probability of the candidate points, and uses linear interpolation between the key feature points to generate a denser set of path points.

6. The intercostal trajectory planning method based on AI reinforcement learning robot according to claim 5, characterized in that: In S103, the position and posture of the probe of the current AI robot in the three-dimensional space are represented by the voxel matrix to determine the next action information of the probe, including: Combined with the voxel matrix, the position and posture of the current AI robot's probe in three-dimensional space are defined as follows: s t ={p t ,R t ,F[x,y,z]}; Where: p t is the current position of the probe, R t is the current posture of the probe, F[x, y, z] is the environment voxel matrix; The next action information of the probe is determined as follows: a t =(Δx,Δy,Δz,ΔφΔψ); Among them, Δx, Δy, Δz are the translation values ​​of the probe, and Δφ, Δψ are the angle adjustment values ​​of the value.

7. The intercostal trajectory planning method based on AI reinforcement learning robot according to claim 6, characterized in that: To determine the parameter changes of the reward function, including: The calculation formula to determine the parameter change value of the reward function is as follows: r t =r c +α1r a +α2r s ; Coverage reward, n t is the number of voxels covered by the current scan, and N is the total number of target voxels; Decay minimizes the reward, d t is the distance between the probe and the target area, R c is the standard distance; r s =1-p t : Shadow avoidance reward, p t is the proportion of the current occluded area; Among them, α1 and α2 are weight coefficients.

8. The intercostal trajectory planning method based on AI reinforcement learning robot according to claim 7, characterized in that: Then, the path planning is optimized through parameter change value learning to complete the objective function update, including: Using dual adversarial deep Q learning to optimize path planning, the network is updated by the following objective function, and the objective function update formula is as follows: Among them, Q(s t+1 , a t+1 ) is the Q value of the current state-action pair, and γ is the discount factor.

9. The intercostal trajectory planning method based on AI reinforcement learning robot according to claim 8, characterized in that: In S105, the probe contact depth is adjusted through force feedback to ensure good acoustic coupling, and then the collected thoracic anatomical image data and contact force data are analyzed in the frequency domain to obtain more accurate target point coordinates, including: Feedback optimization performs frequency domain analysis on the collected thoracic anatomical image data and contact force data, and extracts key feature points through the following filtering equations: The filtering equation is gradually approached to the true value to reduce data jumps and interference. After filtering, more accurate target point coordinates are obtained for path correction and optimization.

10. Intercostal trajectory planning system based on AI reinforcement learning robot, characterized in that: The intercostal trajectory planning method according to any one of claims 1 to 9 is implemented, wherein the intercostal trajectory planning system comprises: An acquisition module, used for receiving instructions for constructing a three-dimensional anatomical model of the intercostal region and anatomical image data of the chest cavity sent by a user terminal; A segmentation module, used to segment the anatomical image data using a deep convolutional image segmentation network to obtain a target area of ​​the anatomical image, and voxelize the target area of ​​the anatomical image using a fixed resolution to generate a three-channel voxel matrix; A path point set module is used to define the surface of the target area of ​​the anatomical image as key feature points, detect the key feature points through a 3D point cloud detection network, and generate a path point set using linear interpolation between the key feature points; The calculation module is used to combine the voxel matrix to represent the position and posture of the current AI robot's probe in three-dimensional space, determine the probe's next action information, determine the parameter change value of the reward function, and then learn to optimize the path planning through the parameter change value to complete the objective function update calculation; The fitting module is used to fit the initially generated path point set through a spline curve to generate a smooth scanning trajectory, and to take path points at equal intervals on the trajectory to initially fit the planned path; The path optimization module is used to adjust the probe contact depth through force feedback, perform frequency domain analysis on the collected thoracic anatomical image data and contact force data, obtain more accurate target point coordinates, and complete the correction and optimization of the planned path.

Citation Information

Cited By

  • Motion track planning system for needle-knife robot

    CN120753786A

  • A motion trajectory planning system for a needle knife robot

    CN120753786B