Lumbar vertebral pedicle screw placement path planning method based on multi-agent reinforcement learning

CN120032001APending Publication Date: 2025-05-23宋丹彤
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510120094.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-25
Publication Date
2025-05-23

Smart Images

  • Figure CN120032001A_ABST
    Figure CN120032001A_ABST
Patent Text Reader

Abstract

The invention discloses a lumbar vertebral pedicle screw placement path planning method based on multi-agent reinforcement learning, and the method comprises the following steps: 1, carrying out the preprocessing of CT image data; step 2, constructing bounding boxes of five lumbar vertebrae, and finally training a segmentation network by using the data; step 3, applying a multi-agent reinforcement learning algorithm to cooperatively optimize a placement path of the screw system in the five sections of lumbar vertebrae, and constructing a reinforcement learning environment and an optimization reward function; and step 4, realizing output of the lumbar vertebral pedicle screw placement path planning scheme. Compared with the prior art, the contact surface area of the screw and a cortical bone area is remarkably increased, so that the stability of postoperative fixation is enhanced; the introduction of a bone mineral density award factor ensures that the screw can avoid an osteoporosis area, the success rate of stable screw placement is improved, the manual analysis time of a doctor on a CT image is remarkably shortened through automatic screw placement planning, the operation complexity and the labor time cost are reduced, and repeated image verification in an operation is not needed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of clinical medicine, and in particular to a lumbar pedicle screw placement path planning method based on multi-agent reinforcement learning. Background Art

[0002] Spinal screw placement is an important technique in spinal surgery, used to treat spinal diseases such as herniated disc, spondylolisthesis, and scoliosis. In recent years, with the advancement of surgical instruments and the development of imaging technology, the accuracy and safety of spinal screw placement have gradually improved. However, traditional spinal screw placement planning mainly relies on the doctor's experience and intraoperative imaging guidance. This method has problems such as strong subjectivity, complex operation, and long operation time. In addition, due to the constraints of the use of clinical surgical tools, traditional technology is difficult to take into account both the stability of the screws and the biomechanical properties of the global arrangement of multiple lumbar vertebrae, resulting in limited surgical results.

[0003] In the current clinical medical field, spinal nail placement surgery mainly relies on the following methods:

[0004] 1. Intraoperative navigation system: Intraoperative CT or 3D navigation system is used to guide the screw placement in real time. This method relies on expensive equipment, and the surgical process is complicated and requires high technical skills of the doctor.

[0005] 2. Robot-assisted surgery: In recent years, some spinal surgery robots (such as Mazor Robotics) can assist doctors in the precise placement of screws. However, this technology and equipment is expensive and cannot be widely used in clinical surgery.

[0006] 3. Empirical planning based on anatomical landmarks: The doctor selects the point where the screw is inserted into the vertebral body and the direction of screw placement based on the patient's CT image data and the doctor's personal surgical experience. This method is highly subjective, and the quality of surgical planning depends on the surgeon's surgical experience, making it difficult to form a standardized operating procedure.

[0007] Disadvantages of traditional technology:

[0008] Traditional intraoperative navigation equipment and robotic systems are mostly only designed for single-segment lumbar screw placement, and cannot meet the needs of setting connecting rods between multiple lumbar vertebrae according to clinical surgical scenarios. In addition, current equipment generally does not include key data such as patient bone density as a consideration for surgical plan design, which may lead to an increase in the failure rate of screw placement surgery.

[0009] In addition, doctors manually plan the screw path or rely on intraoperative imaging assistance, which has the following problems:

[0010] Lack of standardization: The quality of planning varies greatly among different doctors, making it difficult to ensure consistency.

[0011] Complex operation: requires multiple intraoperative imaging verifications, which increases operation time and radiation exposure.

[0012] Unable to adjust dynamically: Traditional automated methods do not incorporate patient anatomical characteristics as a basis for selecting a plan to adjust surgical pin placement planning.

[0013] Ignoring the global aspect: Traditional technology mainly focuses on the screw position of a single lumbar vertebra, while ignoring the biomechanical synergy between multiple lumbar vertebrae.

[0014] Corresponding defects in the effects of traditional technologies: Screw placement in the lumbar spine may reduce the stability of the screws due to the relatively porous bone in the area where the screws pass, increasing the risk of surgical failure or postoperative loosening. For the installation of surgical auxiliary tools, the uneven arrangement of multiple lumbar intervertebral screws may lead to postoperative mechanical imbalance, aggravating the patient's postoperative discomfort and complications. In addition, the operation time supported by traditional technology is relatively long, and the radiation exposure of patients and medical staff is increased. The high cost of surgery supported by high-precision equipment in traditional technology limits the promotion of this technology in small and medium-sized medical institutions. Summary of the invention

[0015] The purpose of the present invention is to provide a lumbar pedicle screw placement path planning method based on multi-agent reinforcement learning that solves the above-mentioned problem.

[0016] In order to achieve the above object, the technical solution adopted by the present invention is: a lumbar pedicle screw placement path planning method based on multi-agent reinforcement learning, the method steps are as follows:

[0017] Step 1: CT image data preprocessing;

[0018] Using a coarse-to-fine network framework based on deep learning, the coarse-grained network is first used to predict the centroid feature points of the lumbar spine, and the lumbar part is automatically extracted from the complete 3D CT image of the patient's spine. Then, the fine-grained network is used to predict the four feature points of the central sagittal plane of each lumbar vertebra and the four feature points of the pedicle position as key data for screw placement path planning.

[0019] Step 2: Use the eight feature points obtained in step 1 to construct bounding boxes of five lumbar vertebrae. Obtain single lumbar vertebrae data through the bounding boxes. Then perform a three-dimensional rigid body rotation operation on each lumbar vertebra to rotate the single lumbar vertebrae CT data and its segmentation mask to a positive horizontal state in space. Finally, use the above data to train the segmentation network.

[0020] Step 3: Use the feature point coordinate data and the binary segmentation mask of the vertebra to locate the initial screw placement plane position and the screw's geometric parameters. Then, based on the bone density distribution inside the vertebra, use the multi-agent reinforcement learning algorithm to collaboratively optimize the placement path of the screw system inside the five lumbar vertebrae, build a reinforcement learning environment and optimize the reward function, so that the optimization route is more in line with clinical medical standards and the result is as close as possible to the result manually planned by the surgeon.

[0021] Step 4: Input CT image data into the optimized trained model to output the lumbar pedicle screw placement path planning solution.

[0022] As a preferred solution, in step 1, during the preprocessing of CT image data, CT images of different patients are corrected and image standardized.

[0023] As a preferred solution, in step three, the reward function includes a screw surface area reward function, a bone density reward function and a linear arrangement reward function.

[0024] As a preferred solution, the screw surface area reward function is designed based on the principle that the larger the contact area, the higher the reward. The specific formula is as follows:

[0025] S=2π×R×L

[0026] Where S is the contact area, R is the radius of the screw, and L is the length of the screw.

[0027] As a preferred solution, the bone density reward function evaluates the regional bone density by reading the HU value of the corresponding area in the original CT image, and takes the sum of the density of the passed areas as a reward.

[0028] As a preferred solution, the linear arrangement reward function evaluates this indicator by calculating the arrangement variance of the screw entry points and rotation angles, and rewards are given based on the design principle that the smaller and more orderly the differences, the higher the reward.

[0029] As a preferred solution, by adjusting the weight of the reward function, the intelligent agent first focuses on optimizing the surface area of ​​the screw and the bone density of the area through which the screw passes. After a single intelligent agent learns to handle basic goals, higher-order global optimization goals are gradually introduced.

[0030] As a preferred solution, when conducting multi-agent reinforcement learning, each lumbar vertebra is regarded as an agent, and learning and training are carried out using centralized training and distributed execution.

[0031] In the centralized training, during the training phase, the performance of each screw on the corresponding lumbar vertebra is integrated, and the placement performance of the entire screw system is analyzed through a shared value network, and a comprehensive evaluation is generated, which is used to guide the training;

[0032] In the distributed execution, during the pinning process, each agent executes independently and can learn from the global information of the system reward feedback to make the best decision.

[0033] As a preferred solution, in step 4, the output of the nail placement path planning solution includes the output of nail placement parameters, three-dimensional visualization and quantification of biomechanical indicators.

[0034] As a preferred solution, the output of the screw placement parameters can output, for each lumbar vertebra, the screw information on the left and right sides of the lumbar vertebra, the radius and length of the screw, and the bone density of the area through which the screw passes. The screw placement plan can be visualized in three dimensions on a three-dimensional model with a controllable observation angle, and the biomechanical indicators can be quantified to provide some numerical evaluation indicators.

[0035] Compared with the prior art, the advantages of the present invention are:

[0036] (1) Accuracy improvement based on clinical needs: The intelligent agent optimized by the reward function can select a better screw placement path, significantly increasing the surface area of ​​contact between the screw and the cortical bone area, thereby enhancing the stability of postoperative fixation; the introduction of the bone density reward factor ensures that the screw can avoid the osteoporotic area and select the high bone density area, thereby improving the success rate of stable screw placement based on clinical surgical standards.

[0037] (2) Collaborative optimization of multiple lumbar vertebrae: The linear arrangement term in the reward function is based on the need for the use of connecting rod tools in clinical surgery. We achieve neat arrangement of ipsilateral screws in multiple lumbar vertebrae by reducing the variance of the entry point and rotation angle of the ipsilateral screws, thereby improving the overall mechanical balance of the spinal surgery screw system. Centralized Training and Decentralized Execution (CTDE) ensures the optimization of the entire screw system through global value evaluation, while ensuring that each intelligent agent can learn the best strategy, thereby achieving balance and coordination between the local and global aspects.

[0038] (3) Improved surgical efficiency and safety: Automated nail placement planning significantly reduces the time doctors spend on manual analysis of CT images, reducing surgical complexity and manual time costs; since the planning plan is completed before surgery, there is no need for repeated image verification during surgery, reducing the radiation risk for patients and doctors.

[0039] (4) Personalized screw placement plan: The reinforcement learning model can adapt to the anatomical structure and pathological characteristics of different patients and provide each patient with a personalized screw placement plan. For special cases (such as patients with severe osteoporosis), the weight of the reward function can be adjusted to prioritize the optimization of bone density and improve the success rate of screw placement.

[0040] (5) Economical and scalable: Compared with the high cost of surgical robot equipment, the implementation of this technology only requires CT images and computing resources. It has high economical and universal applicability and is suitable for promotion in small and medium-sized medical institutions. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 :It is a flow chart of the pedicle screw placement surgical path planning based on multi-agent reinforcement learning of the present invention;

[0042] Figure 2 : is the distribution diagram of the central sagittal plane feature points and pedicle feature points of the present invention;

[0043] Figure 3 : It is a schematic diagram of the construction of the local coordinate system of a single lumbar vertebra and its rigid body rotation effect in space according to the present invention;

[0044] Figure 4 : Schematic diagram of the screw placement plane and screw geometric parameters. DETAILED DESCRIPTION

[0045] The automated path planning for pedicle screw placement of the present invention is intended to realize an end-to-end practical tool for planning the screw placement path and calculating screw parameters, and to complete the entire operation in combination with a robot navigation system, thereby truly realizing the automation of spinal screw placement surgery.

[0046] The present invention will be further described below. The specific embodiments are as follows: a lumbar pedicle screw placement path planning method based on multi-agent reinforcement learning, see Figure 1 , the method steps are as follows:

[0047] Step 1: CT image data preprocessing;

[0048] Using a coarse-to-fine network framework based on deep learning, the coarse-grained network is first used to predict the centroid feature points of the lumbar spine, and the lumbar part is automatically extracted from the complete three-dimensional CT image of the patient's spine. Then, the fine-grained network is used to predict the four feature points of the central sagittal plane of each lumbar vertebra and the four feature points of the pedicle position as the key data for screw placement path planning, such as Figure 2 As shown, the left picture shows the distribution of feature points in the central sagittal plane, and the right picture shows the distribution of feature points in the pedicle;

[0049] The present invention is applied to process three-dimensional medical images. The image data comes from the patient's spine image obtained by the doctor using CT scanning. Different from the traditional network research that only processes two-dimensional medical images, the technology of the present invention can process three-dimensional medical images. Through the deep network learning algorithm, we can accurately find the part where the lumbar spine is located from the overall spine CT data and mark the feature points that need attention (such as the center point of the pedicle).

[0050] Since the CT images of different patients may have differences in resolution or position angle, the CT images of different patients are corrected and standardized during CT image data preprocessing to ensure that the lumbar data format of each patient is consistent, thereby facilitating subsequent calculations and analysis.

[0051] Step 2: Use the eight feature points obtained in step 1 to construct bounding boxes of five lumbar vertebrae. Obtain single lumbar vertebrae data through the bounding boxes. Then perform a three-dimensional rigid body rotation operation on each lumbar vertebra to rotate the single lumbar vertebrae CT data and its segmentation mask to a positive horizontal state in space. The operation process is as follows: Figure 3 As shown, the above data is finally used to train the segmentation network; the purpose of this operation is to reduce the difficulty of network training and thus achieve the effect of improving training accuracy.

[0052] Step 3: Use the feature point coordinate data and the binary segmentation mask of the vertebral body to locate the initial screw placement plane position and the screw geometric parameters, such as Figure 4 As shown, based on the bone density distribution inside the vertebral body, a multi-agent reinforcement learning algorithm is used to collaboratively optimize the placement path of the screw system inside the five lumbar vertebrae, build a reinforcement learning environment and optimize the reward function, so that the optimized route is more in line with clinical medicine standards, ensuring that the results are as close as possible to those manually planned by the surgeon.

[0053] The application of multi-agent reinforcement learning in spinal screw placement planning of the present invention includes a design framework for centralized training and distributed execution, the definition of state space and action space, and a reward function optimization method for dynamic weight adjustment. The reward function is equivalent to the description of surgical rules, telling the agent which behaviors are "correct" and worthy of reward. The reward function includes a screw surface area reward function, a bone density reward function, and a linear arrangement reward function. The reward function is as follows:

[0054] Screw surface area reward function:

[0055] Since the longer and thicker the screw is, the larger the contact area between the screw and the lumbar bone layer is, and the larger the contact area is, the better the stability is. Therefore, the screw surface area reward function is designed based on the principle that the larger the contact area is, the higher the reward is. The specific formula is as follows:

[0056] S=2π×R×L

[0057] Where S is the contact area, R is the radius of the screw, and L is the length of the screw.

[0058] Bone density reward function:

[0059] Since the higher the bone density in an area, the more firmly the screw is fixed, the placement path in the high bone density area is preferentially selected. Therefore, the bone density reward function evaluates the regional bone density by reading the HU value of the corresponding area in the original CT image, and takes the sum of the density of the passed areas as the reward.

[0060] Linear permutation reward function:

[0061] Since the more uniform the arrangement of the ipsilateral screws between different lumbar vertebrae (such as the arrangement of the screw entry points, the difference in the screw rotation angle, etc.), the better the mechanical performance of the spinal screw placement, the linear arrangement reward function evaluates this indicator by calculating the arrangement variance of the screw entry points and the rotation angle, and rewards are given based on the design principle that the smaller and more uniform the difference, the higher the reward.

[0062] During the global design of the reward function, we added a global indicator, namely "linear alignment", to the reward function, which means that even if the agent only focuses on one of its own lumbar vertebrae, the behavior of a single agent will still be affected by the overall performance of the screw placement system, so as to achieve coordinated optimization of multiple lumbar vertebrae. For example, if the screw entry point of a lumbar vertebra deviates too much from the arrangement of the screws on the same side of other lumbar vertebrae, the entire system will be penalized, thus prompting the agent to learn to coordinate and deal with multiple factors and ultimately choose the optimal strategy.

[0063] In order to prevent the intelligent agent from being unable to adapt to complex goals in the early stages of training, the concept of gradual optimization is introduced. By adjusting the weight of the reward function, the intelligent agent is allowed to first focus on optimizing the surface area of ​​the screw and the bone density of the area through which the screw passes. After a single intelligent agent learns to handle basic goals, higher-order global optimization goals are gradually introduced.

[0064] In addition, when performing multi-agent reinforcement learning, the present invention constructs a reinforcement learning environment and designs a virtual "laboratory" to simulate the entire process of a doctor placing screws on a computer. In this environment, each lumbar vertebra is regarded as an "agent" whose task is to find the best screw placement path. Each lumbar vertebra is regarded as an agent. The agent can not only see the key information of the current lumbar vertebra, such as the starting and ending positions of the screws, the boundary line of the lumbar vertebra, and the bone density distribution, but also observe the positions of the screw entry points on the same side of other lumbar vertebrae. Based on the above observation information, the agent can update the position by constantly trying to adjust the angle and length of the screw to find the optimal solution. Actions of the agent: Just like the doctor manually adjusts the position of the screw, the agent also changes the screw placement position by selecting actions such as "rotate the screw 5° to the left".

[0065] Through the use of centralized training and distributed execution for learning and training,

[0066] In the centralized training, during the training phase, the agents responsible for different lumbar vertebrae will "cooperate" with each other, that is, the performance of each screw on the corresponding lumbar vertebra is integrated, and the placement performance of the overall screw system is analyzed through a shared value network, and a comprehensive evaluation is generated, which is used to guide training;

[0067] Decentralized Execution: During the screw placement process, although each agent can only control the screw position of the lumbar vertebra it is responsible for, just like a doctor can only operate the current lumbar vertebra but cannot pay attention to all lumbar vertebrae at the same time, each agent executes independently and can learn from the global information of the system reward feedback to make the best decision.

[0068] Step 4: Input CT image data into the optimized trained model to output the lumbar pedicle screw placement path planning scheme. The output of the screw placement path planning scheme includes the output of screw placement parameters, three-dimensional visualization, and quantification of biomechanical indicators, as follows:

[0069] The output of the screw placement parameters can output, for each lumbar vertebra, the screw information on the left and right sides of the lumbar vertebra, the radius and length of the screws, and the bone density of the area through which the screws pass, and the screw placement plan can be visualized in three dimensions on a three-dimensional model with a controllable observation angle. The doctor can intuitively see the placement position of the screws and the arrangement of the screws in the lumbar spine as a whole, and the biomechanical indicators are quantified, providing some numerical evaluation indicators (such as screw surface area, bone density, and linear arrangement), which can further help the doctor judge the pros and cons of the plan.

[0070] The above is a detailed introduction to a lumbar pedicle screw placement path planning method based on multi-agent reinforcement learning provided by the present invention. This article uses specific examples to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea; at the same time, for those skilled in the art, according to the idea of ​​the present invention, there will be changes in the specific implementation method and application scope, and changes and improvements to the present invention will be possible without exceeding the concept and scope specified in the attached claims. In summary, the content of this specification should not be understood as a limitation on the present invention.

Claims

1. A lumbar pedicle screw placement path planning method based on multi-agent reinforcement learning, characterized in that: The steps are as follows: Step 1: CT image data preprocessing; Using a coarse-to-fine network framework based on deep learning, the coarse-grained network is first used to predict the centroid feature points of the lumbar spine, and the lumbar part is automatically extracted from the complete 3D CT image of the patient's spine. Then, the fine-grained network is used to predict the four feature points of the central sagittal plane of each lumbar vertebra and the four feature points of the pedicle position as key data for screw placement path planning. Step 2: Use the eight feature points obtained in step 1 to construct bounding boxes of five lumbar vertebrae. Obtain single lumbar vertebrae data through the bounding boxes. Then perform a three-dimensional rigid body rotation operation on each lumbar vertebra to rotate the single lumbar vertebrae CT data and its segmentation mask to a positive horizontal state in space. Finally, use the above data to train the segmentation network. Step 3: Use the feature point coordinate data and the binary segmentation mask of the vertebra to locate the initial screw placement plane position and the geometric parameters of the screw. Then, based on the bone density distribution inside the vertebra, use the multi-agent reinforcement learning algorithm to collaboratively optimize the placement path of the screw system inside the five lumbar vertebrae, build a reinforcement learning environment and optimize the reward function. Step 4: Input CT image data into the optimized trained model to output the lumbar pedicle screw placement path planning solution.

2. The method for lumbar pedicle screw placement path planning based on multi-agent reinforcement learning according to claim 1, characterized in that: In step 1, during the preprocessing of CT image data, CT images of different patients are corrected and image standardized.

3. The method for lumbar pedicle screw placement path planning based on multi-agent reinforcement learning according to claim 1, characterized in that: In step three, the reward function includes a screw surface area reward function, a bone density reward function and a linear arrangement reward function.

4. The method for lumbar pedicle screw placement path planning based on multi-agent reinforcement learning according to claim 3 is characterized in that: The screw surface area reward function is designed based on the principle that the larger the contact area, the higher the reward. The specific formula is as follows: S=2π×R×L Where S is the contact area, R is the radius of the screw, and L is the length of the screw.

5. The method for lumbar pedicle screw placement path planning based on multi-agent reinforcement learning according to claim 3, characterized in that: The bone density reward function evaluates the regional bone density by reading the HU value of the corresponding area in the original CT image, and takes the sum of the density of the passed area as the reward.

6. The method for lumbar pedicle screw placement path planning based on multi-agent reinforcement learning according to claim 3, characterized in that: The linear arrangement reward function evaluates this indicator by calculating the arrangement variance of the screw insertion points and rotation angles, and rewards are given based on the design principle that the smaller and more orderly the differences, the higher the reward.

7. The method for lumbar pedicle screw placement path planning based on multi-agent reinforcement learning according to claim 6, characterized in that: By adjusting the weight of the reward function, the intelligent agent can first focus on optimizing the surface area of ​​the screw and the bone density of the area through which the screw passes. After a single intelligent agent learns to handle basic goals, high-order global optimization goals are gradually introduced.

8. The method for lumbar pedicle screw placement path planning based on multi-agent reinforcement learning according to claim 1, characterized in that: When performing multi-agent reinforcement learning, each lumbar vertebra is regarded as an agent, and learning and training are carried out using centralized training and distributed execution. In the centralized training, during the training phase, the performance of each screw on the corresponding lumbar vertebra is integrated, and the placement performance of the entire screw system is analyzed through a shared value network, and a comprehensive evaluation is generated, which is used to guide the training; In the distributed execution, during the pinning process, each agent executes independently and can learn from the global information of the system reward feedback to make the best decision.

9. The method for lumbar pedicle screw placement path planning based on multi-agent reinforcement learning according to claim 1, characterized in that: In step 4, the output of the pin placement path planning solution includes the output of pin placement parameters, three-dimensional visualization, and quantification of biomechanical indicators.

10. The method for lumbar pedicle screw placement path planning based on multi-agent reinforcement learning according to claim 9, characterized in that: The output of the screw placement parameters can output, for each lumbar vertebra, the screw information on the left and right sides of the lumbar vertebra, the radius and length of the screw, and the bone density of the area through which the screw passes, and perform a three-dimensional visualization of the screw placement plan on a three-dimensional model with a controllable observation angle, and quantify the biomechanical indicators to provide some numerical evaluation indicators.

Citation Information

Cited By

  • Preoperative planning quality control method for spinal screw placement, electronic equipment and computer storage medium

    CN122140368A