X-ray image visual prompting system for atrial septal puncture operation
Patent Information
- Application Number
- CN202510611764.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2026-08-28
- Estimated Expiration
- 2045-05-13
AI Technical Summary
[0004]为了解决现有方案用于房间隔穿刺术中需要额外辅助设备,导致需求响应不足,成本高,缺乏面向房间隔穿刺术的专用技术模块设计,缺乏动态调整能力,与实时变化的手术不适配的技术问题,本发明的目的在于提供一种房间隔穿刺手术用X射线影像视觉提示系统,所采用的技术方案具体如下:
[0055]本系统无需额外手术设备,适配术中真实环境,即不依赖三维建模设备、心腔内超声或电磁导航系统,直接基于术中常用的低剂量X射线透视图像进行推理与提示,便于广泛临床落地,特别适合资源受限的手术环境;识别解剖区域结构边界,并基于图像分割基础模型微调,通过视觉引导生成高置信度的候选穿刺点,形成建议穿刺窗,提高穿刺点决策模型定位的可解释性与稳定性;通过模拟解剖结构的运动轨迹,基于注意力机制和运动-语义解耦嵌入模块优化时空模型,实现房间隔穿刺点的动态预测;通过学习多专家穿刺点标注,系统可以预测可靠穿刺点范围而不是单一穿刺点,并将生成的安全穿刺区域概率图与动态视觉提示图进行联动可视化提示,实现术中个性化风险提示与操作精度保障,增强该系统用于临床的可靠性;通过增强现实技术,将安全穿刺区域概率图与动态视觉提示图叠加展示在术者视野中,实现“所见即建议”的辅助操作方式,降低学习门槛,提升年轻术者的操作信心和安全保障,增强现实视觉提示,优化人机协同效率,支持术者做出更精准、更安全的个体化决策。
Smart Images

Figure CN120543797B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of medical imaging technology, specifically to an X-ray imaging visual prompting system for atrial septal puncture surgery. Background Technology
[0002] Atrial septal puncture is a crucial step in many cardiac interventional procedures, such as transcatheter mitral valve repair, left atrial appendage occlusion, and arrhythmia ablation. Its purpose is to safely and precisely introduce a catheter into the left atrium through the atrial septum. Although this procedure is widely used, it presents certain technical challenges, especially in patients with complex anatomy or congenital lesions, where the risks increase and could lead to serious complications such as cardiac perforation and cardiac tamponade. X-ray fluoroscopy, as a traditional imaging method, remains the most commonly used guidance method for atrial septal puncture worldwide (excluding North America) due to its advantages such as immediate imaging, ease of operation, and the absence of additional training in ICE catheter manipulation. Particularly in centers without transesophageal echocardiography guidance, X-ray-guided atrial septal puncture remains the only viable option.
[0003] Traditional TSP (Transseptal Puncture) procedures rely heavily on the surgeon's experience and real-time interpretation of X-ray images. However, this method demands high levels of spatial imagination, image interpretation skills, and hand-eye coordination from the surgeon, resulting in a steep learning curve and limitations on surgical safety and efficiency. Inaccurate judgments can lead to serious complications or even patient death. In recent years, artificial intelligence, particularly deep learning, has made significant progress in medical image recognition and decision support, leading to several assisted puncture techniques. However, most of these rely on additional mechanical devices to achieve 3D imaging and improve puncture positioning accuracy, resulting in high costs and complex procedures. These limitations hinder deployment in grassroots settings or conventional operating rooms, limiting clinical application and implementation. The system lacks sufficient capability and responsiveness to demand. Furthermore, the introduced AI systems are still in their early stages. For example, the previous case CN118319486B proposed an AI-based cardiovascular interventional surgery image guidance system. It uses deep learning methods to analyze various cardiovascular interventional images, constructs a semi-automatic intelligent annotation tool and a cardiovascular interventional surgery image guidance system based on vascular structure segmentation and lesion area segmentation, and performs risk assessment for the surgery. However, it has not established a spatial model across cardiac chamber structures, nor has it involved dedicated algorithm design for puncture window recognition, optimal puncture angle assessment, etc. It lacks a dedicated technical module design for atrial septal puncture, and its design in terms of accuracy control and error tolerance mechanisms is weak. It also lacks image enhancement strategies for the detailed structure of the atrial septum and the ability to dynamically adjust during the operation. Summary of the Invention
[0004] To address the technical problems of existing solutions for atrial septal puncture, which require additional auxiliary equipment, resulting in insufficient responsiveness, high cost, lack of dedicated technical modules for atrial septal puncture, lack of dynamic adjustment capabilities, and incompatibility with real-time surgical changes, the present invention aims to provide an X-ray imaging visual prompting system for atrial septal puncture surgery. The specific technical solution adopted is as follows:
[0005] The data acquisition and preprocessing module is used to: acquire preoperative X-ray videos and perform preprocessing to obtain anatomical structure segmentation masks;
[0006] The data analysis and generation module is used to: construct a puncture point decision model, input the anatomical structure segmentation mask into the puncture point decision model, output a puncture area probability map, and generate candidate puncture points based on the puncture area probability map;
[0007] The data processing module is used to: track the anatomical structure segmentation mask, construct a trajectory set, encode the trajectory set to obtain an embedding vector set, determine the spatial average value, optimize the spatiotemporal model through a trajectory-semantic decoupling mechanism, combine the spatial average value and the anatomical structure segmentation mask to obtain a trajectory consistency loss function, output the prediction result based on the optimized spatiotemporal model, and use the trajectory consistency loss function to determine whether the prediction result matches the candidate puncture point.
[0008] The data prediction module is used to: construct a probability map of the safe puncture area based on multi-expert integration, and generate a dynamic visual cue map based on the prediction results matched with candidate puncture points;
[0009] The data augmentation and decision-making module is used to: overlay dynamic visual cues with a probability map of safe puncture areas using augmented reality technology, and provide real-time feedback from the computing terminal to enable collaborative operation between humans and the system.
[0010] Preferably, preoperative X-ray video is acquired and preprocessed to obtain an anatomical structure segmentation mask, including:
[0011] Acquire preoperative X-ray videos, extract each frame image and standardize its size;
[0012] Based on the identification of anatomical structure region boundaries in each frame of image, the main loss and auxiliary loss are introduced to determine the overall loss function. The basic image segmentation model is used to fine-tune the boundaries of the anatomical structure region to obtain the anatomical structure segmentation mask.
[0013] Preferably, the overall loss function is determined by introducing main loss and auxiliary loss respectively, and the boundaries of the anatomical structure region are fine-tuned using the basic image segmentation model to obtain the anatomical structure segmentation mask, including:
[0014] The formula for calculating the main loss is:
[0015]
[0016] in, Indicates the main loss; Indicates the total number of pixels in the image; Indicates the first 1 pixel; Indicates the first The real label of each pixel; Indicates the first Each pixel corresponds to a mask probability output by the base model based on image segmentation;
[0017] The formula for calculating the auxiliary loss is:
[0018]
[0019] in, Indicates auxiliary loss; To represent a constant, avoid having a denominator of zero;
[0020] The overall loss function is determined, and the corresponding calculation formula is as follows:
[0021]
[0022] in, Represents the total loss function; , These represent the adjustment weights corresponding to the main loss and auxiliary loss, respectively.
[0023] Preferably, a puncture point decision model is constructed, the anatomical structure segmentation mask is input into the puncture point decision model, the puncture area probability map is output, and candidate puncture points are generated based on the puncture area probability map, including:
[0024] The puncture point decision model uses a lightweight U-Net network;
[0025] Anatomical structure segmentation mask is used as a visual cue to generate a location cue map. Input data is determined based on each frame of image and the location cue map. The input data is fed into the puncture point decision model and outputs a puncture area probability map.
[0026] The target area is delineated based on the puncture area probability map, the puncture point decision model is trained, and the preoperative X-ray video is input into the trained puncture point decision model to generate candidate puncture points, forming a suggested puncture window.
[0027] Preferably, an anatomical structure segmentation mask is used as a visual cue to generate a location cue map. Input data is determined based on each frame of the image and the location cue map. This input data is then fed into a puncture point decision model, which outputs a puncture area probability map, including:
[0028] Using an anatomical segmentation mask as a visual cue, a location cue map is generated. The corresponding calculation formula is:
[0029]
[0030] in, Location hint image; , , Segmentation binary mask images representing segmentation masks for the heart contour, the midline of the spine, and the anatomical structure of the coronary sinus electrode catheter, respectively;
[0031] Define each frame as Based on the location hint map, the input data is determined, and the corresponding calculation formula is:
[0032]
[0033] in, Indicates input data;
[0034] Input the data into the puncture point decision model and output a probability map of the puncture area.
[0035] Preferably, the target area is delineated based on the puncture area probability map, and the puncture point decision model is trained, including:
[0036] Based on the probability map of the puncture area, any real puncture point is selected as the center, and a circle is defined with a radius to delineate the candidate area circle as the target area. The corresponding logical expression is:
[0037]
[0038] in, The coordinates of the center of the target region detected by the puncture point decision model are indicated. Represents the coordinates of the actual puncture point; Indicates the radius of the target region.
[0039] Preferably, the anatomical structure segmentation mask is tracked to construct a trajectory set. The trajectory set is then encoded to obtain an embedding vector set. A spatial average value is determined. The spatiotemporal model is optimized using a trajectory-semantic decoupling mechanism. The spatial average value and the anatomical structure segmentation mask are combined to obtain a trajectory consistency loss function. The prediction result is output based on the optimized spatiotemporal model. The trajectory consistency loss function is used to determine whether the prediction result matches the candidate puncture point, including:
[0040] Continuous frame tracking of the anatomical structure segmentation mask is performed to determine the trajectory. All trajectories corresponding to the anatomical structure segmentation mask are obtained to construct a trajectory set. The trajectory set is encoded to obtain an embedding vector set. The spatial average value of the trajectory of any anatomical structure in the preoperative X-ray video is determined.
[0041] The spatiotemporal model is optimized through a trajectory-semantic decoupling mechanism. The semantic center coordinates are determined based on the anatomical structure segmentation mask, and the trajectory consistency loss function is obtained by combining the spatial average value.
[0042] Based on the optimized spatiotemporal model, the prediction results are fused and output through a cross-attention mechanism.
[0043] Set a judgment threshold and use the trajectory consistency loss function to judge whether the prediction result matches the candidate puncture point. If the trajectory consistency loss function is less than the judgment threshold, it means that the prediction result matches the trajectory of the candidate puncture point.
[0044] Preferably, the trajectory is determined, and the corresponding calculation formula is:
[0045]
[0046] in, The first anatomical structure segmentation mask corresponds to the The trajectory of one pixel; Indicates the first The first anatomical structure segmentation mask corresponding to the intra-frame The coordinates of one pixel; This indicates the total number of frames corresponding to the preoperative X-ray video.
[0047] Preferably, the spatiotemporal model is optimized through a trajectory-semantic decoupling mechanism, the semantic center coordinates are determined based on the anatomical structure segmentation mask, and the trajectory consistency loss function is obtained by combining the spatial average value, including:
[0048] Define the anatomical structure segmentation mask as The semantic center coordinates are determined using the following formula:
[0049]
[0050] in, Indicates the coordinates of the semantic center;
[0051] The trajectory consistency loss function is obtained by combining spatial average values, and the corresponding calculation formula is as follows:
[0052]
[0053] in, Represents the trajectory consistency loss function; Indicates an index of anatomical structures; Indicates the first The semantic center coordinates corresponding to each anatomical structure; Indicates the first The spatial average value corresponding to each anatomical structure.
[0054] The present invention has the following beneficial effects:
[0055] This system requires no additional surgical equipment and adapts to real-world surgical environments. It does not rely on 3D modeling equipment, intracardiac ultrasound, or electromagnetic navigation systems; instead, it directly uses commonly used low-dose X-ray fluoroscopy images for reasoning and prompting, facilitating widespread clinical application and making it particularly suitable for resource-constrained surgical environments. It identifies anatomical region boundaries and fine-tunes a basic image segmentation model, generating high-confidence candidate puncture points through visual guidance to form a suggested puncture window, improving the interpretability and stability of the puncture point decision model. By simulating the motion trajectory of anatomical structures, it optimizes the spatiotemporal model based on attention mechanisms and a motion-semantic decoupling embedding module, enabling dynamic atrial septal puncture point detection. Prediction: By learning from the annotations of multiple experts' puncture points, the system can predict the range of reliable puncture points rather than a single puncture point. It also links the generated safe puncture area probability map with a dynamic visual cue map for visualization, enabling personalized risk alerts and ensuring operational accuracy during the procedure, thus enhancing the system's reliability in clinical use. Augmented reality technology overlays the safe puncture area probability map with the dynamic visual cue map in the surgeon's field of vision, achieving a "what you see is what you get" assisted operation method. This lowers the learning threshold, enhances the operational confidence and safety assurance of younger surgeons, optimizes human-machine collaboration efficiency, and supports surgeons in making more accurate and safer individualized decisions. Attached Figure Description
[0056] To more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0057] Figure 1 This is a schematic diagram showing the boundary of the anatomical structure region in an X-ray imaging visual prompting system for atrial septal puncture surgery according to an embodiment of the present invention.
[0058] Figure 2 This is a schematic diagram of the suggested puncture window of an X-ray imaging visual prompting system for atrial septal puncture surgery provided in an embodiment of the present invention;
[0059] Figure 3 This is a schematic diagram of the output results of the data prediction module of an X-ray imaging visual prompting system for atrial septal puncture surgery provided in an embodiment of the present invention. Detailed Implementation
[0060] To further illustrate the technical means and effects adopted by the present invention to achieve its intended purpose, the following, in conjunction with the accompanying drawings and preferred embodiments, details the specific implementation, structure, features, and effects of an X-ray imaging visual prompting system for atrial septal puncture surgery proposed according to the present invention. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. Furthermore, specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.
[0061] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0062] The following describes in detail, with reference to the accompanying drawings, a specific scheme of the X-ray imaging visual prompting system for atrial septal puncture surgery provided by the present invention, the system comprising:
[0063] The data acquisition and preprocessing module is used to: acquire preoperative X-ray videos and perform preprocessing to obtain anatomical structure segmentation masks;
[0064] The data analysis and generation module is used to: construct a puncture point decision model, input the anatomical structure segmentation mask into the puncture point decision model, output a puncture area probability map, and generate candidate puncture points based on the puncture area probability map;
[0065] The data processing module is used to: track the anatomical structure segmentation mask, construct a trajectory set, encode the trajectory set to obtain an embedding vector set, determine the spatial average value, optimize the spatiotemporal model through a trajectory-semantic decoupling mechanism, combine the spatial average value and the anatomical structure segmentation mask to obtain a trajectory consistency loss function, output the prediction result based on the optimized spatiotemporal model, and use the trajectory consistency loss function to determine whether the prediction result matches the candidate puncture point.
[0066] The data prediction module is used to: construct a probability map of the safe puncture area based on multi-expert integration, and generate a dynamic visual cue map based on the prediction results matched with candidate puncture points;
[0067] The data augmentation and decision-making module is used to: overlay dynamic visual cues with a probability map of safe puncture areas using augmented reality technology, and provide real-time feedback from the computing terminal to enable collaborative operation between humans and the system.
[0068] As an optional implementation, in this embodiment, the three major anatomical structures of the heart contour, the midline of the spine, and the coronary sinus electrode catheter are analyzed. Videos of the anatomical structures are obtained based on X-ray imaging technology, and images are extracted to provide a predicted puncture area for atrial septal surgery, thereby improving the safety and success rate of the surgery.
[0069] Please see Figure 1 This illustration shows a schematic diagram of the anatomical structure region boundaries of an X-ray imaging visual prompting system for atrial septal puncture surgery provided by an embodiment of the present invention, wherein red lines indicate the region boundaries of the heart outline; yellow lines indicate the region boundaries of the midline of the spine; and green lines indicate the region boundaries of the coronary sinus electrode catheter.
[0070] Furthermore, the data acquisition and preprocessing module includes:
[0071] Step S11: Acquire preoperative X-ray video, extract each frame image and standardize the size.
[0072] Preferably, in this embodiment, the data sample is a standard preoperative low-dose X-ray fluoroscopy dynamic video during surgery, which records the key moments before surgery and shows the real-time dynamic changes of organs and tissues in the patient's body. By using low-dose X-rays, high-quality image data is obtained while minimizing radiation exposure to patients and medical staff.
[0073] Specifically, standard preoperative low-dose X-ray fluoroscopy dynamic videos of interventional procedures were collected, and each frame of the video was extracted and uniformly cropped and normalized to 512×512 pixels.
[0074] Step S12: Identify the boundaries of the anatomical structure region based on each frame of the image, introduce the main loss and auxiliary loss respectively to determine the overall loss function, and use the basic image segmentation model to fine-tune the boundaries of the anatomical structure region to obtain the anatomical structure segmentation mask.
[0075] Preferably, the basic model for image segmentation includes three modules: an image encoder, a bounding box cue encoder, and a segmentation mask decoder. In order to adapt to the grayscale distribution and weak boundary features of anatomical structures in low-dose X-ray images, a channel separation attention mechanism is added to the segmentation mask decoder during fine-tuning to enhance the model's perception of regions with weak texture, such as the boundaries of the heart contour region.
[0076] Further, step S12 includes:
[0077] The formula for calculating the main loss is:
[0078]
[0079] in, Indicates the main loss; Indicates the total number of pixels in the image; Indicates the first 1 pixel; Indicates the first The real label of each pixel; Indicates the first Each pixel corresponds to a mask probability output by the base model based on image segmentation;
[0080] The formula for calculating the auxiliary loss is:
[0081]
[0082] in, Indicates auxiliary loss; To represent a constant, avoid having a denominator of zero;
[0083] The overall loss function is determined, and the corresponding calculation formula is as follows:
[0084]
[0085] in, Represents the total loss function; , These represent the adjustment weights corresponding to the main loss and auxiliary loss, respectively.
[0086] Specifically, based on the fine-tuning of the basic image segmentation model and expert image annotation, the X-ray image visual cues system used in atrial septal puncture surgery is guided to identify the boundaries of three major anatomical structures—the heart contour, the midline of the spine, and the coronary sinus electrode catheter—in a single frame of X-ray fluoroscopy video. During fine-tuning, a total loss function is used to simultaneously optimize the segmentation accuracy and boundary accuracy of the structural regions, resulting in a more accurate anatomical structure segmentation mask. The main loss uses binary cross-entropy, and the auxiliary loss is a structural consistency loss function, to ensure the consistency and accuracy of the acquired region boundaries with the real boundaries, thus facilitating more precise visual cues for subsequent atrial septal puncture surgery.
[0087] Please see Figure 2 The illustration shows a schematic diagram of the suggested puncture window of an X-ray imaging visual prompting system for atrial septal puncture surgery according to an embodiment of the present invention; wherein the marked box portion is the suggested puncture window.
[0088] Furthermore, the data analysis and generation module includes:
[0089] The puncture point decision model uses a lightweight U-Net network; its unique symmetrical structure enables feature information to be effectively transmitted in the network to improve segmentation accuracy. Furthermore, the lightweight U-Net network optimizes the traditional U-Net structure, reduces the number of parameters, and makes the model concise and efficient while maintaining high segmentation accuracy. In other words, it achieves fast and efficient image segmentation while ensuring accuracy.
[0090] Step S21: Use the anatomical structure segmentation mask as a visual cue to generate a location cue map. Based on each frame of image and the location cue map, determine the input data, input the input data into the puncture point decision model, and output the puncture area probability map.
[0091] Further, step S21 includes:
[0092] Step S211: Using the anatomical structure segmentation mask as a visual cue, generate a location cue map. The corresponding calculation formula is:
[0093]
[0094] in, Location hint image; , , Segmentation binary mask images representing segmentation masks for the heart contour, the midline of the spine, and the anatomical structure of the coronary sinus electrode catheter, respectively;
[0095] Step S212: Define each frame image as Based on the location hint map, the input data is determined, and the corresponding calculation formula is:
[0096]
[0097] in, Indicates input data;
[0098] Step S213: Input the input data into the puncture point decision model and output the puncture area probability map.
[0099] The method involves using an anatomical structure segmentation mask as a visual cue to generate a location cue map. Input data is determined based on each frame of the image and the location cue map, which can accurately obtain the specific location of the input data. After processing by the puncture point decision model, a puncture area probability map can be output to show the probability of puncture in each area of any anatomical structure. This allows for a more intuitive view of which sites have a high success rate for puncture and which have risks or uncertainties.
[0100] Step S22: Delineate the target area based on the puncture area probability map, train the puncture point decision model, and use the trained puncture point decision model to input the preoperative X-ray video to generate candidate puncture points and form a suggested puncture window.
[0101] Preferably, the training of the puncture point decision model also uses binary cross-entropy as the main loss function.
[0102] Further, in step S22, the target area is delineated based on the puncture area probability map, and the puncture point decision model is trained, including:
[0103] Based on the probability map of the puncture area, any real puncture point is selected as the center, and a circle is defined with a radius to delineate the candidate area circle as the target area. The corresponding logical expression is:
[0104]
[0105] in, The coordinates of the center of the target region detected by the puncture point decision model are indicated. Represents the coordinates of the actual puncture point; Indicates the radius of the target region.
[0106] Specifically, the target area is set as the learning objective of the puncture point decision model. The model is trained by combining the main loss function. Then, the preoperative X-ray video is input into the trained puncture point decision model to identify the area with the highest probability in the puncture area probability map, generate a series of candidate puncture points, and delineate suggested puncture windows for the candidate puncture points. This provides reference and guidance for the actual surgical process, effectively improving the accuracy and safety of puncture operation and enhancing the overall surgical outcome.
[0107] Understandably, while low-dose X-rays offer significant advantages in reducing radiation exposure for patients and surgeons, their imaging quality is somewhat compromised, particularly in anatomical structures such as the right lateral border of the heart, where blurred boundaries may occur. Therefore, this application incorporates visual trajectory modeling into each frame of the video, targeting the trajectory or path of anatomical structures to improve the diagnostic accuracy of X-ray fluoroscopy images. Specifically, it performs detailed analysis of cardiac motion and catheter pathways, extracting changes in catheter tip position and cardiac lateral border point changes across multiple frames of the video to form a long-range trajectory-like pattern, capturing the dynamic movement patterns of the heart and catheter. Simulating a target recognition method in dynamic video, the long-range trajectory-like pattern is input into a spatiotemporal model to model the dynamic changes of various anatomical structures and their spatial interactions. This helps improve the spatiotemporal model's ability to detect the relative positional stability of the anatomical structure's surrounding boundaries and predict the puncture point location, significantly enhancing the imaging quality of low-dose X-ray fluoroscopy images in key anatomical structures such as the cardiac contour, providing more reliable and accurate imaging support for clinical diagnosis and interventional procedures.
[0108] Furthermore, the data processing module includes:
[0109] Step S31: Perform continuous frame tracking on the anatomical structure segmentation mask to determine the trajectory, obtain all trajectories corresponding to the anatomical structure segmentation mask to construct a trajectory set, encode the trajectory set to obtain an embedding vector set, and determine the spatial average value of the trajectory of any anatomical structure in the preoperative X-ray video.
[0110] Specifically, in this embodiment, trajectory tracking is performed on the cardiac contour and the coronary sinus electrode catheter, that is, the trajectory of the outer edge of the heart and the trajectory of the tip of the electrode catheter are tracked continuously frame by frame to determine the trajectory and construct a trajectory set.
[0111] Further, in step S31, the trajectory is determined, and the corresponding calculation formula is:
[0112]
[0113] in, The first anatomical structure segmentation mask represents the first anatomical structure segmentation mask. The trajectory of one pixel; Indicates the first The first anatomical structure segmentation mask corresponding to the intra-frame The coordinates of one pixel; This indicates the total number of frames corresponding to the preoperative X-ray video.
[0114] Trajectory encoding is performed on the trajectory set to obtain an embedding vector set, denoted as . It is used to describe the movement patterns of the cardiac periphery and electrode catheters; at the same time, when predicting the puncture area, it should not deviate from the position of the anatomical structure output from the actual trajectory. That is, when using the trajectory of the anatomical structure to predict based on the spatiotemporal model, it should not deviate from the position of the candidate puncture point obtained by the data analysis and generation module, so as to avoid errors in the subsequent puncture surgery. Then, the spatial average value of the corresponding anatomical structure is confirmed by the trajectory confirmed in each frame and the total number of frames.
[0115] Step S32: Optimize the spatiotemporal model through the trajectory-semantic decoupling mechanism, determine the semantic center coordinates based on the anatomical structure segmentation mask, and obtain the trajectory consistency loss function by combining the spatial average value.
[0116] To explain, the trajectory-semantic decoupling mechanism, also known as the motion-semantic decoupling embedding technique, enables the spatiotemporal model to process trajectory motion features and semantic structural features separately. The trajectory motion features are the set of embedding vectors of the trajectory output of the anatomical structure; the semantic structural features are the semantic embeddings of the midline of the spine and the outer edge of the heart generated after fine-tuning based on the image segmentation model, so as to achieve highly robust dynamic structural localization.
[0117] Further, step S32 includes:
[0118] Step S321: Define the anatomical structure segmentation mask as The semantic center coordinates are determined using the following formula:
[0119]
[0120] in, Indicates the coordinates of the semantic center;
[0121] Step S322: Combine the spatial average value to obtain the trajectory consistency loss function, and the corresponding calculation formula is:
[0122]
[0123] in, Represents the trajectory consistency loss function; Indicates an index of anatomical structures; Indicates the first The semantic center coordinates corresponding to each anatomical structure; Indicates the first The spatial average value corresponding to each anatomical structure.
[0124] It can be explained that the introduction of the trajectory consistency loss function can ensure that the predicted puncture point area output by the optimized spatiotemporal model does not deviate too far from the actual outer edge of the heart or the center of gravity of the coronary sinus electrode catheter trajectory in adjacent frames. In other words, the trajectory consistency loss function can ensure that the spatiotemporal model maintains stability and coherence in the time dimension, avoid drastic fluctuations or unreasonable jumps in the prediction results between adjacent frames, and improve the reliability and accuracy of the spatiotemporal model prediction results.
[0125] Step S33: Based on the optimized spatiotemporal model, the prediction results are fused through a cross-attention mechanism. That is, the optimized spatiotemporal model uses a dual-branch structure to process two modalities: trajectory motion features and semantic structure features. Then, the outputs of the two modalities are fused through a cross-attention mechanism to obtain the prediction results. In other words, by analyzing the correlation between trajectory motion features and semantic structure features, the contribution of each modality to the final prediction results is dynamically adjusted to ensure that the optimized spatiotemporal model can capture the most relevant feature information while ignoring irrelevant or noisy data, so as to make full use of the complementarity between different modalities and improve the accuracy and reliability of the prediction results.
[0126] Step S34: Set a judgment threshold and use the trajectory consistency loss function to judge whether the prediction result matches the candidate puncture point. If the trajectory consistency loss function is less than the judgment threshold, it means that the prediction result matches the trajectory of the candidate puncture point.
[0127] The explanation is as follows: the judgment threshold is set according to the specific needs of the actual application. When the final value obtained by the trajectory consistency loss function is less than the judgment threshold, it means that the smaller the Euclidean distance between the semantic center coordinates and the spatial average value, the more reasonable the output prediction result is. This indicates that the prediction result matches the true trajectory, i.e., the trajectory of the candidate puncture point. It can be seen that the prediction result is relatively consistent with the actual observed trajectory in spatial location, verifying the accuracy of the prediction result. Conversely, if the trajectory consistency loss function is greater than the judgment threshold, the Euclidean distance between the semantic center coordinates and the spatial average value is large, indicating that there is a deviation between the prediction result and the true trajectory. The optimized spatiotemporal model needs to be adjusted to ensure that the output prediction result is reasonable.
[0128] Please see Figure 3 This diagram illustrates the output of the data prediction module of an X-ray image visual cues system for atrial septal puncture surgery according to an embodiment of the present invention. Frames 1, 2, and 3 refer to images extracted from any frame of the preoperative X-ray video. The color areas corresponding to the heart contour, spine, and coronary sinus electrode catheter are dynamic visual cues corresponding to the anatomical structures. The ideal puncture area is a probability map of the safe puncture area, and the high-risk puncture area is a high-risk area. Different frame numbers indicate different changes in the movement of the anatomical structures, and the corresponding cues in each frame also change accordingly.
[0129] Understandably, since each expert may have slightly different strategies for selecting puncture points during a transseptal puncture, although these puncture points are generally concentrated in a central area, the actual puncture point location may deviate from this center due to differences in each expert's habits and experience. Therefore, in order to more accurately simulate this puncture decision-making process in actual operation and provide the operator with a more reliable range of puncture point candidates, rather than just predicting a single puncture point location, this study adopts the puncture point annotations from multiple experienced operators on the same pre-puncture X-ray fluoroscopy video, trains an ensemble model simulating multiple experts, and constructs a probability map of safe puncture areas. This map reflects the most likely area for puncture point selection under the experience and habits of different experts; that is, combining historical puncture path data, a probability map of safe puncture areas is constructed based on multi-expert integration to more closely reflect the actual situation of patients.
[0130] Specifically, a regional risk assessment mechanism is introduced for the probability map of safe puncture areas. Heat maps are used to warn high-risk areas such as the atrial septum and the aorta, thus identifying high-risk puncture areas. Then, based on the prediction results confirmed in step S34, a dynamic visual cue map corresponding to the anatomical structure is generated to help the surgeon better identify and avoid potential risk areas during puncture operations, significantly improving the safety and success rate of the surgery.
[0131] As an optional implementation, in this embodiment, in order to achieve real-time feedback of data and spatiotemporal model prediction results, a computing terminal is configured independently of the surgical instruments. Real-time images captured by the camera are input into the computing terminal in real time via wireless transmission. The computing terminal outputs the detected anatomical structures of continuous frames as visual cues, and the prediction results are uploaded to the augmented reality terminal.
[0132] Furthermore, augmented reality (AR) technology is used to overlay dynamic visual cues with a probability map of safe puncture areas, and real-time feedback is provided by a computing terminal to enable collaborative operation between humans and the system. AR technology overlays the acquired dynamic visual cues, probability maps of safe puncture areas, and high-risk puncture areas onto the X-ray fluoroscopic images in the surgeon's field of vision in real time. This allows the surgeon to see various important information more intuitively. In other words, AR technology overlays visual cues onto the real-time X-ray images observed by the surgeon, enabling visualized puncture point markings, warning boundary prompts, and navigation arrow feedback. The computing terminal enhances the output of prediction results, combining real-time feedback with visual cues. This allows the surgeon to make decisions based on terminal feedback, creating a human-centered "intelligent augmented reality" system. This system does not rely on external 3D imaging equipment or navigation platforms, and its lightweight software and visual cues interaction method enable low-threshold deployment, making it widely applicable to routine interatrial septal puncture procedures in both grassroots and large centers.
[0133] Understandably, this system requires no additional surgical equipment and adapts to the real-world surgical environment. It does not rely on 3D modeling equipment, intracardiac ultrasound, or electromagnetic navigation systems; instead, it directly uses commonly used low-dose X-ray fluoroscopy images for reasoning and prompting, facilitating widespread clinical application and making it particularly suitable for resource-constrained surgical environments. It identifies anatomical region boundaries and fine-tunes a basic image segmentation model, generating high-confidence candidate puncture points through visual guidance to form a suggested puncture window, improving the interpretability and stability of the puncture point decision model. By simulating the motion trajectory of anatomical structures, it optimizes the spatiotemporal model based on attention mechanisms and a motion-semantic decoupling embedding module to achieve atrial septal puncture. The system features dynamic prediction; by learning from the annotations of multiple experts' puncture points, it can predict the range of reliable puncture points rather than a single puncture point, and links the generated safe puncture area probability map with the dynamic visual cues for visualization, realizing personalized risk warnings and ensuring operational accuracy during the operation, thus enhancing the reliability of the system in clinical use; through augmented reality technology, the safe puncture area probability map and the dynamic visual cues are overlaid and displayed in the surgeon's field of vision, realizing a "what you see is what you get" assisted operation mode, lowering the learning threshold, improving the operational confidence and safety of young surgeons, and optimizing human-machine collaboration efficiency, supporting surgeons to make more accurate and safer individualized decisions.
[0134] It should be noted that the order of the above embodiments of the present invention is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. The processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0135] The various embodiments in this specification are described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on describing the differences from other embodiments.
Claims
1. An X-ray imaging visual prompting system for atrial septal puncture surgery, characterized in that, The system includes: The data acquisition and preprocessing module is used to: acquire preoperative X-ray videos and perform preprocessing to obtain anatomical structure segmentation masks; The data analysis and generation module is used to: construct a puncture point decision model, input the anatomical structure segmentation mask into the puncture point decision model, output a puncture area probability map, and generate candidate puncture points based on the puncture area probability map; The data processing module is used for: tracking the anatomical structure segmentation mask, constructing a trajectory set, encoding the trajectory set to obtain an embedding vector set, determining the spatial average value, optimizing the spatiotemporal model through a trajectory-semantic decoupling mechanism, combining the spatial average value and the anatomical structure segmentation mask to obtain a trajectory consistency loss function, outputting prediction results based on the optimized spatiotemporal model, and using the trajectory consistency loss function to determine whether the prediction results match the candidate puncture points, including: Continuous frame tracking is performed on the anatomical structure segmentation mask to determine the trajectory. A trajectory set is constructed by acquiring all trajectories corresponding to the anatomical structure segmentation mask. The trajectory set is then encoded to obtain an embedding vector set. The spatial average value of the trajectory of any anatomical structure in the preoperative X-ray video is determined. The formula for determining the trajectory is as follows: ; in, The first anatomical structure segmentation mask corresponds to the The trajectory of one pixel; Indicates the first The first anatomical structure segmentation mask corresponding to the intra-frame The coordinates of one pixel; This indicates the total number of frames corresponding to the preoperative X-ray video; The spatiotemporal model is optimized through a trajectory-semantic decoupling mechanism. Semantic center coordinates are determined based on an anatomical structure segmentation mask, and a trajectory consistency loss function is obtained by combining spatial average values. This function includes: Define the anatomical structure segmentation mask as The semantic center coordinates are determined using the following formula: ; in, Indicates the coordinates of the semantic center; The trajectory consistency loss function is obtained by combining spatial average values, and the corresponding calculation formula is as follows: ; in, Represents the trajectory consistency loss function; Indicates an index of anatomical structures; Indicates the first The semantic center coordinates corresponding to each anatomical structure; Indicates the first The spatial average value corresponding to each anatomical structure; Based on the optimized spatiotemporal model, the prediction results are fused and output through a cross-attention mechanism. Set a judgment threshold and use the trajectory consistency loss function to judge whether the prediction result matches the candidate puncture point. If the trajectory consistency loss function is less than the judgment threshold, it means that the prediction result matches the trajectory of the candidate puncture point. The data prediction module is used to: construct a probability map of the safe puncture area based on multi-expert integration, and generate a dynamic visual cue map based on the prediction results matched with candidate puncture points; The data augmentation and decision-making module is used to: overlay dynamic visual cues with a probability map of safe puncture areas using augmented reality technology, and provide real-time feedback from the computing terminal to enable collaborative operation between humans and the system.
2. The X-ray imaging visual prompting system for atrial septal puncture surgery according to claim 1, characterized in that, Preoperative X-ray videos were acquired and preprocessed to obtain anatomical segmentation masks, including: Acquire preoperative X-ray videos, extract each frame image and standardize its size; Based on the identification of anatomical structure region boundaries in each frame of image, the main loss and auxiliary loss are introduced to determine the overall loss function. The basic image segmentation model is used to fine-tune the boundaries of the anatomical structure region to obtain the anatomical structure segmentation mask.
3. The X-ray imaging visual prompting system for atrial septal puncture surgery according to claim 1, characterized in that, The overall loss function is determined by introducing main loss and auxiliary loss respectively. The boundaries of the anatomical structure regions are fine-tuned using the basic image segmentation model to obtain the anatomical structure segmentation mask, including: The formula for calculating the main loss is: ; in, Indicates the main loss; Indicates the total number of pixels in the image; Indicates the first 1 pixel; Indicates the first The real label of each pixel; Indicates the first Each pixel corresponds to a mask probability output by the base model based on image segmentation; The formula for calculating the auxiliary loss is: ; in, Indicates auxiliary loss; To represent a constant, avoid having a denominator of zero; The overall loss function is determined, and the corresponding calculation formula is as follows: ; in, Represents the total loss function; , These represent the adjustment weights corresponding to the main loss and auxiliary loss, respectively.
4. The X-ray imaging visual prompting system for atrial septal puncture surgery according to claim 1, characterized in that, A puncture point decision model is constructed. An anatomical structure segmentation mask is input into the model, which outputs a puncture region probability map. Candidate puncture points are generated based on this probability map, including: The puncture point decision model uses a lightweight U-Net network; Anatomical structure segmentation mask is used as a visual cue to generate a location cue map. Input data is determined based on each frame of image and the location cue map. The input data is fed into the puncture point decision model and outputs a puncture area probability map. The target area is delineated based on the puncture area probability map, the puncture point decision model is trained, and the preoperative X-ray video is input into the trained puncture point decision model to generate candidate puncture points, forming a suggested puncture window.
5. The X-ray imaging visual prompting system for atrial septal puncture surgery according to claim 4, characterized in that, Anatomical structure segmentation masks are used as visual cues to generate location cue maps. Input data is determined based on each frame of the image and the location cue map. This input data is then fed into the puncture point decision model, which outputs a puncture region probability map, including: Using an anatomical segmentation mask as a visual cue, a location cue map is generated. The corresponding calculation formula is: ; in, Location hint image; , , Segmentation binary mask images representing segmentation masks for the heart contour, the midline of the spine, and the anatomical structure of the coronary sinus electrode catheter, respectively; Define each frame as Based on the location hint map, the input data is determined, and the corresponding calculation formula is: ; in, Indicates input data; Input the data into the puncture point decision model and output a probability map of the puncture area.
6. The X-ray imaging visual prompting system for atrial septal puncture surgery according to claim 4, characterized in that, The target region is delineated based on the puncture area probability map, and a puncture point decision model is trained, including: Based on the probability map of the puncture area, any real puncture point is selected as the center, and a circle is defined with a radius to delineate the candidate area circle as the target area. The corresponding logical expression is: ; in, The coordinates of the center of the target region detected by the puncture point decision model are indicated. Represents the coordinates of the actual puncture point; Indicates the radius of the target region.
Citation Information
Patent Citations
An artificial intelligence-based image guidance system for cardiovascular interventional surgery
CN118319486B
Patient rehabilitation training data acquisition method and system based on visual identification
CN119446391A
Positioning puncture surgery robot system with reconfigurable working space
CN119856983A