Transform-based atrial septum puncture needle tracking method, system and program product
Through the Transformer-based method, the position of the puncture needle is tracked in real time, which solves the problems of low tracking accuracy and insufficient real-time performance in the prior art, and improves the safety and efficiency of the surgery.
Patent Information
- Application Number
- CN202411968607.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-23
AI Technical Summary
The prior art is difficult to accurately track the position of the puncture needle when processing complex ultrasound images, resulting in increased surgical risks and prolonged operation time.
The atrial septum puncture needle tracking method based on Transformer is used to track the position of the puncture needle in real time through steps such as data preprocessing, feature extraction, Transformer encoding and position prediction.
It improves the tracking accuracy and robustness of the puncture needle, meets the needs of real-time tracking, and reduces surgical risks and operating time.
Smart Images

Figure CN120032080A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of medical image processing and computer-assisted puncture, and more specifically, to a Transformer-based atrial septum puncture needle tracking method, system and program product. Background Art
[0002] Transseptal Puncture (TSP) is an interventional cardiology procedure used to treat a variety of congenital heart diseases. In this procedure, the doctor uses a puncture needle to pass through the esophagus through the atrial septum (the thin wall between the left and right atria of the heart) into the right atrium under the guidance of transesophageal echocardiography (TEE), thereby establishing a channel between the left and right atria. Transesophageal echocardiography (TEE) has become the preferred imaging method for guiding transseptal puncture due to its clear imaging and high degree of visualization. However, due to factors such as the low signal-to-noise ratio of ultrasound images and the puncture needle tip being easily obstructed by tissue, it is very difficult to manually track the trajectory of the puncture needle tip, which increases the risk and operation time of the operation.
[0003] Existing puncture needle tracking methods mainly rely on image processing techniques, such as image segmentation, feature extraction, and template matching. However, these methods have the following shortcomings when processing complex ultrasound images:
[0004] 1. Poor robustness: Ultrasound images have large noise and complex tissue structures, and traditional image processing methods are difficult to accurately extract the features of puncture needles. Due to the presence of image noise, traditional image processing methods often find it difficult to effectively segment the contour of the puncture needle, resulting in inaccurate tracking results.
[0005] 2. Weak generalization ability: Template matching methods that rely on specific image features are difficult to adapt to the anatomical differences between different patients. The cardiac structure and tissue characteristics of different patients may vary greatly, and traditional template matching methods often have difficulty adapting to these differences, resulting in unstable tracking results.
[0006] 3. Lack of real-time performance: The multi-stage image processing process has a large amount of computation and is difficult to meet the needs of real-time tracking. Traditional image processing methods usually require multiple stages of processing, such as preprocessing, feature extraction, matching, etc. These processing steps have a large amount of computation and are difficult to achieve real-time tracking.
[0007] The technical solutions disclosed in the invention patent applications with publication numbers CN 116051538 A, CN 116912494 A and CN 118396936 A are mainly focused on the image analysis, segmentation or classification of echocardiograms, and are respectively applied to the semantic segmentation of the left ventricle, the fusion of different image information for the classification of cardiac parts, and the image segmentation of the four-chamber heart structure. The purpose of their invention is to improve the segmentation accuracy and analysis effect of echocardiogram images to assist diagnosis. Although they use the Transformer model, the focus is on the segmentation and classification of static images, that is, the processing of a single ultrasound image. These technologies extract and fuse multi-scale features of images through multiple modules (such as Swin Transformer, convolution fusion module, etc.) to perform left ventricle segmentation or classification of cardiac structures. The key technology of these solutions is how to efficiently extract information of fixed structures in images and improve segmentation accuracy, such as multi-scale feature extraction based on Transformer, semantic segmentation of images, etc. The goal is to improve the image processing and analysis process, but it cannot process dynamically changing structural information.
[0008] Since the position of the puncture needle changes dynamically during the operation, how to accurately and in real time track the position of the puncture needle from the ultrasound image, overcome technical obstacles such as low signal-to-noise ratio, tissue occlusion and real-time requirements, and develop a method and system for dynamic tracking of the atrial septal puncture needle to improve the tracking accuracy and real-time performance of the puncture needle during surgery is still a technical problem that needs to be solved urgently. Summary of the invention
[0009] In view of the problems existing in the prior art, the present invention proposes a Transformer-based atrial septal puncture needle tracking method, system and program product, which aims to track the position of the puncture needle in real time through transesophageal echocardiography, improve the tracking accuracy, robustness and real-time performance of the puncture needle during surgery, and ensure the safety and efficiency of atrial septal puncture.
[0010] To achieve the above objectives, in a first aspect, the present invention provides a Transformer-based atrial septal puncture needle tracking method, comprising:
[0011] Step S1, data preprocessing: performing preprocessing including image cropping, denoising and image enhancement on the acquired real-time TEE image sequence;
[0012] Step S2, feature extraction: extracting deep features from the preprocessed TEE image using a pretrained convolutional neural network; the convolutional neural network automatically learns semantic information in the image, including features of the puncture needle and surrounding tissues;
[0013] Step S3, Transformer encoding: inputting the extracted deep feature sequence into the Transformer encoder; the Transformer encoder is a sequence model based on a self-attention mechanism, which learns the contextual relationship of the puncture needle in time and space;
[0014] Step S4, position prediction: using a Transformer decoder to decode the encoded features and predict the position coordinates of the puncture needle tip in the next frame of TEE image; the Transformer decoder predicts the position of the puncture needle tip in the next frame of image based on the previous feature sequence and context information;
[0015] Step S5, trajectory generation: connecting the continuously predicted puncture needle tip positions to generate a complete puncture needle tip trajectory; the trajectory is used to guide the doctor to accurately track the position of the puncture needle during the operation.
[0016] Furthermore, in step S2, the convolutional neural network is a ResNet18 model, and the last fully connected layer is replaced by a feature map.
[0017] Furthermore, in step S3, in the Transformer encoder, the feature vector of each position is processed by a multi-head self-attention mechanism; the multi-head self-attention mechanism simultaneously considers the relationship between multiple positions in the input sequence.
[0018] Furthermore, in step S3, for position i, its self-attention vector is obtained by calculating the similarity between the feature vectors of all positions in the input sequence; the similarity is calculated using a dot product attention mechanism or a bilinear attention mechanism, and the formula is expressed as follows:
[0019]
[0020] Among them, $Q$ represents the query vector, $K$ represents the key vector, $V$ represents the value vector, $ $ represents the dimension of the key vector; by calculating the similarity, the attention distribution of each position is obtained, and then it is weighted and summed with the value vector to obtain the self-attention vector of position i.
[0021] Furthermore, in step S4, the self-attention mechanism is also used in the Transformer decoder, including several layers of Transformer decoding layers; in a decoding layer, in addition to considering the feature vector output by the encoder, it is also necessary to consider the feature vector output by the previous decoding layer.
[0022] Furthermore, in step S5, a fully connected layer is used to map the output of the Transformer decoder to a two-dimensional coordinate (x, y) representing the position of the puncture needle tip.
[0023] In a second aspect, the present invention provides a Transformer-based atrial septal puncture needle tracking system, which uses the Transformer-based atrial septal puncture needle tracking method as described above to dynamically track the atrial septal puncture needle, including:
[0024] The image preprocessing module performs preprocessing including image cropping, denoising and image enhancement on the acquired real-time TEE image sequence;
[0025] Feature extraction module, which uses a pre-trained convolutional neural network to extract deep features from pre-processed TEE images;
[0026] Transformer encoding module, which inputs the extracted deep feature sequence into the Transformer encoder;
[0027] The position prediction module uses the Transformer decoder to decode the encoded features, and the position prediction head predicts the position coordinates of the puncture needle tip in the next frame of TEE image;
[0028] The trajectory generation module connects the continuously predicted puncture needle tip positions to generate a complete puncture needle tip trajectory.
[0029] In a final aspect, the present invention provides a computer program product. When the computer program product is run on a computer, the computer is enabled to execute the Transformer-based transseptal puncture needle tracking method as described above.
[0030] Compared with the prior art, the present invention has the following technical effects:
[0031] 1. Improved accuracy: Compared with traditional image processing methods such as image segmentation, feature extraction, and template matching, Transformer can better capture the semantic information of the puncture needle and surrounding tissue, thereby improving tracking accuracy. Transformer's self-attention mechanism can learn the contextual relationship of the puncture needle in time and space, further improving tracking accuracy.
[0032] 2. Enhanced robustness: Traditional image processing methods are easily affected by factors such as noise and tissue occlusion when processing complex ultrasound images, resulting in unstable tracking results. The Transformer-based method can improve the robustness of tracking by learning a large amount of training data, and can better adapt to the anatomical differences of different patients and the presence of image noise.
[0033] 3. Improved real-time performance: Traditional image processing methods usually require multiple stages of processing, which requires a large amount of calculation and is difficult to meet the needs of real-time tracking. The Transformer-based method can improve real-time performance through parallel computing and other technologies, and can track the position of the puncture needle in real-time TEE image sequences.
[0034] 4. Strong scalability: The Transformer-based method can flexibly adapt to different ultrasound devices and image acquisition methods, and has good scalability. By adjusting the parameters of the model and the format of the input data, the method can be applied to different types of atrial septal puncture and other interventional surgeries. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 It is a schematic diagram of the software and hardware collaborative system of the present invention.
[0036] Figure 2 The figure is a flow chart of the atrial septal puncture needle tracking method of the present invention.
[0037] Figure 3 It is a schematic diagram of the effect of an embodiment of the present invention. DETAILED DESCRIPTION
[0038] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments, but they are not intended to limit the present invention.
[0039] In the following detailed description, many specific details are set forth to provide a more thorough understanding of the present invention. However, it is obvious to those skilled in the art that well-known algorithms and models are not shown in detail to avoid obscuring the subject matter of the present invention.
[0040] In addition, the order of execution of actions, steps, etc. in the devices and methods shown in the claims, specifications and drawings can be implemented in any order as long as there is no special explicit limitation on the order and the output of the previous processing is not used in the subsequent processing.
[0041] Example 1
[0042] See also Figure 1 and Figure 2 This embodiment provides a method for tracking a transseptal puncture needle based on Transformer, and the specific implementation method is as follows:
[0043] Step S1: Data preprocessing:
[0044] - Image cropping: Crop out non-critical areas in the TEE image and keep only the area containing the puncture needle and surrounding tissues.
[0045] - Denoising: Use filters or denoising algorithms to process images to reduce noise interference in the image.
[0046] - Image enhancement: Use image enhancement techniques, such as contrast adjustment, brightness adjustment, etc., to improve the visualization effect of the image.
[0047] Step S2, feature extraction: Use a pre-trained convolutional neural network (CNN) to extract deep features from the pre-processed TEE image. CNN can automatically learn the semantic information in the image, including the features of the puncture needle and surrounding tissue.
[0048] Use the pre-trained CNN model to extract features. Suppose the input image is (I) and the CNN model is (f_{CNN}), then the feature extraction process is:
[0049] [ ]
[0050] Where (F) is the extracted feature map.
[0051] Step S3, Transformer encoding: Input the extracted deep feature sequence into the Transformer encoder. Transformer is a sequence model based on the self-attention mechanism, which can effectively learn the contextual relationship in the sequence. In this method, the Transformer encoder can learn the contextual relationship of the puncture needle in time and space, and improve the tracking accuracy.
[0052] First, the feature map (F) is converted into a feature sequence (S), and then the Transformer encoder (T_{enc}) is applied:
[0053] [ ]
[0054] [ ]
[0055] Among them, (Z) is the encoded feature sequence.
[0056] Transformer encoder formula
[0057] Suppose the input sequence is (X), the query, key, and value matrices are (Q, K, V), and the self-attention mechanism is calculated as:
[0058] [ ]
[0059] [ ]
[0060] in( ) is the learned weight matrix and (h) is the number of attention heads.
[0061] The principle of Transformer is mainly based on the self-attention mechanism, which learns the contextual relationship in the sequence by paying self-attention to the information at each position in the sequence. Specifically, Transformer consists of an encoder and a decoder, where the encoder is used to extract the features of the input sequence and the decoder is used to generate the output sequence.
[0062] In the encoder, the feature vector of each position is processed by a multi-head self-attention mechanism. The multi-head self-attention mechanism can simultaneously consider the relationship between multiple positions in the input sequence. Specifically, for position i, its self-attention vector is obtained by calculating the similarity between the feature vectors of all positions in the input sequence. The similarity can be calculated using the dot product attention mechanism or the bilinear attention mechanism.
[0063] The formula is as follows:
[0064]
[0065] Among them, $Q$ represents the query vector, $K$ represents the key vector, $V$ represents the value vector, $ $ represents the dimension of the key vector. By calculating the similarity, the attention distribution of each position can be obtained, and then it is weighted summed with the value vector to obtain the self-attention vector of position i.
[0066] Step S4, position prediction: Use the Transformer decoder to decode the encoded features and predict the position coordinates of the puncture needle tip in the next frame of TEE image. The Transformer decoder can predict the position of the puncture needle tip in the next frame of image based on the previous feature sequence and context information.
[0067] In the decoder, the prediction result of the next position can be generated by inputting the feature vector output by the encoder into the decoder. The self-attention mechanism is also used in the decoder, but in the decoder, in addition to considering the feature vector output by the encoder, it is also necessary to consider the feature vector output by the previous decoder. In this way, contextual information can be used to generate more accurate prediction results.
[0068] The formula is as follows:
[0069]
[0070] Among them, $Decoder(i)$ represents the output of the decoder at position i, $Q$ represents the query vector, $K$ represents the key vector, $V$ represents the value vector, and $FeedForward$ represents the feedforward neural network. The decoder generates the final prediction result through a multi-layer stacked self-attention mechanism and feedforward neural network.
[0071] Using the Transformer decoder ( ) to predict the position of the puncture needle tip (P):
[0072] [ ]
[0073] Transformer encoder formula
[0074] The output of the decoder is calculated based on the output of the encoder (Z) and the previous output (Y):
[0075] [ ]
[0076] [ ]
[0077] Among them, (Linear) is the final fully connected layer, which is used to output position prediction.
[0078] Step S5, trajectory generation: Connect the continuously predicted puncture needle tip positions to generate a complete puncture needle tip trajectory. The trajectory can be used to guide the doctor to accurately track the position of the puncture needle during the operation, thereby improving the safety and efficiency of the operation.
[0079] The predicted position ( ) are linked in chronological order:
[0080] [ ]
[0081] Where (T) is the total number of time steps.
[0082] Through the above transformer encoding and position prediction steps, the puncture needle tip can be tracked in the TEE image sequence. The data preprocessing, feature extraction, transformer encoding and position prediction steps in the specific implementation can be adapted to different atrial septal punctures and other interventional surgeries by adjusting the parameters of the model and the format of the input data.
[0083] A preferred embodiment is provided below. Assumption: We use a simplified 2D TEE image sequence, and the size of each frame is 256x256 pixels. The puncture needle tip is clearly visible in the image and has been preprocessed (cropping, denoising, and enhancement).
[0084] Model design:
[0085] CNN feature extractor: Use a pre-trained ResNet18 model and replace the last fully connected layer with a feature map with an output dimension of 512.
[0086] Transformer Encoder: Contains 2 Transformer Encoder layers, each layer contains 8 attention heads, and the hidden layer dimension is 512.
[0087] Transformer decoder: Contains 2 layers of Transformer Decoder layer, each layer contains 8 attention heads, and the hidden layer dimension is 512.
[0088] Position prediction head: A fully connected layer is used to map the decoder output to 2D coordinates (x, y) representing the position of the puncture needle tip.
[0089] Training process:
[0090] Data preparation: A TEE image sequence dataset containing annotations of the puncture needle tip trajectory was collected.
[0091] Data enhancement: Randomly rotate, scale, translate, and so on the image to increase the diversity of the data.
[0092] Model training: Use Adam optimizer, learning rate 0.0001, batch size 32, and train for 100 epochs.
[0093] Loss function: Use the mean squared error (MSE) loss function to calculate the distance between the predicted position and the true position.
[0094] Reasoning process:
[0095] Input: Input a frame of TEE image.
[0096] Feature Extraction: Use CNN feature extractor to extract image features.
[0097] Transformer encoding: Input the features into the Transformer encoder to obtain the encoded feature sequence.
[0098] Transformer decoding: The encoded feature sequence is input into the Transformer decoder, and combined with the prediction results at the previous moment, the position of the puncture needle tip at the current moment is predicted.
[0099] Trajectory generation: See Figure 3 , connect the consecutive predicted position points to form a complete trajectory.
[0100] From the above method steps and test results, it can be seen that the present invention tracks the position of the puncture needle in real time through transesophageal echocardiography (TEE), overcomes the technical obstacles of low signal-to-noise ratio, tissue occlusion and high real-time requirements, to ensure the accuracy and safety of atrial septal puncture, improves the tracking accuracy and real-time performance of the puncture needle during surgery, and reduces surgical risks.
[0101] Example 2
[0102] This embodiment provides a Transformer-based atrial septal puncture needle tracking system, which is used in conjunction with a hardware system such as an atrial septal puncture needle and a TEE imaging system, and uses the Transformer-based atrial septal puncture needle tracking method described in Example 1 to dynamically track the atrial septal puncture needle, including:
[0103] The image preprocessing module performs preprocessing including image cropping, denoising and image enhancement on the acquired real-time TEE image sequence;
[0104] Feature extraction module, which uses a pre-trained convolutional neural network to extract deep features from pre-processed TEE images;
[0105] Transformer encoding module, which inputs the extracted deep feature sequence into the Transformer encoder;
[0106] The position prediction module uses the Transformer decoder to decode the encoded features, and the position prediction head predicts the position coordinates of the puncture needle tip in the next frame of TEE image;
[0107] The trajectory generation module connects the continuously predicted puncture needle tip positions to generate a complete puncture needle tip trajectory.
[0108] The above-mentioned Transformer-based transseptal puncture needle tracking method can be embodied in the form of a computer program product or a software functional unit. If the above-mentioned Transformer-based transseptal puncture needle tracking method is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Therefore, the essence of this technical solution or the part that contributes to the prior art or the part of this technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including several instructions for an electronic system (which can be a personal computer, a server, or a network system, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program codes.
[0109] Those skilled in the art will appreciate that the units, i.e., algorithm steps, of the various examples described in conjunction with this embodiment can be implemented in electronic hardware or in a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0110] Those skilled in the art should understand that those skilled in the art can implement variations by combining the prior art and the above embodiments, which will not be described in detail here. Such variations do not affect the essential content of the present invention, and will not be described in detail here.
[0111] The above describes the preferred embodiments of the present invention. It should be understood that the present invention is not limited to the above-mentioned specific embodiments, and the systems and structures that are not described in detail should be understood to be implemented in a common manner in the art; any technician familiar with the art can use the above-disclosed methods and technical contents to make many possible changes and modifications to the technical solutions of the present invention without departing from the scope of the technical solutions of the present invention, or modify them into equivalent embodiments of equivalent changes, which does not affect the essential content of the present invention. Therefore, any simple modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solutions of the present invention are still within the scope of protection of the technical solutions of the present invention.
Claims
1. A method for tracking atrial septal puncture needle based on Transformer, characterized in that: include: Step S1, data preprocessing: performing preprocessing including image cropping, denoising and image enhancement on the acquired real-time TEE image sequence; Step S2, feature extraction: extracting deep features from the preprocessed TEE image using a pretrained convolutional neural network; the convolutional neural network automatically learns semantic information in the image, including features of the puncture needle and surrounding tissues; Step S3, Transformer encoding: inputting the extracted deep feature sequence into the Transformer encoder; the Transformer encoder is a sequence model based on a self-attention mechanism, which learns the contextual relationship of the puncture needle in time and space; Step S4, position prediction: using a Transformer decoder to decode the encoded features and predict the position coordinates of the puncture needle tip in the next frame of TEE image; the Transformer decoder predicts the position of the puncture needle tip in the next frame of image based on the previous feature sequence and context information; Step S5, trajectory generation: connecting the continuously predicted puncture needle tip positions to generate a complete puncture needle tip trajectory; the trajectory is used to guide the doctor to accurately track the position of the puncture needle during the operation.
2. The method for tracking a transseptal puncture needle based on Transformer according to claim 1, characterized in that: In step S2, the convolutional neural network is a ResNet18 model, and the last fully connected layer is replaced by a feature map.
3. The method for tracking a transseptal puncture needle based on Transformer according to claim 2, characterized in that: In step S3, in the Transformer encoder, the feature vector of each position is processed by a multi-head self-attention mechanism; the multi-head self-attention mechanism simultaneously considers the relationship between multiple positions in the input sequence.
4. The method for tracking a transseptal puncture needle based on Transformer according to claim 3, characterized in that: In step S3, for position i, its self-attention vector is obtained by calculating the similarity between the feature vectors of all positions in the input sequence; the similarity is calculated using a dot product attention mechanism or a bilinear attention mechanism, and the formula is as follows: Among them, $Q$ represents the query vector, $K$ represents the key vector, $V$ represents the value vector, $ $ represents the dimension of the key vector; by calculating the similarity, the attention distribution of each position is obtained, and then it is weighted and summed with the value vector to obtain the self-attention vector of position i.
5. The method for tracking a transseptal puncture needle based on Transformer according to claim 3, characterized in that: In step S4, the self-attention mechanism is also used in the Transformer decoder, including several layers of Transformer decoding layers; in a decoding layer, in addition to considering the feature vector output by the encoder, it is also necessary to consider the feature vector output by the previous decoding layer.
6. The method for tracking a transseptal puncture needle based on Transformer according to claim 1, characterized in that: In step S5, a fully connected layer is used to map the output of the Transformer decoder to a two-dimensional coordinate (x, y) representing the position of the puncture needle tip.
7. Transformer-based atrial septal puncture needle tracking system, characterized by: Dynamically tracking the atrial septal puncture needle using the Transformer-based atrial septal puncture needle tracking method according to any one of claims 1 to 6 comprises: The image preprocessing module performs preprocessing including image cropping, denoising and image enhancement on the acquired real-time TEE image sequence; Feature extraction module, which uses a pre-trained convolutional neural network to extract deep features from pre-processed TEE images; Transformer encoding module, which inputs the extracted deep feature sequence into the Transformer encoder; The position prediction module uses the Transformer decoder to decode the encoded features, and the position prediction head predicts the position coordinates of the puncture needle tip in the next frame of TEE image; The trajectory generation module connects the continuously predicted puncture needle tip positions to generate a complete puncture needle tip trajectory.
8. A computer program product, characterized in that When the computer program product is run on a computer, the computer is enabled to execute the Transformer-based transseptal puncture needle tracking method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Transform echocardiogram-based left ventricle segmentation method
CN116051538A
Reliable echocardiogram segmentation method fusing relative position information
CN116912494A
Ultrasonic cardiogram analysis method and system, terminal and computer readable storage medium
CN118396936A