Multi-modal fusion spinal region puncture positioning method and system based on deep learning
By employing a multimodal fusion method based on deep learning, the problem of insufficient ultrasound positioning accuracy in complex patient populations was solved, enabling precise puncture path planning and real-time safety correction, thereby improving the success rate and safety of spinal surgery.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-24
- Publication Date
- 2026-04-03
AI Technical Summary
Existing technologies suffer from reduced accuracy in ultrasound localization for complex patient groups, such as elderly patients with spinal osteophytes or obese patients with excessive subcutaneous fat layers, making it difficult to meet the needs for precise puncture. This is especially true in minimally invasive spinal surgery, where puncture path planning is complex and carries risks such as nerve and vascular damage.
A deep learning-based multimodal fusion method is used to construct an enhanced real-time spinal model through image segmentation, target recognition, 3D reconstruction, and ultrasound image fusion. This model recommends puncture paths and parameters, corrects puncture needle deviation in real time, improves puncture success rate, and reduces the probability of adverse events.
It has improved the success rate of spinal puncture in complex patients, reduced the probability of adverse events, enhanced the medical skills of junior physicians, improved the imbalance of medical resources, and promoted the development of digital medical surgical instruments.
Smart Images

Figure CN121786453A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method, system, storage medium, and electronic device for spinal region puncture localization based on deep learning and multimodal fusion. Background Technology
[0002] Spinal puncture is a common and crucial procedure in clinical diagnosis and treatment. It is performed precisely by experienced physicians in numerous diagnostic and treatment processes, including the diagnosis of neurological diseases, pain management, and anesthesia administration. This procedure falls under the category of minimally invasive surgery. Its core principle involves using a specialized spinal needle, under strict aseptic conditions, to precisely insert it into the subarachnoid or epidural space of the lumbar vertebral segment. Cerebrospinal fluid samples are collected for biochemical and cytological analysis to aid in disease diagnosis, or anesthetic, analgesic, or therapeutic drugs are administered, serving various medical purposes such as clinical diagnosis, disease treatment, or surgical anesthesia. Precise localization of the epidural space is particularly critical throughout the procedure. Its accuracy directly affects the effective avoidance of unplanned dural injury, nerve damage, cerebrospinal fluid leakage, and side effects caused by drug injection deviation. Therefore, extremely high precision is required. The successful implementation of this technique typically relies heavily on the operator's accumulated clinical experience, deep understanding of spinal anatomy, and proficiency in spinal ultrasound imaging and image interpretation. Therefore, rich clinical experience and mastery of spinal ultrasound techniques are two core factors for improving the success rate of puncture. Simultaneously, maintaining a high degree of caution and strictly adhering to operating procedures are crucial to correctly inserting the needle while minimizing the risk of puncturing the dura mater or damaging nerve roots, the spinal cord, or other neural structures. However, for special patient groups such as those with anatomical abnormalities due to scoliosis, elderly patients with severe spinal degenerative changes, or obese patients with excessive subcutaneous fat, regardless of the operator's extensive clinical experience or expertise in using spinal ultrasound for localization, significant challenges arise, including difficulties in anatomical structure identification and complex puncture path planning.
[0003] Furthermore, with the rapid development of minimally invasive spinal surgery techniques, percutaneous transforaminal endoscopic discectomy plays an increasingly important role in the surgical treatment of lumbar disc herniation, lumbar spinal stenosis, and other lumbar spine diseases, and has become a crucial technique in the field of minimally invasive lumbar spine surgery. This procedure requires precise puncture at the spinal segment corresponding to the affected intervertebral disc, targeting the intervertebral foramen region. An operative channel is gradually established using a puncture needle to reach the lesion within the spinal canal. The endoscopic system is then used to clearly observe and remove the herniated disc tissue to relieve nerve compression. During this complex procedure, the surgeon must perform real-time, precise positioning under X-ray guidance, which demands a high level of spatial positioning ability, anatomical knowledge, and clinical experience. Precise positioning during puncture and surgical manipulation is crucial for reducing the incidence of adverse events such as nerve and vascular injury, minimizing surgical trauma, shortening surgical time, and promoting rapid postoperative recovery. Therefore, continuously exploring and improving the accuracy of spinal region positioning has significant clinical value for improving the overall success rate of spinal surgery and reducing the incidence of adverse events.
[0004] Ultrasound examination, as a medical imaging instrument with significant advantages such as portability, no ionizing radiation hazard, and the ability to provide immediate examination results, has been widely adopted in various clinical departments and its technology has been continuously upgraded. Extensive clinical practice has proven its ability to accurately locate the vertebral bodies of the spine and guide puncture procedures, providing crucial real-time imaging support for clinical operations. However, in actual clinical applications, it has been found that in complex patient groups, such as elderly patients with severe spinal osteophytes or obese patients with thick subcutaneous fat layers leading to reduced ultrasound penetration, ultrasound images are prone to reduced resolution and blurred anatomical structures. This means that the accuracy of ultrasound localization still needs further improvement, and the precision and reliability of its puncture guidance can significantly decrease, making it difficult to meet the precise localization needs of complex cases. Summary of the Invention This invention provides a deep learning-based multimodal fusion method, system, storage medium, and electronic device for spinal region puncture localization, which can accurately provide puncture points and recommended puncture paths, improve the success rate of spinal canal puncture in complex patients, and reduce the probability of adverse events.
[0005] This invention provides a deep learning-based multimodal fusion method for spinal region puncture localization, comprising: Obtain a spinal image containing the puncture site, and input the spinal image into a trained image segmentation model to obtain the image segmentation result; The image segmentation result is input into a trained target recognition model to perform target recognition and obtain target region information; Two-dimensional slice data is obtained based on the spinal image. Each pixel of the two-dimensional slice data is converted into a three-dimensional voxel, and the surface geometry model is extracted to obtain a three-dimensional reconstruction model. A continuous sequence of spinal ultrasound images is acquired, and the spinal ultrasound image sequence and the three-dimensional reconstruction model are input into a GAN model for fusion to obtain an enhanced real-time spinal model image. Recommend puncture path and puncture parameters based on puncture type and corresponding target area.
[0006] Furthermore, according to the above-mentioned deep learning-based multimodal fusion spinal region puncture localization method, the image segmentation model is PSPNet, and the processing procedure of the image segmentation model includes: The spine image is input into the backbone network of PSPNet for feature extraction to obtain an initial feature map; The feature map is input into the pyramid pooling module for pooling to obtain multiple pyramid feature maps; The upsampled pyramid feature map and the initial feature map are concatenated to obtain a fused feature map; The fused feature map is input into a convolutional layer to obtain the image segmentation result; the image segmentation result includes the target region, the avoidance region, the punctureable region, and the danger warning region.
[0007] Furthermore, according to the above-mentioned deep learning-based multimodal fusion spinal region puncture localization method, the image segmentation model processing further includes: Multiple consecutive frames of spine images are input into the Transformer encoder to obtain multi-scale feature representations; Arrange the multi-scale feature representations of multiple consecutive frames in chronological order to form a temporal feature sequence; The temporal feature sequence is input into a Transformer encoder for temporal processing and contextual information fusion to obtain enhanced features; The enhanced features are input into a decoder, which includes multiple decoding layers. A classification head is applied to the last decoding layer to obtain the image segmentation result.
[0008] Furthermore, according to the above-mentioned deep learning-based multimodal fusion spinal region puncture localization method, the target recognition model is the CenterNet model, and the target region information output by the CenterNet model includes: the center point heatmap, size, and center point offset of the target region.
[0009] Furthermore, according to the above-mentioned deep learning-based multimodal fusion spinal region puncture localization method, each pixel of the two-dimensional slice data is converted into a three-dimensional voxel, and a surface geometric model is extracted to obtain a three-dimensional reconstruction model, including: The two-dimensional slices are preprocessed; The preprocessed 2D slices are stacked into 3D voxels; Extract isosurfaces to construct a three-dimensional reconstruction model.
[0010] Furthermore, according to the aforementioned deep learning-based multimodal fusion spinal region puncture localization method, the puncture path and puncture parameters are recommended based on the puncture type and the corresponding target region, including: If the puncture type is subarachnoid puncture, based on the lumbar segment where the cauda equina terminates in the enhanced real-time spinal model image, the recommended puncture point is the segment below the cauda equina. At the same time, the puncture path is recommended based on whether the midline and paramidline paths pass through avoidance areas, and suggestions are given for puncture point, puncture angle and puncture depth. If the puncture type is epidural puncture, then based on the enhanced real-time spinal model image, the target area is the dura mater. The puncture path is recommended based on whether the midline path and para-midline path pass through the avoidance area. At the same time, suggestions are given for the puncture point, puncture angle and puncture depth. If the puncture type is percutaneous endoscopic discectomy, the spinal segment to be operated on is predicted based on the enhanced real-time spinal model image, and the puncture path is recommended based on whether the puncture path passes through the avoidance area.
[0011] Furthermore, according to the above-mentioned deep learning-based multimodal fusion spinal region puncture localization method, the method further includes: After recommending the puncture path and puncture parameters based on the puncture type and corresponding target area, the procedure includes: During the puncture, the trajectory of the puncture needle is tracked in real time, and based on the enhanced real-time spinal model image, real-time safety boundary correction and early warning are performed for deviations of the puncture needle.
[0012] This invention also provides a deep learning-based multimodal fusion spinal region puncture localization system, comprising: An image acquisition and segmentation module is used to acquire a spinal image containing the puncture site, and input the spinal image into a trained image segmentation model to obtain the image segmentation result; The target recognition module is used to input the image segmentation result into the trained target recognition model to perform target recognition and obtain target region information; The three-dimensional reconstruction module is used to obtain two-dimensional slice data based on the spine image, convert each pixel of the two-dimensional slice data into a three-dimensional voxel, and extract the surface geometry model to obtain a three-dimensional reconstruction model. The fusion module is used to acquire a continuous sequence of spinal ultrasound images, and input the spinal ultrasound image sequence and the three-dimensional reconstruction model into the GAN model for fusion to obtain an enhanced real-time spinal model image; the puncture path and puncture parameter recommendation module is used to recommend puncture paths and puncture parameters according to the puncture type and the corresponding target area.
[0013] The present invention also provides a computer-readable storage medium storing a plurality of instructions adapted for loading by a processor to execute any of the above-described deep learning-based multimodal fusion spinal region puncture localization methods.
[0014] The present invention also provides an electronic device, including a processor and a memory, wherein the processor is electrically connected to the memory, the memory is used to store instructions and data, and the processor is used in the steps of the deep learning-based multimodal fusion spinal region puncture localization method described above.
[0015] This invention provides a deep learning-based multimodal fusion method, system, storage medium, and electronic device for spinal region puncture localization. The invention segments and identifies spinal images, constructs a three-dimensional reconstruction model of the spine, and recommends puncture paths and parameters based on the puncture type and corresponding target region. This invention can accurately provide puncture points and recommended puncture paths, improving the success rate of spinal canal punctures in complex patients, reducing the probability of adverse events, improving the medical skills of junior physicians, addressing the imbalance of medical resources among different hospitals, and promoting the development of new digital medical surgical instruments. Attached Figure Description
[0016] The technical solution and other beneficial effects of the present invention will become apparent from the following detailed description of specific embodiments of the invention, in conjunction with the accompanying drawings.
[0017] Figure 1 A flowchart of a deep learning-based multimodal fusion spinal region puncture localization method provided in an embodiment of the present invention.
[0018] Figure 2 This is a schematic diagram of the structure of the PSPNet model provided in an embodiment of the present invention.
[0019] Figure 3 This is a flowchart illustrating the processing of the CenterNet model provided in an embodiment of the present invention.
[0020] Figure 4 This is a schematic diagram of the epidural / subarachnoid puncture path provided in an embodiment of the present invention.
[0021] Figure 5 This is a schematic diagram of the transforaminal puncture path provided in an embodiment of the present invention.
[0022] Figure 6 This is a schematic diagram of the structure of a deep learning-based multimodal fusion spinal region puncture localization system provided in an embodiment of the present invention.
[0023] Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0024] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0025] This invention provides a deep learning-based multimodal fusion spinal region puncture localization method, system, storage medium, and electronic device. The deep learning-based multimodal fusion spinal region puncture localization system provided in this invention can be integrated into an electronic device, such as a terminal or server. The terminal can include a tablet computer, laptop computer, personal computer (PC), microprocessor box, or other devices.
[0026] Please see Figure 1 , Figure 1 The flowchart illustrates a deep learning-based multimodal fusion spinal region puncture localization method provided in this embodiment of the invention, which is applied in electronic devices. This deep learning-based multimodal fusion spinal region puncture localization method includes the following steps: S1. Obtain a spinal image containing the puncture site, and input the spinal image into the trained image segmentation model to obtain the image segmentation result.
[0027] The spinal images are CT / MRI images.
[0028] In one embodiment, the image segmentation model is PSPNet (Pyramid Scene Parsing Network). Figure 2 This is a schematic diagram of the structure of the PSPNet model provided in an embodiment of the present invention, as shown below. Figure 2 As shown, the image segmentation model's processing steps include: S111, the spine image is input into the backbone network of PSPNet for feature extraction to obtain the initial feature map.
[0029] Specifically, the backbone network is a pre-trained CNN (Convolutional Neural Network), which includes: convolutional layers (for extracting image features), activation layers (such as ReLU, introducing non-linearity), pooling layers (for reducing feature dimensionality), and fully connected layers (for classification or regression tasks). Convolutional layers extract features by learning local features in the image (such as edges and textures). Activation layers transform input values into non-linear outputs using non-linear functions, enabling the network to learn and simulate more complex function mappings. Pooling layers reduce the spatial dimensionality of the feature map, thereby reducing the number of parameters and computational cost while retaining the most important information, helping to prevent overfitting. Fully connected layers integrate the local features extracted by the convolutional and pooling layers and then output the result.
[0030] The training process of the above CNN includes: The spinal image data was cleaned, image names were standardized, and capitalization and whitespace were corrected. Spinal images from multiple batches of databases were merged to form a dataset. The dataset was then divided in a 7:3 ratio, with 70% of the images used for the training set and 30% for the test set. A loss function (cross-entropy) was designed. Let p(x) be the target distribution and q(x) be the prediction distribution. Import the training set, initialize the network weights, and set the training parameters (learning rate, batch size, number of training epochs, etc.). Import the training set into the CNN model. Validate the trained CNN model using a validation set and adjust the training parameters, such as the learning rate and batch size, based on the performance on the validation set. To prevent overfitting, training can be stopped when the performance on the validation set no longer improves.
[0031] S112, input the feature map into the pyramid pooling module for pooling to obtain multiple pyramid feature maps.
[0032] S113, the upsampled pyramid feature map and the initial feature map are concatenated to obtain the fused feature map.
[0033] S114, the fused feature map is input into the convolutional layer to obtain the image segmentation result; the image segmentation result includes the target region (dura mater category, intervertebral foramen category), avoidance region (spinous process category, transverse process category, articular process category, lamina category, visceral category, blood vessel category, quadratus lumborum category), punctureable region (erector spinae category, interspinous ligament category, ligamentum flavum category) and danger warning region (visceral category, blood vessel category, spinal cord category).
[0034] To compensate for the shortcomings of CNNs in long-range dependency modeling and multi-scale feature fusion, this invention proposes using a Transformer encoder to address these deficiencies. The image segmentation model's processing may further include: S121, input multiple consecutive frames of spine images into the Transformer encoder to obtain multi-scale feature representations; S122, arrange the multi-scale feature representations of multiple consecutive frames in chronological order to form a temporal feature sequence; S123, input the temporal feature sequence into the Transformer encoder for temporal processing and contextual information fusion to obtain enhanced features; S124, the enhanced features are input into the decoder, which consists of multiple decoding layers. A classification head is applied to the last decoding layer to obtain the image segmentation result.
[0035] S2, input the image segmentation result into the trained target recognition model to perform target recognition and obtain target region information.
[0036] In one embodiment, the target recognition model is the CenterNet model. Figure 3 This is a flowchart of the CenterNet model processing provided in an embodiment of the present invention. The target recognition process includes: S21, Preprocessing and Feature Extraction.
[0037] Preprocess the input image segmentation results: resize the input image to a fixed size (e.g., 512×512). Backbone network feature extraction: Use deep convolutional networks (such as Hourglass, ResNet, DLA, etc.) to extract multi-scale feature maps; Feature pyramid construction: Generate feature maps of different resolutions to capture targets of different sizes.
[0038] S22, Key Point Heatmap Prediction.
[0039] Center point heatmap generation: Output a heatmap, with a size that is 1 / 4 of the original image (e.g., 128×128). Each channel corresponds to an object category, and the peak position in the heatmap represents the center of the corresponding object category; Gaussian kernel encoding: The true center point position is encoded using a Gaussian kernel to generate training labels; S23, Target Attribute Regression.
[0040] Size regression: For each detected center point, regress the width and height of the target; Center point offset regression: Due to the position quantization error caused by downsampling, the center point offset is regressed in detail.
[0041] S24, Post-processing and Output.
[0042] Peak point extraction: Local peak points are extracted using 3×3 max pooling on the heatmap; Threshold filtering: Filtering detection results based on confidence thresholds; Top-K selection: Retain the K detection results with the highest confidence; Bounding box reconstruction: Combine center point position, size regression value and offset to reconstruct the final bounding box.
[0043] The target recognition loss function adopts the center point detection loss function in CenterNet, and the loss function is as follows:
[0044] in, The center point heatmap loss is used to locate the center point of the target area; The bounding box size loss is used to predict the size of the target region; This is the offset loss, used to correct the precise location of the center point of the target area in the heatmap. and The hyperparameters are used to adjust the weights of the loss in each part.
[0045] Using the CenterNet model, the center point of the lesion region and its corresponding bounding box are accurately predicted during forward propagation. Subsequently, the neural network weights are optimized through backpropagation. Among them, the first The input features of a layer neural network are The number of layers in a deep learning feature extraction network is This embodiment combines center point detection and size prediction, updating the target recognition model parameters through forward and backward propagation to gradually improve the detection accuracy of target regions. The trained target recognition model is then used to identify the target region in each image.
[0046] S3 obtains two-dimensional slice data based on the spine image, converts each pixel of the two-dimensional slice data into a three-dimensional voxel, and extracts the surface geometry model to obtain a three-dimensional reconstruction model.
[0047] In one embodiment, step S3 includes: S31, preprocess the two-dimensional slices; S32, stack the pre-processed two-dimensional slices into three-dimensional voxels; S33, extract isosurfaces to construct a three-dimensional reconstruction model.
[0048] In puncture navigation, the construction of a 3D visualization model is crucial. This embodiment employs a voxel-based reconstruction method, integrating continuous CT / MRI 2D slice data to convert each pixel into a 3D voxel. Isosurfaces are extracted using algorithms such as the Marching Cubes Algorithm, ultimately constructing a high-precision 3D anatomical model. This model clearly displays the target area (such as the epidural space and intervertebral foramen) and avoidance structures (blood vessels, internal organs, etc.), and enables multi-angle interactive observation through visualization tools such as VTK / OpenGL. This provides doctors with an intuitive reference for puncture path planning, significantly improving the safety and accuracy of puncture surgery.
[0049] S4. Acquire a continuous sequence of spinal ultrasound images. Input the spinal ultrasound image sequence and the 3D reconstruction model into the GAN model for fusion to obtain an enhanced real-time spinal model image.
[0050] In one embodiment, step S4 includes the following steps: S41, acquire a continuous sequence of spinal ultrasound images; Specifically, step S41 includes: S411, Data Preprocessing and Feature Extraction.
[0051] (1) Ultrasound image processing workflow Image acquisition: The patient's spine is scanned in real time using an ultrasound probe to acquire continuous spinal ultrasound sequences. (2) Identification of transverse processes and spinous processes: Pre-trained deep networks are used to detect transverse and spinous process feature points in ultrasound images, and the spatial coordinates of these bony landmarks are labeled. The detected key points are converted into geometric feature vectors containing information on location, shape, and spatial relationships.
[0052] S412, 3D reconstruction model processing.
[0053] It automatically segments important anatomical structures such as the dura mater, intervertebral foramen, transverse process, and spinous process, generates binary masks and contour information of various structures, and unifies images of different modalities into the same spatial coordinate system.
[0054] S42, the spinal ultrasound image sequence and the three-dimensional reconstruction model are input into the GAN model for fusion to obtain an enhanced real-time spinal model image.
[0055] During spinal puncture, even slight changes in the patient's position (such as prone or lateral decubitus) can cause displacement of X-ray or ultrasound images. The greatest value of GAN fusion lies in using ultrasound feature points to lock the precise position and orientation of the 3D model on the current scanning plane in real time, achieving real-time anatomical posture correction.
[0056] Step S42 includes: GAN model construction The GAN model network structure is an encoder-decoder architecture, using a U-Net-like structure and including skip connections. The GAN model has a dual-path input: path one is the real-time ultrasound image and its extracted transverse / spinous process feature maps; path two is the 3D reconstruction model. At the bottleneck layer, the features of the two modalities are deeply fused, specifically using the transverse and spinous processes as reference points. These "transverse and spinous processes" serve as key geometric anchor points. The generator is activated to generate a fused image, ensuring that the transverse and spinous process features on it are spatially precisely aligned with their corresponding features on the reference image. The proposed puncture area is then reconstructed to obtain a real-time spinal model image. This allows for real-time ultrasound tracking of the current body position and correction of potential positional errors in the 3D model, while simultaneously improving the anatomical recognition accuracy of the ultrasound image.
[0057] The GAN model consists of a generator and a discriminator, with defined loss functions for the generator and discriminator, respectively. The generator fuses images with similar features, and its loss function measures the difference between the generated and target images. The discriminator is a binary classification model that accepts real images and images generated by the generator and attempts to distinguish between them. The model trains the generator and discriminator alternately, allowing them to learn competitively, improving the generator's ability to generate realistic images and the discriminator's classification accuracy. The training process of a GAN is essentially a binary minimax game between the generator and discriminator. During training, one model is fixed while the parameters of the other are updated, and the two models are trained iteratively. The generator continuously improves its generation ability through the discriminator's discrimination, while the discriminator continuously improves its discrimination ability by learning the distribution of the data. Through continuous adversarial training, the probability of the discriminator correctly judging the source of the training samples is maximized. Ultimately, the generator and discriminator reach an equilibrium, meaning the discriminator can no longer distinguish whether its input comes from real or generated samples.
[0058] S5 recommends puncture path and puncture parameters based on puncture type and corresponding target area.
[0059] Figure 4 This is a schematic diagram of the epidural / subarachnoid puncture path provided in an embodiment of the present invention. Figure 5This is a schematic diagram of the transforaminal puncture path provided in an embodiment of the present invention. Specifically, if the puncture type is lumbar puncture or subarachnoid anesthesia, the recommended puncture point is the lower segment below the cauda equina in the enhanced real-time spinal model image. At the same time, the puncture path is recommended based on whether the midline path and paramidline path pass through the avoidance area, and suggestions are given for the puncture point (distance from the spinous process), puncture angle (angle with the coronal plane), and puncture depth. If the puncture type is epidural puncture, then based on the enhanced real-time spinal model image, the target area is the dura mater. The puncture path is recommended based on whether the midline path and para-midline path pass through the avoidance area. At the same time, suggestions are given for the puncture point, puncture angle and puncture depth. If the puncture type is percutaneous endoscopic discectomy, the spinal segment to be operated on is predicted based on the enhanced real-time spinal model image, and the puncture path is recommended based on whether the puncture path passes through the avoidance area.
[0060] Furthermore, a puncture operation model is constructed, with the target area label as the target (T), the puncture needle label as the treatment method (M), and the puncture needle reaching the target position as the output result (Y), while identifying potential confounding factors (C). A causal graph is used to represent the causal relationship between the target, the puncture needle, and the result, and confounding factors are explicitly identified in the causal graph, considering their potential impact on the causal relationship. The influence of confounding factors is controlled using statistical methods (such as regression analysis, propensity score matching, etc.). The puncture operation model is trained on data with added puncture needle labels.
[0061] Furthermore, during the puncture process, the trajectory of the puncture needle is tracked in real time, and based on the enhanced real-time spinal model image, real-time safety boundary correction and early warning are performed for deviations of the puncture needle.
[0062] Specifically, the current spatial coordinates (x, y) of the puncture needle tip are captured in real time from the ultrasound probe. p ,y p ,z p ) and its instantaneous forward direction vector ( x , y , z If the puncture needle deviates from the target area, the needle's trajectory will be displayed in real time during this process. Based on the needle body, a hypothetical forward direction will be established. If the needle is outside the safe boundary of the target area, a warning will be issued; if it approaches internal organs, a red danger warning will be issued. Along the real-time forward vector ( The system projects the future distance to predict the needle insertion trajectory segment. When the predicted needle insertion trajectory segment approaches the danger warning area (<5mm), a red danger warning is triggered, and remedial measures are suggested according to the guidelines. Based on the method described in the above embodiments, this embodiment will further describe it from the perspective of a deep learning-based multimodal fusion spinal region puncture positioning system. This deep learning-based multimodal fusion spinal region puncture positioning system can be implemented as an independent entity or integrated into an electronic device. The electronic device can be a terminal, server, or other devices. The terminal can include a tablet computer, laptop computer, personal computer (PC), microprocessor box, or other devices.
[0063] Please see Figure 6 , Figure 6 This invention specifically describes a deep learning-based multimodal fusion spinal region puncture localization system provided in an embodiment of the invention, which is applied in electronic devices. This deep learning-based multimodal fusion spinal region puncture localization system may include: An image acquisition and segmentation module is used to acquire a spinal image containing the puncture site, and input the spinal image into a trained image segmentation model to obtain the image segmentation result; The target recognition module is used to input the image segmentation result into the trained target recognition model to perform target recognition and obtain target region information; The three-dimensional reconstruction module is used to obtain two-dimensional slice data based on the spine image, convert each pixel of the two-dimensional slice data into a three-dimensional voxel, and extract the surface geometry model to obtain a three-dimensional reconstruction model. The fusion module is used to acquire a continuous sequence of spinal ultrasound images, and input the spinal ultrasound image sequence and the three-dimensional reconstruction model into the GAN model for fusion to obtain an enhanced real-time spinal model image. The puncture path and puncture parameter recommendation module is used to recommend puncture paths and puncture parameters based on the puncture type and the corresponding target area.
[0064] In specific implementation, the above modules and / or units can be implemented as independent entities, or they can be arbitrarily combined and implemented as the same or several entities. For the specific implementation of the above modules and / or units, please refer to the previous method embodiments. For the specific beneficial effects that can be achieved, please also refer to the beneficial effects in the previous method embodiments, which will not be repeated here.
[0065] In addition, this embodiment of the invention also provides an electronic device, which may be a computer, tablet computer, or other similar device. This electronic device can implement the steps of any embodiment of the deep learning-based multimodal fusion spinal region puncture localization method provided in this embodiment of the invention. Therefore, it can achieve the beneficial effects that any deep learning-based multimodal fusion spinal region puncture localization method provided in this embodiment of the invention can achieve, as detailed in the preceding embodiments, and will not be repeated here.
[0066] Figure 7 A specific structural block diagram of an electronic device provided in an embodiment of the present invention is shown. This electronic device can be used to implement the deep learning-based multimodal fusion spinal region puncture localization method provided in the above embodiments. The electronic device 500 can be a terminal, server, or other device. The terminal can include a tablet computer, laptop computer, personal computer (PC), microprocessor box, or other devices.
[0067] The memory 520 can be used to store software programs and modules, such as the program instructions / modules corresponding to those in the above embodiments. The processor 580 executes various functional applications and data processing by running the software programs and modules stored in the memory 520. The memory 520 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 520 may further include memory remotely located relative to the processor 580, and these remote memories can be connected to the electronic device 500 via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0068] The input unit 530 can be used to receive input numeric or character information, and to generate a keyboard and mouse related to user settings and function control. Display unit 540 can be used to display information input by the user or information provided to the user, as well as various graphical user interfaces, which can be composed of graphics, text, icons, video, and any combination thereof. Display unit 540 may include display panel 541, which may optionally be configured in the form of LCD (Liquid Crystal Display), OLED (Organic Light-Emitting Diode), or other similar forms.
[0069] Audio circuitry 560, speaker 561, and microphone 562 provide an audio interface between the user and electronic device 500. Audio circuitry 560 converts received audio data into electrical signals and transmits them to speaker 561, where speaker 561 converts them into sound signals for output. Conversely, microphone 562 converts collected sound signals into electrical signals, which are then received by audio circuitry 560, converted back into audio data, and processed by processor 580. The audio data is then transmitted via RF circuitry 510 to, for example, another terminal, or output to memory 520 for further processing. Audio circuitry 560 may also include an earphone jack to facilitate communication between external headphones and electronic device 500.
[0070] Electronic device 500, through transmission module 570 (e.g., Wi-Fi module), can help users receive requests, send information, etc., providing users with wireless broadband internet access. Although transmission module 570 is shown in the figure, it is understood that it is not an essential component of electronic device 500 and can be omitted as needed without changing the essence of the invention.
[0071] The processor 580 is the control center of the electronic device 500. It connects to various parts of the phone via various interfaces and lines, and performs various functions and processes data of the electronic device 500 by running or executing software programs and / or modules stored in the memory 520, and by calling data stored in the memory 520, thereby providing overall monitoring of the electronic device. Optionally, the processor 580 may include one or more processing cores; in some embodiments, the processor 580 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may also not be integrated into the processor 580.
[0072] Electronic device 500 also includes a power supply 590 (such as a battery) that supplies power to various components. In some embodiments, the power supply may be logically connected to processor 580 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The power supply 590 may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components.
[0073] Although not shown, the electronic device 500 also includes cameras (such as front-facing cameras and rear-facing cameras), Bluetooth modules, etc., which will not be described in detail here. Specifically, in this embodiment, the display unit of the electronic device is a touch screen display, and the mobile terminal also includes a memory and one or more programs, wherein one or more programs are stored in the memory and configured to be executed by one or more processors. One or more programs contain instructions for performing the following operations: Obtain a spinal image containing the puncture site, and input the spinal image into a trained image segmentation model to obtain the image segmentation result; The image segmentation result is input into a trained target recognition model to perform target recognition and obtain target region information; Two-dimensional slice data is obtained based on the spinal image. Each pixel of the two-dimensional slice data is converted into a three-dimensional voxel, and the surface geometry model is extracted to obtain a three-dimensional reconstruction model. A continuous sequence of spinal ultrasound images is acquired, and the spinal ultrasound image sequence and the three-dimensional reconstruction model are input into a GAN model for fusion to obtain an enhanced real-time spinal model image. Recommend puncture path and puncture parameters based on puncture type and corresponding target area.
[0074] In practice, the above modules can be implemented as independent entities or combined in any way to be implemented as the same or several entities. For the specific implementation of the above modules, please refer to the previous method implementation examples, which will not be repeated here.
[0075] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by instructions, or by instructions controlling related hardware. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor. Therefore, embodiments of the present invention provide a storage medium storing multiple instructions that can be loaded by a processor to execute the steps of any embodiment of the deep learning-based multimodal fusion spinal region puncture localization method provided by the present invention.
[0076] The computer-readable storage medium may include: read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
[0077] Since the instructions stored in the storage medium can execute the steps in any embodiment of the deep learning-based multimodal fusion spinal region puncture localization method provided in the embodiments of the present invention, the beneficial effects that any deep learning-based multimodal fusion spinal region puncture localization method provided in the embodiments of the present invention can achieve can be realized. For details, please refer to the previous embodiments, which will not be repeated here.
[0078] The foregoing has provided a detailed description of a deep learning-based multimodal fusion spinal region puncture localization method, system, storage medium, and electronic device provided by embodiments of the present invention. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. At the same time, for those skilled in the art, there will be changes in specific implementation methods and application scope based on the ideas of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.
Claims
1. A multimodal fusion spinal region puncture localization method based on deep learning, characterized in that, The method includes: Obtain a spinal image containing the puncture site, and input the spinal image into a trained image segmentation model to obtain the image segmentation result; The image segmentation result is input into a trained target recognition model to perform target recognition and obtain target region information; Two-dimensional slice data is obtained based on the spinal image. Each pixel of the two-dimensional slice data is converted into a three-dimensional voxel, and the surface geometry model is extracted to obtain a three-dimensional reconstruction model. A continuous sequence of spinal ultrasound images is acquired, and the spinal ultrasound image sequence and the three-dimensional reconstruction model are input into a GAN model for fusion to obtain an enhanced real-time spinal model image. Recommend puncture path and puncture parameters based on puncture type and corresponding target area.
2. The deep learning-based multimodal fusion spinal region puncture localization method according to claim 1, characterized in that, The image segmentation model is PSPNet, and the processing steps of the image segmentation model include: The spine image is input into the backbone network of PSPNet for feature extraction to obtain an initial feature map; The feature map is input into the pyramid pooling module for pooling to obtain multiple pyramid feature maps; The upsampled pyramid feature map and the initial feature map are concatenated to obtain a fused feature map; The fused feature map is input into a convolutional layer to obtain the image segmentation result; the image segmentation result includes the target region, the avoidance region, the punctureable region, and the danger warning region.
3. The deep learning-based multimodal fusion spinal region puncture localization method according to claim 1, characterized in that, The image segmentation model's processing also includes: Multiple consecutive frames of spine images are input into the Transformer encoder to obtain multi-scale feature representations; Arrange the multi-scale feature representations of multiple consecutive frames in chronological order to form a temporal feature sequence; The temporal feature sequence is input into a Transformer encoder for temporal processing and contextual information fusion to obtain enhanced features; The enhanced features are input into a decoder, which includes multiple decoding layers. A classification head is applied to the last decoding layer to obtain the image segmentation result.
4. The deep learning-based multimodal fusion spinal region puncture localization method according to claim 1, characterized in that, The target recognition model is the CenterNet model, and the target region information output by the CenterNet model includes: a heatmap of the center point of the target region, its size, and the center point offset.
5. The deep learning-based multimodal fusion spinal region puncture localization method according to claim 1, characterized in that, Each pixel of the two-dimensional slice data is converted into a three-dimensional voxel, and a surface geometry model is extracted to obtain a three-dimensional reconstructed model, including: The two-dimensional slices are preprocessed; The preprocessed 2D slices are stacked into 3D voxels; Extract isosurfaces to construct a three-dimensional reconstruction model.
6. The deep learning-based multimodal fusion spinal region puncture localization method according to claim 1, characterized in that, Based on the puncture type and the corresponding target area, the recommended puncture path and puncture parameters include: If the puncture type is subarachnoid puncture, based on the lumbar segment where the cauda equina terminates in the enhanced real-time spinal model image, the recommended puncture point is the segment below the cauda equina. At the same time, the puncture path is recommended based on whether the midline and paramidline paths pass through avoidance areas, and suggestions are given for puncture point, puncture angle and puncture depth. If the puncture type is epidural puncture, then based on the enhanced real-time spinal model image, the target area is the dura mater. The puncture path is recommended based on whether the midline path and para-midline path pass through the avoidance area. At the same time, suggestions are given for the puncture point, puncture angle and puncture depth. If the puncture type is percutaneous endoscopic discectomy, the spinal segment to be operated on is predicted based on the enhanced real-time spinal model image, and the puncture path is recommended based on whether the puncture path passes through the avoidance area.
7. The deep learning-based multimodal fusion spinal region puncture localization method according to claim 1, characterized in that, After recommending the puncture path and puncture parameters based on the puncture type and corresponding target area, the procedure includes: During the puncture, the trajectory of the puncture needle is tracked in real time, and based on the enhanced real-time spinal model image, real-time safety boundary correction and early warning are performed for deviations of the puncture needle.
8. A deep learning-based multimodal fusion spinal region puncture localization system, characterized in that, include: An image acquisition and segmentation module is used to acquire a spinal image containing the puncture site, and input the spinal image into a trained image segmentation model to obtain the image segmentation result; The target recognition module is used to input the image segmentation result into the trained target recognition model to perform target recognition and obtain target region information; The three-dimensional reconstruction module is used to obtain two-dimensional slice data based on the spine image, convert each pixel of the two-dimensional slice data into a three-dimensional voxel, and extract the surface geometry model to obtain a three-dimensional reconstruction model. The fusion module is used to acquire a continuous sequence of spinal ultrasound images, and input the spinal ultrasound image sequence and the three-dimensional reconstruction model into the GAN model for fusion to obtain an enhanced real-time spinal model image. The puncture path and puncture parameter recommendation module is used to recommend puncture paths and puncture parameters based on the puncture type and the corresponding target area.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a plurality of instructions adapted for loading by a processor to execute the deep learning-based multimodal fusion spinal region puncture localization method according to any one of claims 1 to 7.
10. An electronic device, characterized in that, The method includes a processor and a memory, the processor being electrically connected to the memory, the memory being used to store instructions and data, and the processor being used to execute the steps in the deep learning-based multimodal fusion spinal region puncture localization method according to any one of claims 1 to 7.