Ultrasonic endoscope navigation system and method based on deep learning
Through parallel encoder architecture and multimodal fusion technology, the problems of anatomical structure recognition accuracy and high learning threshold in ultrasonic endoscopic image analysis are solved, and efficient and real-time ultrasonic endoscopic image analysis and multimodal information fusion are achieved, which promotes the popularization of ultrasonic endoscopic technology in primary medical institutions.
Patent Information
- Application Number
- CN202510806559.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-17
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-17
AI Technical Summary
Existing ultrasound endoscopic image analysis depends on operator experience, long learning curves, limited anatomical structure recognition accuracy, and difficult to popularize in primary medical institutions. Inadequate multimodal information fusion leads to a decrease in model robustness.
The parallel encoder architecture is adopted to combine CNN and Transformer branches to extract local and global features, introduce bidirectional LSTM network to analyze timing correlation, integrate MaskScoring R-CNN and elastic registration algorithm for multimodal alignment, design lightweight decoder and mixed loss functions, and enhance the generalization ability of the model through adversarial training.
Significantly improve the accuracy of anatomical structure segmentation and video anatomy coherence, lower learning thresholds, enhance the model's adaptability to dynamic scenarios and cross-modal registration robustness, and meet clinical real-time needs.
Smart Images

Figure CN120339273A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of medical image analysis, and particularly to an endoscopic ultrasound navigation system and method based on deep learning. Background Art
[0002] Pancreatic cancer is characterized by insidious onset and strong invasiveness. By the time of diagnosis, it has often progressed to the advanced stage, and the 5-year survival rate is less than 10%. Early diagnosis is the key to improving the prognosis. However, existing screening methods such as tumor markers and imaging examinations have low sensitivity to early lesions. Endoscopic ultrasound, with its advantages of high resolution and close-range exploration, can detect abnormal pancreatic echoes with a diameter ≤ 5 mm and has become a core technology for early pancreatic cancer screening. However, the analysis of endoscopic ultrasound images highly depends on the operator's experience and spatial imagination. Its scanning sections are variable and the anatomical structures are complex, resulting in a long learning curve and insufficient popularity in grass-roots hospitals. According to statistics, only about 10% of endoscopic physicians globally can proficiently operate endoscopic ultrasound, which greatly limits its clinical application.
[0003] Currently, although deep learning-based medical image analysis has made progress in fields such as gastroscopy and colonoscopy, there are significant limitations in the analysis of endoscopic ultrasound images: traditional CNNs rely on local convolution operations and are difficult to model the long-range spatial dependence relationships between organs and blood vessels in endoscopic ultrasound images, resulting in limited accuracy in anatomical structure recognition; the dynamic videos of endoscopic ultrasound require temporal correlation analysis between consecutive frames, while existing single-frame segmentation models lack coherent modeling of the overall structure of the video and are prone to misjudgment due to image jitter or perspective mutation; tumor compression or invasion often leads to anatomical structure deformation, and models relying solely on a single endoscopic ultrasound modality are prone to performance degradation due to data distribution shift and need to fuse multi-modal information such as CT / MRI to improve robustness. Therefore, based on the above problems, the present invention proposes an endoscopic ultrasound navigation system and method based on deep learning. Summary of the Invention
[0004] Technical Objectives In order to solve the above problems, the objective of the present invention is to provide an endoscopic ultrasound navigation system and method based on deep learning, aiming to solve the key bottlenecks in the analysis of endoscopic ultrasound images, assist endoscopic physicians in quickly identifying key anatomical structures and lesions in endoscopic ultrasound images, reduce the operation learning threshold, improve the early diagnosis efficiency of tumors in the biliary and pancreatic systems such as pancreatic cancer, and promote the popularization and application of endoscopic ultrasound technology in grass-roots medical institutions.
[0005] Technical Solutions To achieve the above object, the present invention provides an endoscopic ultrasound navigation system and method based on deep learning. The system adopts a parallel encoder architecture, extracts local detail features through a CNN branch, and the Transformer branch uses the self-attention mechanism to capture global context, and combines a channel attention module to achieve feature adaptive fusion; introduces a bidirectional LSTM network and a dynamic weight adjustment mechanism to analyze the spatio-temporal correlation between consecutive frames and optimize the temporal coherence of video segmentation; integrates MaskScoring R-CNN and an elastic registration algorithm to achieve cross-modal alignment and feature complementarity between endoscopic ultrasound and CT / MRI images, and uses adversarial training to generate synthetic data to enhance the generalization ability of the model; designs a lightweight decoder and a hybrid loss function, retains detail information through multi-level skip connections, compresses the model parameter quantity, and ensures that the single-frame processing delay is less than 50 ms to meet the clinical real-time requirements.
[0006] In a first aspect, the present invention provides an endoscopic ultrasound navigation system based on deep learning, including: A parallel encoder module for extracting local features of endoscopic ultrasound images through a CNN branch and capturing multi-scale global context information through a Transformer branch adopting a pyramid structure; the Transformer branch includes a non-overlapping patch embedding layer and multi-stage downsampling; A channel attention fusion module for adaptively weighted fusion of the extracted features; A decoder module for generating an anatomical structure segmentation mask based on the fusion features through multi-level skip connections; A temporal processing module for receiving the frame-by-frame segmentation results output by the decoder module and realizing the temporal coherence analysis of the endoscopic ultrasound video through a bidirectional LSTM network combined with an Attention mechanism; A multi-modal fusion module for registering and fusing endoscopic ultrasound images; the multi-modal fusion module includes an adversarial elastic registration sub-module, which uses a generator to predict a non-linear deformation field and a discriminator to distinguish real and synthetic registration images, and at the same time combines mutual information maximization constraints to ensure anatomical structure consistency; A spatio-temporal attention fusion module for dynamically allocating spatial and temporal attention weights according to the segmentation confidence of the current frame.
[0007] Further, the resolution of the feature map output by the Transformer branch in the 8-fold downsampling stage is aligned with that of the CNN branch, and the spatial dimension of the feature map is matched through bilinear interpolation.
[0008] Further, the upsampling operation of the decoder module sequentially performs 2x, 2x, and 4x upsampling using transposed convolutional layers, and the loss function is jointly optimized using Dice Loss and cross-entropy loss.
[0009] Furthermore, the multimodal fusion module further includes the following sub-modules: The Mask Scoring R-CNN module is used to segment soft tissue targets from CT / MRI images; The elastic transformation module realizes cross-modal registration of endoscopic ultrasound and CT / MRI based on B-spline interpolation and mutual information optimization algorithm; The multimodal feature enhancement module is used to input the registered fusion image into the parallel encoder module to improve the recognition accuracy of anatomical structure variation cases.
[0010] Furthermore, the spatio-temporal attention fusion module captures the global spatial correlation between organs and blood vessels in a single-frame image using the self-attention layer of the Transformer branch, extracts the temporal evolution features between adjacent frames through a bidirectional LSTM network, and dynamically assigns spatial and temporal attention weights according to the segmentation confidence of the current frame. The fusion formula is:
[0011] In the formula, is the dynamic fusion output; is the dynamic weight parameter; is the global spatial attention map; is the temporal attention weight; is the slope coefficient, determined by fitting the confidence decay curve of tumor compression cases; is the segmentation confidence of the current frame; is the confidence threshold.
[0012] By jointly modeling the spatial features and temporal dependencies within endoscopic ultrasound video frames, it effectively fuses local details and global context information, significantly improves the accuracy of anatomical structure segmentation and the coherence of video parsing, while suppressing misidentifications caused by image jitter or perspective mutations, and enhancing the model's adaptability to dynamic scenarios.
[0013] Furthermore, the total loss function of the adversarial elastic registration module is: In the formula, is the output of the total loss function; , , and are the balance factors of each loss; is the adversarial loss; is the mutual information loss; and are the balance coefficients of the total variation and the L2 regularization term; is the total variation; is the L2 regularization term; is the segmentation confidence of the current frame; is the confidence threshold; is the KL divergence regularization term.
[0014] Through non-linear deformation field prediction and adversarial training optimization, break through the limitations of traditional rigid registration, achieve high-precision cross-modal alignment of endoscopic ultrasound and CT / MRI images, significantly improve the registration robustness of complex anatomical variation cases, and enhance the generalization ability of the model under different medical center data distributions.
[0015] In a second aspect, the present invention provides a method for annotating anatomical structures based on endoscopic ultrasound video that does not involve the diagnosis and treatment of diseases, including: Filter invalid frames in the endoscopic ultrasound video based on a pre-trained CNN model and perform standardization processing on the valid frames; Perform anatomical structure segmentation on a single-frame image through a Transformer-CNN network; Combine a bidirectional LSTM network to correct the temporal consistency of the segmentation results of consecutive frames.
[0016] Furthermore, the standardization processing includes gray-scale normalization, CLAHE contrast enhancement, fixed-size cropping, and metadata anonymization. The metadata anonymization specifically means cropping the region of interest or replacing the patient privacy information in the image with a black mask.
[0017] In a third aspect, the present invention also provides a deep learning-based endoscopic ultrasound navigation method that does not involve the diagnosis and treatment of diseases. The method is based on the system described in the first aspect above and includes: Extract local features of the endoscopic ultrasound image and capture multi-scale global context information; Perform adaptive weighted fusion on the extracted features and generate an anatomical structure segmentation mask based on the fused features; Receive the per-frame segmentation results output and perform temporal coherence analysis of the endoscopic ultrasound video through a bidirectional LSTM network; Register and fuse the endoscopic ultrasound images.
[0018] In a fourth aspect, the present invention also provides a training device for a multi-modal endoscopic ultrasound navigation model, including: An adversarial training module for generating synthetic training data of endoscopic ultrasound and CT / MRI through a generative adversarial network; A transfer learning module that initializes the weights of the Transformer-CNN network using the parameters of a pre-trained IDEA gastroscope navigation model; A hybrid loss function for joint optimization by combining Dice Loss, cross-entropy loss, and temporal consistency loss.
[0019] In a fifth aspect, the present invention further provides a computer device, including a management platform and a memory. The management platform is connected to the memory. The memory is used to store a computer program, and the management platform is used to execute the computer program stored in the memory, so that the computer device implements the aforementioned method for endoscopic ultrasound navigation based on deep learning.
[0020] In a sixth aspect, the present invention further provides a computer-readable storage medium. A computer program is stored in the computer-readable storage medium. When the computer program is executed by a management platform, the aforementioned method for endoscopic ultrasound navigation based on deep learning is implemented.
[0021] The present invention realizes the feature fusion of local details and global context through a parallel encoder architecture, combines a dynamic spatio-temporal attention mechanism to enhance the video temporal coherence, introduces a multi-modal adversarial elastic registration network to solve the non-linear alignment problem of cross-modal images, and supplements it with a lightweight decoder design to compress the computational complexity. The system effectively overcomes the bottlenecks in traditional endoscopic ultrasound image analysis, such as fragmented local features, insufficient temporal dependence, sensitivity to anatomical variations, and limited real-time performance. It significantly improves the anatomical structure segmentation accuracy, cross-modal registration robustness, and dynamic scene adaptability, provides a highly reliable artificial intelligence-assisted tool for early screening of pancreatic cancer, reduces the learning threshold and clinical application cost of endoscopic ultrasound technology, and promotes its popularization in primary medical institutions.
[0022] Beneficial effects By implementing the endoscopic ultrasound navigation system and method based on deep learning provided by the present invention, the following technical effects are achieved: (1) Through the collaborative design of the CNN and Transformer branches in this application, both local feature extraction and global context modeling are considered, overcoming the limitation of feature expression caused by inductive bias in traditional single-branch models in medical image segmentation, and realizing the efficient recognition and accurate positioning of multi-scale anatomical structures in endoscopic ultrasound images.
[0023] (2) Based on the design of multi-level skip connections and a hybrid loss function, while retaining high-resolution detailed features, the number of model parameters is compressed, the computational complexity is significantly reduced, and the single-frame processing delay is ensured to meet the clinical real-time requirements, providing seamless navigation assistance for endoscopic physicians.
[0024] (3) By jointly modeling the spatial features and temporal dependence within endoscopic ultrasound video frames, local details and global context information are effectively fused, significantly improving the accuracy of anatomical structure segmentation and the coherence of video analysis. At the same time, misidentifications caused by image jitter or sudden perspective changes are suppressed, enhancing the model's adaptability to dynamic scenes.
[0025] (4) Through the prediction and adversarial training optimization of the non-linear deformation field, break through the limitations of traditional rigid registration, achieve high-precision cross-modal alignment of endoscopic ultrasound and CT / MRI images, significantly improve the registration robustness of complex anatomical variation cases, and enhance the generalization ability of the model under different medical center data distributions. Brief Description of the Drawings
[0026] To make the above-described deep learning-based endoscopic ultrasound navigation system and method of the present invention more clearly understandable, the following will briefly introduce the drawings required for the specific implementation of the present invention. Obviously, the drawings in the following description are only some embodiments of the present invention, and those skilled in the art can obtain other drawings based on these drawings without creative efforts.
[0027] Figure 1 Represents the technical roadmap of this application; Figure 2 Represents the schematic diagram of the IDEA-EUS model. Detailed Description of the Preferred Embodiments
[0028] To facilitate the understanding of the embodiments of the present invention, first, the abbreviations and key terms that may be involved in the embodiments of the present invention will be explained or defined. For abbreviations or key terms that are not defined, they are all commonly understood by those skilled in the art.
[0029] EUS: Endoscopic ultrasound; IDEA-EUS: Deep learning-based endoscopic ultrasound navigation system model; MRI: Magnetic resonance imaging; MR: Magnetic resonance.
[0030] Example 1: A deep learning-based endoscopic ultrasound navigation system and method are provided, and the technical roadmap is as Figure 1 shown. Specifically as follows.
[0031] Construct an endoscopic ultrasound navigation image recognition model based on the Transformer-CNN network, and the construction process specifically includes; 1. Data collection, the research objects include the clinical data of patients undergoing endoscopic ultrasound examination, and the data includes endoscopic ultrasound images and video data, among which pathological data is collected for tumor diagnosis cases.
[0032] 2. Data preprocessing, including data anonymization, data filtering, and data normalization and standardization. The data preprocessing process includes: Step 1. Anonymize the image in one of the following two ways: ① Remove the metadata overlay by cropping the region of interest of the endoscopic ultrasound; ②By editing specific sub-regions of the image, replace potentially confounding image features, patient names, medical record numbers, procedure timestamps, or procedure locations with black blocks. For example, if an image is marked with unique markers or text to highlight the diagnostic features of a potential bias model, these images are considered to have confounding image features, and if these features cannot be removed without damaging the image, these images will not be included in the final database.
[0033] Step 2: Data filtering, including: ①Use the ordinary white light endoscopy mode images and endoscopic ultrasound mode images for the training and testing of CNN1 to distinguish the endoscopic ultrasound mode; ②Use the ineffective endoscopic ultrasound images and clear endoscopic ultrasound images to train and test CNN2, and eliminate the ineffective images in further data collection. The continuous frame images extracted from the endoscopic ultrasound video are filtered by the above model and then data-labeled. The images of the same patient are grouped in the same dataset.
[0034] Step 3: Data normalization and standardization to improve the contrast and details of the images and make the images clearer. Use linear standardization to map the gray values of the images to the range between 0 and 255. For the images extracted from the endoscopic ultrasound video, they are obtained by interval sampling at a frequency of one image per 5 frames, and then the images are cropped using fixed length and width pixel values (224×224) to obtain endoscopic ultrasound images that can be used to train the model.
[0035] 3. Data labeling: Use the preprocessed endoscopic ultrasound and CT / MRI images for data labeling so that the model can accurately identify and segment different anatomical structures and lesions in the images.
[0036] Use Labelme software to label the abdominal organs, main vascular structures, and space-occupying lesions that can be recognized in the endoscopic ultrasound path, which are divided into 23 labels, including: liver, gallbladder, pancreatic head, pancreatic neck, pancreatic body, pancreatic tail, spleen, kidney, adrenal gland, main pancreatic duct, common bile duct, abdominal aorta, celiac trunk, common hepatic artery, splenic artery, superior mesenteric artery, inferior vena cava, portal vein, splenic vein, superior mesenteric vein, solid tumor, cystic lesion, and enlarged lymph node. The same image can have multiple labels.
[0037] 4. Image recognition model construction: In the stage of constructing the deep learning model for image recognition, use the labeled and segmented endoscopic ultrasound image data to train the model to achieve automatic recognition of anatomical structures and lesions in the endoscopic ultrasound images. Use the Transformer-CNN network to construct an efficient and accurate image recognition model. Through training and optimization on a large amount of labeled data, the model will learn the feature representations of different structures and lesions in the images and be able to accurately recognize and segment unlabeled images.
[0038] ① Model construction: Use Nvidia A100 to build a Transformer-CNN network under the PyTorch deep learning framework to automatically identify anatomical structures in endoscopic ultrasound images.
[0039] The Transformer-CNN network is a deep learning model for medical image segmentation, which can effectively utilize multi-scale information to improve the accuracy of segmentation. This model adopts a parallel encoder structure, with a CNN to extract local information and another encoder using a Transformer to extract global information. Then, the information extracted by the two encoders is fused, and a decoder is used to generate the final segmentation result.
[0040] The specific implementation process is as follows: The Transformer-CNN network is a new parallel encoder architecture proposed considering the unique advantages of CNN and Transformer. Traditional segmentation models rely only on a single encoder branch to integrate global and local information in the image, while this method combines two independent branches to capture semantic information from the input image respectively. The overall architecture follows a paradigm similar to the U-Net encoder-decoder, consisting of an encoder, a channel attention module, skip connections, and a decoder. This parallel encoder combines the advantages of ResNet and Transformer. A pyramid structure is introduced in the Transformer component to capture global features at multiple resolutions. In addition, a channel attention module is used to enhance the expressive power of the parallel encoder, enrich the extracted features, and provide guidance for the subsequent decoding process. Skip connections and decoder modules are also used to estimate the final segmentation mask. It mainly consists of three parts: Parallel Transformer-CNN Encoder, which introduces a pyramid structure in the Transformer to obtain global feature maps of different scales. To extract more comprehensive and diverse feature representations, the parallel encoder is divided into three stages. In the first stage, it operates on a 2D input image of size H×W×3 using a patch embedding layer with a patch size of 4, ensuring non-overlapping patches. Then the resulting feature map is processed through the Transformer layer to obtain its global information. Entering the second stage, to retain the detailed information in the patch embedding layer, the patch size is reduced to 2. This will generate a feature map with a resolution of H / 8×W / 8, which is then further processed through the Transformer layer. In the third stage, the patch size remains at 2. Using a Transformer encoder with a pyramid structure, the feature maps are downsampled in the order of factors 4, 8, and 16. To eliminate the spatial dimension differences of the multi-scale feature maps, a bilinear interpolation operation is performed on the feature map output by the Transformer branch during the 8-fold downsampling stage to strictly align its resolution with the feature map output by the CNN branch. This downsampling process is essential for generating feature maps of different scales, allowing a wider receptive field and capturing the hierarchical representation of the input image. For the other branch, ResNet is used as the backbone to capture the local details of the image, and it is also downsampled by factors 4, 8, and 16 respectively to ensure that the resolution of the local information extracted by ResNet is consistent with that extracted by the Transformer.
[0041] Channel Attention Module, which is used to combine the local and global information obtained from the CNN and Transformer branches, enabling valuable information to be passed from the parallelized encoder to the decoder, thus obtaining pixel-level segmentation results. Since simple convolution operations are not sufficient to effectively fuse local and global features, the information content of the feature channels is used to weight them. By activating the channels that contribute to the segmentation results and suppressing the irrelevant channels, effective feature fusion is achieved using the local and global information of the encoder.
[0042] Decoder. The Transformer-CNN network achieves a semantically rich multi-scale feature representation by implementing a parallel encoder and a channel attention module. Similar to TransUNet, the Transformer-CNN network uses skip connections to link low-resolution features to high-resolution features, which are then passed to the decoder to generate the final segmentation mask. The decoder uses convolutional layers to extract combined multi-scale features, including 3×3 convolution, batch normalization, and ReLU layers. The upsampling operation is performed using transposed convolutional layers, obtaining 2x, 2x, and 4x upsampling in sequence. Dice loss and Cross-Entropy are used as loss functions to train the Transformer-CNN network.
[0043] ② Model validation and testing. An endoscopic ultrasound image dataset is constructed for model validation and testing. The dataset contains endoscopic ultrasound images from different patients, including normal images and lesion images. The images in the dataset have been segmented and annotated, with segmentation labels for different anatomical structures and lesions. Dice and JAC are used to evaluate the segmentation effect of the model. The higher the Dice coefficient and JAC coefficient values, the better the segmentation effect.
[0044] Build a real-time navigation model IDEA-EUS based on endoscopic ultrasound video. The principle is as Figure 2 shown, and the construction process specifically includes: 1. Model construction. Since the Transformer-CNN network can only be used for image processing, using the Transformer-CNN network alone to segment endoscopic images will make the segmentation of the entire endoscopic video lack coherence and it is difficult to utilize the information between frames. Therefore, an LSTM network is proposed to solve this problem. This network mainly consists of a Transformer-CNN network part, a Bi-LSTM part, and an Attention part. Samples are taken from the endoscopic video, and the sampled endoscopic ultrasound images are respectively used by the Transformer-CNN network for feature extraction. Then, the learned features are used as the input for each time point of the Bi-LSTM, and the Attention mechanism is used for weight learning to achieve the optimal segmentation effect.
[0045] 2. Model details. It is proposed to use the Transformer-CNN network to extract features from the input endoscopic ultrasound video frames, that is, to process these video frames into the form of feature vectors that the Bi-LSTM can receive and process.
[0046] Bi-LSTM consists of a forward propagation LSTM and a backward propagation LSTM. In LSTM, the memory controller is generally used to decide which information to forget and retain, and the input and output of information are realized through the three structures of input gate, forget gate and output gate. The traditional LSTM network only considers the information of the forward propagation process, thus ignoring the future information. In Bi-LSTM, the input at the current moment depends not only on the previous video frame, but also on the subsequent video frames, fully considering the timing information before and after the video frame. The Attention mechanism is a brain signal processing mechanism similar to human vision. By calculating the weights of the feature vectors output from the Bi-LSTM network at different times, some important features are highlighted, so that the entire network model can show better performance.
[0047] 3. Model training and optimization. During the training process, the parameters of the Transformer-CNN network are frozen. The labeled ultrasound endoscopic images and labels are used as the input and output of the model training. During the training, the training data is enhanced by random rotation, image shrinkage and enlargement, and blurring. The network is optimized by optimization methods such as Adam or SGD.
[0048] 4. Model performance evaluation. In terms of performance evaluation, DICE and JAC are used to evaluate the segmentation effect of Transformer-CNN and LSTM network models on anatomical structures and lesions. The model inference time, number of floating-point operations, and number of multiplication and accumulation operations are used to evaluate the running speed and computational complexity of the model. Finally, a stable real-time navigation model IDEA-EUS based on ultrasound endoscopic video is formed. IDEA-EUS can identify anatomical parts such as organs and blood vessels displayed in the current picture, as well as abnormal structures in real time in ultrasound endoscopic video.
[0049] The IDEA-EUS model is optimized by combining multimodal data. The optimization process includes: Due to huge space-occupying lesions or tumor invasion, the normal anatomical structure of the abdominal cavity will have a large variation, which is relatively common in the population suitable for ultrasound endoscopy. In order to cope with the differences in individual anatomical structures, the IDEA-EUS model is optimized by combining multimodal imaging data, and a method based on CT / MRI and ultrasound endoscopy image fusion is proposed to construct a multimodal anatomical structure recognition model.
[0050] 1. Data collection: During the collection of ultrasound endoscopic video data, if the patient has undergone CT or MRI examination within one month before or after, the corresponding CT / MRI image data will be collected at the same time.
[0051] 2. Data preprocessing and annotation, including: Abdominal CT, using a fixed number of slices, intercept the part of the abdominal CT that is most relevant to the video acquisition site of endoscopic ultrasound. Use ITK-SNAP software to label abdominal organs, major vascular structures, and space-occupying lesions, with the same labels as endoscopic ultrasound. Adjust to the appropriate window width and window level using an adaptive method; normalize the image pixel values to the range of 0-255. According to convex optimization theory and data probability distribution, perform centering preprocessing on the image by removing the mean, normalize using the Z-Score method, and standardize based on the mean and standard deviation of the data to make the data approximately follow a Gaussian distribution. Perform image enhancement and contrast-limited adaptive histogram equalization operations on the grayscale and ground truth masks.
[0052] Abdominal MRI, using a fixed number of slices, intercept the part of the abdominal MRI that is most relevant to the video acquisition site of endoscopic ultrasound. Use TK-SNAP software to label abdominal organs, major vascular structures, and space-occupying lesions, with the same labels as CT. Adjust to the appropriate window width and window level using an adaptive method; normalize the image pixel values to the range of 0-255. According to convex optimization theory and data probability distribution, perform centering preprocessing on the image by removing the mean, normalize using the Z-Score method, and standardize based on the mean and standard deviation of the data to make the data approximately follow a Gaussian distribution. Perform image enhancement and contrast-limited adaptive histogram equalization operations on the grayscale and ground truth masks.
[0053] 3. Model construction: Construct an anatomical structure recognition model under endoscopic ultrasound video based on multi-modal imaging information, including the following steps: ① Based on Mask Scoring R-CNN, segment soft tissue target objects from endoscopic ultrasound, CT / MRI respectively; ② To establish the correspondence between images, extract feature points from the segmented soft tissue target regions. Feature point extraction methods include corner detection, edge detection, and regional centroid; ③ Since the acquisition methods of endoscopic ultrasound and CT / MRI images are different, there are differences in shape, size, position, etc. To match the two images, one of the images needs to be elastically transformed. Elastic transformation functions include affine transformation with B-spline interpolation, biomechanical model constraints, and finite element simulation; ④ To evaluate the effect of image registration, define mutual information as the objective function to measure the correlation of information between the two images; ⑤ Use the gradient descent method to minimize the objective function and solve the optimization problem of the elastic transformation function parameters to complete image registration; ⑥ According to the registration result, fuse the endoscopic ultrasound and CT / MRI images together to obtain a fused image; ⑦ Based on the fused image, use the Transformer-CNN network to achieve segmentation on anatomical structures and lesions.
[0054] Example 2: Based on the foregoing embodiments, a dynamic spatio-temporal attention fusion mechanism is added. This mechanism uses the self-attention layer of Transformer to capture the global spatial correlation between organs and blood vessels in a single-frame image. At the same time, it extracts the temporal evolution features between adjacent frames through bidirectional LSTM, and adaptively assigns attention weights according to the segmentation confidence of the current frame to suppress misidentifications caused by sudden changes in perspective or image blur.
[0055] First, use the Transformer encoder to generate a global spatial attention map for each frame of endoscopic ultrasound image , and the formula is
[0056] In the formula, is the global spatial attention map; is a function; and are the query and key matrices; is the dimensional scaling factor; is the Sigmoid function; is the learnable weight matrix; is the local feature output by the CNN branch; Input the Transformer features of 5 consecutive frames into the bidirectional LSTM to output the temporal attention weights , and the formula is:
[0057] In the formula, is the temporal attention weight; is the activation function; is the hidden state of the forward LSTM at the (t-1)th frame; is the hidden state of the forward LSTM at the tth frame; is the hidden state of the forward LSTM at the (t+1)th frame; is the hidden state of the backward LSTM at the (t-1)th frame; is the hidden state of the backward LSTM at the tth frame; is the hidden state of the backward LSTM at the (t+1)th frame; and are the learnable parameters; According to the segmentation confidence of the current frame, dynamically fuse the spatial and temporal attention:
[0058] In the formula, is the output of the dynamic fusion; is a dynamic weight parameter, which is used to dynamically adjust the contribution ratio of spatial and temporal attention according to the confidence level; is the slope coefficient, which is used to control the transition smoothness of the fusion weight and is determined by fitting the confidence attenuation curve of the tumor compression case; is the segmentation confidence of the current frame, reflecting the reliability of spatial features; is the confidence threshold, which is used to determine whether to prioritize spatial or temporal features.
[0059] Verification shows that in the case of obtaining an average error similar to that of the above embodiment, on a test set containing 200 pancreatic tumor cases, the dynamic spatiotemporal attention fusion mechanism increases the Dice coefficient from 89.1% of the baseline model to 92.3%, and the anatomical structure misrecognition rate of video clips is reduced from 12.5% to 6.8%. The single-frame processing delay is 38ms, which can meet clinical real-time requirements. The results show that this mechanism significantly improves the limitations of traditional single-frame segmentation models in dynamic scenes by jointly modeling the spatial features and temporal dependencies within ultrasound endoscopic video frames. It can effectively enhance the accuracy and continuity of anatomical structure segmentation, reduce misrecognition caused by sudden changes in perspective or image blur, and maintain real-time processing capabilities, providing endoscopists with smooth and coherent navigation feedback. In addition, its dynamic weight allocation strategy improves the model's adaptability to complex anatomical variations and ensures robustness in highly dynamic inspection scenarios.
[0060] Embodiment 3: On the basis of the above-mentioned embodiment, a multimodal adversarial elastic registration network is added. The network uses a generator to predict the nonlinear deformation field, uses a discriminator to distinguish between real and synthetic registration images, and combines mutual information maximization constraints to ensure the consistency of the anatomical structure after deformation, thereby improving the accuracy and generalization ability of multimodal fusion.
[0061] The generator takes ultrasound endoscopic images and CT images as input and outputs the deformation field, the formula is:
[0062] In the formula, is the deformation field; For the generator; This is an endoscopic ultrasound image; It is a CT image; are the generator parameters, and the network structure is U-Net+residual block; The discriminator distinguishes between real registration pairs and synthetic registration pairs, and the adversarial loss function is:
[0063] In the formula, To combat the loss, a discriminator is used to distinguish between real and synthetic registration pairs; is the mathematical expectation, representing the calculation of the loss under the true data distribution.
[0064] The mutual information loss adopts neural network-based probability density estimation to avoid the curse of dimensionality problem of the traditional histogram method and maximize the mutual information between the deformed endoscopic ultrasound image and the CT image:
[0065] In the formula, is the mutual information loss, maximizing the mutual information between the deformed endoscopic ultrasound image and the CT image to maintain anatomical consistency; is the number of training samples; is the joint probability and marginal probability estimator parameterized by the neural network; is the counter; is the spatial transformation operation based on the displacement field; is the regularization term weight, used to control the smoothness of the displacement field; is the total variation regularization term of the displacement field, used to constrain the local smoothness of the deformation field.
[0066] The total loss function integrates the adversarial loss, mutual information constraint, and deformation field regularization: In the formula, is the output of the total loss function; , , and are the balance factors of each loss, used to adjust the weights between the adversarial loss, mutual information loss, and displacement field regularization term; and are the balance coefficients of the total variation and L2 regularization term; is the total variation; is the L2 regularization term; is the segmentation confidence of the current frame, reflecting the reliability of the registration result; is the confidence threshold, used to dynamically adjust the regularization strength; is the KL divergence regularization term, used to constrain the displacement field distribution output by the generator.
[0067] Suppose 50 cases of multimodal imaging data of pancreatic cancer patients are selected, each case including endoscopic ultrasound dynamic video, abdominal CT, and MRI images. Among them, 30 cases are normal cases of local anatomical structures, and 20 cases have significant deformation of the pancreatic morphology due to tumor invasion. The cross-modal registration of endoscopic ultrasound and CT is performed using the multimodal adversarial elastic registration network, and the performance differences with the traditional B-spline registration algorithm are compared.
[0068] The effects of the multimodal adversarial elastic registration network are shown in Table 1.
[0069] Table 1. Summary of the effects of the multi-modal adversarial elastic registration network
[0070] According to the experimental table, by manually annotating 10 anatomical landmark points in endoscopic ultrasound and CT images and calculating the Euclidean distance of the corresponding points after registration, the TRE of the multi-modal adversarial elastic registration network dropped to 1.2 mm in cases of anatomical variation, proving that its elastic deformation field can effectively compensate for non-linear displacement; by measuring the overlap degree of the pancreatic regions segmented by endoscopic ultrasound and CT, the mean value of the multi-modal adversarial elastic registration network reached 0.91, with a standard deviation of 4.3%, indicating a significant improvement in its generalization ability among different cases and medical centers; the multi-modal adversarial elastic registration network accelerates convergence through adversarial training, reducing the number of iterations by 40% and shortening the registration time per case to 1.5 seconds, which can meet the clinical real-time requirements. Experimental verification shows that it significantly improves the registration accuracy of endoscopic ultrasound and CT / MRI images. Especially in cases with severe anatomical structure variation, it can effectively compensate for non-linear displacement, ensuring the reliability of multi-modal feature fusion. At the same time, the adversarial training strategy enhances the generalization ability of the model to the data distribution of different medical centers, reduces the dependence on the amount of labeled data, and provides a technical basis for the precise navigation of complex cases.
[0071] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable non-transitory storage media containing computer-usable program code.
[0072] The present invention can provide computer program instructions to the management platform of a general computer, a special computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed through the management platform of the computer or other programmable data processing devices generate a device for implementing the system.
[0073] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing devices to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions of the system.
[0074] These computer program instructions can also be loaded onto a computer or other programmable data processing devices, so that a series of operation steps are executed on the computer or other programmable devices to generate computer-implemented processing, and thus the instructions executed on the computer or other programmable devices provide steps for implementing the functions of the system.
Claims
1. An endoscopic ultrasound navigation system based on deep learning, characterized in that, Including: A parallel encoder module, which is used to extract local features of endoscopic ultrasound images through a CNN branch and capture multi-scale global context information through a Transformer branch adopting a pyramid structure; the Transformer branch includes a non-overlapping patch embedding layer and multi-stage downsampling; A channel attention fusion module, which is used to adaptively weight and fuse the extracted features; A decoder module, which generates an anatomical structure segmentation mask based on the fused features through multi-level skip connections; A temporal processing module, which is used to receive the frame-by-frame segmentation results output by the decoder module and realize the temporal coherence analysis of endoscopic ultrasound videos through a bidirectional LSTM network combined with an Attention mechanism; A multi-modal fusion module, which is used to register and fuse endoscopic ultrasound images; the multi-modal fusion module includes an adversarial elastic registration sub-module, which uses a generator to predict a non-linear deformation field, and a discriminator to distinguish between real and synthetic registered images, and at the same time combines mutual information maximization constraints to ensure anatomical structure consistency; A spatio-temporal attention fusion module, which is used to dynamically allocate spatial and temporal attention weights according to the segmentation confidence of the current frame.
2. The system according to claim 1, wherein: The resolution of the feature map output by the Transformer branch in the 8-fold downsampling stage is aligned with the CNN branch, and the spatial dimension of the feature map is matched through bilinear interpolation.
3. The system according to claim 1, wherein: The upsampling operation of the decoder module uses transposed convolutional layers to perform 2x, 2x, and 4x upsampling in sequence, and the loss function is jointly optimized using Dice Loss and cross-entropy loss.
4. The system according to claim 1, wherein: The spatio-temporal attention fusion module uses the self-attention layer of the Transformer branch to capture the global spatial correlation between organs and blood vessels in a single-frame image, extracts the temporal evolution features between adjacent frames through a bidirectional LSTM network, and dynamically allocates spatial and temporal attention weights according to the segmentation confidence of the current frame. The fusion formula is: In the formula, is the dynamic fusion output; is the dynamic weight parameter; is the global spatial attention map; is the temporal attention weight; is the slope coefficient, determined by fitting the confidence decay curve of tumor compression cases; is the segmentation confidence of the current frame; is the confidence threshold.
5. The system according to claim 1, wherein: The total loss function of the adversarial elastic registration module is: In the formula, is the output of the total loss function; , , and are the balance factors of each loss; is the adversarial loss; is the mutual information loss; and are the balance coefficients of the total variation and the L2 regularization term; is the total variation; is the L2 regularization term; is the segmentation confidence of the current frame; is the confidence threshold; is the KL divergence regularization term.
6. An anatomical structure annotation method based on endoscopic ultrasound video that does not involve the diagnosis and treatment of diseases, characterized in that, Including: Filtering invalid frames in the endoscopic ultrasound video based on a CNN model and normalizing the valid frames; Performing anatomical structure segmentation on a single-frame image through a Transformer-CNN network; Performing temporal consistency correction on the continuous frame segmentation results through a bidirectional LSTM network.
7. The method according to claim 6, wherein: The normalization process further includes metadata anonymization, specifically by cropping the region of interest or replacing the patient privacy information in the image with a black mask.
8. A deep learning-based endoscopic ultrasound navigation method not involving the diagnosis and treatment of diseases, wherein: The implementation of the method is based on the system according to any one of claims 1-5: The method includes: Extracting local features of endoscopic ultrasound images and capturing multi-scale global context information; Adaptive weighted fusion is performed on the extracted features, and an anatomical structure segmentation mask is generated based on the fused features; The output frame-by-frame segmentation results are received, and the temporal coherence analysis of the endoscopic ultrasound video is realized through a bidirectional LSTM network; The endoscopic ultrasound images are registered and fused.
9. A computer device, comprising a management platform and a memory, the management platform being connected to the memory, the memory being used for storing computer programs, characterized in that: The management platform is used to execute the computer program stored in the memory, so that the computer device executes the method described in claim 8.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is run, it executes the method described in claim 8.
Citation Information
Patent Citations
Video object recognition method, device and equipment and storage medium
CN109815931A
Remote sensing image semantic segmentation method and system
CN114359734A
Image processing method and device, terminal equipment and storage medium
CN115311219A
Four-axis fusion method based on CNN and Transform
CN116188928A
Medical image analyzing and processing system based on image analysis
CN118485643A
Cited By
Ultrasonic image discrimination perception pre-training method based on cooperative training framework
CN121213578A
Focus determination method and device based on multi-modal data
CN121236040A
Human body medical image classification method based on ultrasonic image and CT image
CN121505368A
4D ultrasonic quality control method, device and program product
CN121904565A
Deep learning-based ultrasonic endoscope navigation system and method for biliopancreatic system
CN122199879A