Multi-modal medical image analysis system, medium and equipment for gastrointestinal disease identification
By using a multimodal medical image analysis system, feature fusion is performed using self-attention mechanism and multi-head attention strategy to generate a dense deformation field for localization of gastrointestinal lesions. This solves the problems of insufficient imaging quality and poor data stability in existing technologies, and achieves high-precision real-time localization of gastrointestinal lesions.
Patent Information
- Application Number
- CN202511841297.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-09
- Publication Date
- 2026-01-06
- Estimated Expiration
- 2045-12-09
AI Technical Summary
Existing gastrointestinal endoscopic imaging technologies suffer from insufficient image quality, significant environmental interference, and poor data stability in gastrointestinal mucosal image analysis, making it difficult to meet the needs of real-time localization of gastrointestinal lesions.
A multimodal medical image analysis system was adopted. Multi-scale texture and structural features were extracted through modal feature encoding and position awareness modules. The correspondence between features between modalities was learned by combining dynamic feature alignment module. Feature fusion was performed using self-attention mechanism and multi-head attention strategy to generate dense deformation field for registration. Finally, the gastrointestinal lesions were located using YOLOv5 model.
It improves the accuracy of locating gastrointestinal lesions, enhances data stability and anti-interference ability, and can accurately identify key structural features in complex backgrounds to meet real-time detection needs.
Smart Images

Figure CN121280431A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of medical image processing technology, and particularly relates to a multimodal medical image analysis system, medium and device for identifying gastrointestinal diseases. Background Technology
[0002] The statements in this section are merely background information related to the present invention and do not necessarily constitute prior art.
[0003] In the field of biomedical testing, analyzing blood oxygen saturation using gastrointestinal mucosal images is a key method for the early detection of gastrointestinal and systemic diseases. As an important digestive and absorptive tissue, the blood oxygen saturation of the gastrointestinal mucosa, such as arterial and venous blood oxygen saturation, directly reflects local metabolic status and systemic circulatory function. By analyzing the blood oxygen levels in the blood vessels of the gastrointestinal mucosa, doctors can indirectly assess gastrointestinal diseases such as gastritis, gastric ulcers, colorectal cancer, inflammatory bowel disease, and gastrointestinal polyps, as well as systemic diseases such as hypertension, coronary heart disease, stroke risk, chronic kidney disease, and Alzheimer's disease.
[0004] Currently, commonly used imaging techniques mainly include color imaging in gastrointestinal endoscopy, optical coherence tomography (OCT), and laser Doppler flowmetry (LDF). Color imaging in gastrointestinal endoscopy uses different wavelengths of light (such as 570nm and 605nm) to analyze differences in vascular color to calculate blood oxygen saturation; OCT combined with flow imaging (OCT-A) assesses microvascular oxygenation status; and LDF measures the correlation between blood flow velocity and blood oxygen concentration. However, these traditional methods, which calculate blood oxygen saturation based on the spectral absorption differences between chromohyhemoglobin and deoxyhemoglobin, have several drawbacks: Limited imaging quality: The complex internal environment of the gastrointestinal tract (such as mucus interference and intestinal peristalsis) leads to insufficient resolution and contrast of gastrointestinal mucosal images, which in turn seriously affects the accuracy of the data; Significant environmental interference: External factors such as the intensity of the endoscope light source and the reflection of the mucosal surface can easily cause deviations in spectral data, making the test results unstable; Poor data stability: It requires manual collection of a large number of samples for verification, and the whole process is complex and time-consuming, making it difficult to meet the needs of real-time detection. Summary of the Invention
[0005] To address at least one of the technical problems mentioned above, this invention provides a multimodal medical image analysis system, medium, and device for identifying gastrointestinal diseases. These systems have strong anti-interference capabilities and high data stability, and can meet the real-time requirements for locating gastrointestinal lesions.
[0006] To achieve the above objectives, the present invention adopts the following technical solution: A first aspect of the present invention provides a multimodal medical image analysis system for identifying gastrointestinal diseases, comprising the following steps: The modal feature encoding and position awareness module is used to extract multi-scale texture and structural features from color images and OCT depth images of gastrointestinal endoscopy, respectively, and to fuse the extracted multi-scale texture and structural features to obtain a cross-modal feature sequence. The dynamic feature alignment module is used to align features between modalities based on feature sequences with added positional encoding, combined with a cross-attention layer, by sharing a query key-value matrix, and to learn the spatial correspondence between different modalities in vascular morphology and gastrointestinal wall layer structure. The feature extraction module is used to extract multi-scale global context feature sequences and local detail feature sequences based on aligned cross-modal vascular and gastrointestinal wall interlayer structural features, and fuse multi-scale and local detail features to obtain fused multi-modal features; The dynamic deformation field generation and calibration module is used to convert the fused multimodal features into a displacement field, then upsample the displacement field to the original image resolution to generate a dense deformation field, introduce physiological signals, and adjust the dense deformation field based on the physiological signals to obtain the adjusted dense deformation field. The multi-stage registration module is used for coarse and fine registration based on the adjusted dense deformation field to obtain the registered color images and OCT depth images of the gastrointestinal endoscope.
[0007] Furthermore, in the modal feature encoding and position awareness module, the extraction of cross-modal vascular features based on the acquired gastrointestinal endoscopy color image and gastrointestinal endoscopy OCT depth image includes inputting the gastrointestinal endoscopy color image and gastrointestinal endoscopy OCT depth image into the corresponding CNN encoder respectively, extracting multi-scale texture features and structural features, and fusing the extracted multi-scale texture features and structural features to obtain a cross-modal feature sequence.
[0008] Furthermore, in the modal feature encoding and location awareness module, based on the added encoded feature sequence, the spatial correspondence of different modal data in terms of vascular morphology and gastrointestinal wall structure is learned. This includes combining a cross-attention layer, using gastrointestinal endoscopy image tokens as queries and gastrointestinal endoscopy OCT image tokens as keys, calculating attention weights, and combining the attention weights to learn the spatial correspondence of different modal data in terms of vascular morphology and gastrointestinal wall structure.
[0009] Furthermore, in the feature extraction module, a multi-head global attention mechanism is used to capture global context information at different scales for each head, resulting in a multi-scale global context feature sequence. This multi-scale global context feature sequence is then combined with interlayer structure tokens from the gastrointestinal endoscopy OCT depth image and a local variability attention mechanism to extract local detail feature sequences.
[0010] Furthermore, in the feature extraction module, each head captures global contextual information at different scales, specifically: heads 1 and 2 capture the main trunks of large blood vessels, heads 3 and 4 focus on the edges of gastrointestinal lesions, and heads 5 and 6 capture microvascular details.
[0011] Furthermore, in the feature extraction module, when fusing global contextual information and local detail features to obtain fused multimodal features, a pyramid feature fusion strategy is adopted. The bottom layer superimposes global vascular distribution features with local microvascular texture features, the middle layer fuses global gastrointestinal specific reference region relative position features with local gastrointestinal lesion edge gradient features, and the top layer combines global spectral features with local vascular bifurcation detail features.
[0012] Furthermore, in the dynamic deformation field generation and calibration module, the dynamic adjustment of the dense deformation field parameters is expressed as follows: , in, For the adjusted dense deformation field parameters, The parameters of the dense deformation field before adjustment. The weighting coefficients are influenced by physiological signals. S is the activation value of the gated recurrent unit after encoding the physiological signal, and S is the output state vector.
[0013] Furthermore, the system also includes a gastrointestinal lesion localization module, which inputs the registered gastrointestinal endoscopy color image into the backbone network of the YOLOv5 model, extracts the two-dimensional texture features of the gastrointestinal lesions, performs depth slice feature extraction on the registered gastrointestinal endoscopy OCT depth image, obtains the three-dimensional structural features of the gastrointestinal lesion region, fuses the two-dimensional texture features and the three-dimensional structural features of the gastrointestinal lesion region and inputs them into the YOLOv5 detection head, outputs the candidate boxes and confidence scores of the gastrointestinal lesions, and determines the center coordinates and range of the gastrointestinal lesion based on the candidate boxes.
[0014] A second aspect of the present invention provides a computer-readable storage medium.
[0015] A computer-readable storage medium having a computer program stored thereon, which is implemented when executed by a processor: Multi-scale texture and structural features were extracted from color images and OCT depth images of gastrointestinal endoscopy, and the extracted multi-scale texture and structural features were fused to obtain a cross-modal feature sequence. Based on the feature sequences with added position encoding, combined with the cross-attention layer, the spatial correspondence between different modal data in terms of vascular morphology and gastrointestinal wall layer structure is learned by aligning intermodal features through a shared query key-value matrix; Based on the aligned cross-modal vascular and gastrointestinal wall interlayer structural features, multi-scale global context feature sequences and local detail feature sequences are extracted, and multi-scale and local detail features are fused to obtain fused multi-modal features; The fused multimodal features are converted into a displacement field, and then the displacement field is upsampled to the original image resolution to generate a dense deformation field. Physiological signals are introduced, and the dense deformation field is adjusted based on the physiological signals to obtain the adjusted dense deformation field. Coarse and fine registration were performed based on the adjusted dense deformation field to obtain the registered color images and OCT depth images of the gastrointestinal endoscope.
[0016] A third aspect of the present invention provides a computer device.
[0017] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, implements: Multi-scale texture and structural features were extracted from color images and OCT depth images of gastrointestinal endoscopy, and the extracted multi-scale texture and structural features were fused to obtain a cross-modal feature sequence. Based on the feature sequences with added position encoding, combined with the cross-attention layer, the spatial correspondence between different modal data in terms of vascular morphology and gastrointestinal wall layer structure is learned by aligning intermodal features through a shared query key-value matrix; Based on the aligned cross-modal vascular and gastrointestinal wall interlayer structural features, multi-scale global context feature sequences and local detail feature sequences are extracted, and multi-scale and local detail features are fused to obtain fused multi-modal features; The fused multimodal features are converted into a displacement field, and then the displacement field is upsampled to the original image resolution to generate a dense deformation field. Physiological signals are introduced, and the dense deformation field is adjusted based on the physiological signals to obtain the adjusted dense deformation field. Coarse and fine registration were performed based on the adjusted dense deformation field to obtain the registered color images and OCT depth images of the gastrointestinal endoscope.
[0018] Compared with the prior art, the beneficial effects of the present invention are: 1. This invention utilizes a self-attention mechanism to perform global correlation modeling on the feature sequences of multimodal images. By calculating the attention weights of any two pixel features in the feature sequence (such as feature vectors of key structures like vascular branch nodes, gastrointestinal lesion edge contours, and microvascular textures), it accurately quantifies the correlation strength of different regions in terms of spatial location and structural morphology. This allows global structural features such as the continuity of vascular trunks and branches, and the distribution relationship between gastrointestinal lesions and surrounding blood vessels to be incorporated into the feature learning scope. This breaks through the limitations of traditional local feature extraction methods, establishes feature dependencies at the full-image scale, and improves the accuracy of gastrointestinal lesion localization. 2. This invention targets vulnerable regions in gastrointestinal mucosal images (such as gastrointestinal mucosal reflections, shadows caused by vascular intersections, and low-contrast microvascular regions), employing a multi-head attention multi-subspace learning strategy. Different attention subspaces focus on local detail features (such as the edge texture of microvessels and the boundary gradient of gastrointestinal lesions) and global contextual features (such as the overall distribution trend of the vascular network and the relative positional relationship between gastrointestinal lesions and specific reference regions of the gastrointestinal tract). Through the complementary fusion of multi-subspace features, the interference of local noise (such as pixel value anomalies caused by reflections and feature blurring caused by shadows) on global registration is effectively weakened, ensuring that key structural features can still be accurately identified and enhanced in complex backgrounds.
[0019] Advantages of additional aspects of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description
[0020] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.
[0021] Figure 1 This is a block diagram of a multimodal medical image analysis system for identifying gastrointestinal diseases provided in an embodiment of the present invention. Detailed Implementation
[0022] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0023] It should be noted that the following detailed description is illustrative and intended to provide further explanation of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0024] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of exemplary embodiments according to the invention. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise. Furthermore, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0025] To address the complex vascular distribution, lesion morphology, and susceptibility to intestinal peristalsis in gastrointestinal mucosal images, this invention leverages the advantages of the Transformer self-attention mechanism to extract global image features, significantly improving registration robustness. First, the self-attention mechanism is used to perform global correlation modeling on the feature sequences of multimodal images. By calculating the attention weights of any two pixel features in the feature sequence (such as feature vectors of key structures like vascular branch nodes, gastrointestinal lesion edges, and microvascular textures), the correlation strength of different regions in spatial location and structural morphology is precisely quantified. This incorporates global structural features such as the continuity of vascular trunks and branches, and the distribution relationship between gastrointestinal lesions and surrounding blood vessels into the feature learning scope, overcoming the limitations of traditional local feature extraction methods and establishing feature dependencies at the full-image scale. Meanwhile, for special regions in gastrointestinal mucosal images that are easily disturbed (such as gastrointestinal mucosal reflection, shadows caused by vascular intersections, and low-contrast microvascular regions), a multi-head attention multi-subspace learning strategy is adopted: different attention subspaces focus on local detail features (such as the edge texture of microvessels and the boundary gradient of gastrointestinal lesions) and global context features (such as the overall distribution trend of vascular networks and the relative positional relationship between gastrointestinal lesions and specific reference regions of the gastrointestinal tract). Through the complementary fusion of multi-subspace features, the interference of local noise (such as pixel value abnormalities caused by reflection and feature blurring caused by shadows) on global registration is effectively weakened, ensuring that key structural features can still be accurately identified and enhanced in complex backgrounds.
[0026] Example 1 like Figure 1 As shown, this embodiment provides a multimodal medical image analysis system for gastrointestinal disease identification, including the following steps: The image acquisition module is used to acquire color images and OCT depth images of gastrointestinal endoscopes. In the image acquisition module, multispectral gastrointestinal endoscopy is used to acquire color and grayscale images, covering wavelength bands such as 550nm and 600nm; a gastrointestinal endoscopy optical coherence tomography image acquisition module is introduced to acquire depth information of the gastrointestinal wall.
[0027] In this embodiment, when acquiring color images of gastrointestinal endoscopy, the endoscopy parameters are set to high resolution (e.g., 4K or higher) and short exposure time (e.g., 1 / 1000 second) to ensure that the images are clear, complete, and can accurately reflect the spectral information of the gastrointestinal mucosa.
[0028] The image preprocessing module is used to denoise and enhance the contrast of the acquired color images and OCT depth images of gastrointestinal endoscopy. In the image preprocessing module, the median filtering algorithm is used to remove mucus interference noise in the image, and the histogram equalization method is used to enhance the contrast of the image. For gastrointestinal endoscopy OCT images, wavelet transform is used to denoise and preserve the image details (such as the interlayer structure of the gastrointestinal wall and microvascular texture).
[0029] The cross-modal feature encoding and position-aware module is used to input the color image of the gastrointestinal endoscope and the depth image of the gastrointestinal endoscope OCT to the corresponding CNN encoder, respectively, to extract multi-scale texture features and structural features, and to fuse the extracted multi-scale texture features and structural features to obtain a cross-modal feature sequence. In this embodiment, the CNN encoder adopts a lightweight design, such as MobileNetV3, which reduces the computational load while maintaining the feature resolution (512×512). The extracted multi-scale texture features include features such as blood vessel bifurcation, gastrointestinal lesion edges, and microvascular textures, while the structural features include the thickness of the gastrointestinal wall layers. The extracted multi-scale texture features and structural features are converted into sequence tokens through linear projection, with each token corresponding to a 16×16 pixel region, forming a feature sequence of dimension N×D, where N is the number of tokens and D is the dimension of the feature vector. In this embodiment, D=512.
[0030] The positional encoding addition module is used to add positional encoding to cross-modal feature sequences to obtain feature sequences with added positional encoding. In this embodiment, a hybrid positional coding strategy is adopted. Sinusoidal positional codes are added to the tokens of the gastrointestinal endoscopy images to capture global spatial relationships, and learnable positional codes are added to the tokens of the gastrointestinal endoscopy OCT images to adapt to nonlinear changes in the depth direction. The specific hybrid positional coding strategy is expressed as follows: , in, Indicates the position code at the 1st i The first feature token, the first j The values in each dimension are used to assign spatial location information to the features of gastrointestinal endoscopy images or gastrointestinal endoscopy OCT images, helping the model to perceive the spatial distribution relationship of the features; This is a learned location embedding function designed for gastrointestinal endoscopic OCT images. The input is the index i of a feature token, and the output is the learned location encoding vector corresponding to that token. Since gastrointestinal endoscopic OCT is a tomographic imaging technique of the gastrointestinal wall (with complex inter-layer structures), "learned encoding" can more flexibly capture the locational correlations of its features. i Indexed by token j For dimensional indexing.
[0031] The dynamic feature alignment module is used to align features between modalities based on feature sequences with added positional encoding, combined with a cross-attention layer, by sharing a query key-value matrix, and to learn the spatial correspondence between different modalities in vascular morphology and gastrointestinal wall layer structure. In this embodiment, a cross-attention layer is added after the encoder, using gastrointestinal endoscopy image tokens as query Q and gastrointestinal endoscopy OCT image tokens as key values (KV) to calculate attention weights: , After obtaining the attention weights, the spatial correspondence between different modalities of data in terms of vascular morphology and gastrointestinal wall interlayer structure is learned in the following way: two-dimensional structural features such as vascular bifurcation and gastrointestinal lesion edges are captured using gastrointestinal endoscopy image tokens (Q), and three-dimensional structural features such as gastrointestinal wall interlayer thickness and vascular depth are provided using gastrointestinal endoscopy OCT image tokens (KV); the attention weights will assign higher weights to features in K that match the structure of Q (such as the specific interlayer boundary in OCT corresponding to the vascular bifurcation point in gastrointestinal endoscopy), and realize the spatial correlation mapping of intermodal features through weighted summation of V; at the same time, the intermodal feature difference loss (such as L1 loss) is minimized through backpropagation, and the parameters of the QKV matrix are iteratively optimized so that the aligned feature sequence achieves the optimal balance in terms of vascular morphological continuity (such as trunk-branch connection) and interlayer structural consistency (such as the gastrointestinal mucosal muscle layer depth corresponding to the gastrointestinal lesion area).
[0032] The feature extraction module is used to obtain a multi-scale global context feature sequence by combining aligned cross-modal vascular features with a multi-head global attention mechanism. Each head captures global context information at different scales. Combining the multi-scale global context feature sequence with interlayer structure tokens of gastrointestinal endoscopy OCT depth images and a local variability attention mechanism, a local detail feature sequence is extracted. The multi-scale and local detail features are fused to obtain the fused multimodal features. The feature extraction module specifically includes a global information extraction module, a local information extraction module, and a feature fusion module. The global information extraction module employs a multi-head global attention encoder, with each head independently capturing global context at different scales. Input: a sequence of key structural features aligned between two modalities (dimension N×D, where N is the number of tokens and D=512, including alignment features such as gastrointestinal mucosal vascular texture and interlayer structure of gastrointestinal endoscopy OCT). Output: a multi-scale global context feature sequence (dimension maintained at N×D, each token incorporating structural association information at the corresponding scale across the entire image, such as heads 1-2 outputting global distribution features of large blood vessel trunks, heads 3-4 outputting global contour features of gastrointestinal lesion edges, and heads 5-6 outputting global network features of microvessels). Specifically, heads 1 and 2 capture the main trunks of large blood vessels, heads 3 and 4 focus on the edges of gastrointestinal lesions, and heads 5 and 6 capture microvascular details. In this embodiment, the receptive fields of heads 1 and 2 are 64×64, the receptive fields of heads 3 and 4 are 32×32, and the receptive fields of heads 5 and 6 are 16×16. After obtaining the global features, gradient vanishing is mitigated through residual connections and layer normalization (LayerNorm), ensuring the effective transmission of global features.
[0033] This invention targets vulnerable regions in gastrointestinal mucosal images (such as gastrointestinal mucosal reflections, shadows caused by vascular intersections, and low-contrast microvascular regions). It employs a multi-head attention multi-subspace learning strategy: different attention subspaces focus on local detail features (such as the edge texture of microvessels and the boundary gradient of gastrointestinal lesions) and global contextual features (such as the overall distribution trend of the vascular network and the relative positional relationship between gastrointestinal lesions and specific reference regions of the gastrointestinal tract). Through the complementary fusion of multi-subspace features, the interference of local noise (such as pixel value anomalies caused by reflections and feature blurring caused by shadows) on global registration is effectively weakened, ensuring that key structural features can still be accurately identified and enhanced in complex backgrounds.
[0034] The local information extraction module is used to introduce deformable attention into the Transformer decoder. Inputs include: a multi-scale global contextual feature sequence output from the global information extraction module (as query tokens, focusing on local regions requiring refinement, such as vascular bifurcation and low-contrast microvessels), and inter-layer structure tokens from the original gastrointestinal endoscopy OCT images (as a reference feature library). Outputs include: a sub-pixel aligned local detail feature sequence (dimension N×D, each query token corresponding to a weighted feature of 8 reference points, highlighting the fine correlation of local structures, such as OCT inter-layer boundary matching features at vascular bifurcation). For example, for vascular bifurcation regions in gastrointestinal mucosal images, the model automatically samples the inter-layer boundary tokens of the surrounding gastrointestinal endoscopy OCT images to achieve sub-pixel alignment. The sampling formula is: , in, The predicted offset. As weight, In this embodiment, the number of reference points is... =8, The feature vector corresponding to the final sampling position in the reference feature library; The feature fusion module is used to associate and fuse the multi-scale global features output by the global attention encoder with the local detail features output by the local deformation perception module using a pyramid feature fusion strategy to obtain the fused feature vector. In this embodiment, the pyramid feature fusion strategy includes fusing the multi-scale global features output by the global attention encoder with the local detail features output by the local deformation perception module through a skip connection: the bottom layer (4× downsampling) superimposes the global vascular distribution features with the local microvascular texture features to enhance the integrity of the vascular network; the middle layer (2× downsampling) fuses the relative position features of global gastrointestinal lesions and specific reference regions of the gastrointestinal tract with the edge gradient features of local gastrointestinal lesions to improve the accuracy of structural localization; the top layer (original image) combines global spectral features with local vascular bifurcation detail features to optimize the recognition of microstructures; during the fusion process, element-level addition and channel attention weighting (assigning higher weights to key structural feature channels) are used to ensure the effective association between global and local features.
[0035] The dynamic deformation field generation and calibration module is used to convert the fused multimodal features into a displacement field, then upsample the displacement field to the original image resolution to generate a dense deformation field, introduce physiological signals, and adjust the dense deformation field based on the physiological signals to obtain the adjusted dense deformation field. In this embodiment, during the dense deformation field process, the attention weights output by the decoder are converted into a displacement field (Δx, Δy) through a two-layer fully connected network, with each token corresponding to a two-dimensional displacement vector. The displacement field is then upsampled to the original image resolution (512×512) using bilinear interpolation to form the dense deformation field. , represented as: , The dense deformation field generated by this formula is used for geometric correction of gastrointestinal endoscopic OCT images: This represents the horizontal and vertical offset of each pixel (x, y) in the image. Pixel positions are adjusted through pixel resampling (e.g., bilinear interpolation) to correct structural misalignment between the image and the gastrointestinal endoscopy image caused by intestinal peristalsis (e.g., ±15° deformation, local stretching). This provides corrected image data for subsequent registration steps to locate gastrointestinal lesions. This deformation field directly acts on the gastrointestinal endoscopy OCT image to correct for local stretching (e.g., ±15° deformation) caused by intestinal peristalsis.
[0036] In this embodiment, physiological signals are introduced, and the dense deformation field is adjusted based on the physiological signals to obtain the adjusted dense deformation field, including: Physiological signals such as heart rate (HR) and blood pressure (BP) are introduced as auxiliary inputs and encoded into a state vector s through a gated recurrent unit (GRU). The deformation field parameter θ is dynamically adjusted using the following formula: , , in, The dense deformation field parameters before adjustment (including all pixels) , Offset matrix); The weighting coefficient for the influence of physiological signals is set to 0.1~0.3 (an empirical value is set to control the adjustment range of physiological signals on the deformation field and avoid over-correction). The activation value is the result of encoding physiological signals into the gated recurrent unit (GRU), and S is the output state vector (integrating temporal features of heart rate and blood pressure). The ReLU function ensures that the activation value is non-negative, avoiding reverse adjustment. The adjusted dense deformation field parameters incorporate physiological signal feedback, such as appropriately reducing the local deformation amplitude when the heart rate is too fast, thereby improving the stability of the deformation field.
[0037] A multi-stage registration module is used to perform coarse and fine registration based on the adjusted dense deformation field to obtain registered color images and OCT depth images of gastrointestinal endoscopy. In this embodiment, the coarse registration stage is as follows: the encoder is fixed (to avoid low-level feature shift), and only the parameters are optimized; an optimizer is used (the learning rate is set to 1e-4), with mean squared error (MSE) as the loss function (calculating the pixel value difference between the corrected gastrointestinal endoscopy OCT image and the gastrointestinal endoscopy image); iterates for 200 rounds until MSE < 0.05, to achieve global structural alignment (such as matching the main trunk of large blood vessels and the approximate area of gastrointestinal lesions).
[0038] Fine-tuning stage: Unfreeze the CNN encoder (allowing fine-tuning of low-level features), introduce Structural Similarity (SSIM) loss (measuring image structural similarity, weight 0.6) and Edge-Preserving Loss (weight 0.4, avoiding edge blurring through edge gradient penalty); use AdamW optimizer (learning rate set to 5e-5, weight decay 1e-5) to focus on optimizing the alignment accuracy of blood vessel bifurcation points (through keypoint matching loss) and the interlayer boundaries of gastrointestinal endoscopy OCT (through interlayer distance loss); iterate 50-100 times until SSIM>0.92, achieving accurate alignment of fine structures such as gastrointestinal lesion edges and microvessels.
[0039] The visual disc localization module is used to fuse the registered color images of gastrointestinal endoscopy and the depth images of gastrointestinal endoscopy OCT to obtain the localization results of gastrointestinal lesions; In this embodiment, the registered color images of gastrointestinal endoscopy are input into the backbone network (CSPDarknet) of the YOLOv5 model to extract the two-dimensional texture features of the gastrointestinal tract, including the grayscale difference between the gastrointestinal lesions and the surrounding tissues, and surface morphological features such as edge contours. Depth slice feature extraction is performed on the registered gastrointestinal endoscopy OCT depth images to obtain the three-dimensional structural features of the gastrointestinal lesion region, such as the infiltration depth of the gastrointestinal lesion and the thickness distribution of each layer of the gastrointestinal wall. The two-dimensional texture features of the gastrointestinal lesion and the three-dimensional structural features of the gastrointestinal lesion region are fused and input into the YOLOv5 detection head to output the candidate boxes and confidence scores of the gastrointestinal lesions. The center coordinates and range of the gastrointestinal lesion are determined based on the candidate boxes. In this embodiment, the non-maximum suppression algorithm (NMS threshold is set to 0.5 in this embodiment) is used to remove duplicate detection boxes, and the shape (irregular or near-circular) and location prior information (usually located in a specific segment of the gastrointestinal tract) of the gastrointestinal lesion are combined for verification and correction, and finally the accurate center coordinates and range of the gastrointestinal lesion are output.
[0040] Specifically, when fusing the two-dimensional texture features and three-dimensional structural features of gastrointestinal lesions, a convolutional attention mechanism can be used to fuse the two types of features and dynamically allocate weights to enhance the role of depth features. For example, when distinguishing between gastrointestinal lesions and adjacent normal tissues (such as inflammatory areas or densely vascularized areas), the three-dimensional structural features of gastrointestinal endoscopy OCT are preferentially used to improve the discriminative power. The final output of the gastrointestinal lesion localization results is used to define the region of interest (ROI): delineate the key areas for subsequent gastrointestinal mucosal vessel segmentation and spectral analysis, reduce the interference of irrelevant background areas in the gastrointestinal mucosal image, and improve the efficiency and accuracy of subsequent analysis.
[0041] The multi-task learning module is used to learn gastrointestinal mucosal vessel segmentation results and blood oxygen saturation based on the localization results of gastrointestinal lesions. In this embodiment, multi-task learning is used to enhance the synergy between vascular structure recognition (assisting in the localization of gastrointestinal lesions) and blood oxygen calculation (assisting in disease assessment). The specific content is as follows: gastrointestinal mucosal vessel segmentation and blood oxygen saturation prediction are combined into a multi-task learning problem. Synergistic optimization is achieved by designing a dual-tower architecture of "shared encoder + task-specific decoder": the improved UNet encoder is used as a shared feature extraction layer, which includes 5 downsampling stages. Multi-scale features are extracted through residual blocks and skip connections. The shallow features capture details such as the edge of the blood vessels, the middle features depict the morphology and spectral characteristics of the blood vessels, and the deep features aggregate global blood vessel distribution information. These features serve both tasks at the same time, avoiding redundant extraction.
[0042] The vessel segmentation task outputs a binary vessel mask through an upsampling branch, while the blood oxygenation prediction task outputs arterial / venous blood oxygenation values through a fully connected branch. The vessel segmentation task is optimized using a cross-entropy loss function (weighted cross-entropy is used to address class imbalance, with vessel pixel weights set to 5 and background to 1), and the blood oxygenation prediction task is optimized using a mean squared error loss function (the prediction errors for arterial and venous blood oxygenation are calculated separately). The total loss function for multi-task learning is: , in, It is the cross-entropy loss of blood vessel segmentation. This is the mean squared error loss in blood oxygen saturation prediction. and It is the weighting coefficient.
[0043] The blood oxygen saturation learning task is based on a deep learning spectral analysis model that directly learns the mapping relationship of blood oxygen saturation from multispectral images, rather than relying solely on the traditional optical density ratio and Beer-Lambert's law.
[0044] Specifically, convolutional neural networks (CNNs) or Transformer networks are used for end-to-end blood oxygen saturation prediction. The model input is the grayscale value of a multispectral image, and the output is the blood oxygen saturation value of arteries and veins.
[0045] Formula: Blood oxygen saturation The calculation formula is: ,in, It is the optical density ratio. and These are empirical parameters.
[0046] Furthermore, a dynamic calibration mechanism is introduced to adjust the blood oxygen saturation calculation model based on individual patient differences and real-time physiological status.
[0047] Based on physiological parameters such as heart rate and blood pressure, adjust empirical parameters in real time. and To ensure the calculated results more closely match actual blood oxygen saturation values, the calculated optical density ratio is filtered using methods such as median filtering or Gaussian filtering to remove noise interference. Simultaneously, a tissue scattering model is established to correct for optical density measurement deviations caused by scattering from gastrointestinal mucosal tissue.
[0048] By combining physiological parameters such as heart rate (HR), blood pressure (BP), and respiratory rate (RR), empirical parameters are dynamically adjusted. a and b Adjust using the following formula: , , Dynamically calibrated blood oxygen saturation for: , in, This is the current measured heart rate, which is the number of times the heart beats per minute. The baseline heart rate refers to the heart rate of the human body in a "resting, basal metabolic" state (such as the resting heart rate), serving as a reference benchmark for physiological states. This is the currently measured systolic blood pressure, the pressure inside the arteries during cardiac contraction. The baseline systolic blood pressure is the reference value for systolic blood pressure under basal body conditions. This is the currently measured diastolic blood pressure, the pressure within the arteries during cardiac diastole. The baseline diastolic blood pressure is the reference value for diastolic blood pressure under basal conditions in the human body. a and b These are empirical parameters before dynamic adjustment. and These are dynamically adjusted empirical parameters. This is the heart rate-related adjustment factor. It is the blood pressure-related adjustment coefficient. This dynamic adjustment mechanism can significantly improve the accuracy of blood oxygen saturation calculation, making it closer to the true value.
[0049] Example 2 This embodiment provides a computer-readable storage medium storing a computer program that, when executed by a processor, performs the following steps: Multi-scale texture and structural features were extracted from color images and OCT depth images of gastrointestinal endoscopy, and the extracted multi-scale texture and structural features were fused to obtain a cross-modal feature sequence. Based on the feature sequences with added position encoding, combined with the cross-attention layer, the spatial correspondence between different modal data in terms of vascular morphology and gastrointestinal wall layer structure is learned by aligning intermodal features through a shared query key-value matrix; Based on the aligned cross-modal vascular and gastrointestinal wall interlayer structural features, multi-scale global context feature sequences and local detail feature sequences are extracted, and multi-scale and local detail features are fused to obtain fused multi-modal features; The fused multimodal features are converted into a displacement field, and then the displacement field is upsampled to the original image resolution to generate a dense deformation field. Physiological signals are introduced, and the dense deformation field is adjusted based on the physiological signals to obtain the adjusted dense deformation field. Coarse and fine registration were performed based on the adjusted dense deformation field to obtain the registered color images and OCT depth images of the gastrointestinal endoscope.
[0050] Example 3 This embodiment provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it performs the following steps: Multi-scale texture and structural features were extracted from color images and OCT depth images of gastrointestinal endoscopy, and the extracted multi-scale texture and structural features were fused to obtain a cross-modal feature sequence. Based on the feature sequences with added position encoding, combined with the cross-attention layer, the spatial correspondence between different modal data in terms of vascular morphology and gastrointestinal wall layer structure is learned by aligning intermodal features through a shared query key-value matrix; Based on the aligned cross-modal vascular and gastrointestinal wall interlayer structural features, multi-scale global context feature sequences and local detail feature sequences are extracted, and multi-scale and local detail features are fused to obtain fused multi-modal features; The fused multimodal features are converted into a displacement field, and then the displacement field is upsampled to the original image resolution to generate a dense deformation field. Physiological signals are introduced, and the dense deformation field is adjusted based on the physiological signals to obtain the adjusted dense deformation field. Coarse and fine registration were performed based on the adjusted dense deformation field to obtain the registered color images and OCT depth images of the gastrointestinal endoscope.
[0051] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of hardware embodiments, software embodiments, or embodiments combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage and optical storage) containing computer-usable program code.
[0052] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1A device that provides the functions specified in one or more boxes.
[0053] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0054] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0055] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.
[0056] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A multi-modality medical image analysis system for gastrointestinal disorder identification, characterized in that, The method comprises the following steps: a modal feature encoding and position perception module is used to extract multi-scale texture features and structure features of the gastrointestinal endoscope color image and the gastrointestinal endoscope OCT depth image respectively, and the multi-scale texture features and the structure features are fused to obtain a cross-modal feature sequence; a dynamic feature alignment module is used to align the features based on the added position encoding, combine a cross-attention layer, align the modal features through a shared query key-value matrix, and learn the spatial correspondence of different modal data on the blood vessel morphology and the gastrointestinal wall interlayer structure; a feature extraction module is used to extract a multi-scale global context feature sequence and a local detail feature sequence based on the aligned cross-modal blood vessel and gastrointestinal wall interlayer structure features, and fuse the multi-scale and local detail features to obtain a fused multi-modal feature; a dynamic deformation field generation and calibration module is used to convert the fused multi-modal feature into a displacement field, upsample the displacement field to the original image resolution, generate a dense deformation field, introduce a physiological signal, and adjust the dense deformation field based on the physiological signal to obtain an adjusted dense deformation field; a multi-stage registration module is used to perform coarse registration and fine registration based on the adjusted dense deformation field to obtain the registered gastrointestinal endoscope color image and the gastrointestinal endoscope OCT depth image.
2. The multi-modality medical image analysis system for gastrointestinal disorder identification of claim 1, wherein, In the modal feature encoding and position perception module, the cross-modal blood vessel features are extracted based on the obtained gastrointestinal endoscope color image and the gastrointestinal endoscope OCT depth image, which comprises inputting the gastrointestinal endoscope color image and the gastrointestinal endoscope OCT depth image into corresponding CNN encoders respectively, extracting multi-scale texture features and structure features, and fusing the extracted multi-scale texture features and structure features to obtain a cross-modal feature sequence.
3. The multi-modality medical image analysis system for gastrointestinal disorder identification of claim 1, wherein, In the modal feature encoding and position perception module, the spatial correspondence of different modal data on the blood vessel morphology and the gastrointestinal wall interlayer structure is learned based on the encoded feature sequence, which comprises combining a cross-attention layer, taking the gastrointestinal endoscope image tokens as the query and the gastrointestinal endoscope OCT image tokens as the key value, calculating the attention weight, and learning the spatial correspondence of different modal data on the blood vessel morphology and the gastrointestinal wall interlayer structure based on the attention weight.
4. The multi-modality medical image analysis system for gastrointestinal disorder identification of claim 1, wherein, In the feature extraction module, a multi-head global attention mechanism is combined, each head captures global context information of different scales to obtain a multi-scale global context feature sequence, a local variability attention mechanism is combined based on the multi-scale global context feature sequence and the interlayer structure tokens of the gastrointestinal endoscope OCT depth image, and a local detail feature sequence is extracted.
5. The multi-modality medical image analysis system for gastrointestinal disorder identification of claim 4, wherein, In the feature extraction module, each head captures global context information of different scales, which specifically comprises: head 1 and head 2 capture large blood vessel trunks, head 3 and head 4 focus on gastrointestinal lesion edges, and head 5 and head 6 capture microvessel details.
6. The multi-modality medical image analysis system for gastrointestinal disorder identification of claim 1, wherein, In the feature extraction module, when the global context information and the local detail features are fused to obtain the fused multi-modal features, a pyramid feature fusion strategy is adopted, the global blood vessel distribution features are superimposed with the local micro blood vessel texture features at the bottom layer, the global gastrointestinal specific reference region relative position features are fused with the local gastrointestinal lesion edge gradient features at the middle layer, and the global spectral features are combined with the local blood vessel bifurcation detail features at the high layer.
7. The multi-modality medical image analysis system for gastrointestinal disorder identification of claim 1, wherein, In the dynamic deformation field generation and calibration module, the dynamic adjustment of the dense deformation field parameters is represented as: , wherein, is the adjusted dense deformation field parameter, is the unadjusted dense deformation field parameter, is the physiological signal influence weight coefficient, is the activation value of the encoded physiological signal by the gated recurrent unit, and S is the output state vector.
8. The multi-modality medical image analysis system for gastrointestinal disorder identification of claim 1, wherein, The system further comprises a gastrointestinal lesion positioning module. The gastrointestinal endoscope color image after registration is input into the backbone network of the YOLOv5 model to extract two-dimensional texture features of the gastrointestinal lesion, the gastrointestinal endoscope OCT depth image after registration is subjected to depth slice feature extraction to obtain three-dimensional structure features of the gastrointestinal lesion region, and the two-dimensional texture features of the gastrointestinal lesion and the three-dimensional structure features of the gastrointestinal lesion region are input into the detection head of the YOLOv5 after being fused to output gastrointestinal lesion candidate boxes and confidence. The gastrointestinal lesion center coordinates and range are determined according to the gastrointestinal lesion candidate boxes.
9. A computer readable storage medium having stored thereon a computer program, characterized in that, The program is executed by the processor to implement the following steps: Multi-scale texture features and structure features of the gastrointestinal endoscope color image and the gastrointestinal endoscope OCT depth image are extracted respectively, and the extracted multi-scale texture features and structure features are fused to obtain a cross-modal feature sequence; Based on the feature sequence after adding the position encoding, the cross attention layer is combined to align the features between the modalities by using the shared query key-value matrix, and the spatial correspondence of the different modal data on the blood vessel morphology and the gastrointestinal wall interlayer structure is learned; Based on the aligned cross-modal blood vessel and gastrointestinal wall interlayer structure features, multi-scale global context feature sequences and local detail feature sequences are extracted, and the multi-scale and local detail features are fused to obtain fused multi-modal features; The fused multi-modal features are converted into a displacement field, the displacement field is upsampled to the original image resolution, a dense deformation field is generated, a physiological signal is introduced, the dense deformation field is adjusted based on the physiological signal, and an adjusted dense deformation field is obtained. Based on the adjusted dense deformation field, coarse registration and fine registration are performed to obtain the gastrointestinal endoscope color image and the gastrointestinal endoscope OCT depth image after registration.
10. A computer device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor executes the program to implement the following steps: Multi-scale texture features and structure features of the gastrointestinal endoscope color image and the gastrointestinal endoscope OCT depth image are extracted respectively, and the extracted multi-scale texture features and structure features are fused to obtain a cross-modal feature sequence; Based on the feature sequence after adding the position encoding, the cross attention layer is combined to align the features between the modalities by using the shared query key-value matrix, and the spatial correspondence of the different modal data on the blood vessel morphology and the gastrointestinal wall interlayer structure is learned; Based on the aligned cross-modal blood vessel and gastrointestinal wall interlayer structure features, multi-scale global context feature sequences and local detail feature sequences are extracted, and the multi-scale and local detail features are fused to obtain fused multi-modal features; The fused multi-modal features are converted into a displacement field, and the displacement field is up-sampled to the resolution of the original image to generate a dense deformation field. A physiological signal is introduced, and the dense deformation field is adjusted based on the physiological signal to obtain an adjusted dense deformation field. Based on the adjusted dense deformation field, coarse registration and fine registration are performed to obtain registered gastrointestinal endoscope color images and gastrointestinal endoscope OCT depth images.
Citation Information
Patent Citations
Multi-focus joint segmentation method in retina OCT image based on hybrid network
CN118657800A
Spatial registration and image fusion method for deep learning based on SS-OCT and TOF data
CN120219452A
Vascular dynamics video registration quantification method and device based on vascular skeleton
CN120279070A
Building method of multi-modal three-dimensional medical image segmentation and registration model, and application thereof
US20250238932A1
Systems and methods for automated widefield optical coherence tomography angiography
WO2017218738A1
Cited By
Medical image feature extraction method and system based on morphological structure double-path interaction
CN122091115A