A Multimodal Endoscopic Feature Fusion System for Assessing Inflammatory Bowel Disease Activity
By using a multimodal endoscopic feature fusion system, combined with EUS microarray multimodal endoscopy and white light endoscopy, the consistency and accuracy issues of IBD activity assessment under endoscopy have been resolved, achieving more accurate assessment of inflammatory bowel disease activity and supporting precision diagnosis and treatment.
Patent Information
- Application Number
- CN202511211603.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-28
- Publication Date
- 2026-01-30
- Estimated Expiration
- 2045-08-28
AI Technical Summary
Existing endoscopic assessments of IBD activity suffer from poor consistency and difficulty in assessing the depth of inflammation affecting the intestinal wall and transmural changes around the intestine. AI assessments are mostly limited to single-modal optics, resulting in low accuracy and hindering the development of precision diagnosis and treatment.
A multimodal endoscopic feature fusion system was adopted, combining EUS microarray multimodal endoscope and white light endoscope. Through image preprocessing, white light surface feature extraction and ultrasound transmural feature extraction, the mucosal features were fused using a multimodal endoscopic IBD activity assessment model to improve the accuracy of assessment.
It improves the accuracy of IBD activity assessment, enhances the consistency between endoscopic activity assessment and clinical, laboratory, and histological activity, enables real-time assessment of treatment efficacy, and guides the optimization of treatment strategies.
Smart Images

Figure CN120747066B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of medical information processing technology, and more specifically, to an inflammatory bowel disease activity assessment system based on multimodal endoscopic feature fusion. Background Technology
[0002] Inflammatory bowel disease (IBD) has an unknown etiology and primarily affects young adults. Its characteristics include a chronic, recurrent, and highly disabling nature, seriously jeopardizing patients' health. With significant advancements in the development of novel drugs such as biologics and innovative treatment strategies, the need for precision medicine is increasing. Endoscopic remission of IBD has become a crucial basis for patient treatment planning. However, endoscopic assessment of IBD activity is limited by:
[0003] (1) There is poor consistency among endoscopists;
[0004] (2) Optical endoscopy, as the main assessment method, can only observe the surface of the intestinal mucosa and is difficult to assess the depth of the intestinal wall affected by inflammation and the transmural changes around the intestine.
[0005] Currently, there are still discrepancies between endoscopic activity scores and clinical and histological activity scores, which hinders the development of precision diagnosis and treatment of IBD.
[0006] In recent years, the development of artificial intelligence (AI) technology has made it possible for researchers to accurately and efficiently assess IBD activity. Di Ruscio et al. constructed a model for assessing and predicting the histological healing of UC under white light endoscopy using image classification; Bossuyt P et al. constructed a computer algorithm to automatically assess the red density of the mucosal surface in patients with moderate to severe UC, thereby predicting mucosal healing. However, existing AI research is mostly limited to single-modal optical assessment and is largely limited to static image assessment. Summary of the Invention
[0007] To overcome the shortcomings of the prior art, such as low accuracy of endoscopic activity assessment of IBD and reliance on single-modal optics in AI assessment, this invention provides an inflammatory bowel disease activity assessment system based on multimodal endoscopic feature fusion.
[0008] To solve the above-mentioned technical problems, the technical solution of the present invention is as follows:
[0009] A system for assessing the activity of inflammatory bowel disease based on multimodal endoscopic feature fusion, comprising:
[0010] A multimodal data receiving module is used to receive dynamic sequence data of a first EUS microarray multimodal endoscope; wherein, the dynamic sequence data of the first EUS microarray multimodal endoscope includes a first EUS endoscope image and a corresponding first white light endoscope image;
[0011] The image preprocessing module is used to perform noise reduction processing on the dynamic sequence data of the first EUS microarray multimodal endoscope to determine the withdrawal observation segment; wherein, the withdrawal observation segment includes a first white light effective observation image and a first ultrasound effective observation image;
[0012] The white light surface feature extraction module is used to carry a trained white light surface feature extraction model and use the white light surface feature extraction model to extract mucosal surface features from the first white light effective observation image to determine the first white light endoscopic mucosal features that reflect the state of the intestinal wall mucosal surface.
[0013] The ultrasound transmural feature extraction module is used to carry a trained ultrasound transmural feature extraction model and use the ultrasound transmural feature extraction model to perform structural recognition and measurement on the first ultrasound effective observation image to determine the first ultrasound mucosal features reflecting the transmural structural state of the intestine.
[0014] The multimodal endoscopic IBD activity assessment module is used to carry a trained multimodal dynamic endoscopic IBD activity assessment model, and to determine the IBD activity based on the mucosal characteristics under the first white light endoscope and the mucosal characteristics under the first ultrasound using the multimodal dynamic endoscopic IBD activity assessment model.
[0015] Compared with the prior art, the beneficial effects of the technical solution of the present invention are:
[0016] This application discloses an inflammatory bowel disease (IBD) activity assessment system based on multimodal endoscopic feature fusion. The system uses an image preprocessing module to denoise the dynamic sequence data of the first EUS microarray multimodal endoscopy to overcome noise interference from noisy images in the dynamic sequence data. Mucosal features under white light and ultrasound are extracted using a white light surface feature extraction module and an ultrasound transmural feature extraction module, respectively. Finally, the system fuses multimodal endoscopic features using a multimodal dynamic endoscopic IBD activity assessment model in the multimodal endoscopic IBD activity assessment module to determine IBD activity, improving the accuracy of IBD activity assessment and enhancing the consistency between endoscopic activity assessment and clinical, laboratory, and histological activity. This overcomes the limitations of previous disease assessments based on single-modality, static images. Attached Figure Description
[0017] Figure 1This is a schematic diagram of the inflammatory bowel disease activity assessment system described in the embodiments of this application.
[0018] Figure 2 This is a schematic diagram of the training process of the multimodal dynamic endoscopic IBD activity assessment model described in the embodiments of this application.
[0019] Figure 3 This is a schematic diagram of the training process of the ultrasonic transmural feature extraction model in the embodiments of this application.
[0020] Figure 4 This is a schematic diagram of the training process of the lesion recognition model described in the embodiments of this application.
[0021] Figure 5 This is a schematic diagram of the training process of the white light surface feature extraction model described in the embodiments of this application.
[0022] Figure 6 This is a schematic diagram of the hardware entity of the electronic device described in the embodiments of this application. Detailed Implementation
[0023] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms are interchangeable where appropriate; this is merely a way of distinguishing objects with the same attributes in the description of embodiments of this application. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, so that a process, method, system, product, or apparatus that comprises a series of units is not necessarily limited to those units, but may include other units not explicitly listed or inherent to those processes, methods, products, or apparatuses. The term "determine" broadly covers a wide variety of actions, including acquiring, calculating, processing, deriving, investigating, searching (e.g., searching in a table, database, or other data structure), probing, and similar actions; it may also include receiving (e.g., receiving information), accessing (e.g., accessing data in memory), and similar actions; it may also include generating, creating, establishing, and similar actions; and parsing, selecting, choosing, and similar actions, etc. Definitions of other terms will be given in the following description.
[0024] It should be noted that when one element is considered to be "connected" to another element, it can be directly connected to the other element or connected to the other element through an intermediary element. Furthermore, in the following embodiments, "connection" should be understood as "electrical connection," "communication connection," etc., if there is transmission of electrical signals or data between the connected objects.
[0025] The accompanying drawings are for illustrative purposes only and should not be construed as limiting the scope of this patent.
[0026] To better illustrate this embodiment, some parts in the accompanying drawings may be omitted, enlarged, or reduced, and do not represent the actual product dimensions;
[0027] It will be understood by those skilled in the art that certain well-known structures and their descriptions may be omitted in the accompanying drawings.
[0028] The technical solution of the present invention will be further described below with reference to the accompanying drawings and embodiments.
[0029] This embodiment provides an inflammatory bowel disease activity assessment system based on multimodal endoscopic feature fusion, as shown in Figure 1, including:
[0030] The multimodal data receiving module 110 is used to receive the first EUS microarray multimodal endoscope dynamic sequence data; wherein, the first EUS microarray multimodal endoscope dynamic sequence data includes a first EUS endoscope image and a corresponding first white light endoscope image.
[0031] The image preprocessing module 120 is used to perform noise reduction processing on the dynamic sequence data of the first EUS microarray multimodal endoscope to determine the withdrawal observation segment; wherein, the withdrawal observation segment includes a first white light effective observation image and a first ultrasound effective observation image;
[0032] The white light surface feature extraction module 130 is used to carry a trained white light surface feature extraction model and use the white light surface feature extraction model to extract mucosal surface features from the first white light effective observation image to determine the mucosal features under the first white light endoscope.
[0033] The ultrasound transmural feature extraction module 140 is used to carry a trained ultrasound transmural feature extraction model and use the ultrasound transmural feature extraction model to perform structural recognition and measurement on the first ultrasound effective observation image to determine the first ultrasound mucosal features reflecting the transmural structural state of the intestine.
[0034] The multimodal endoscopic IBD activity assessment module 150 is used to carry a trained multimodal dynamic endoscopic IBD activity assessment model, and to determine the IBD activity based on the mucosal characteristics under the first white light endoscope and the mucosal characteristics under the first ultrasound using the multimodal dynamic endoscopic IBD activity assessment model.
[0035] Those skilled in the art should understand that EUS (Endoscopic Ultrasonography) involves inserting a miniature ultrasound scanning probe into the human body through the biopsy channel of an electronic endoscope. While the electronic endoscope observes the mucosal surfaces of internal organs, the miniature ultrasound scanning imaging system can acquire tomographic images of the walls of internal organs. Unlike transabdominal ultrasound, EUS is placed directly against the intestinal wall mucosa, reducing interference from air and other impurities in the abdominal cavity. This allows for better assessment of the depth and extent of intestinal wall inflammation and accurate evaluation of periintestinal fat and abscess formation.
[0036] It should be noted that this embodiment constructs a system for real-time assessment of IBD patient activity based on multimodal data from ultrasound and optical endoscopy. The system extracts mucosal features under white light and ultrasound (i.e., the first white light endoscopic mucosal feature reflecting the intestinal wall mucosal surface state and the first ultrasound mucosal feature reflecting the intestinal transmural structural state) using the white light surface feature extraction model in the white light surface feature extraction module and the ultrasound transmural feature extraction model in the ultrasound transmural feature extraction module, respectively. The system then determines the multimodal IBD activity using the multimodal dynamic endoscopic IBD activity assessment model in the multimodal endoscopic IBD activity assessment module. This overcomes the limitations of existing technologies that rely on single-modality, static images for disease assessment, resulting in high accuracy. In clinical practice, it improves the consistency between endoscopic activity assessment and clinical, laboratory, and histological activity, enabling real-time evaluation of the patient's response to the current treatment plan, guiding treatment strategy optimization, and providing a new auxiliary tool for precision diagnosis and treatment of IBD.
[0037] It should also be noted that the first EUS endoscopic image and the first white light endoscopic image can be dynamic sequence images, each sequence containing a fixed number of time steps (frames); or they can be non-dynamic sequence images, using a fixed number of images to be equivalent to dynamic sequence images.
[0038] In some preferred embodiments, determining the activity of IBD based on the first white light endoscopic mucosal characteristics and the first ultrasound-guided mucosal characteristics includes:
[0039] The first white light endoscopic mucosal features and the first white light effective observation image, the first ultrasound-guided mucosal features and the first ultrasound-guided effective observation image are respectively encoded into three-dimensional vectors to establish a three-dimensional vector sequence; wherein, the three-dimensional vector includes the modality type and time step determined based on the first white light effective observation image or the first ultrasound-guided effective observation image, and the mucosal feature vector determined based on the first white light endoscopic mucosal features or the first ultrasound-guided mucosal features;
[0040] The activity of the IBD is determined by using the three-dimensional vector sequence as input to the multimodal dynamic endoscopic IBD activity assessment model.
[0041] In some specific implementations, the three-dimensional vector is represented as [modality type, time step, mucosal feature vector]. The modality type can be represented using one-hot encoding (white light or EUS); the time step is the corresponding image frames arranged in a time sequence; and the mucosal feature vector is the first white light endoscopic mucosal feature determined by the modality type through numerical mapping, or the first ultrasound-guided mucosal feature extracted through a pre-trained convolutional neural network (such as ResNet).
[0042] In some optional embodiments, the training process of the multimodal dynamic endoscopic IBD activity assessment model, as shown in Figure 2, includes:
[0043] Acquire dynamic sequence data of the third EUS microarray multimodal endoscope; wherein, the dynamic sequence data of the third EUS microarray multimodal endoscope includes the third white light effective observation image and the corresponding mucosal features under the third white light endoscope, as well as the third EUS effective observation image and the corresponding mucosal features under the third ultrasound.
[0044] Using pathological examination results as the gold standard, the IBD pathological activity of the third EUS microarray multimodal endoscopy dynamic sequence data was labeled as the true label, and the true label was smoothed.
[0045] A multimodal endoscopic brainstem dysplasia (IBD) activity training dataset is established based on the dynamic sequence data of the third EUS microarray multimodal endoscope and the smoothed real labels; wherein, the dynamic sequence data of the third EUS microarray multimodal endoscope is encoded in three-dimensional vector form;
[0046] An initial multimodal dynamic endoscopic IBD activity assessment model is established, and the model is iteratively trained on the multimodal endoscopic IBD activity training dataset. Cross-entropy is used as the loss function, and the network parameters are updated through backpropagation based on the cross-entropy loss until the training ends.
[0047] It should be noted that smoothing the true labels helps reduce overfitting of the multimodal dynamic endoscopic IBD activity assessment model to the training data and enhances the model's generalization ability. As a non-limiting example, label smoothing can be achieved by assigning a higher probability (e.g., 0.9) to the true labels and distributing the remaining probabilities equally among the other categories (e.g., 0.1).
[0048] Exemplary, the third white light endoscopic mucosal features and the third ultrasound-guided mucosal features can be extracted from the third white light effective observation image or the third EUS effective observation image by the trained white light surface feature extraction model and ultrasound transmural feature extraction model, respectively.
[0049] In some specific implementations, during the training of the multimodal dynamic endoscopic IBD activity assessment model, Adam or AdamW is used as the optimizer, and the learning rate and other hyperparameters are set. Cross-entropy loss is calculated by comparing the predicted results with the true labels; gradients are calculated through backpropagation, and the optimizer is used to update the model parameters until the model converges or reaches the preset number of training rounds.
[0050] Furthermore, model performance is validated using five-fold cross-validation or other cross-validation methods, and metrics such as accuracy, precision, recall, and F1-score are used to evaluate model performance. Based on the feedback from cross-validation results, hyperparameters of the model, such as learning rate, batch size, number of encoder layers, and number of attention heads, are optimized to achieve optimal model performance.
[0051] In some preferred embodiments, the noise reduction processing based on the dynamic sequence data of the first EUS microarray multimodal endoscope includes:
[0052] Based on the Farneback dense optical flow algorithm, the forward optical flow vector and the reverse optical flow vector between adjacent frames of the first white light endoscope image are calculated;
[0053] Based on the reverse optical flow vector, reverse verification is performed to eliminate interfering optical flow, determine the optical flow residual vector, and then calculate the mirror withdrawal speed.
[0054] The optical flow residual vector is converted into HSV channel values to construct the corresponding HSV image;
[0055] The recurrent neural network is used to identify the continuous HSV images, determine the state of the first white light endoscope image (such as colonoscopy withdrawal observation, insertion, sliding, aspiration, flushing, biopsy, etc.), and remove the segments that are not in the observation state from the first white light endoscope image and the corresponding first EUS endoscope image, and determine the first white light effective observation image and the first ultrasound effective observation image.
[0056] It should be noted that the above preferred embodiment calculates the dense optical flow of adjacent frames, eliminates interfering optical flow through reverse verification, converts the verified optical flow into an HSV image, and sends the sequence into a recurrent neural network to obtain the state of the image sequence, thereby avoiding interference from non-observed segments and overcoming noise interference from multimodal data.
[0057] Furthermore, the image preprocessing module is also used to perform image segmentation and scaling on the first white light endoscope image before noise reduction processing, including:
[0058] The first white light endoscope image is decoded into a frame image. The frame image is then segmented using a U-net-based image semantic segmentation network to remove interference information from the frame image while retaining the mucosal portion of the frame image, resulting in a segmented frame image.
[0059] The segmented frame image is scaled using region interpolation, and the scaled segmented frame image is then re-encoded into the first white light endoscope image.
[0060] In some specific implementation processes, due to differences in endoscope models, brands, etc., there may be significant differences between different original video images. By decoding, segmenting, and recoding the endoscope images, the negative impact of image format differences on model recognition can be removed, and interference information other than image information can be eliminated / reduced, as follows:
[0061] Video decoding: The first white light endoscope image is decoded into an image (i.e., a frame image) at a rate of 25 frames per second;
[0062] Background segmentation: The image region in the decoded image of the first white light endoscope image is segmented by the U-net network, which is compatible with video images of different sizes and models of white light endoscopes, to remove interference information other than image information.
[0063] Image scaling: The size of the cropped endoscopic image is about 1200×1000. The endoscopic image is reduced to 360×360 (or other sizes) by region interpolation to ensure the real-time performance of subsequent recognition.
[0064] Re-encoding: The scaled, segmented frame image is re-encoded into the first white light endoscope image.
[0065] In some preferred embodiments, the training process of the ultrasonic transmural feature extraction model, as shown in Figure 3, includes:
[0066] Acquire a second endoscopic ultrasound image for training;
[0067] Using expert experience, the second endoscopic ultrasound image is used to annotate the intestinal wall mucosal structure (mucosal layer, submucosa, muscularis propria, etc.), and then the center line of each intestinal wall mucosal structure is extracted. The layer thickness of each intestinal wall mucosal structure is measured perpendicular to the center line to quantify the characteristic parameters of the mucosal structure.
[0068] Based on the annotation results (i.e. the above-mentioned annotation results on structure) and the quantification results, an IBD ultrasound structure training dataset was established for learning to delineate the layers of intestinal wall mucosal structure and measure structure.
[0069] The initial ultrasound transmural feature extraction model is constructed based on D-LinkNet or ResNet, and then the ultrasound transmural feature extraction model is transferred to the IBD ultrasound structure training dataset for transfer learning.
[0070] As a non-limiting example, the characteristic parameters of the mucosal structure include, but are not limited to, the degree of uniformity of the mucosal structure layers, the degree of thickening, and the mucosal thickness.
[0071] It should be understood that the trained ultrasound transmural feature extraction model can automatically delineate the contours of mucosal layers such as the intestinal wall mucosa, submucosa, and muscularis propria in ultrasound images (i.e., EUS endoscopic images) and quantify feature parameters such as layer thickness.
[0072] In some specific implementation processes, the annotation results of the intestinal wall mucosal structure are reflected by the coordinate values of the boundary lines of each layer of the intestinal wall.
[0073] In some specific implementation processes, Halcon software was used to extract the centerline of each intestinal wall mucosal structure and measure the layer thickness of each intestinal wall mucosal structure, and then visualized the results.
[0074] In some specific implementation processes, the second endoscopic ultrasound images are obtained by retrospectively collecting data from patients with ulcerative colitis and Crohn's disease who have previously visited the hospital.
[0075] Furthermore, the second endoscopic ultrasound image is also labeled with periintestinal structures, so that the ultrasound transmural feature extraction model can learn the three-dimensional features of the periintestinal tract.
[0076] In some alternative embodiments, the ultrasound transmural feature extraction module is also used to output an ultrasound layer diagram of the intestinal wall and / or intestinal wall structure.
[0077] In some specific implementations, a Transformer-based decoder is incorporated into the ultrasound transmural feature extraction module to reconstruct the feature vectors output by the ultrasound transmural feature extraction model into an ultrasound layer diagram. The ultrasound layer diagram is marked with color blocks for each intestinal wall mucosal structure and with numerical values for the corresponding layer thickness.
[0078] In some optional embodiments, the ultrasound transmural feature extraction module is further equipped with a lesion recognition model for determining lesion features, and generates a schematic diagram of the lesion marked with the lesion based on the first effective ultrasound observation image and the lesion features; the lesion features are also used to perform feature fusion with the first ultrasound-guided mucosal features and then input as new first ultrasound-guided mucosal features into the multimodal dynamic endoscopic IBD activity assessment model; wherein, the training process of the lesion recognition model, as shown in Figure 4, includes:
[0079] Using expert experience and combined with the quantitative results of characteristic parameters of various intestinal wall mucosal structures, the lesion formation sites (lesions affecting the intestinal wall) in the second endoscopic ultrasound images in the IBD ultrasound structure training dataset are determined, and the second endoscopic ultrasound images are labeled with label values to mark the lesions (such as the location of stenotic intestinal segments and abscess formation) and lesion types (such as different lesion types such as abscess, stenosis, inflammation, fibrosis, etc.), forming an IBD ultrasound lesion training dataset for learning to identify lesion sites and lesion types;
[0080] The initial lesion identification model is constructed based on UNet++, and the lesion identification model is then transferred to the IBD ultrasound lesion training dataset for transfer learning.
[0081] It should be noted that UNet++ is a neural network improved upon the original UNet architecture, primarily optimizing skip connections. Compared to natural images, medical image processing (such as lesion segmentation or anomaly recognition) requires higher accuracy. UNet++ achieves this through finer-grained feature fusion and a dense skip connection mechanism, replacing skip connections with dense convolution blocks. This allows features of different resolutions to be fused and passed multiple times through a layer-by-layer refinement path, reducing the semantic gap between deep and shallow features. Furthermore, it enhances the network's segmentation capabilities for complex target regions, making it particularly suitable for extracting complex morphological features from medical images.
[0082] It should be understood that variants of UNet++ can also be applied to the above embodiments, and this disclosure does not limit this.
[0083] As a non-limiting example, those skilled in the art may also use R-CNN, ViT, or other model frameworks to construct the lesion identification model.
[0084] In some specific implementation processes, the lesion identification model is trained based on the TensorFlow deep learning framework, and the accuracy of the model results is evaluated using the IoU (Intersection over Union) value.
[0085] In some specific implementation processes, the lesion features are reconstructed into a schematic diagram of the lesion by the decoder mounted in the ultrasound transmural feature extraction module, and the corresponding lesions are marked with color blocks.
[0086] In some specific implementation processes, the lesion features are fused with the original first ultrasound-guided mucosal features through a multi-head attention mechanism to obtain new first ultrasound-guided mucosal features.
[0087] In some other specific implementations, a new first ultrasound-guided mucosal feature is obtained by splicing the lesion features with the original first ultrasound-guided mucosal features.
[0088] Based on the original first ultrasound mucosal features, first white light endoscopy mucosal features, and lesion features, the accuracy of the activity output by the multimodal endoscopic IBD activity assessment module will be further improved, bringing it closer to pathology-based IBD activity.
[0089] In some specific implementations, the second endoscopic ultrasound image is augmented to expand the number of samples in the training set (the IBD ultrasound structure training dataset and the IBD ultrasound lesion training dataset).
[0090] In some optional embodiments, the training process of the white light surface feature extraction model, as shown in Figure 5, includes:
[0091] Acquire a second white light endoscopic image that is cross-modal aligned with the second endoscopic ultrasound image; wherein the second white light endoscopic image is associated with vascular status information, bleeding status information and / or mucosal integrity information;
[0092] Using expert experience, combined with the vascular status information, the bleeding status information and / or the mucosal integrity information, the endoscopic mucosal surface status reflected by the second white light endoscope image is graded and evaluated to determine the surface feature evaluation result corresponding to the second white light endoscope image;
[0093] IBD white light training dataset is constructed based on the second white light endoscope image and the corresponding surface feature evaluation results;
[0094] The initial white light surface feature extraction model is constructed based on any one of Vision Transformer, EfficientNet, ResNet or Inception. The white light surface feature extraction model is then iteratively trained on the IBD white light training dataset using the K-fold cross-validation method. Training stops when the validation loss value does not decrease for a preset number of rounds.
[0095] As a non-limiting example, the surface feature assessment result may be an activity assessment result based on imaging, or it may be an assessment result of the mucosal surface coating condition, boundary clarity assessment result, or color normality assessment result, etc.
[0096] In some specific implementation processes, two or more expert physicians, based on the endoscopic features (vessels, bleeding, ulcer depth, ulcer size, stenosis) proposed by the Mayo Endoscopic Score, the Ulcerative Colitis Endoscopic Activity Index Score (UCEIS), and / or the Simple Endoscopic Score for Crohn's Disease (SES-CD) scoring criteria, uniformly assess the endoscopic mucosal surface features of the second white light endoscopic images at different times for different IBD patients. Combining information on vascular status, bleeding status, and mucosal integrity, a total score of the overall white light endoscopic mucosal features for the IBD patient (i.e., using image-based activity as the surface feature assessment result) is determined as the label for training the white light surface feature extraction model. Specifically, if the original expert physicians cannot reach a consensus, a new expert physician is introduced to make a comprehensive judgment.
[0097] In addition, to improve the accuracy of the white light surface feature extraction model in assessing mucosal features under white light endoscopy, the model performance can be verified by combining the clinical manifestations (such as abdominal pain, frequency of diarrhea, hematochezia, etc.), laboratory indicators (such as fecal calcitonin, C-reactive protein, blood routine analysis, etc.) and pathological histological activity scores of specific patients during the endoscopic examination.
[0098] In some specific implementation processes, the white light surface feature extraction model is iteratively trained based on the five-fold cross-validation method. Training is stopped when the validation loss value does not decrease for five consecutive rounds.
[0099] Exemplarily, during the training of the white light surface feature extraction model, a classification head is connected to the output of the model to map the category (denoted as the predicted value) to which the feature vector output by the model belongs. During training, the parameters of the classification head are frozen, and the white light surface feature extraction model is updated based on the loss function calculated from the predicted value and the label value (i.e., the surface feature evaluation result) until the end of training when the classification head is discarded.
[0100] As a non-limiting example, the classification head sequentially includes a global average pooling layer, an FC layer, a BN (BatchNorm) layer, a ReLU layer, and another FC layer.
[0101] As another non-limiting example, the classification head includes an FC layer and a Softmax layer.
[0102] It should be noted that the white light surface feature extraction module will determine the category of the first white light valid observation image based on the preset confidence cut-off value.
[0103] In some preferred embodiments, the white light surface feature extraction module is further used to weight the retraction speed of the first white light effective observation image (i.e., the image after removing non-observation segments) with the corresponding mucosal features under the first white light endoscope, determine the distribution of different feature mucosa of the effective observation intestinal segment and the corresponding mucosa proportion, and generate a colonoscopy schematic diagram reflecting the distribution of mucosal features of different effective observation intestinal segments.
[0104] In the above preferred embodiment, the observation status and observation duration reflected by the retraction speed can be used to correct the distribution results of mucosal features under colonoscopy.
[0105] In some specific implementation processes, the classification results of the mucosal characteristics of the effective intestinal segments will be marked with preset colors on the colonoscopy diagram.
[0106] It should be noted that, based on the visualized colonoscopy diagram, users can retrieve colonoscopy diagrams related to previous white light endoscopy images to generate follow-up comparison results of disease characteristics under colonoscopy, so as to assess the temporal changes in mucosal characteristics of specific patients' colonoscopy manifestations.
[0107] Furthermore, before constructing the IBD white light training dataset, data augmentation is performed on the second white light endoscope image, including but not limited to random erasure, rotation, and flipping, to expand the sample.
[0108] It should be noted that the second white light endoscopic image is a fragment of the image during the observation phase after the endoscope is withdrawn.
[0109] In some preferred embodiments, the multimodal dynamic endoscopic IBD activity assessment model includes multiple Transformer encoders stacked sequentially and an output layer; wherein,
[0110] Each Transformer encoder includes a multi-head self-attention unit, a first residual connection and layer normalization unit, an FNN (Feed-Forward Neural Network) unit, and a second residual connection and layer normalization unit. The multi-head self-attention unit is used to compute attention weights in parallel based on the received input and then concatenate them to perform a linear transformation into the corresponding weighted feature vector. The first residual connection and layer normalization unit is used to perform residual connection and layer normalization on the weighted feature vectors received by the multi-head self-attention unit and their output, and then use this as the input to the FNN unit. The FNN unit includes two fully connected layers for feature mapping, employing ReLU activation and linear activation, respectively. The second residual connection and layer normalization unit is used to perform residual connection and layer normalization on the output of the FNN unit and its input (i.e., the normalized multi-head self-attention output), and then use this as the input to the next Transformer encoder or the output layer.
[0111] The output layer includes two fully connected layers (Dense) and a Softmax layer. The first fully connected layer includes a linear transformation layer and a ReLU activation function, and the second fully connected layer is a linear transformation layer. The Softmax layer transforms the output of the fully connected layer into a category probability distribution that reflects the activity of the multimodal endoscopic IBD.
[0112] In some specific implementations, the multimodal dynamic endoscopic IBD activity assessment model is equipped with 6 Transformer encoders. The output of each encoder layer is used as the input of the next encoder layer. The above steps are repeated until all Transformer encoders have completed their calculations.
[0113] It should be understood that the multi-head self-attention mechanism can achieve linear transformation to generate query (Q), key (K), and value (V) vectors. Through... Calculate attention weights, use multiple heads (e.g., 8 heads) to compute attention in parallel, and finally concatenate and linearly transform to output the weighted feature vector.
[0114] In the output layer, the fully connected layer has a two-layer structure. The first layer is a linear transformation + ReLU activation, and the second layer is a linear transformation (without activation function). This transforms the final output vector of the Transformer encoder through the two fully connected layers, and then the output of the fully connected layer is transformed into the corresponding class probability distribution through the Softmax layer.
[0115] In some preferred embodiments, the system further includes a visualization module for visually displaying the colonoscopy diagram under white light, the ultrasound layer diagram and lesion diagram under ultrasound, and the activity level (e.g., in image or report form).
[0116] In some examples, a computer program is provided, including computer-readable code, which, when executed in a computer device, allows a processor in the computer device to perform actions for implementing some or all of the modules / methods in the system.
[0117] This embodiment provides a computer-readable storage medium storing at least one instruction, at least one program, code set, or instruction set. The at least one instruction, at least one program, code set, or instruction set is loaded and executed by a processor, causing the processor to execute some or all of the modules / methods in the system provided in this embodiment.
[0118] It is understood that the storage medium can be transient or non-transient. Exemplarily, the storage medium includes, but is not limited to, various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0119] By way of example, the processor may be a central processing unit (CPU), a microprocessor unit (MPU), a digital signal processor (DSP), or a field programmable gate array (FPGA), etc.
[0120] By way of example, the read-only memory includes, but is not limited to, MASK ROM, PROM, EPROM, EEPROM, Flash, etc.
[0121] By way of example, the random access memory includes, but is not limited to, DRAM, SRAM, SDRAM, DDR SDRAM, etc.
[0122] In some examples, a computer program product is provided, which can be implemented by hardware, software, or a combination thereof. As a non-limiting example, the computer program product can be embodied in the storage medium, or it can be embodied in a software product, such as an SDK (Software Development Kit).
[0123] As a non-limiting example, a computer program product is provided, comprising a computer program or computer-executable instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer program or computer-executable instructions from the computer-readable storage medium, and executes the computer-executable instructions, causing the electronic device to perform some or all of the modules / methods described in the embodiments of this application.
[0124] This embodiment also proposes an electronic device, including a memory and a processor. The memory stores at least one instruction, at least one program, code set, or instruction set. When the processor executes the at least one instruction, at least one program, code set, or instruction set, it implements some or all of the modules / methods in the system described in the foregoing embodiments.
[0125] In some examples, a hardware entity of the electronic device is provided, referring to Figure 6, including: a processor, a memory, and a communication interface; wherein, the processor typically controls the overall operation of the electronic device; the communication interface is used to enable the electronic device to communicate with other terminals or servers via a network; the memory is configured to store instructions and applications executable by the processor, and can also cache data to be processed or already processed (including but not limited to image data, audio data, voice communication data, and video communication data) to be processed by the processor and various modules in the electronic device, and can be implemented by flash memory (FLASH), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or random access memory (RAM).
[0126] A processor may include one or more processing elements. Therefore, a processor may include one or more integrated circuits (ICs) configured to perform the functions of the processor. Furthermore, each integrated circuit may include circuitry (e.g., a first circuit, a second circuit, and other circuitry, etc.) configured to perform the functions of the processor.
[0127] Furthermore, data can be transferred between the processor, communication interface, and memory via a bus, which can include any number of interconnected buses and bridges, connecting various circuits of one or more processors and memories together.
[0128] The same or similar labels correspond to the same or similar parts;
[0129] The terms used to describe positional relationships in the accompanying drawings are for illustrative purposes only and should not be construed as limiting this application.
[0130] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other.
[0131] In different specific implementations, the methods or systems described in this application can be implemented in software, hardware, or a combination thereof. Furthermore, the order of the method steps can be changed, and various elements can be added, reordered, combined, omitted, or modified.
[0132] Obviously, the above embodiments of this application are merely examples for clearly illustrating this application, and are not intended to limit the implementation of this application, nor are they intended to limit this application. For those skilled in the art, other variations or modifications can be made based on the above description. The separate structural / functional modules or units can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part. The structure and function of the separate components can be implemented as a combined structure or component. It is neither necessary nor possible to exhaustively describe all embodiments here. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the scope of protection of the claims of this application.
Claims
1. A system for assessing inflammatory bowel disease activity based on multi-modal endoscopic feature fusion, characterized in that, The method comprises the following steps: a multi-modal data receiving module for receiving first EUS microarray multi-modal endoscopic dynamic sequence data; wherein the first EUS microarray multi-modal endoscopic dynamic sequence data comprises first EUS endoscopic images and corresponding first white light endoscopic images; an image preprocessing module for performing noise reduction processing according to the first EUS microarray multi-modal endoscopic dynamic sequence data, and determining a retreat observation segment; wherein the retreat observation segment comprises first white light effective observation images and first ultrasonic effective observation images; a white light surface feature extraction module for loading a trained white light surface feature extraction model, and using the white light surface feature extraction model to extract mucosal surface features from the first white light effective observation images, and determining first white light endoscopic mucosal features reflecting the state of the intestinal wall mucosal surface; an ultrasonic transmural feature extraction module for loading a trained ultrasonic transmural feature extraction model, and using the ultrasonic transmural feature extraction model to perform structure recognition and measurement on the first ultrasonic effective observation images, and determining first ultrasonic mucosal features reflecting the state of the intestinal transmural structure; a multi-modal endoscopic IBD activity assessment module for loading a trained multi-modal dynamic endoscopic IBD activity assessment model, and using the multi-modal dynamic endoscopic IBD activity assessment model to determine the activity of IBD according to the first white light endoscopic mucosal features and the first ultrasonic mucosal features; the multi-modal dynamic endoscopic IBD activity assessment model comprises a plurality of Transformer encoders stacked in sequence and an output layer; wherein each Transformer encoder comprises a multi-head self-attention mechanism unit, a first residual connection and layer normalization unit, an FNN unit, and a second residual connection and layer normalization unit; the multi-head self-attention mechanism unit is used to calculate attention weights in parallel based on the received input, and linearly transform the concatenated weighted feature vectors into corresponding weighted feature vectors; the first residual connection and layer normalization unit is used to perform residual connection and layer normalization on the input and output weighted feature vectors received by the multi-head self-attention mechanism unit, and use the result as the input of the FNN unit; the FNN unit comprises two fully connected layers for feature mapping, which respectively use ReLU activation and linear activation; the second residual connection and layer normalization unit is used to perform residual connection and layer normalization on the output of the FNN unit and the input, and use the result as the input of the next Transformer encoder or the output layer; and the output layer comprises two fully connected layers and a Softmax layer, the first fully connected layer comprises a linear transformation layer and a ReLU activation function, and the second fully connected layer is a linear transformation layer; the Softmax layer converts the output of the fully connected layer into a category probability distribution reflecting the multi-modal endoscopic IBD activity.
2. The system for evaluating the activity of inflammatory bowel disease based on multi-modal endoscopic feature fusion according to claim 1, characterized in that, determining the activity of IBD according to the first white light endoscopic mucosal features and the first ultrasonic mucosal features comprises: The first white light endoscopic mucosal feature and the first white light effective observation image, the first ultrasonic mucosal feature and the first ultrasonic effective observation image are respectively coded as three-dimensional vectors to establish a three-dimensional vector sequence; wherein the three-dimensional vector includes a modal type determined based on the first white light effective observation image or the first ultrasonic effective observation image, a time step, and a mucosal feature vector determined based on the first white light endoscopic mucosal feature or the first ultrasonic mucosal feature; The three-dimensional vector sequence is used as an input of the multi-modal dynamic endoscopic IBD activity evaluation model to determine the activity of the IBD.
3. The system for evaluating the activity of inflammatory bowel disease based on multi-modal endoscopic feature fusion according to claim 2, characterized in that, The training process of the multi-modal dynamic endoscopic IBD activity evaluation model includes: Obtaining third EUS microarray multi-modal endoscopic dynamic sequence data; wherein the third EUS microarray multi-modal endoscopic dynamic sequence data includes third white light effective observation images and corresponding third white light endoscopic mucosal features, and third EUS effective observation images and corresponding third ultrasonic mucosal features; Taking the pathological examination result as a standard, the third EUS microarray multi-modal endoscopic dynamic sequence data is labeled with IBD pathological activity as a true label, and the true label is smoothed; Based on the third EUS microarray multi-modal endoscopic dynamic sequence data and the smoothed true label, a multi-modal endoscopic IBD activity training data set is established; wherein the third EUS microarray multi-modal endoscopic dynamic sequence data is encoded in the form of a three-dimensional vector; An initial multi-modal dynamic endoscopic IBD activity evaluation model is established, and the multi-modal dynamic endoscopic IBD activity evaluation model is iteratively trained on the multi-modal endoscopic IBD activity training data set, cross-entropy is used as a loss function, and network parameters are updated through back propagation according to the cross-entropy loss until the training is completed.
4. The system for evaluating the activity of inflammatory bowel disease based on multi-modal endoscopic feature fusion according to any one of claims 1-3, characterized in that, The training process of the ultrasonic transmural feature extraction model includes: Obtaining a second ultrasonic endoscopic image for training; Using expert experience, the second ultrasonic endoscopic image is labeled with data related to the intestinal wall mucosal structure, and the center line of each intestinal wall mucosal structure is extracted, and the hierarchical thickness of each intestinal wall mucosal structure is measured perpendicular to the center line to quantify the feature parameters of the mucosal structure; Based on the labeling result and the quantization result, an IBD ultrasonic structure training data set for learning to outline the hierarchical structure of the intestinal wall mucosa and the feature quantization is established; Based on D-LinkNet or ResNet, an initial ultrasonic transmural feature extraction model is constructed, and the ultrasonic transmural feature extraction model is migrated on the IBD ultrasonic structure training data set.
5. The system for evaluating the activity of inflammatory bowel disease based on multi-modal endoscopic feature fusion according to claim 4, characterized in that, The ultrasonic transmural feature extraction module also carries a lesion recognition model for determining lesion features, and generates a lesion schematic diagram marked with lesions based on the first ultrasonic effective observation image and the lesion features; the lesion features are also used for feature fusion with the first ultrasonic mucosal features as new first ultrasonic mucosal features input into the multi-modal dynamic endoscopic IBD activity evaluation model; The training process of the lesion recognition model includes: The lesion formation site in the second endoscopic ultrasound image in the IBD ultrasound structure training data set is determined by expert experience, and the second endoscopic ultrasound image is labeled with a label value about the lesion and the lesion type to form an IBD ultrasound lesion training data set for learning to identify the lesion site and the lesion type; An initial lesion identification model is constructed, and the lesion identification model is trained on the IBD ultrasound lesion training data set.
6. The system for evaluating the activity of inflammatory bowel disease based on multi-modal endoscopic feature fusion according to claim 4, characterized in that, The training process of the white light surface feature extraction model includes: Obtain a second white light endoscopic image that is cross-modality aligned with the second endoscopic ultrasound image; wherein the second white light endoscopic image is associated with blood vessel state information, bleeding state information, and / or mucosal integrity information; The mucosal surface state reflected by the second white light endoscopic image is evaluated by combining the blood vessel state information, the bleeding state information, and / or the mucosal integrity information, and the surface feature evaluation result corresponding to the second white light endoscopic image is determined by expert experience. An IBD white light training data set is constructed based on the second white light endoscopic image and the corresponding surface feature evaluation result. An initial white light surface feature extraction model is constructed based on any one of Vision Transformer, EfficientNet, ResNet, or Inception, and the white light surface feature extraction model is iteratively trained on the IBD white light training data set based on the K-fold cross-validation method. When the validation loss value does not decrease for a predetermined number of rounds, the training is stopped.
7. The system for evaluating the activity of inflammatory bowel disease based on multi-modal endoscopic feature fusion according to any one of claims 1-3, characterized in that, The noise reduction processing based on the first EUS microarray multi-modality endoscopic dynamic sequence data includes: Based on the Farneback dense optical flow algorithm, the forward optical flow vector and the reverse optical flow vector between adjacent frame images of the first white light endoscopic image are calculated; Based on the reverse optical flow vector, interference optical flow is excluded by reverse verification to determine the optical flow residual vector, thereby calculating the withdrawal speed; The optical flow residual vector is converted into HSV channel values to construct a corresponding HSV image; The state of the first white light endoscopic image is determined by recognizing consecutive HSV images using a recurrent neural network, and the first white light endoscopic image and the corresponding first EUS endoscopic image are removed from the non-observation state, thereby determining the first white light effective observation image and the first ultrasound effective observation image.
8. The system for evaluating the activity of inflammatory bowel disease based on multi-modal endoscopic feature fusion according to claim 7, characterized in that, The image preprocessing module is also used to perform image segmentation and scaling on the first white light endoscopic image before noise reduction processing, including: The first white light endoscopic image is decoded into frame images, and the frame images are segmented using a U-net-based image semantic segmentation network to remove interference information in the frame images and retain the mucosal portion in the frame images, thereby obtaining segmented frame images; The segmented frame images are scaled based on the region interpolation method, and the scaled segmented frame images are re-encoded into the first white light endoscopic image.
9. The system for evaluating the activity of inflammatory bowel disease based on multi-modal endoscopic feature fusion according to claim 7, characterized in that, The white light surface feature extraction module is further configured to weight the withdrawal speed of the first white light effective observation image and the corresponding first white light endoscopic mucosa feature, determine the distribution of different characteristic mucosa of the effective observation intestinal segment and the corresponding mucosa proportion, and generate a colonoscopy schematic diagram reflecting the mucosa feature distribution of different effective observation intestinal segments.
Citation Information
Patent Citations
Tumor classification model training method based on multiple modes
CN118521834A
Ucerative colitis virtual biopsy method and system based on multi-modal image
CN120015247A