A Prediction System for Inflammatory Bowel Disease Treatment Based on Multimodal Data

The inflammatory bowel disease (IBD) efficacy prediction system, which integrates multimodal data fusion, comprehensively assesses the clinical, pathological, endoscopic, and imaging characteristics of IBD patients, addressing the shortcomings in the accuracy and timeliness of efficacy prediction in existing technologies and providing more accurate diagnostic and treatment support.

CN120708939BActive Publication Date: 2025-12-02SUN YAT SEN UNIV +1

Patent Information

Application Number
CN202511211606.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-28
Publication Date
2025-12-02
Estimated Expiration
2045-08-28

AI Technical Summary

Technical Problem

Existing technologies are insufficient to predict the efficacy of biologics in patients with inflammatory bowel disease by comprehensively analyzing multi-level and multi-dimensional characteristics, resulting in insufficient accuracy and timeliness in diagnosis and treatment decisions.

Method used

The system employs a multimodal data acquisition module, a text feature extraction module, an endoscopic intestinal mucosal feature recognition module, an image feature recognition module, and a pathological feature extraction module. Combined with the Transformer model, it performs multimodal data fusion to achieve a comprehensive evaluation of clinical, pathological, endoscopic, and imaging features.

Benefits of technology

It improves the accuracy and timeliness of predicting the treatment efficacy of inflammatory bowel disease, providing a more effective support tool for the diagnosis and treatment decisions of IBD patients.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120708939B_ABST
    Figure CN120708939B_ABST
Patent Text Reader

Abstract

This invention discloses a system for predicting the efficacy of inflammatory bowel disease based on multimodal data, relating to the field of medical information processing. The system includes: a multimodal data acquisition module for receiving multimodal data; a text feature extraction module for determining text feature vectors using a text feature extraction model; an endoscopic intestinal mucosal feature recognition module for determining intestinal mucosal features using an endoscopic intestinal mucosal feature recognition model; an image feature recognition module for determining imaging features using an image feature recognition model; a pathological feature extraction module for determining pathological features using a pathological recognition model; and an efficacy prediction module for predicting the efficacy of biologics based on multimodal data features using a Transformer model and fully connected layers. Compared to existing technologies, this invention achieves efficacy prediction of biologics based on the fusion of clinical, pathological, endoscopic, and imaging multimodal features, improving the accuracy and clinical rationality of diagnostic and treatment decisions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of medical information processing technology, and more specifically, to a system for predicting the efficacy of inflammatory bowel disease based on multimodal data. Background Technology

[0002] Inflammatory bowel disease (IBD) is a chronic, relapsing inflammatory bowel disease. IBD has a protracted and recurring course, is difficult to predict, and exhibits complex disease progression, severely impacting patients' quality of life. Biologics, targeting specific molecules or receptors in the inflammatory process, are used to treat IBD patients who do not respond well to or cannot tolerate conventional drugs, aiming to induce and maintain remission, reduce the incidence of IBD-related hospitalizations and surgeries, and improve patients' quality of life. However, approximately one-third of patients experience primary non-response to biologic therapy, and 30%-50% eventually develop secondary non-response during treatment. Ineffective biologic therapy not only increases the financial burden on patients but also exposes them to risks such as severe infections and autoimmune reactions due to its potent immunosuppressive effects. Therefore, accurately predicting the efficacy of biologics in IBD patients is crucial for improving their prognosis and quality of life.

[0003] The clinical manifestations and drug sensitivity of IBD patients exhibit significant individual heterogeneity. Clinicians typically need to conduct a comprehensive assessment based on demographic, laboratory, histological, imaging, and even transcriptomic information to achieve precise treatment. However, previous studies have largely relied on single-modality data such as clinical or laboratory indicators and imaging features, failing to achieve a comprehensive analysis of the multi-level and multi-dimensional characteristics of IBD patients. This approach cannot fully reflect the biological characteristics of the disease and is insufficient for effectively predicting clinical outcomes. Summary of the Invention

[0004] To overcome the limitations of the prior art in predicting the efficacy of inflammatory bowel disease, this invention provides a multimodal data-based system for predicting the efficacy of inflammatory bowel disease.

[0005] To solve the above-mentioned technical problems, the technical solution of the present invention is as follows:

[0006] A system for predicting the treatment efficacy of inflammatory bowel disease based on multimodal data, comprising:

[0007] A multimodal data acquisition module is used to receive multimodal data; wherein, the multimodal data includes clinical information, demographic information, laboratory test information, imaging information, tissue WSI images, and endoscopic images;

[0008] The text feature extraction module is used to load a trained text feature extraction model and use the text feature extraction model to determine text feature vectors based on at least one of the clinical information, the demographic information and the laboratory test information.

[0009] An endoscopic intestinal mucosal feature recognition module is used to carry a trained endoscopic intestinal mucosal feature recognition model and to determine intestinal mucosal features based on the endoscopic images using the endoscopic intestinal mucosal feature recognition model; wherein, the intestinal mucosal feature mapping effectively observes the clinical disease activity and corresponding distribution area information of the intestinal segment;

[0010] An image feature recognition module is used to carry a trained image feature recognition model and use the image feature recognition model to extract features based on the imaging information to determine imaging features; wherein, the imaging features map lesion area information;

[0011] The pathological feature extraction module is used to load a trained pathological recognition model and use the pathological recognition model to determine pathological features based on the tissue WSI image; wherein, the pathological features map tissue structure, pathological activity, morphological information and therapeutic type;

[0012] The efficacy prediction module is used to carry a Transformer model and a fully connected layer, and to determine the efficacy prediction results of the biological agent based on the text feature vector, the intestinal mucosal features, the imaging features and the pathological features.

[0013] Compared with the prior art, the beneficial effects of the technical solution of the present invention are:

[0014] This application discloses a multimodal data-based system for predicting the efficacy of inflammatory bowel disease (IBD) treatment. The system includes a multimodal data acquisition module, a text feature extraction module, an endoscopic intestinal mucosal feature recognition module, an image feature recognition module, a pathological feature extraction module, and a treatment efficacy prediction module. These modules aggregate multimodal and multiscale information, and then the treatment efficacy prediction module performs multimodal data fusion to achieve efficacy prediction of biological agents based on clinical, pathological, endoscopic, and imaging multimodal feature fusion. Compared to existing technologies, this application improves the accuracy, timeliness, and rationality of diagnosis and treatment decisions, providing an effective support tool for predicting treatment efficacy, exploring disease progression, and developing treatment plans for IBD patients. Attached Figure Description

[0015] Figure 1 This is a schematic diagram of the inflammatory bowel disease efficacy prediction system described in the embodiments of this application.

[0016] Figure 2This is a schematic diagram of the process by which the endoscopic intestinal mucosal feature recognition module determines intestinal mucosal features in the embodiments of this application.

[0017] Figure 3 This is a flowchart illustrating the training process of the pathological recognition model described in this application embodiment.

[0018] Figure 4 This is a schematic diagram of the hardware entity of the electronic device provided in the embodiments of this application. Detailed Implementation

[0019] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms are interchangeable where appropriate; this is merely a way of distinguishing objects with the same attributes in the description of embodiments of this application. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, so that a process, method, system, product, or apparatus that comprises a series of units is not necessarily limited to those units, but may include other units not explicitly listed or inherent to those processes, methods, products, or apparatuses. The term "determine" broadly covers a wide variety of actions, including acquiring, calculating, processing, deriving, investigating, searching (e.g., searching in a table, database, or other data structure), probing, and similar actions; it may also include receiving (e.g., receiving information), accessing (e.g., accessing data in memory), and similar actions; it may also include generating, creating, establishing, and similar actions; and parsing, selecting, choosing, and similar actions, etc. Definitions of other terms will be given in the following description.

[0020] It should be noted that when one element is considered to be "connected" to another element, it can be directly connected to the other element or connected to the other element through an intermediary element. Furthermore, in the following embodiments, "connection" should be understood as "electrical connection," "communication connection," etc., if there is transmission of electrical signals or data between the connected objects.

[0021] The accompanying drawings are for illustrative purposes only and should not be construed as limiting the scope of this patent.

[0022] To better illustrate this embodiment, some parts in the accompanying drawings may be omitted, enlarged, or reduced, and do not represent the actual product dimensions;

[0023] It will be understood by those skilled in the art that certain well-known structures and their descriptions may be omitted in the accompanying drawings.

[0024] The technical solution of the present invention will be further described below with reference to the accompanying drawings and embodiments.

[0025] This embodiment provides a system for predicting the efficacy of treatment for inflammatory bowel disease based on multimodal data, as shown in Figure 1, including:

[0026] A multimodal data acquisition module is used to receive multimodal data; wherein, the multimodal data includes clinical information, demographic information, laboratory test information, imaging information, tissue WSI images, and endoscopic images;

[0027] The text feature extraction module is used to load a trained text feature extraction model and use the text feature extraction model to determine text feature vectors based on at least one of the clinical information, the demographic information and the laboratory test information.

[0028] An endoscopic intestinal mucosal feature recognition module is used to carry a trained endoscopic intestinal mucosal feature recognition model and to determine intestinal mucosal features based on the endoscopic images using the endoscopic intestinal mucosal feature recognition model; wherein, the intestinal mucosal feature mapping effectively observes the clinical disease activity and corresponding distribution area information of the intestinal segment;

[0029] An image feature recognition module is used to carry a trained image feature recognition model and use the image feature recognition model to extract features based on the imaging information to determine imaging features; wherein, the imaging features map lesion area information;

[0030] The pathological feature extraction module is used to load a trained pathological recognition model and use the pathological recognition model to determine pathological features based on the tissue WSI image; wherein, the pathological features map tissue structure, morphological information, pathological activity and therapeutic type.

[0031] The efficacy prediction module is used to carry a Transformer model and a fully connected layer, and to determine the efficacy prediction results of the biological agent based on the text feature vector, the intestinal mucosal features, the imaging features and the pathological features.

[0032] The inflammatory bowel disease (IBD) efficacy prediction system disclosed in this embodiment utilizes a multimodal data acquisition module, a text feature extraction module, an endoscopic intestinal mucosal feature recognition module, an image feature recognition module, and a pathological feature extraction module. It aggregates multimodal and multiscale information from clinical, pathological, endoscopic, and imaging data. Then, through the efficacy prediction module, it performs multimodal data fusion to comprehensively evaluate the individual physical examination of a specific IBD patient, simulating the logic of evidence-based diagnosis and treatment. It realizes the efficacy prediction of biological agents based on the fusion of clinical, pathological, endoscopic, and imaging multimodal features, improving the accuracy, timeliness, and clinical rationality of diagnosis and treatment decisions. It provides an effective support tool for predicting the efficacy of IBD patients, exploring the course of the disease, and formulating treatment plans.

[0033] As a non-limiting example, the clinical information includes, but is not limited to, clinical symptoms, bowel movements, and clinical medications, and the laboratory test information includes, but is not limited to, hemoglobin and erythrocyte sedimentation rate at different time points.

[0034] It should be noted that the clinical disease activity reflects the quantitative results of the imaging manifestations of IBD disease progression; the distribution area information represents the corresponding anatomical structure location.

[0035] In some preferred embodiments, referring to FIG2, determining the intestinal mucosal characteristics based on the endoscopic images includes:

[0036] Based on the Farneback dense optical flow algorithm, the forward optical flow vector and the reverse optical flow vector between adjacent frames of the endoscopic image are calculated;

[0037] Based on the reverse optical flow vector, reverse verification is performed to eliminate interfering optical flow, and the optical flow residual vector is determined, thereby calculating the retraction observation speed.

[0038] The optical flow residual vector is converted into HSV (Hue, Saturation, Value) channel values ​​to construct the corresponding HSV image;

[0039] The recurrent neural network is used to identify continuous HSV images, determine the state of the endoscopic images (such as colonoscopy withdrawal, insertion, sliding, aspiration, flushing, biopsy, etc.), and remove segments that are not in the observation state to determine the endoscopic image observation segments.

[0040] Using the endoscopic intestinal mucosal feature recognition model, the intra-abdominal bronchoscopic motility (IBD) is determined based on the endoscopic image observation segment. The IBD endoscopic motility is then weighted with the withdrawal observation speed to determine the intestinal mucosal features that reflect the motility distribution and corresponding mucosal proportion of different effective observation intestinal segments in the endoscopic image observation segment.

[0041] In some optional embodiments, referring to FIG2, before determining the optical flow residual vector, the endoscopic image undergoes a first preprocessing, including:

[0042] The endoscopic image is decoded into a frame image, and the frame image is segmented using an image semantic segmentation network to remove interference information from the frame image while retaining the mucosal portion of the frame image, thus obtaining the segmented frame image.

[0043] The segmented frame image is scaled using region interpolation, and the scaled segmented frame image is then re-encoded into the endoscopic image.

[0044] It should be noted that due to differences in endoscope models, brands, etc., there may be significant differences between different original video images. By decoding, segmenting, and recoding the endoscope images, the negative impact of image format differences on model recognition can be removed, and interference information other than image information can be eliminated / reduced.

[0045] In some specific implementations, the image semantic segmentation network is built based on Unet or its improved model.

[0046] In some specific implementation processes, the size of the endoscopic image is around 1200×1000. The endoscopic image is reduced to 360×360 through region interpolation to ensure the real-time performance of image recognition.

[0047] In some optional embodiments, the training process of the endoscopic intestinal mucosal feature recognition model includes:

[0048] Acquire training endoscopic images in the state of withdrawal observation;

[0049] Using expert experience, the vascular status information, bleeding status information and mucosal ulcer status reflected in the training endoscopic images are graded and assessed to determine the grading assessment results.

[0050] Based on the grading assessment results regarding the vascular status information, bleeding status information, and ulcer status information, the corresponding activity score of the training endoscopic images is determined;

[0051] After data augmentation of the training endoscopic images, an endoscopic image training sample set is established based on the training endoscopic images and activity scores.

[0052] The initial endoscopic intestinal mucosal feature recognition model was constructed based on ResNet and iteratively trained on the endoscopic image training sample set.

[0053] As a non-limiting example, the data augmentation includes, but is not limited to, preprocessing the training endoscopic images by randomly erasing, rotating, or flipping them.

[0054] In some specific implementation processes, based on the endoscopic characteristics proposed by the Ulcerative Colitis Endoscopic Index of Severity (UCEIS) and the Mayo Endoscopic Activity Score (MES), three expert physicians conduct graded assessments of the mucosa.

[0055] In some optional embodiments, the endoscopic intestinal mucosal feature recognition module is also used to mark the activity of the effectively observed intestinal segment with a preset color on the colonoscopy diagram after determining the intestinal mucosal features, and output the visualization results of the activity under IBD endoscopy.

[0056] It is understandable that, based on the visualization results of IBD endoscopic activity, it is possible to retrieve several endoscopic image data associated with a specific patient, and then generate follow-up comparison results of colonoscopy activity to assess the temporal changes in the colonoscopy performance of a specific patient.

[0057] Furthermore, to fully utilize the temporal changes in colonoscopy findings in specific patients, an LSTM (Long Short-Term Memory) model can be used for feature extraction.

[0058] In some preferred embodiments, determining pathological features based on the tissue WSI image using the pathological identification model includes:

[0059] The tissue WSI image is cut into pixel blocks of a specified pixel size, and the Otsu algorithm is used to filter the background of the pixel blocks;

[0060] Each group of pixel blocks corresponding to the background-filtered tissue WSI image is used as input to the pathological recognition model to determine the pathological features of the tissue WSI image.

[0061] In some alternative embodiments, referring to Figure 3, the training process of the pathology recognition model includes:

[0062] Images of IBD intestinal tissue obtained after formalin fixation, paraffin embedding, and HE staining were processed by image segmentation and background filtering to obtain training inflammatory bowel disease pathomic pixel blocks.

[0063] For each WSI image of the colonic tissue, at least a portion of the training inflammatory bowel disease pathomic pixel blocks are selected, and the pathological activity, tissue structure, morphological information and treatment type of each training inflammatory bowel disease tissue pathological image pixel block are labeled to construct a pathological dataset.

[0064] The pathology dataset is divided into A subset of pathological data, One subset of the pathological data is used as the training set, and the remaining one subset of the pathological data is used as the test set;

[0065] Based on multilayer perceptron construction An initial pathology identification model, let The pathological identification model mentioned above is in The model is trained on each training set to update its parameters; wherein the training inflammatory bowel disease pathomic pixel blocks are used as input to the pathological identification model, and the treatment outcome type is used as the output label of the pathological identification model.

[0066] Training is stopped when the average correlation between each tissue structure, pathological activity, morphological information and efficacy calculated by the pathological identification model on a validation set containing at least a portion of the training data of the corresponding training set does not improve over multiple consecutive periods.

[0067] Verify on the test set respectively. The stability of the pathological identification models is assessed, and the model with the best performance is selected as the pathological identification model for online use.

[0068] It should be understood that because WSIs have huge pixel sizes, organizing WSI images involves image segmentation and using the Otsu algorithm to filter the white background of the image, which can filter out background information that does not need to be analyzed.

[0069] In some specific implementations, WSI images are divided into 224×224 pixel blocks, and 10,000 small pixel blocks are selected in each tissue WSI image.

[0070] In some specific implementation processes, a pathological identification model was built using Python based on the TensorFlow deep learning framework. For preprocessed intestinal pathomic images (WSI images), the segmented WSI images were randomly divided into 5 parts, with 4 parts randomly selected as the training set and the remaining part as the test set, resulting in 5 combinations. Models were then built for each combination to verify model stability, and the optimal model was selected. Training was stopped when the average correlation of the efficacy type calculated on the validation set (containing 10% of the training data) did not improve within 50 consecutive cycles. After training, the constructed pathological identification model could output a corresponding confidence score for each segmented small slice, i.e., the feature evaluation and identification result of that small slice.

[0071] In some preferred embodiments, the training process of the image feature recognition model includes:

[0072] Acquire training radiographic images and resample them to a specified spatial resolution using linear interpolation;

[0073] Based on the image pixels, the resampled training image is standardized using Z-scores to determine the standardized image corresponding to the training image. This process is represented as follows:

[0074]

[0075] In the formula, Represents the original pixel value. Represents the average pixel value. The standard deviation of pixel values. This represents the standardized pixel value;

[0076] Using expert experience, based on the principle of minimum irregular curves, lesion regions (i.e., regions of interest (VOI)) are annotated on the standardized imaging images corresponding to the training imaging images, and the lesion regions are reconstructed in three dimensions according to the annotation results to form the corresponding training image VOI.

[0077] The initial image feature recognition model is constructed based on a 3D CNN (3D Convolutional Neural Network). The training image is used as the input of the image feature recognition model, and the corresponding VOI of the training image is used as the output label of the image feature recognition model. The image feature recognition model is then trained iteratively.

[0078] It should be understood that for three-dimensional MRI (Magnetic Resonance Imaging) data, the image feature recognition model based on 3D CNN after training can better capture the spatial information of three-dimensional lesion data. It can directly perform convolution operations on the volume data through three-dimensional convolution kernels to achieve high-level feature extraction of lesions.

[0079] In some specific implementation processes, the IBD imaging expert physician first determines a set of pixel intensities of training images as standardized reference points, converts the pixel intensities of the images into Z-scores to achieve standardization, calculates the Z-scores of the training images, and then applies these reference points to the standardization of new images.

[0080] In some alternative embodiments, determining the imaging features includes:

[0081] The imaging information is received and resampled to a specified spatial resolution; wherein the imaging information includes three-dimensional image data;

[0082] The resampled image information is standardized by Z-score and used as input to the trained image feature recognition model. High-level features are extracted from the image information using a three-dimensional convolution kernel as the image features.

[0083] As a non-limiting example, the training radiographic images and / or three-dimensional image data may be magnetic resonance images.

[0084] It should be noted that due to differences in the models, brands, and parameters of the equipment used to acquire training imaging images and / or 3D image data, especially the differences in the spatial resolution of magnetic resonance images, the learning results and recognition accuracy of the image feature recognition model may be poor. This embodiment uses a resampling method to ensure that the model can learn consistent spatial information from different data sources.

[0085] In some specific implementation processes, the training imaging images are resampled to Spatial resolution is used to form a new image sequence.

[0086] In some preferred embodiments, the training process of the text feature extraction model includes:

[0087] Acquire training text information; wherein, the training text information includes training clinical information, training laboratory test information, and training demographic information;

[0088] A text feature extraction regular expression library is constructed using expert experience as the text feature extraction model, and corresponding regular expressions are configured.

[0089] Based on the text feature extraction model, the training text information is vectorized to obtain the training text feature vector.

[0090] The training text feature vectors are evaluated using expert experience, and the text feature extraction model is optimized based on the evaluation results of the training text feature vectors.

[0091] In some specific implementation processes, three expert physicians from the IBD diagnosis and treatment alliance jointly determine the possible descriptive methods for each text prognostic-related feature (such as drug type, drug dosage, stool condition, erythrocyte sedimentation rate, C-reactive protein, and X-ray showing intestinal shortening, etc.), and then construct a regular expression library for text feature extraction. The raw clinical information in the database is automatically vectorized using regular expressions, and the extracted vector labels are jointly reviewed by the three expert physicians. After expert review and approval, based on the results of the reviewed clinical information, the accuracy of the text feature extraction model in extracting vector labels is evaluated using metrics such as precision, recall, and F1 score, and the regular expression model is continuously optimized.

[0092] In some specific implementations, the text feature vector includes clinical information features. (Based on clinical information extraction) and laboratory test characteristics (Based on information extracted from laboratory tests).

[0093] In some preferred embodiments, determining the predicted efficacy of the biologics includes:

[0094] The text feature vector, the intestinal mucosal features, the imaging features, and the pathological features are concatenated to obtain multimodal features;

[0095] The multimodal features are used as input to the efficacy prediction module. The self-attention mechanism of the Transformer model is used to extract multimodal high-level features. The multimodal high-level features are mapped to the efficacy prediction results of the biologics through multiple fully connected layers, and the corresponding prediction probabilities are output using the Sigmoid activation function.

[0096] It should be noted that this embodiment employs a multimodal feature pre-fusion strategy to promote the integrity of feature learning and facilitate the full fusion of information and exploration of interrelationships among various modalities. Furthermore, the use of vector concatenation facilitates the reallocation of computing resources, ensuring the timeliness and efficiency of prediction results output.

[0097] In some specific implementation processes, the clinical information features are analyzed in the efficacy prediction module. Laboratory test characteristics Intestinal mucosal characteristics Imaging features and pathological features Vectors are concatenated to form multimodal features. As input to the Transformer model, it learns the correlation and complementarity between different modalities of data through the self-attention mechanism based on the Transformer model.

[0098] More specifically, the Transformer model includes multiple Transformer coding layers.

[0099] In some specific implementation processes, the efficacy prediction module visualizes the efficacy prediction results of the biological agent using different colors and text based on the prediction probability, including but not limited to: using a red warning box to indicate that the efficacy prediction for a specific patient is "poor", using a yellow warning box to indicate that the efficacy prediction for a specific patient is "pending", and using green text to indicate that the efficacy prediction for a specific patient is "good".

[0100] In some specific implementation processes, the weights of each feature vector will be fed back to the user (such as medical staff) in the form of numerical values ​​or bar charts, so that the user can see the efficacy prediction ideas and provide a reference for the user's information evaluation.

[0101] In some examples, a computer program is provided, including computer-readable code, which, when executed in a computer device, allows a processor in the computer device to perform actions for implementing some or all of the modules / methods in the system.

[0102] This embodiment provides a computer-readable storage medium storing at least one instruction, at least one program, code set, or instruction set. The at least one instruction, at least one program, code set, or instruction set is loaded and executed by a processor, causing the processor to execute some or all of the modules / methods in the system provided in this embodiment.

[0103] It is understood that the storage medium can be transient or non-transient. Exemplarily, the storage medium includes, but is not limited to, various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0104] By way of example, the processor may be a central processing unit (CPU), a microprocessor unit (MPU), a digital signal processor (DSP), or a field programmable gate array (FPGA), etc.

[0105] By way of example, the read-only memory includes, but is not limited to, MASK ROM, PROM, EPROM, EEPROM, Flash, etc.

[0106] By way of example, the random access memory includes, but is not limited to, DRAM, SRAM, SDRAM, DDR SDRAM, etc.

[0107] In some examples, a computer program product is provided, which can be implemented by hardware, software, or a combination thereof. As a non-limiting example, the computer program product can be embodied in the storage medium, or it can be embodied in a software product, such as an SDK (Software Development Kit).

[0108] As a non-limiting example, a computer program product is provided, comprising a computer program or computer-executable instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer program or computer-executable instructions from the computer-readable storage medium, and executes the computer-executable instructions, causing the electronic device to perform some or all of the modules / methods described in the embodiments of this application.

[0109] This embodiment also proposes an electronic device, including a memory and a processor. The memory stores at least one instruction, at least one program, code set, or instruction set. When the processor executes the at least one instruction, at least one program, code set, or instruction set, it implements some or all of the modules / methods in the system described in the embodiment.

[0110] In some examples, a hardware entity of the electronic device is provided, referring to Figure 4, including: a processor, a memory, and a communication interface; wherein, the processor typically controls the overall operation of the electronic device; the communication interface is used to enable the electronic device to communicate with other terminals or servers via a network; the memory is configured to store instructions and applications executable by the processor, and can also cache data to be processed or already processed (including but not limited to image data, audio data, voice communication data, and video communication data) to be processed by the processor and various modules in the electronic device, and can be implemented by flash memory (FLASH), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or random access memory (RAM).

[0111] A processor may include one or more processing elements. Therefore, a processor may include one or more integrated circuits (ICs) configured to perform the functions of the processor. Furthermore, each integrated circuit may include circuitry (e.g., a first circuit, a second circuit, and other circuitry, etc.) configured to perform the functions of the processor.

[0112] Furthermore, data can be transferred between the processor, communication interface, and memory via a bus, which can include any number of interconnected buses and bridges, connecting various circuits of one or more processors and memories together.

[0113] The same or similar labels correspond to the same or similar parts;

[0114] The terms used to describe positional relationships in the accompanying drawings are for illustrative purposes only and should not be construed as limiting this application.

[0115] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other.

[0116] In different specific implementations, the methods or systems described in this application can be implemented in software, hardware, or a combination thereof. Furthermore, the order of the method steps can be changed, and various elements can be added, reordered, combined, omitted, or modified.

[0117] Obviously, the above embodiments of this application are merely examples for clearly illustrating this application, and are not intended to limit the implementation of this application, nor are they intended to limit this application. For those skilled in the art, other variations or modifications can be made based on the above description. The separate structural / functional modules or units can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part. The structure and function of the separate components can be implemented as a combined structure or component. It is neither necessary nor possible to exhaustively describe all embodiments here. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the scope of protection of the claims of this application.

Claims

1. A system for predicting the treatment efficacy of inflammatory bowel disease based on multimodal data, characterized in that, include: A multimodal data acquisition module is used to receive multimodal data; wherein, the multimodal data includes clinical information, demographic information, laboratory test information, imaging information, tissue WSI images, and endoscopic images; The text feature extraction module is used to load a trained text feature extraction model and use the text feature extraction model to determine text feature vectors based on at least one of the clinical information, the demographic information and the laboratory test information. An endoscopic intestinal mucosal feature recognition module is used to carry a trained endoscopic intestinal mucosal feature recognition model and to determine intestinal mucosal features based on the endoscopic images using the endoscopic intestinal mucosal feature recognition model; wherein, the intestinal mucosal feature mapping effectively observes the clinical disease activity and corresponding distribution area information of the intestinal segment; An image feature recognition module is used to carry a trained image feature recognition model and use the image feature recognition model to extract features based on the imaging information to determine imaging features; wherein, the imaging features map lesion area information; The pathological feature extraction module is used to load a trained pathological recognition model and use the pathological recognition model to determine pathological features based on the tissue WSI image; wherein, the pathological features map pathological activity, tissue structure, morphological information and therapeutic type. The efficacy prediction module is used to carry a Transformer model and a fully connected layer, and to determine the efficacy prediction results of the biological agent based on the text feature vector, the intestinal mucosal features, the imaging features and the pathological features. The determination of intestinal mucosal characteristics based on the endoscopic images includes: Based on the Farneback dense optical flow algorithm, the forward optical flow vector and the reverse optical flow vector between adjacent frames of the endoscopic image are calculated; Based on the reverse optical flow vector, reverse verification is performed to eliminate interfering optical flow, and the optical flow residual vector is determined, thereby calculating the retraction observation speed. The optical flow residual vector is converted into HSV channel values ​​to construct the corresponding HSV image; The recurrent neural network is used to identify continuous HSV images, determine the state of the endoscopic images, and remove segments that are not in the observation state to determine the endoscopic image observation segments. Using the endoscopic intestinal mucosal feature recognition model, the intra-abdominal bronchoscopic motility (IBD) is determined based on the endoscopic image observation segment. The IBD endoscopic motility is then weighted with the withdrawal observation speed to determine the intestinal mucosal features that reflect the motility distribution and corresponding mucosal proportion of different effective observation intestinal segments in the endoscopic image observation segment.

2. The inflammatory bowel disease treatment efficacy prediction system based on multimodal data according to claim 1, characterized in that, Before determining the optical flow residual vector, the endoscopic image undergoes a first preprocessing step, including: The endoscopic image is decoded into a frame image, and the frame image is segmented using an image semantic segmentation network to remove interference information from the frame image while retaining the mucosal portion of the frame image, thus obtaining the segmented frame image. The segmented frame image is scaled using region interpolation, and the scaled segmented frame image is then re-encoded into the endoscopic image.

3. The inflammatory bowel disease treatment efficacy prediction system based on multimodal data according to claim 1, characterized in that, The training process of the endoscopic intestinal mucosal feature recognition model includes: Acquire training endoscopic images in the state of withdrawal observation; Using expert experience, the vascular status information, bleeding status information and mucosal ulcer status reflected in the training endoscopic images are graded and assessed to determine the grading assessment results. Based on the grading assessment results regarding the vascular status information, the bleeding status information, and the ulcer status information, the corresponding activity score of the training endoscopic images is determined; After data augmentation of the training endoscopic images, an endoscopic image training sample set is established based on the training endoscopic images and activity scores. The initial endoscopic intestinal mucosal feature recognition model was constructed based on ResNet and iteratively trained on the endoscopic image training sample set.

4. The inflammatory bowel disease treatment efficacy prediction system based on multimodal data according to claim 1, characterized in that, The process of determining pathological features based on the tissue WSI image using the pathological recognition model includes: The tissue WSI image is cut into pixel blocks of a specified pixel size, and the Otsu algorithm is used to filter the background of the pixel blocks; Each group of pixel blocks corresponding to the background-filtered tissue WSI image is used as input to the pathological recognition model to determine the pathological features of the tissue WSI image.

5. The inflammatory bowel disease treatment efficacy prediction system based on multimodal data according to claim 4, characterized in that, The training process of the pathology identification model includes: WSI images of intestinal tissue after formalin fixation, paraffin embedding, and HE staining were obtained, and then the images were segmented and the background filtered to obtain training inflammatory bowel disease pathomic pixel blocks. For each WSI image of the intestinal tissue, at least a portion of the training inflammatory bowel disease pathomic pixel blocks are selected, and the pathological activity, tissue structure, morphological information and therapeutic effects are labeled for each training inflammatory bowel disease pathomic pixel block to construct a pathological dataset. The pathology dataset is divided into A subset of pathological data, One subset of the pathological data is used as the training set, and the remaining one subset of the pathological data is used as the test set; Based on multilayer perceptron construction An initial pathology identification model, let The pathological identification model mentioned above is in The model is trained on each training set to update its parameters; wherein the training inflammatory bowel disease pathomic pixel blocks are used as input to the pathological identification model, and the treatment outcome type is used as the output label of the pathological identification model. Training is stopped when the average correlation between each tissue structure, pathological activity, morphological information and efficacy calculated by the pathological identification model on a validation set containing at least a portion of the training data of the corresponding training set does not improve over multiple consecutive periods. Verify on the test set respectively. The stability of the pathological identification models is assessed, and the model with the best performance is selected as the pathological identification model for online use.

6. The inflammatory bowel disease treatment efficacy prediction system based on multimodal data according to claim 1, characterized in that, The training process of the image feature recognition model includes: Acquire training radiographic images and resample them to a specified spatial resolution using linear interpolation; Based on the image pixels, the resampled training image is standardized using Z-scores to determine the standardized image corresponding to the training image. This process is represented as follows: In the formula, Represents the original pixel value. Represents the average pixel value. The standard deviation of pixel values. This represents the standardized pixel value; Using expert experience, based on the principle of minimum irregular curve, the lesion region is labeled on the standardized imaging image corresponding to the training imaging image, and the lesion region is reconstructed in three dimensions according to the labeling results to form the corresponding training image VOI; The initial image feature recognition model is constructed based on 3D CNN. The training image is used as the input of the image feature recognition model, and the corresponding VOI of the training image is used as the output label of the image feature recognition model. The image feature recognition model is trained iteratively.

7. The inflammatory bowel disease treatment efficacy prediction system based on multimodal data according to claim 6, characterized in that, The determination of imaging features includes: The imaging information is received and resampled to a specified spatial resolution; wherein the imaging information includes three-dimensional image data; The resampled image information is standardized by Z-score and used as input to the trained image feature recognition model. High-level features are extracted from the image information using a three-dimensional convolution kernel as the image features.

8. The inflammatory bowel disease treatment efficacy prediction system based on multimodal data according to claim 1, characterized in that, The training process of the text feature extraction model includes: Acquire training text information; wherein, the training text information includes training clinical information, training laboratory test information, and training demographic information; A text feature extraction regular expression library is constructed using expert experience as the text feature extraction model, and corresponding regular expressions are configured. Based on the text feature extraction model, the training text information is vectorized to obtain the training text feature vector. The training text feature vectors are evaluated using expert experience, and the text feature extraction model is optimized based on the evaluation results of the training text feature vectors.

9. A system for predicting the efficacy of inflammatory bowel disease treatment based on multimodal data according to any one of claims 1-8, characterized in that, The determination of the predicted efficacy of the biological agent includes: The text feature vector, the intestinal mucosal features, the imaging features, and the pathological features are concatenated to obtain multimodal features; The multimodal features are used as input to the efficacy prediction module. The self-attention mechanism of the Transformer model is used to extract multimodal high-level features. The multimodal high-level features are mapped to the efficacy prediction results of the biologics through multiple fully connected layers, and the corresponding prediction probabilities are output using the Sigmoid activation function.

Citation Information

Patent Citations

  • Artificial intelligence auxiliary system for digestive endoscopic images of inflammatory bowel diseases

    CN111524124A

  • Multi-modal prediction model construction method and system for analyzing liver cancer recurrence data

    CN117612711A

Cited By

  • Anorectal disease multi-modal data feature analysis method and system

    CN122436260A

  • A method and system for feature analysis of multimodal data of anorectal diseases

    CN122436260B