Inflammatory bowel disease curative effect prediction system based on multi-modal data
The inflammatory bowel disease efficacy prediction system based on multimodal data fusion uses multimodal data acquisition and feature recognition modules, combined with the Transformer model, to solve the problems of insufficient accuracy and timeliness of inflammatory bowel disease efficacy prediction in existing technologies, and achieve more accurate diagnosis and treatment decision support.
Patent Information
- Application Number
- CN202511211606.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-28
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2045-08-28
AI Technical Summary
Existing technologies make it difficult to effectively predict the efficacy of biological agents in patients with inflammatory bowel disease through single-modality data, and are unable to achieve comprehensive analysis of multi-level and multi-dimensional characteristics, resulting in insufficient accuracy and timeliness in diagnosis and treatment decisions.
A multimodal data acquisition module, a text feature extraction module, an endoscopic intestinal mucosal feature recognition module, an imaging feature recognition module, and a pathological feature extraction module are used. Through multimodal data fusion and the Transformer model, the efficacy of biological agents is predicted and the clinical-pathological-endoscopic-imaging characteristics of IBD patients are comprehensively evaluated.
It improves the accuracy and timeliness of diagnosis and treatment decisions for IBD patients, provides effective support for efficacy prediction and treatment plan formulation, and enhances the rationality and accuracy of predictions.
Smart Images

Figure CN120708939A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of medical information processing technology, and more specifically, to an inflammatory bowel disease efficacy prediction system based on multimodal data. Background Art
[0002] Inflammatory bowel disease (IBD) is a chronic, relapsing inflammatory disease of the intestine. The course of IBD is protracted and recurrent, unpredictable, and complex, severely impacting patients' quality of life. Biologics target specific molecules or receptors in the inflammatory process and are used to treat IBD patients who respond poorly to or cannot tolerate traditional drugs to induce and maintain remission, reducing the incidence of IBD-related hospitalizations and surgeries and improving patients' quality of life. However, approximately one-third of patients experience primary non-response to biologic treatment, and 30%-50% of patients ultimately develop secondary non-response during treatment. Ineffective biologic treatment not only increases the financial burden on patients, but also exposes them to risks such as severe infection and autoimmune reactions due to their potent immunosuppressive effects. Therefore, accurately predicting the efficacy of biologics in IBD patients is key to improving their prognosis and quality of life.
[0003] IBD patients exhibit significant individual heterogeneity in their clinical manifestations and drug sensitivity. Clinicians typically need to conduct comprehensive assessments based on demographic, laboratory, histological, imaging, and even transcriptome data to achieve precise treatment. However, previous studies have often relied on single-modality data, such as clinical or laboratory indicators or imaging features, failing to comprehensively analyze the multifaceted and multidimensional characteristics of IBD patients, fully reflecting the biological characteristics of the disease and making it difficult to effectively predict clinical efficacy outcomes. Summary of the Invention
[0004] In order to overcome the limitations of the above-mentioned prior art in predicting the efficacy of inflammatory bowel disease, the present invention provides an inflammatory bowel disease efficacy prediction system based on multimodal data.
[0005] In order to solve the above technical problems, the technical solutions of the present invention are as follows: A system for predicting the therapeutic effect of inflammatory bowel disease based on multimodal data, comprising: A multimodal data acquisition module, configured to receive multimodal data, wherein the multimodal data includes at least two of clinical information, demographic information, laboratory examination information, imaging information, tissue WSI images, and endoscopic images; a text feature extraction module, configured to carry a trained text feature extraction model and utilize the text feature extraction model to determine a text feature vector based on at least one of the clinical information, the demographic information, and the laboratory examination information; An endoscopic intestinal mucosal feature recognition module, configured to carry a trained endoscopic intestinal mucosal feature recognition model and utilize the model to determine intestinal mucosal features based on the endoscopic image; wherein the intestinal mucosal feature map effectively observes the clinical disease activity of the intestinal segment and the corresponding distribution area information; An image feature recognition module, configured to carry a trained image feature recognition model and utilize the image feature recognition model to extract features based on the imaging information to determine imaging features; wherein the imaging features map lesion area information; A pathology feature extraction module, configured to carry a trained pathology recognition model and utilize the pathology recognition model to determine pathology features based on the tissue WSI image; wherein the pathology features map tissue structure, pathology activity, morphology information, and therapeutic effect type; The efficacy prediction module is used to carry the Transformer model and the fully connected layer, and determine the efficacy prediction result of the biological agent based on at least two of the text feature vector, the intestinal mucosal features, the imaging features and the pathological features.
[0006] Compared with the prior art, the beneficial effects of the technical solution of the present invention are: The present application discloses a system for predicting the efficacy of inflammatory bowel disease based on multimodal data, including a multimodal data acquisition module, a text feature extraction module, an endoscopic intestinal mucosal feature recognition module, an image feature recognition module, a pathological feature extraction module and an efficacy prediction module. The multimodal data acquisition module, the text feature extraction module, the endoscopic intestinal mucosal feature recognition module, the image feature recognition module and the pathological feature extraction module aggregate multimodal and multiscale information, and then the efficacy prediction module performs multimodal data fusion to achieve biological agent efficacy prediction based on clinical-pathological-endoscopic-imaging multimodal feature fusion. Compared with the existing technology, the present application improves the accuracy, timeliness and rationality of diagnosis and treatment decisions, and provides an effective support tool for efficacy prediction, disease course exploration and treatment plan formulation for IBD patients. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] Figure 1 Schematic diagram of the structure of the inflammatory bowel disease efficacy prediction system described in the examples of this application.
[0008] Figure 2 Schematic diagram of the process of determining intestinal mucosal features by the endoscopic intestinal mucosal feature recognition module described in the embodiment of the present application.
[0009] Figure 3 Schematic diagram of the training process of the pathology recognition model described in the embodiments of this application.
[0010] Figure 4 A schematic diagram of the hardware entity of the electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0011] The terms "first", "second" etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequential order. It should be understood that the terms used in this way can be interchangeable in appropriate circumstances, and this is merely a way of distinguishing the objects of the same attribute when describing them in the embodiments of the present application. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, so that the process, method, system, product or equipment comprising a series of units need not be limited to those units, but may include other units that are not clearly listed or inherent to these processes, methods, products or equipment. The term "determine" widely covers various actions, may include obtaining, calculating, computing, processing, deriving, investigating, searching (for example, searching in a table, a database or other data structure), ascertaining and similar actions, may also include receiving (for example, receiving information), accessing (for example, accessing data in a memory) and similar actions, may also include generating, creating, establishing and similar actions, and parsing, selecting, selecting and similar actions etc. The relevant definitions of other terms will be provided in the following description.
[0012] It should be noted that when an element is considered to be "connected" to another element, it can be directly connected to the other element or connected to the other element through an intervening element. In addition, the "connection" in the following embodiments should be understood as "electrical connection", "communication connection", etc., if there is transmission of electrical signals or data between the connected objects.
[0013] The accompanying drawings are for illustrative purposes only and are not to be construed as limiting this patent; In order to better illustrate this embodiment, some parts in the drawings may be omitted, enlarged, or reduced, and do not represent the actual product size; It is understandable to those skilled in the art that some well-known structures and descriptions thereof may be omitted in the drawings.
[0014] The technical solution of the present invention is further described below with reference to the accompanying drawings and embodiments.
[0015] This embodiment provides a system for predicting the therapeutic effect of inflammatory bowel disease based on multimodal data, as shown in FIG1 , including: A multimodal data acquisition module, configured to receive multimodal data, wherein the multimodal data includes at least two of clinical information, demographic information, laboratory examination information, imaging information, tissue WSI images, and endoscopic images; a text feature extraction module, configured to carry a trained text feature extraction model and utilize the text feature extraction model to determine a text feature vector based on at least one of the clinical information, the demographic information, and the laboratory examination information; An endoscopic intestinal mucosal feature recognition module, configured to carry a trained endoscopic intestinal mucosal feature recognition model and utilize the model to determine intestinal mucosal features based on the endoscopic image; wherein the intestinal mucosal feature map effectively observes the clinical disease activity of the intestinal segment and the corresponding distribution area information; An image feature recognition module, configured to carry a trained image feature recognition model and utilize the image feature recognition model to extract features based on the imaging information to determine imaging features; wherein the imaging features map lesion area information; A pathology feature extraction module, configured to carry a trained pathology recognition model and utilize the pathology recognition model to determine pathology features based on the tissue WSI image; wherein the pathology features map tissue structure, morphological information, pathology activity, and therapeutic effect type; The efficacy prediction module is used to carry the Transformer model and the fully connected layer, and determine the efficacy prediction result of the biological agent based on at least two of the text feature vector, the intestinal mucosal features, the imaging features and the pathological features.
[0016] The inflammatory bowel disease efficacy prediction system disclosed in this embodiment utilizes a multimodal data acquisition module, a text feature extraction module, an endoscopic intestinal mucosal feature recognition module, an image feature recognition module, and a pathological feature extraction module to aggregate multimodal and multi-scale information such as clinical, pathological, endoscopic, and imaging information. Multimodal data fusion is then performed through the efficacy prediction module to comprehensively evaluate individual physical examinations of specific IBD patients, simulate evidence-based diagnosis and treatment logic, and realize biological agent efficacy prediction based on clinical-pathological-endoscopic-imaging multimodal feature fusion, thereby improving the accuracy, timeliness, and clinical rationality of diagnosis and treatment decisions, and providing an effective support tool for efficacy prediction, disease course exploration, and treatment plan formulation for IBD patients.
[0017] As a non-limiting example, the clinical information includes but is not limited to clinical symptoms, stool conditions, clinical medications, etc., and the laboratory test information includes but is not limited to hemoglobin and erythrocyte sedimentation rate at different time points, etc.
[0018] It should be noted that the clinical disease activity reflects the quantitative results of the imaging manifestations of IBD disease progression; the distribution area information represents the corresponding anatomical structure location.
[0019] In some preferred embodiments, referring to FIG. 2 , determining the intestinal mucosal characteristics based on the endoscopic image includes: Based on the Farneback dense optical flow algorithm, the forward optical flow vector and the reverse optical flow vector between adjacent frames of the endoscopic image are calculated; Performing reverse verification based on the reverse optical flow vector to eliminate interference optical flow, determining the optical flow residual vector, and thus calculating the mirror withdrawal observation speed; Convert the optical flow residual vector into HSV (Hue, Saturation, Value) channel values and construct the corresponding HSV image; Using a recurrent neural network to identify the continuous HSV images, determine the state of the endoscopic image (such as colonoscopy withdrawal, endoscope insertion, endoscope sliding, water suction, water flushing, biopsy, etc.), and eliminate the segments belonging to the non-observation state to determine the endoscopic image observation segment; The endoscopic intestinal mucosal feature recognition model is used to determine the IBD endoscopic activity based on the endoscopic image observation segment, and the IBD endoscopic activity is weighted with the endoscopic observation withdrawal speed to determine the intestinal mucosal features that reflect the activity distribution of different effective observation intestinal segments in the endoscopic image observation segment and the corresponding mucosal proportion.
[0020] In some optional embodiments, referring to FIG. 2 , before determining the optical flow residual vector, the endoscopic image is subjected to a first preprocessing, including: Decoding the endoscopic image into a frame image, performing image segmentation on the frame image using an image semantic segmentation network, removing interference information in the frame image, retaining the mucosal portion of the frame image, and obtaining a segmented frame image; The segmented frame image is scaled based on a regional interpolation method, and the scaled segmented frame image is re-encoded into the endoscopic image.
[0021] It should be noted that due to differences in endoscope models and brands, there may be significant differences between different video raw images. By decoding, segmenting, and re-encoding endoscopic images, the negative impact of image format differences on model recognition can be removed, and interference information other than image information can be eliminated / reduced.
[0022] In some specific implementations, the image semantic segmentation network is constructed based on Unet or an improved model thereof.
[0023] In some specific implementations, the size of the endoscopic image is about 1200×1000, and the endoscopic image is reduced to 360×360 by regional interpolation to ensure real-time image recognition.
[0024] In some optional embodiments, the training process of the endoscopic intestinal mucosal feature recognition model includes: Acquiring training endoscopic images in a withdrawn observation state; Using expert experience to perform a graded assessment of the vascular status information, bleeding status information, and mucosal ulcer status reflected in the training endoscopic images, and determine a graded assessment result; determining a corresponding activity score of the training endoscopic image according to the graded evaluation results of the vascular status information, the bleeding status information, and the ulcer status information; After data enhancement is performed on the training endoscopic images, an endoscopic image training sample set is established based on the training endoscopic images and the activity scores; The initial endoscopic intestinal mucosal feature recognition model is constructed based on ResNet, and iterative training is performed on the endoscopic image training sample set.
[0025] As a non-limiting example, the data enhancement includes but is not limited to pre-processing the training endoscopic images such as random erasing, rotation, and flipping.
[0026] In some specific implementation processes, three expert physicians performed graded assessments of the mucosa based on the endoscopic features proposed by the Ulcerative Colitis Endoscopic Index of Severity (UCEIS) and the Mayo Endoscopic Activity Score (MES).
[0027] In some optional embodiments, the endoscopic intestinal mucosal feature recognition module is further used to mark the activity of the effectively observed intestinal segment with a preset color on the colonoscopic schematic diagram after determining the intestinal mucosal features, and output the IBD endoscopic activity visualization results.
[0028] It can be understood that based on the visualization results of IBD endoscopic activity, it is possible to retrieve several endoscopic image data associated with a specific patient, and then generate colonoscopic activity follow-up comparison results to evaluate the temporal changes in the colonoscopic manifestations of a specific patient.
[0029] Furthermore, in order to make full use of the temporal changes in the colonoscopic manifestations of specific patients, the LSTM (Long Short-Term Memory) model can be used for feature extraction.
[0030] In some preferred embodiments, determining pathological features based on the tissue WSI image using the pathology recognition model includes: Cutting the tissue WSI image into pixel blocks of a specified pixel size, and performing background filtering on the pixel blocks using the Otsu algorithm; Each group of pixel blocks corresponding to the tissue WSI image after background filtering is used as input of the pathology recognition model to determine the pathological features of the tissue WSI image.
[0031] In some optional embodiments, referring to FIG3 , the training process of the pathology recognition model includes: Obtain WSI (Whole Slide Image) images of IBD intestinal tissue after formalin fixation, paraffin embedding, and HE staining, perform image segmentation and background filtering, and obtain inflammatory bowel disease pathology pixel blocks for training; For each of the colon tissue WSI images, selecting at least a portion of the training inflammatory bowel disease pathology pixel blocks, annotating each of the training inflammatory bowel disease tissue pathology image pixel blocks with pathological activity, tissue structure, morphological information, and therapeutic effect type, to construct a pathology dataset; The pathology dataset is divided into A subset of pathological data, The pathological data subsets are used as training sets, and the remaining pathological data subset is used as a test set; Based on multi-layer perceptron An initial pathology recognition model, The pathology recognition model is The training sets are used to train the model parameters respectively; wherein the training inflammatory bowel disease pathology pixel blocks are used as the input of the pathology recognition model, and the efficacy outcome type is used as the output label of the pathology recognition model; When the average correlation of each tissue structure, pathological activity, morphological information, and therapeutic effect calculated by the pathology recognition model on a validation set containing at least a portion of the training data of the corresponding training set does not improve over multiple consecutive cycles, training is stopped; Verify on the test set The stability of the pathology recognition models is evaluated, and the model with the best performance is selected as the pathology recognition model for online use.
[0032] It should be understood that due to the huge pixel size of WSIs, the WSI images were cut and the Otsu algorithm was used to filter the white background of the image to filter out the background information that does not need to be analyzed.
[0033] In some specific implementations, the WSIs image is divided into 224×224 pixel blocks, and 10,000 small pixel blocks are selected from each tissue WSI image.
[0034] In some specific implementations, a pathology recognition model is constructed using Python, based on the TensorFlow deep learning framework. Preprocessed intestinal pathology omics images (WSI images) are randomly divided into five parts, four of which are randomly selected as training sets, and the remaining one as a test set. Five combinations are obtained, and models are constructed for each of these combinations to verify model stability and select the optimal model. Training is terminated when the average correlation between efficacy types calculated on a validation set containing 10% of the training data does not improve over 50 consecutive cycles. After training, the constructed pathology recognition model outputs a corresponding confidence level for each segmented small slice, representing the feature evaluation and recognition result for that small slice.
[0035] In some preferred embodiments, the training process of the image feature recognition model includes: Obtain training imaging images and resample them to the specified spatial resolution based on linear interpolation; According to the image pixels, the resampled training imaging image is subjected to Z-score normalization to determine the normalized imaging image corresponding to the training imaging image. The process is expressed as follows: Where, represents the original pixel value, represents the mean pixel value, represents the standard deviation of pixel values, Represents the normalized pixel value; Using expert experience and based on the minimum irregular curve principle, the standardized imaging image corresponding to the training imaging image is annotated with the lesion region (i.e., volume of interest (VOI) annotation), and the lesion region is three-dimensionally reconstructed based on the annotated result to form the corresponding training imaging VOI; An initial image feature recognition model is constructed based on a 3D CNN (3D Convolutional Neural Network), the training imaging image is used as the input of the image feature recognition model, the corresponding training image VOI is used as the output label of the image feature recognition model, and the image feature recognition model is iteratively trained.
[0036] It should be understood that for three-dimensional MRI (Magnetic Resonance Imaging) data, the trained image feature recognition model based on 3D CNN can better capture the spatial information of three-dimensional lesion data, and directly perform convolution operations on the volume data through the three-dimensional convolution kernel to achieve high-level feature extraction of the lesions.
[0037] In some specific implementations, an IBD imaging expert physician first determines the pixel intensities of a set of training images as reference points for standardization, converts the pixel intensities of the images into Z-scores to achieve standardization, calculates the Z-scores of the training images, and then applies these reference points to the standardization of new images.
[0038] In some optional embodiments, determining the imaging features includes: receiving the imaging information and resampling it to a specified spatial resolution; wherein the imaging information includes three-dimensional image data; The resampled imaging information is Z-score normalized and used as the input of the trained imaging feature recognition model. High-level features are extracted from the imaging information as the imaging features through a three-dimensional convolution kernel.
[0039] As a non-limiting example, the training imaging images and / or three-dimensional image data may be magnetic resonance images.
[0040] It should be noted that due to differences in the models, brands, and parameters of the equipment used to acquire the radiological and / or three-dimensional image data used for training, especially differences in the spatial resolution of magnetic resonance images, the learning results and recognition accuracy of the image feature recognition model may be poor. This embodiment uses a resampling method to ensure that the model can learn consistent spatial information from different data sources.
[0041] In some implementations, the training imaging images are resampled to Spatial resolution, forming a new image sequence.
[0042] In some preferred embodiments, the training process of the text feature extraction model includes: Acquiring training text information; wherein the training text information includes training clinical information, training laboratory examination information, and training demographic information; Utilizing expert experience to construct a text feature extraction regular library as the text feature extraction model, and configuring corresponding regular expressions; Vectorizing the training text information based on the text feature extraction model to obtain a training text feature vector; The training text feature vector is evaluated using expert experience, and the text feature extraction model is optimized according to the evaluation result of the training text feature vector.
[0043] In some specific implementations, three expert physicians from the IBD Diagnosis and Treatment Alliance jointly determined possible descriptions for each text's prognostic features (such as medication type, dosage, stool condition, erythrocyte sedimentation rate, C-reactive protein, and X-ray findings of intestinal shortening). This approach then constructed a regular expression library for text feature extraction. The raw clinical information in the database was automatically vectorized using regular expressions, and the extracted vector labels were reviewed by the three expert physicians. After the expert review, the accuracy of the vector labels extracted by the text feature extraction model was evaluated based on the reviewed clinical information using metrics such as precision, recall, and F1 score, and the regularization model was continuously optimized.
[0044] In some specific implementations, the text feature vector includes clinical information features (Based on clinical information extraction) and laboratory test characteristics (Based on laboratory test information extraction).
[0045] In some preferred embodiments, determining the predicted result of the biological agent efficacy includes: splicing at least two of the text feature vector, the intestinal mucosal feature, the imaging feature, and the pathological feature to obtain a multimodal feature; The multimodal features are used as the input of the efficacy prediction module, the self-attention mechanism of the Transformer model is used to extract the multimodal high-level features, the multimodal high-level features are mapped to the biological agent efficacy prediction results through multiple fully connected layers, and the corresponding prediction probability is output using the Sigmund activation function.
[0046] It should be noted that this embodiment employs a multimodal feature pre-fusion strategy to promote the integrity of feature learning and facilitate the full fusion of information and exploration of relationships between modalities. Furthermore, the use of vector concatenation facilitates the redistribution of computing resources, ensuring the immediacy and effectiveness of prediction output.
[0047] In some specific implementations, in the efficacy prediction module, the clinical information features , laboratory examination characteristics , intestinal mucosal characteristics , imaging features and pathological characteristics Vectors are spliced to form multimodal features , as the input of the Transformer model, the correlation and complementarity between different modal data are learned based on the self-attention mechanism of the Transformer model.
[0048] More specifically, the Transformer model includes multiple Transformer encoding layers.
[0049] In some specific implementation processes, the efficacy prediction module visualizes the biological agent efficacy prediction results in different colors and texts based on the prediction probability, including but not limited to: a red warning box to indicate that the efficacy prediction for a specific patient is "poor", a yellow prompt box to indicate that the efficacy prediction for a specific patient is "pending", and green text to indicate that the efficacy prediction for a specific patient is "good".
[0050] In some specific implementation processes, the weights of each eigenvector will be fed back to users (such as medical staff) in the form of numerical values or bar graphs, so that users can view the efficacy prediction ideas and provide a reference for users to make information evaluation.
[0051] In some examples, a computer program is provided, comprising a computer-readable code. When the computer-readable code is run in a computer device, a processor in the computer device executes the code to implement part or all of the modules / methods in the system.
[0052] This embodiment provides a computer-readable storage medium, on which is stored at least one instruction, at least one program, code set, or instruction set. The at least one instruction, at least one program, code set, or instruction set is loaded and executed by a processor, so that the processor executes some or all of the modules / methods in the system provided in the embodiments of the present application.
[0053] It is understood that the storage medium may be transient or non-transient. Exemplarily, the storage medium includes, but is not limited to, a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, among other media capable of storing program code.
[0054] Exemplarily, the processor may be a central processing unit (CPU), a microprocessor (MPU), a digital signal processor (DSP), or a field programmable gate array (FPGA).
[0055] Exemplarily, the read-only memory includes but is not limited to MASK ROM, PROM, EPROM, EEPROM, Flash, etc.
[0056] Exemplarily, the random access memory includes but is not limited to DRAM, SRAM, SDRAM, DDR SDRAM, etc.
[0057] In some examples, a computer program product is provided, which can be implemented in hardware, software, or a combination thereof. As a non-limiting example, the computer program product can be embodied as the storage medium, or as a software product, such as an SDK (Software Development Kit).
[0058] As a non-limiting example, a computer program product is provided, comprising a computer program or computer-executable instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer program or computer-executable instructions from the computer-readable storage medium and executes the computer-executable instructions, causing the electronic device to perform some or all of the modules / methods in the system described in the embodiments of the present application.
[0059] This embodiment also proposes an electronic device, including a memory and a processor, wherein the memory stores at least one instruction, at least one program, code set or instruction set, and when the processor executes the at least one instruction, at least one program, code set or instruction set, it implements part or all of the modules / methods in the system described in the embodiment.
[0060] In some examples, a hardware entity of the electronic device is provided, see Figure 4, including: a processor, a memory and a communication interface; wherein the processor generally controls the overall operation of the electronic device; the communication interface is used to enable the electronic device to communicate with other terminals or servers through a network; the memory is configured to store instructions and applications executable by the processor, and can also cache data to be processed or processed by the processor and various modules in the electronic device (including but not limited to image data, audio data, voice communication data and video communication data), and can be implemented by flash memory (FLASH), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM) or random access memory (RAM).
[0061] A processor may include one or more processing elements. Thus, a processor may include one or more integrated circuits (ICs) configured to perform the functions of the processor. Furthermore, each integrated circuit may include circuits (e.g., a first circuit, a second circuit, and other circuits) configured to perform the functions of the processor.
[0062] Furthermore, data may be transmitted between the processor, the communication interface and the memory via a bus, which may include any number of interconnected buses and bridges, connecting various circuits of one or more processors and memories.
[0063] The same or similar reference numerals correspond to the same or similar components; The terms used in the drawings to describe positional relationships are for illustrative purposes only and are not to be construed as limiting the present application. It should be noted that, unless there is any conflict, the embodiments and features in the embodiments of this application can be combined with each other.
[0064] In different specific implementations, the method or system described in this application can be implemented in software, hardware or a combination thereof. In addition, the order of the steps of the method can be changed, and various elements can be added, reordered, combined, omitted, modified, etc.
[0065] Obviously, the above embodiments of the present application are merely examples for clearly illustrating the present application, and are not intended to limit the implementation methods of the present application, and are not intended to limit the present application. For those skilled in the art, other different forms of changes or modifications can be made based on the above description. Each discrete structural / functional module or unit can be integrated together to form an independent part, or each module can exist alone, or two or more modules can be integrated to form an independent part, and the structure and function of the discrete components can be implemented as a combined structure or component. It is not necessary and impossible to enumerate all the implementation methods here. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present application should be included in the scope of protection of the claims of the present application.
Claims
1. A system for predicting the efficacy of inflammatory bowel disease based on multimodal data, characterized in that: include: A multimodal data acquisition module, configured to receive multimodal data, wherein the multimodal data includes at least two of clinical information, demographic information, laboratory examination information, imaging information, tissue WSI images, and endoscopic images; a text feature extraction module, configured to carry a trained text feature extraction model and utilize the text feature extraction model to determine a text feature vector based on at least one of the clinical information, the demographic information, and the laboratory examination information; An endoscopic intestinal mucosal feature recognition module, configured to carry a trained endoscopic intestinal mucosal feature recognition model and utilize the model to determine intestinal mucosal features based on the endoscopic image; wherein the intestinal mucosal feature map effectively observes the clinical disease activity of the intestinal segment and the corresponding distribution area information; An image feature recognition module, configured to carry a trained image feature recognition model and utilize the image feature recognition model to extract features based on the imaging information to determine imaging features; wherein the imaging features map lesion area information; A pathology feature extraction module, configured to carry a trained pathology recognition model and utilize the pathology recognition model to determine pathology features based on the tissue WSI image; wherein the pathology features map pathology activity, tissue structure, morphology information, and therapeutic effect type; The efficacy prediction module is used to carry the Transformer model and the fully connected layer, and determine the efficacy prediction result of the biological agent based on at least two of the text feature vector, the intestinal mucosal features, the imaging features and the pathological features.
2. The inflammatory bowel disease efficacy prediction system based on multimodal data according to claim 1, characterized in that: Determining the intestinal mucosal characteristics based on the endoscopic image includes: Based on the Farneback dense optical flow algorithm, the forward optical flow vector and the reverse optical flow vector between adjacent frames of the endoscopic image are calculated; Performing reverse verification based on the reverse optical flow vector to eliminate interference optical flow, determining the optical flow residual vector, and thus calculating the mirror withdrawal observation speed; Convert the optical flow residual vector into HSV channel value and construct the corresponding HSV image; Using a recurrent neural network to identify the continuous HSV images, determine the state of the endoscopic image, and eliminate segments belonging to a non-observation state to determine the endoscopic image observation segment; The endoscopic intestinal mucosal feature recognition model is used to determine the IBD endoscopic activity based on the endoscopic image observation segment, and the IBD endoscopic activity is weighted with the endoscopic observation withdrawal speed to determine the intestinal mucosal features that reflect the activity distribution of different effective observation intestinal segments in the endoscopic image observation segment and the corresponding mucosal proportion.
3. The inflammatory bowel disease efficacy prediction system based on multimodal data according to claim 2, characterized in that: Before determining the optical flow residual vector, the endoscopic image is subjected to a first preprocessing step, including: Decoding the endoscopic image into a frame image, performing image segmentation on the frame image using an image semantic segmentation network, removing interference information in the frame image, retaining the mucosal portion of the frame image, and obtaining a segmented frame image; The segmented frame image is scaled based on a regional interpolation method, and the scaled segmented frame image is re-encoded into the endoscopic image.
4. The inflammatory bowel disease efficacy prediction system based on multimodal data according to claim 2, characterized in that: The training process of the endoscopic intestinal mucosal feature recognition model includes: Acquiring training endoscopic images in a withdrawn observation state; Using expert experience to perform a graded assessment of the vascular status information, bleeding status information, and mucosal ulcer status reflected in the training endoscopic images, and determine a graded assessment result; determining a corresponding activity score of the training endoscopic image according to the graded evaluation results of the vascular state information, the bleeding state information, and the ulcer state information; After data enhancement is performed on the training endoscopic images, an endoscopic image training sample set is established based on the training endoscopic images and the activity scores; The initial endoscopic intestinal mucosal feature recognition model is constructed based on ResNet, and iterative training is performed on the endoscopic image training sample set.
5. The inflammatory bowel disease efficacy prediction system based on multimodal data according to claim 1, characterized in that: Determining pathological features according to the tissue WSI image using the pathology recognition model includes: Cutting the tissue WSI image into pixel blocks of a specified pixel size, and performing background filtering on the pixel blocks using the Otsu algorithm; Each group of pixel blocks corresponding to the tissue WSI image after background filtering is used as input of the pathology recognition model to determine the pathological features of the tissue WSI image.
6. The inflammatory bowel disease efficacy prediction system based on multimodal data according to claim 5, characterized in that: The training process of the pathology recognition model includes: Obtain WSI images of intestinal tissue after formalin fixation, paraffin embedding, and HE staining, perform image segmentation and background filtering, and obtain inflammatory bowel disease pathology pixel blocks for training; For each of the intestinal tissue WSI images, selecting at least a portion of the training inflammatory bowel disease pathology pixel blocks, annotating each of the training inflammatory bowel disease pathology pixel blocks with pathological activity, tissue structure, morphological information, and therapeutic effect, and constructing a pathology dataset; The pathology dataset is divided into A subset of pathological data, The pathological data subsets are used as training sets, and the remaining pathological data subset is used as a test set; Based on multi-layer perceptron An initial pathology recognition model, The pathology recognition model is The training sets are used to train the model parameters respectively; wherein the training inflammatory bowel disease pathology pixel blocks are used as the input of the pathology recognition model, and the efficacy outcome type is used as the output label of the pathology recognition model; When the average correlation of each tissue structure, pathological activity, morphological information, and therapeutic effect calculated by the pathology recognition model on a validation set containing at least a portion of the training data of the corresponding training set does not improve over multiple consecutive cycles, training is stopped; Verify on the test set The stability of the pathology recognition models is evaluated, and the model with the best performance is selected as the pathology recognition model for online use.
7. The inflammatory bowel disease efficacy prediction system based on multimodal data according to claim 1, characterized in that: The training process of the image feature recognition model includes: Obtain training imaging images and resample them to the specified spatial resolution based on linear interpolation; According to the image pixels, the resampled training imaging image is subjected to Z-score normalization to determine the normalized imaging image corresponding to the training imaging image. The process is expressed as follows: Where, represents the original pixel value, represents the mean pixel value, represents the standard deviation of pixel values, Represents the normalized pixel value; Using expert experience, based on the minimum irregular curve principle, the standardized imaging image corresponding to the training imaging image is annotated with the lesion area, and the lesion area is three-dimensionally reconstructed according to the annotated result to form the corresponding training imaging VOI; An initial image feature recognition model is constructed based on 3D CNN, the training imaging image is used as the input of the image feature recognition model, the corresponding training image VOI is used as the output label of the image feature recognition model, and the image feature recognition model is iteratively trained.
8. The inflammatory bowel disease efficacy prediction system based on multimodal data according to claim 7, characterized in that: The determination of imaging features includes: receiving the imaging information and resampling it to a specified spatial resolution; wherein the imaging information includes three-dimensional image data; The resampled imaging information is Z-score normalized and used as the input of the trained imaging feature recognition model. High-level features are extracted from the imaging information as the imaging features through a three-dimensional convolution kernel.
9. The inflammatory bowel disease efficacy prediction system based on multimodal data according to claim 1, characterized in that: The training process of the text feature extraction model includes: Acquiring training text information; wherein the training text information includes training clinical information, training laboratory examination information, and training demographic information; Utilizing expert experience to construct a text feature extraction regular library as the text feature extraction model, and configuring corresponding regular expressions; Vectorizing the training text information based on the text feature extraction model to obtain a training text feature vector; The training text feature vector is evaluated using expert experience, and the text feature extraction model is optimized according to the evaluation result of the training text feature vector.
10. The inflammatory bowel disease efficacy prediction system based on multimodal data according to any one of claims 1 to 9, characterized in that: The method of determining the predicted results of biological agent efficacy includes: splicing at least two of the text feature vector, the intestinal mucosal feature, the imaging feature, and the pathological feature to obtain a multimodal feature; The multimodal features are used as the input of the efficacy prediction module, the self-attention mechanism of the Transformer model is used to extract the multimodal high-level features, the multimodal high-level features are mapped to the biological agent efficacy prediction results through multiple fully connected layers, and the corresponding prediction probability is output using the Sigmund activation function.
Citation Information
Patent Citations
Artificial intelligence auxiliary system for digestive endoscopic images of inflammatory bowel diseases
CN111524124A
Multi-modal prediction model construction method and system for analyzing liver cancer recurrence data
CN117612711A
Transformer substation safe operation monitoring method and system based on OpenCV
CN118485973A
Disease and pest detection method and system based on dynamic scene
CN120107234A