A method, device, medium, and product for predicting ulcerative colitis relapse
Patent Information
- Application Number
- CN202610986725.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-03
- Publication Date
- 2026-09-25
AI Technical Summary
现有的多模态融合方法多采用在同一时间截面进行静态特征拼接或简单的早期/晚期融合策略,这种静态对齐方式难以反映不同模态之间的时间错位关系与动态演变依赖,限制了对疾病复发全过程的准确建模与预测
本申请提供了一种溃疡性结肠炎复发预测方法、设备、介质及产品,分别对临床文本数据、结构化指标数据和医学图像数据进行独立的时序建模,能够针对不同模态数据的固有特性(如文本的语义性、指标的连续性、图像的空间性)提取具有时间依赖性的特征表示,有效克服了传统基于静态数据或固定窗口建模方法容易忽略关键历史信息或平滑短期病情变化的缺陷,更真实地反映了疾病随时间推移的演化过程。通过时间轴对齐和特征空间对齐,消除了不同模态因采样频率、观测时间不一致以及特征维度差异带来的融合障碍;随后,引入跨模态注意力机制进行动态融合,能够在不同的时间阶段,自适应地学习和调整各模态特征的贡献权重,从而有效捕捉疾病复发过程中(如从生化异常到影像改变再到症状加重)模态间的互补信息与时序依赖关系,构建出高度一致且信息丰富的患者动态表征。通过引入时间感知模型进行时序建模,并结合全局池化操作,不仅完整保留了多模态时序特征的动态关联,还能自适应地聚焦于与疾病复发高度相关的关键时间步,过滤冗余噪声。最终通过线性投影将复杂的全局时序表征精准映射为疾病复发状态的预测结果,实现了对长周期、动态演变特征的精准建模。本申请通过构建“多模态时序表征-跨模态时序融合-疾病状态预测”的三层架构,能够有效处理真实世界中医疗数据的非均匀采样与高度异构性问题,实现多模态动态时序对齐与深度融合,从而为临床长期疾病管理提供精准、可靠的数据预测支持。
Smart Images

Figure CN122822334A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence, and in particular to a method, device, medium, and product for predicting the recurrence of ulcerative colitis. Background Technology
[0002] Ulcerative colitis (UC) is an inflammatory bowel disease characterized by chronic inflammation of the colonic mucosa, with a typical course of alternating relapses and remissions. Dynamic monitoring of disease activity and relapse prediction are crucial for long-term clinical management. Because UC relapse risk is influenced by multiple factors, related clinical data are typically multimodal (e.g., clinical text in electronic medical records, structured indicators from laboratory tests, and medical images from endoscopy and pathology), strongly time-dependent, and highly heterogeneous. Therefore, effectively processing and analyzing such multimodal time-series data to achieve accurate relapse risk prediction is an important research direction in the field of medical artificial intelligence data processing.
[0003] In recent years, deep learning-based data processing methods have made some progress in multimodal information fusion and disease state prediction tasks. However, in relapse prediction tasks such as ulcerative colitis, which have long-term and dynamically evolving characteristics, existing multimodal time-series data processing methods still face the following technical bottlenecks: First, existing methods struggle to effectively handle the irregularities and sparsity of time-series data. Real-world long-term clinical follow-up data spans several years, and due to the influence of patient behavior and disease fluctuations, data collection often exhibits significant temporal irregularities and sparsity. Existing time-series modeling methods mostly rely on regular time intervals or fixed sliding windows for feature extraction and modeling. This approach is ill-suited to accurately depicting the non-uniformly sampled disease evolution processes in the real world, easily leading to the neglect of crucial historical time-series information or the over-smoothing of short-term dramatic changes in disease condition, thereby reducing the accuracy of time-series feature representation.
[0004] Secondly, existing methods struggle to address the issues of temporal heterogeneity and asynchronous evolution across multiple modalities. In clinical practice, the sampling frequency and clinical significance of data from different modalities vary significantly, and disease recurrence often exhibits a progressive evolution from abnormal biochemical indicators to changes in imaging morphology, and finally to worsening clinical symptoms. Existing multimodal fusion methods mostly employ static feature stitching at the same time point or simple early / late fusion strategies. This static alignment approach fails to reflect the temporal misalignment and dynamic evolutionary dependencies between different modalities, limiting the accurate modeling and prediction of the entire disease recurrence process.
[0005] In summary, to address the shortcomings of existing technologies, there is an urgent need for an intelligent prediction method for UC recurrence to improve the accuracy and robustness of UC recurrence prediction. Summary of the Invention
[0006] The purpose of this application is to provide a method, device, medium, and product for predicting the recurrence of ulcerative colitis, which can improve the accuracy and robustness of predicting the recurrence of ulcerative colitis.
[0007] To achieve the above objectives, this application provides the following solution: In a first aspect, this application provides a method for predicting the recurrence of ulcerative colitis, including: Acquire multimodal time-series data of patients; the multimodal time-series data includes clinical text data, structured index data, and medical image data; Temporal modeling is performed on multimodal temporal data respectively, and corresponding temporal feature representations are extracted; the temporal feature representations include: text temporal embedding, structured variable temporal embedding sequence, and image temporal embedding sequence; The extracted temporal feature representations are aligned in time axis and feature space; and a dynamic cross-modal attention mechanism is used for dynamic fusion to obtain the fused multimodal temporal embedding sequence. Based on the fused multimodal temporal embedding sequence, temporal modeling and global pooling are performed using a time-aware model to obtain the prediction results.
[0008] Secondly, this application provides a device for predicting the recurrence of ulcerative colitis, comprising: The data acquisition layer is used to acquire patients' multimodal time-series data, which includes clinical text data, structured indicator data, and medical image data. The multimodal temporal representation layer is used to perform temporal modeling on multimodal temporal data and extract corresponding temporal feature representations; the temporal feature representations include: text temporal embedding, structured variable temporal embedding sequence, and image temporal embedding sequence; A cross-modal temporal fusion layer is used to align the extracted temporal feature representations in terms of time axis and feature space; and a dynamic cross-modal attention mechanism is used for dynamic fusion to obtain the fused multimodal temporal embedding sequence. The disease state prediction layer is used to obtain prediction results by performing temporal modeling and global pooling based on the fused multimodal temporal embedded sequence through a time-aware model.
[0009] Thirdly, this application provides a computer device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the ulcerative colitis recurrence prediction method.
[0010] Fourthly, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the described method for predicting recurrence of ulcerative colitis.
[0011] Fifthly, this application provides a computer program product, including a computer program that, when executed by a processor, implements the described method for predicting recurrence of ulcerative colitis.
[0012] According to the specific embodiments provided in this application, this application has the following technical effects: This application provides a method, device, medium, and product for predicting recurrence of ulcerative colitis. It independently performs temporal modeling on clinical text data, structured indicator data, and medical image data, extracting time-dependent feature representations based on the inherent characteristics of different data modalities (such as the semantic nature of text, the continuity of indicators, and the spatial nature of images). This effectively overcomes the shortcomings of traditional static data-based or fixed-window modeling methods, which tend to ignore key historical information or smooth short-term disease changes, thus more realistically reflecting the disease's evolution over time. By aligning the time axis and feature space, it eliminates fusion barriers caused by inconsistencies in sampling frequency, observation time, and feature dimension differences between different modalities. Subsequently, a cross-modal attention mechanism is introduced for dynamic fusion, which adaptively learns and adjusts the contribution weights of each modality's features at different time stages. This effectively captures complementary information and temporal dependencies between modalities during the disease recurrence process (such as from biochemical abnormalities to imaging changes to symptom exacerbation), constructing a highly consistent and information-rich dynamic patient representation. By introducing a time-aware model for temporal modeling and combining it with global pooling, this approach not only fully preserves the dynamic correlation of multimodal temporal features but also adaptively focuses on key time steps highly correlated with disease recurrence, filtering out redundant noise. Finally, linear projection accurately maps the complex global temporal representation to the predicted disease recurrence state, achieving precise modeling of long-term, dynamically evolving characteristics. This application constructs a three-layer architecture of "multimodal temporal representation - cross-modal temporal fusion - disease state prediction," effectively addressing the non-uniform sampling and high heterogeneity of real-world medical data. It achieves multimodal dynamic temporal alignment and deep fusion, thereby providing accurate and reliable data prediction support for long-term clinical disease management. Attached Figure Description
[0013] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0014] Figure 1 This is a schematic flowchart of a method for predicting recurrence of ulcerative colitis in one embodiment of this application; Figure 2 This is a schematic diagram of the architecture of a method for predicting the recurrence of ulcerative colitis in one embodiment of this application. Detailed Implementation
[0015] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0016] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0017] In one exemplary embodiment, such as Figure 1 and Figure 2 As shown, a method for predicting recurrence of ulcerative colitis is provided. This method is executed by a computer device, specifically by a terminal or server alone, or by both a terminal and a server. In this embodiment, the method is applied to a server as an example, and includes the following steps S101 to S104. Wherein: S101, acquire the patient's multimodal time-series data; the multimodal time-series data includes clinical text data, structured indicator data, and medical image data; S102, perform time series modeling on the multimodal time series data respectively, and extract the corresponding time series feature representations; the time series feature representations include: text time series embedding, structured variable time series embedding sequence and image time series embedding sequence; S102 specifically includes: S21, semantic features are extracted from the clinical text data using a pre-trained language model, and a self-attention mechanism is used to simultaneously fuse the semantic features of multiple texts at the same time point. At the same time, temporal position encoding and gating mechanisms are introduced to model cross-temporal dependencies, resulting in text temporal embeddings; and the text temporal embeddings are used as text temporal features. The pre-trained language model is trained based on unstructured text (such as chief complaint, present medical history, follow-up records, etc.) in electronic medical records; preferably, the pre-trained language model is a clinical text pre-trained model (BioClinicalBERT). The process of determining text temporal embedding is as follows: (1) For each time step t∈{1,2,...,N} (N is the time step), this time point contains a set of multiple types of clinical texts Xt={x t1 ,x t2 ,...,x tk (k is the number of text types). A clinical text pre-trained model is used to independently extract semantic features for each text type. The input is a single-type text x. ti The output is the semantic embedding of this type of text. ti .
[0018] (2) A unified representation is obtained by fusing multiple text embeddings at the same time point through a self-attention mechanism. : ;in, : Real number field (representing features consisting of real numbers); : Feature dimension after self-attention fusion (model hyperparameter, a fixed value set manually, such as 256 / 512 / 768). : Represents the unified features obtained through fusion It is A dimensional real vector; (3) Constructing a temporal sequence of text semantic features Adding timing position coding P yields Modeling cross-temporal dependencies through temporal attention : ;in, This is a temporal sequence of text semantic features after incorporating temporal position encoding; (4) Introduce a gating mechanism: ;in, The gating coefficient vector at time step t (controlling the fusion weights of historical time series features); The Sigmoid activation function (maps values to the [0,1] interval to achieve a gating effect); It is a learnable gating weight matrix; For learnable gated bias vectors; (5) The text features after gating and fusion are: ;in, Let be the gated fused text features at time step t. This is the Hadamard product (element-level multiplication, where corresponding values are multiplied). The temporal feature vector at time step t is the output of the temporal attention. (6) Text temporal embedding is as follows: Where N is the total number of time steps (the maximum number of times a patient is followed up). For N rows A real matrix of columns (each row corresponds to a text semantic feature at a time step). For text temporal feature dimensions (model hyperparameters); S22, Based on the structured index data, a long-sequence time-series prediction model is used to capture long-distance time dependencies, learn the dynamic pattern of index changes over time, and obtain a structured variable time-series embedding sequence; and the structured variable time-series embedding sequence is used as a form time-series feature; For structured data such as laboratory test indicators: Continuous variables are standardized; a long-sequence modeling network is constructed to capture long-distance time dependencies; and dynamic patterns of indicator changes over time are learned. The Informer long-sequence time series prediction model is used as the core time series modeling model to specifically handle the long-sequence dependencies of structured variables.
[0019] Specifically, input structured indicator data, denoted as time series. , This represents the set of laboratory indicators at time step t. Output a structured variable time-series embedding sequence. .in, The dimensions of single-step structured indicators (such as the total number of blood routine and biochemical indicators). Structured time series data can be represented as an N-row / (d_{form}\) real matrix.
[0020] S23, Based on the medical image data, spatial pathological features are extracted using a visual transformer model, temporal position coding is introduced to construct a temporal structure, and the pathological evolution relationship is captured through a cross-temporal attention mechanism to obtain an image temporal embedding sequence; and the image temporal embedding sequence is used as the image temporal feature. S23 targets medical images (endoscopy, ultrasound, pathology) acquired at different time points; specifically, the visual transformer model is the Medical Vision Transformer (MedViT); the visual transformer model is used as the core spatial feature extraction model to specifically capture spatial pathological features in medical images.
[0021] The medical image data at time step t is input using a medical vision transformer. (H and W represent the image height and width, respectively, and 3 represents the number of RGB channels); output the spatial pathological feature embedding of the corresponding image. : .in, The dimension embedded for spatial pathological features; By introducing a temporal attention mechanism, the evolution of lesion morphology is captured, and an image temporal sequence is constructed: the spatial features of a single image at N time steps are combined. By integrating, the basic image time series is obtained. : ; Adding temporal position embedding: Introducing temporal position coding The temporally enhanced image embedding sequence is obtained as follows: ; By capturing the pathological evolution relationships of medical images at different time steps through cross-temporal attention, the formula is the same as that for the text modality: ; Obtain image temporal embedding sequence : ; S103, the extracted temporal feature representations are aligned in time axis and feature space; and a dynamic cross-modal attention mechanism is used for dynamic fusion to obtain a fused multimodal temporal embedding sequence; thus, multimodal temporal features can be aligned and collaboratively modeled in a unified time frame to construct a consistent dynamic representation of the patient. S103 specifically includes: S31, the extracted temporal feature representation is mapped to a unified time axis using a timestamp-based attention pooling or interpolation strategy to obtain an equal-length temporal representation; To avoid significant differences in sampling frequency and observation time among different modes, each mode is first mapped to a unified time axis. Let the shared time series be... For the original temporal embeddings of each modality: (1) Text temporal embedding: ; (2) Structured variable temporal embedding sequence: ; (3) Image temporal embedding sequence: ; By using timestamp-based attention pooling or interpolation strategies, each modality is mapped to a unified time axis, resulting in an equal-length temporal representation: ; in, Equal-length time series representations corresponding to different modes; This process preserves the original time information while aligning the multimodal data in the time dimension, providing a foundation for subsequent fusion.
[0022] S32 performs linear projection of the equal-length time series representation of each mode using a learnable parameter matrix.
[0023] To avoid differences in representation dimension and distribution space between features from different modalities, they need to be mapped to a unified feature space before cross-modal fusion. Therefore, a linear projection is performed on the aligned temporal features of each modality: ; in, The learnable parameter matrix is projected, and the modal features are unified as follows: ; This step effectively alleviates the problem of inconsistent representations between modalities and provides a unified semantic space for cross-modal interaction.
[0024] S33, based on the aligned temporal feature representation of each time step, a cross-modal attention head is used to learn the weights of each modality; For the t-th time step, based on multimodal features Modal weights are learned through Temporal Cross-Modal Attention (TCMA): ; in, ; S34, based on the weights, adaptively model the importance of different modalities at each time step, dynamically adjust the weights of each modality, and obtain the fused multimodal temporal embedding sequence.
[0025] The final fused multimodal temporal embedding is obtained. : ; Calculations at all time steps yield a unified, fused multimodal temporal embedding sequence. : ; This mechanism can dynamically adjust the contribution of each modality at different time stages, thereby better capturing the complementary information and temporal dependencies between modalities during disease evolution.
[0026] S104, based on the fused multimodal temporal embedded sequence, uses a time-aware model for temporal modeling and global pooling to obtain the prediction result.
[0027] The time-aware model is a time-aware Transformer model. The time-aware Transformer model preserves the dynamic correlation of multimodal time series and can adaptively focus on key time steps related to recurrence.
[0028] The prediction process of the time-aware Transformer model is as follows: ①Location encoding: Add Time-aware timing position coding Relative temporal information of encoding time steps : ; ②Time series modeling: The input is a Time-aware Transformer encoder, which models the long-short-term dependencies across time steps through temporal self-attention, capturing dynamic patterns related to relapse in multimodal time series. Global average pooling is applied to the full-temporal latent features output by the Transformer to compress and obtain a single patient's global temporal representation. : ; ③ Recurrence prediction: By using linear projection and activation functions, global temporal features are... Mapped to relapse status prediction value : ; This leads to the prediction results of the patient's relapse status. ,in Predicted as a relapse, The prediction is that there will be no recurrence.
[0029] Based on the same inventive concept, this application also provides an ulcerative colitis recurrence prediction device for implementing the aforementioned ulcerative colitis recurrence prediction method. The solution provided by this device is similar to the solution described in the above method; therefore, the specific limitations of one or more embodiments of the ulcerative colitis recurrence prediction device provided below can be found in the limitations of the ulcerative colitis recurrence prediction method described above, and will not be repeated here.
[0030] In one exemplary embodiment, a device for predicting recurrence of ulcerative colitis is provided, comprising: The data acquisition layer is used to acquire patients' multimodal time-series data, which includes clinical text data, structured indicator data, and medical image data. The Multimodal Temporal Representation Layer is used to perform temporal modeling on multimodal temporal data and extract corresponding temporal feature representations. The temporal feature representations include: text temporal embeddings, structured variable temporal embedding sequences, and image temporal embedding sequences. The Temporal Multimodal Fusion Layer is used to align the extracted temporal feature representations in terms of time axis and feature space; and a dynamic cross-modal attention mechanism is used for dynamic fusion to obtain the fused multimodal temporal embedding sequence. The Disease State Prediction Layer is used to obtain prediction results based on the fused multimodal temporal embedded sequences, through temporal modeling and global pooling.
[0031] In one exemplary embodiment, a computer device is provided, which may be a server or a terminal. The computer device includes a processor, memory, input / output interfaces (I / O), and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is connected to the system bus via the I / O interfaces. The processor of the computer device provides computing and control capabilities. The memory of the computer device includes non-volatile storage media and internal memory. The non-volatile storage media stores an operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The I / O interfaces of the computer device are used for exchanging information between the processor and external devices. The communication interface of the computer device is used for communicating with external terminals via a network connection. When the computer program is executed by the processor, it implements a method for predicting recurrence of ulcerative colitis.
[0032] In one exemplary embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.
[0033] In one exemplary embodiment, a computer-readable storage medium is provided storing a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.
[0034] In one exemplary embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.
[0035] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.
[0036] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments described above. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM).
[0037] The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.
[0038] In this application, all actions to acquire signals, information, or data are carried out in compliance with the relevant data protection laws and policies of the country where the location is situated, and with the authorization granted by the owner of the relevant device.
[0039] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0040] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A method for predicting recurrence of ulcerative colitis, characterized in that, include: Acquire multimodal time-series data of patients; The multimodal time-series data includes clinical text data, structured indicator data, and medical image data; Temporal modeling is performed on multimodal time series data respectively, and corresponding temporal feature representations are extracted; The temporal feature representations include: text temporal embedding, structured variable temporal embedding sequences, and image temporal embedding sequences; The extracted temporal feature representations are aligned in time axis and feature space; and a dynamic cross-modal attention mechanism is used for dynamic fusion to obtain the fused multimodal temporal embedding sequence. Based on the fused multimodal temporal embedding sequence, temporal modeling and global pooling are performed using a time-aware model to obtain the prediction results.
2. The method for predicting recurrence of ulcerative colitis according to claim 1, characterized in that, The step of performing time series modeling on multimodal time series data and extracting corresponding time series feature representations specifically includes: Semantic features are extracted from the clinical text data using a pre-trained language model, and a self-attention mechanism is used to simultaneously fuse semantic features of multiple texts at the same time point. At the same time, temporal position encoding and gating mechanisms are introduced to model cross-temporal dependencies, resulting in text temporal embedding. Based on the structured index data, a long-sequence time-series prediction model is used to capture long-distance time dependencies, learn the dynamic pattern of index changes over time, and obtain a structured variable time-series embedding sequence. Based on the medical image data, spatial pathological features are extracted using a visual transformer model, temporal position encoding is introduced to construct a temporal structure, and the pathological evolution relationship is captured through a cross-temporal attention mechanism to obtain an image temporal embedding sequence.
3. The method for predicting recurrence of ulcerative colitis according to claim 2, characterized in that, The process involves extracting semantic features from the clinical text data using a pre-trained language model, simultaneously fusing semantic features from multiple text types at the same time point using a self-attention mechanism, and introducing temporal position encoding and gating mechanisms for cross-temporal dependency modeling to obtain text temporal embeddings. Specifically, this includes: Using formula Determine the temporal embedding of text ; Where N is the total number of time steps, For N rows A real matrix of columns, For text temporal features, Let be the gated fused text features at time step t. , For Hadama accumulation, Let be the temporal feature vector at time step t, output by the temporal attention. Let be the semantic features of the text at time step t. Let be the gating coefficient vector at time step t.
4. The method for predicting recurrence of ulcerative colitis according to claim 1, characterized in that, The step of aligning the extracted temporal feature representations in terms of time axis and feature space specifically includes: Attention pooling or interpolation strategies based on timestamps map the extracted temporal feature representations to a unified time axis, resulting in equal-length temporal representations. Linear projection is performed on the equal-length time series representations of each mode using a learnable parameter matrix.
5. The method for predicting recurrence of ulcerative colitis according to claim 1, characterized in that, The process employs a dynamic cross-modal attention mechanism for dynamic fusion to obtain a fused multimodal temporal embedding sequence, specifically including: Based on the aligned temporal feature representation at each time step, a cross-modal attention head is used to learn the weights of each modality; Based on the weights, the importance of different modalities at each time step is adaptively modeled, and the weights of each modality are dynamically adjusted to obtain the fused multimodal temporal embedding sequence.
6. The method for predicting recurrence of ulcerative colitis according to claim 1, characterized in that, The time-aware model is a time-aware Transformer model.
7. A device for predicting recurrence of ulcerative colitis, characterized in that, include: The data acquisition layer is used to acquire patients' multimodal time-series data; The multimodal time-series data includes clinical text data, structured indicator data, and medical image data; The multimodal temporal representation layer is used to perform temporal modeling on multimodal temporal data and extract corresponding temporal feature representations. The temporal feature representations include: text temporal embedding, structured variable temporal embedding sequences, and image temporal embedding sequences; A cross-modal temporal fusion layer is used to align the extracted temporal feature representations in terms of time axis and feature space; and a dynamic cross-modal attention mechanism is used for dynamic fusion to obtain the fused multimodal temporal embedding sequence. The disease state prediction layer is used to obtain prediction results by performing temporal modeling and global pooling based on the fused multimodal temporal embedded sequence through a time-aware model.
8. A computer device, comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor executes the computer program to implement the method for predicting recurrence of ulcerative colitis as described in any one of claims 1-6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the method for predicting recurrence of ulcerative colitis as described in any one of claims 1-6.
10. A computer program product, comprising a computer program, characterized in that, When executed by a processor, the computer program implements the method for predicting recurrence of ulcerative colitis as described in any one of claims 1-6.