A step-by-step multimodal data fusion system for early diagnosis and staging of cholangiocarcinoma

By using a step-by-step multimodal data fusion system, which extracts image and non-image features through a self-attention mechanism and a pre-trained network, an interpretable cholangiocarcinoma diagnostic report is generated. This solves the problems of insufficient multimodal information fusion and uninterpretable AI models in existing technologies, and achieves highly accurate and robust early diagnosis and staging of cholangiocarcinoma.

CN122177411APending Publication Date: 2026-06-09BEILUN DISTRICT PEOPLES HOSPITAL OF NINGBO CITY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEILUN DISTRICT PEOPLES HOSPITAL OF NINGBO CITY
Filing Date
2026-02-26
Publication Date
2026-06-09

AI Technical Summary

Technical Problem

Existing technologies for the early diagnosis and staging of cholangiocarcinoma suffer from insufficient multimodal information fusion, and AI models lack interpretability and robustness, resulting in inaccurate diagnostic results and a lack of clinical trust.

Method used

A step-by-step multimodal data fusion system is adopted, including data preprocessing, feature extraction, feature fusion, sharing modules and distributed diagnostic framework. Image features are extracted through self-attention mechanism and pre-trained ResNet18 network, non-image features are extracted by combining multilayer perceptron network, and interpretable techniques such as Grad-CAM and SHAP are used to generate diagnostic reports.

Benefits of technology

It achieves highly accurate and interpretable early diagnosis and staging of cholangiocarcinoma, improving the reliability and robustness of diagnostic results and meeting clinical needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122177411A_ABST
    Figure CN122177411A_ABST
Patent Text Reader

Abstract

This invention provides a step-by-step multimodal data fusion system for the early diagnosis and staging of cholangiocarcinoma. It employs a self-attention mechanism to weightedly fuse high-level semantic feature vectors extracted from medical imaging data, clinical feature vectors extracted from clinical laboratory data, and biomarker feature vectors extracted from liquid biopsy biomarker data. This dynamically and deeply mines the nonlinear complementary and synergistic relationships among the three data sets, generating a more informative and discriminative global fusion feature vector. Furthermore, by identifying a distributed diagnostic framework for cholangiocarcinoma-staging-metastasis prediction, it simulates the progressive diagnostic logic of clinicians—first qualitative, then quantitative, and finally risk assessment—resulting in outputs that better meet clinical needs. Simultaneously, it combines a shared fully connected layer with a distributed framework composed of various classifiers, and utilizes step-by-step decision-making to isolate error propagation, thereby enhancing the system's overall anti-interference capability and generalization ability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of clinical laboratory diagnostic technology, and more specifically, to a stepwise multimodal data fusion system for the early diagnosis and staging of cholangiocarcinoma. Background Technology

[0002] Cholangiocarcinoma (CCA) is a malignant tumor originating from the epithelial cells of the bile ducts. It is classified into intrahepatic, extrahepatic, and extra-hepatic carcinomas and is considered a fatal malignant tumor of the hepatobiliary system. Most CCA cases are asymptomatic, characterized by high invasiveness and extremely poor prognosis. Furthermore, the bile ducts are located deep within the body, making early symptoms insidious or atypical. The vast majority of patients are diagnosed at an advanced stage, missing the optimal window for surgical treatment, resulting in a very low five-year survival rate. Currently, tumor markers CEA and CA19-9 lack sufficient specificity and accuracy, making it difficult to predict early-stage CCA. Moreover, current medical practice does not recommend routine examinations of CT, MRI, and PET / CT in asymptomatic individuals; ERCP and other invasive procedures, including percutaneous cholangiography and endoscopic ultrasound, are also not recommended. Therefore, early and accurate diagnosis and staging of CCA are crucial for improving patient prognosis and developing personalized treatment plans.

[0003] Artificial intelligence (AI) has achieved great success in various medical research fields, including radiology, pathology, gastroenterology, and ophthalmology. AI has significant advantages in analyzing massive amounts of data and identifying disease patterns and trends, such as histological, cytological, serum biomarker, and radiological methods. Currently, the clinical diagnosis and staging of cholangiocarcinoma mainly relies on the comprehensive assessment of multimodal information, but existing methods have significant limitations: (1) Existing AI models for auxiliary diagnosis are simple in terms of multimodal information fusion methods. They usually use simple feature vector splicing or post-decision fusion methods, which fail to deeply explore the correlation between modes. (2) Most existing deep learning-based AI systems are like "black boxes". Their decision-making process is not transparent, they cannot provide convincing judgment basis for clinicians, lack interpretability, and are difficult to gain the trust of clinicians. (3) It is not robust enough when facing complex situations and data noise in the real world. Summary of the Invention

[0004] The technical problem to be solved by this invention is how to provide an early diagnosis and staging system for cholangiocarcinoma that can simulate clinical reasoning processes, deeply integrate multimodal information, have strong interpretability, and be highly robust to real-world data.

[0005] This invention provides a step-by-step multimodal data fusion system for the early diagnosis and staging of cholangiocarcinoma, comprising: The data preprocessing module is used to acquire and preprocess medical imaging data, clinical laboratory data, and liquid biopsy biomarker data; The feature extraction module is used to extract high-level semantic feature vectors from medical imaging data, clinical feature vectors from clinical laboratory data, and biomarker feature vectors from liquid biopsy biomarker data. The feature fusion module uses a self-attention mechanism to weightedly fuse high-level semantic feature vectors, clinical feature vectors, and biomarker feature vectors to generate a global fused feature vector. The shared module converts the globally fused feature vector into a hidden representation vector through a fully connected layer; The bile duct cancer identification module takes the hidden representation vector as input, calculates the bile duct cancer identification probability through dot product operation and nonlinear probability mapping, and outputs the identification result of whether the bile duct cancer identification probability is positive. The tumor staging module is connected to the bile duct cancer identification module. When the identification result is positive, it takes the hidden representation vector as input, calculates the probability of late-stage bile duct cancer through dot product operation and nonlinear probability mapping, and outputs the tumor staging result based on the probability of late-stage bile duct cancer. The lymph node metastasis prediction module is connected to the tumor staging module. When the staging result is late stage, it takes the hidden representation vector as input, calculates the lymph node metastasis probability through dot product operation and nonlinear probability, and outputs the lymph node metastasis prediction result based on the metastasis probability.

[0006] Compared with existing technologies, this application has the following advantages: By employing a self-attention mechanism to weightedly fuse high-level semantic feature vectors extracted from medical imaging data, clinical feature vectors extracted from clinical laboratory data, and biomarker feature vectors extracted from liquid biopsy biomarker data, it dynamically and deeply mines the nonlinear complementary and synergistic relationships among the three types of data, generating a more informative and discriminative global fusion feature vector; furthermore, by identifying a distributed diagnostic framework for cholangiocarcinoma-staging-metastasis prediction, it simulates the progressive diagnostic logic of clinicians—first qualitative, then quantitative, and finally risk assessment—resulting in outputs that better meet clinical needs; simultaneously, it combines a shared fully connected layer with a distributed framework composed of various classifiers, and utilizes step-by-step decision-making to isolate error propagation, thereby improving the overall anti-interference capability and generalization performance of the system.

[0007] In one possible implementation, the medical imaging data includes computed tomography images and / or magnetic resonance cholangiopancreatography images in the plain scan, arterial phase, portal venous phase, and delayed phase; the clinical laboratory data includes at least one of carbohydrate antigen 19-9, alkaline phosphatase, total bilirubin, ALT, AST, and platelet count; and the liquid biopsy biomarker data includes the expression level of exosomal microRNA and / or the mutation status of circulating free deoxyribonucleic acid.

[0008] In one possible implementation, the preprocessing of the data preprocessing module specifically includes: Medical imaging data were resampled and intensity normalized, and window width and window level were used to adjust the medical imaging data of the liver or bile duct region. Numerical variables in clinical laboratory data and liquid biopsy biomarker data are normalized, and categorical variables in clinical laboratory data and liquid biopsy biomarker data are converted into numerical vectors using one-hot encoding.

[0009] Compared with existing technologies, by introducing specific preprocessing techniques such as window width and level adjustment for medical images and normalization and encoding of non-image data, the problem of inconsistent scales of multi-source data and the inability to directly fuse them has been solved. Furthermore, the technology has effectively eliminated differences in equipment and protocols, improved data quality, and provided a stable foundation for subsequent feature fusion.

[0010] In one possible implementation, the feature extraction module includes an image feature extraction unit, a clinical feature extraction unit, and a biomarker feature extraction unit; The image feature extraction unit is used to extract high-level semantic feature vectors from medical image data to characterize the morphological characteristics, enhancement characteristics, and relationships with surrounding tissues of lesions. The image feature extraction unit uses a ResNet18 pre-trained on a natural image dataset as its backbone network, removes the classification layer at the end of the ResNet18, and applies global average pooling to the last convolutional layer to form a convolutional neural network. The standardized medical image data is input into the convolutional neural network, and the output is a high-level semantic feature vector. ; The clinical feature extraction unit is used to extract clinical feature vectors from clinical laboratory data to characterize the patient's overall health status and the pathophysiological state of the biliary system. ; The biomarker feature extraction unit is used to extract biomarker feature vectors from liquid biopsy biomarker data to characterize the integration patterns and functional states of key molecular markers. ; Both the clinical feature extraction unit and the biomarker feature extraction unit employ a multilayer perceptron network to output clinical feature vectors from the standardized clinical laboratory data and liquid biopsy biomarker data, respectively. and biomarker feature vectors .

[0011] In one possible implementation, the feature fusion module specifically performs the following operations: Step A1, divide the dimensions as follows: , , High-level semantic feature vector Clinical feature vector and biomarker feature vectors Through a learnable weight matrix and bias Linear projection to a unified dimension The expression is: , ; Step A2, for each projected feature vector Generate query vectors Key vector Sum value vector The expressions are as follows: ; ; ; In the formula, , , It is a learnable matrix; Step A3, calculate the modes With all modes Attention weights between The expression is: ; Step A4, for the value vector Perform a weighted summation to obtain the modalities. context vector The expression is: ; Step A5: Perform average pooling on the context vectors of all modalities to obtain the global fused feature vector. : .

[0012] Compared with existing technologies, by employing pre-trained ResNet18 to extract image features and MLP to extract non-image features, and by using detailed mathematical formulas to define the self-attention fusion process, the problems of insufficient disclosure of feature extraction and fusion mechanisms and insufficient creative support are solved. This further realizes that the technical solution is highly feasible and creatively specific, and protects the technical effects of the core algorithm details from feature extraction to deep fusion.

[0013] In one possible implementation, the sharing module will globally fuse feature vectors. The hidden representation vector is obtained through a shared fully connected layer. The expression is: ; In the formula, This represents the learnable weight matrix. , This represents a learnable bias vector. This indicates a modified linear unit activation function, introducing nonlinearity, expressed as: ; This represents the hidden representation vector. ; Represents the hidden representation vector Dimensions.

[0014] Compared with existing technologies, by clearly defining the calculation formula of the shared fully connected layer and ensuring that each classifier has an independent output layer, the problems of unclear sharing mechanism and ambiguous classifier relationships are solved. Furthermore, an efficient architecture that reveals the separation of shared feature representation learning and task-specific decision-making is achieved, thus protecting the core technical effect of lightweight model design.

[0015] In one possible implementation, the bile duct cancer identification module specifically performs the following operations: Step B1: A pre-trained bile duct cancer identification classifier receives the hidden representation vector output by the fully connected layer. Output the probability of bile duct cancer identification. The expression is: ; ; In the formula, , representing the first weight vector specific to the bile duct cancer identification task, the first weight vector Used for hidden representation vectors The characteristics of the lesion were weighted according to morphological, clinical biochemical, and molecular evidence. , indicating the first bias specific to the bile duct cancer identification task; This indicates that the dot product operation is used on the hidden representation vector. The overall score is obtained by weighted summation of the characteristics of lesion morphology, clinical biochemical abnormalities, and molecular abnormalities. Represents a nonlinear probability mapping function; Step B2, based on the preset probability threshold Probability of cholangiocarcinoma identification The recognition result is represented as: .

[0016] In one possible implementation, the tumor staging module specifically performs the following operations: Step C1: When the output recognition result is positive, the hidden representation vector output by the fully connected layer is received using a pre-trained tumor staging classifier. Output the probability of bile duct cancer staging. The expression is: ; ; In the formula, , representing the second weight vector specific to the tumor staging task, the second weight vector Used for hidden representation vectors The tumor burden and invasion characteristics, biliary obstruction characteristics, and progressive biological characteristics were weighted together. This indicates a second bias specific to the tumor staging task; This indicates that the dot product operation is used on the hidden representation vector. The total score is obtained by weighted summation of features related to tumor burden and invasion, biliary obstruction, and progressive biological characteristics. Represents a nonlinear probability mapping function; Step C2, based on the preset probability threshold Probability of staging of bile duct cancer The phased results are represented as follows: .

[0017] In one possible implementation, the lymph node metastasis prediction module specifically performs the following operations: Step D1: When the staging result is late stage, a pre-trained lymph node metastasis prediction classifier is used to receive the hidden representation vector output by the fully connected layer. Output lymph node metastasis probability The expression is: ; ; In the formula, , representing the third weight vector specific to the lymph node metastasis prediction task, the third weight vector It is configured as a sparse vector, with a large number of non-zero weights concentrated on the hidden representation vector. Feature dimensions of coded lymph node status characteristics and metastatic molecular characteristics; , representing a learnable bias scalar specific to the lymph node metastasis prediction task; This indicates that the dot product operation is used on the hidden representation vector. The comprehensive score is obtained by weighted summation of the feature dimensions of the encoded lymph node status characteristics and metastatic molecular characteristics. Step D2, based on the preset probability threshold and the probability of lymph node metastasis The output transition prediction result is represented as follows: .

[0018] Compared with existing technologies, by providing specific probability calculation formulas and decision threshold rules for each classifier, the problems of the classification decision process being like a "black box" and having unclear scope of rights are solved. Furthermore, the mathematical transformation process from features to diagnostic conclusions is fully disclosed, and each diagnostic step has clear and verifiable judgment criteria.

[0019] In one possible implementation, an interpretive module is also included, the interpretive module being configured to: Based on the bile duct cancer identification probability output by the bile duct cancer identification module, a gradient-weighted activation map-based technique is used to generate a heat map identifying the lesion area on medical imaging data. Based on SHAP value analysis technology, the contribution of each feature in the clinical laboratory data and liquid biopsy biomarker data to the identification results, staging results and / or lymph node metastasis prediction results is analyzed, and a feature importance ranking chart is generated.

[0020] Compared with existing technologies, by introducing Grad-CAM and SHAP, two interpretability technologies, the problem of opaque decision-making processes and lack of clinical trust caused by the black box in the background technology is solved. It further enables the provision of visual lesion localization for images, quantitative contribution analysis for data, and the generation of interpretable diagnostic reports. Attached Figure Description

[0021] Fig. 1 This is a system block diagram of this application; Fig. 2 A heatmap generated from a CT image; Fig. 3 To generate a feature importance ranking graph. Detailed Implementation

[0022] First, those skilled in the art should understand that these embodiments are merely used to explain the technical principles of the embodiments of this application and are not intended to limit the scope of protection of the embodiments of this application. Those skilled in the art can make adjustments as needed to adapt to specific application scenarios.

[0023] In the description of the embodiments of this application, it should be noted that, unless otherwise explicitly specified and limited, the terms "connected" and "linked" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium. Those skilled in the art can understand the specific meaning of the above terms in the embodiments of this application based on the specific circumstances.

[0024] In the embodiments of this application, unless otherwise expressly specified and limited, "above" or "below" the second feature can mean that the first feature is in direct contact with the second feature, or that the first feature is in indirect contact with the second feature through an intermediate medium. Furthermore, "above," "on top of," and "over" the second feature can mean that the first feature is directly above or diagonally above the second feature, or simply that the first feature is at a higher horizontal level than the second feature. "Below," "below," and "under" the second feature can mean that the first feature is directly below or diagonally below the second feature, or simply that the first feature is at a lower horizontal level than the second feature.

[0025] The present application will now be described in further detail with reference to the accompanying drawings and specific embodiments.

[0026] See Figs. 1-3 As shown in the figure, this application discloses a step-by-step multimodal data fusion system for early diagnosis and staging of cholangiocarcinoma, including a data preprocessing module, a feature extraction module, a feature fusion module, and a sharing module connected in sequence, and further including a cholangiocarcinoma identification module, a tumor staging module, and a lymph node metastasis prediction module connected to the sharing module respectively; wherein: The data preprocessing module is used to acquire and preprocess medical imaging data, clinical laboratory data, and liquid biopsy biomarker data. In this embodiment of the application, the medical imaging data includes: acquiring multi-phase computed tomography (CT) images and / or magnetic resonance cholangiopancreatography (MRCP) images of the patient's upper abdomen, including plain scan, arterial phase, portal venous phase, and delayed phase; the images are input in DICOM format. The clinical trial data includes the patient's blood test results, including key indicators such as carbohydrate antigen 19-9 (CA19-9), alkaline phosphatase (ALP), total bilirubin (TBil), alanine aminotransferase (ALT), aspartate aminotransferase (AST), and platelet count. The liquid biopsy biomarker data includes: obtaining liquid biopsy analysis reports from patient plasma, including the expression level (relative quantitative value) of exosomal microRNAs (such as miR-21), and the fusion mutation status (presence / absence) of specific genes (such as FGFR2) in circulating cell-free DNA (cfDNA).

[0027] The preprocessing of the medical imaging data specifically includes: 1A) Using a bilinear interpolation algorithm, adjust the spatial resolution of computed tomography (CT) images and / or magnetic resonance cholangiopancreatography (MRI) images to a pre-defined target size. In this application embodiment, the preset size is 224×224 pixels; the expression is: ; In the formula, This refers to the acquired computed tomography (CT) scan images and / or magnetic resonance cholangiopancreatography (MRI) images. , Represents the interpolation function. This represents the resampled image; 2A) Resampled image The intensity is normalized to the [0, 1] interval to eliminate intensity differences caused by different scanning devices, protocols, and doses. The expression is as follows: ; In the formula, This indicates the intensity normalized image in coordinates. Pixel value at that location, and These represent the minimum and maximum pixel values ​​in the image, respectively. 3A) To highlight the liver and bile duct regions, a liver window is applied to optimize the display of all images, with a specific "liver window" applied to CT images. The window width (WW) determines the contrast range of the displayed image, and the window level (WL) determines the center value of this range. Through this transformation, the image grayscale values ​​are focused within the typical HU value range containing liver and bile duct tissues. The specific expression is: ; In the formula, , This represents the pixel value after window width and window level adjustments. This represents the scaling factor that maps HU values ​​to the grayscale of the image display; window width. Window position To maximize the contrast of the target tissues (liver, bile ducts) and suppress the display of irrelevant tissues (such as bones, fat) (optimized for liver / bile ducts, can be adjusted to WW=200HU depending on the device).

[0028] Preprocessing of clinical laboratory data and liquid biopsy biomarker data specifically includes: 1B) Median imputation is performed on missing values ​​in numerical variables (such as CA19-9, ALP, etc.) in clinical laboratory data and liquid biopsy biomarker data, and categorical variables or labeled missing values ​​in clinical laboratory data and liquid biopsy biomarker data are individually labeled as new categories. 2B) Standardization processing for numerical variables, involving normalization or standardization; The categorical variables are converted into numerical vectors using one-hot encoding. The categorical variables include gender (male, female, unknown) and FGFR2 fusion status (existence / non-existence). After one-hot encoding, the gender is encoded as

[100] (male),

[010] (female), and

[001] (unknown); the FGFR2 fusion status is encoded as

[10] (existence) and

[01] (non-existence).

[0029] The feature extraction module is used to extract high-level semantic feature vectors from medical imaging data to characterize the morphological characteristics, enhancement characteristics and surrounding tissue relationships of lesions; to extract clinical feature vectors from clinical laboratory data to characterize the overall health status and pathophysiological state of the biliary system of patients; and to extract biomarker feature vectors from liquid biopsy biomarker data to characterize the integration pattern and functional status information of key molecular markers. In this embodiment, to automatically learn and extract deep visual patterns related to the diagnosis of cholangiocarcinoma from medical image data, a ResNet18 model pre-trained on a large natural image dataset (such as ImageNet) is used as the backbone network. The global average pooling layer and fully connected classification layer at the end of the ResNet18 model are removed, retaining the general feature encoder. A global average pooling operation is then performed after the last convolutional layer, forming the convolutional neural network used in this application to process medical image data. The standardized medical image is input into the convolutional neural network model, and the feature map output by the last convolutional layer... , Indicates length, Indicates width. In this embodiment of the application, the number of channels is indicated. , After global average pooling operation for each feature channel The expression for spatial averaging is: ; Finally, a 512-dimensional high-level semantic feature vector was obtained. .

[0030] To map structured non-imaging data (clinical indicators and biomarkers) into dense, fusionable vector representations, this application employs a multilayer perceptron network (MLP) to extract clinical feature vectors from standardized clinical laboratory data and biomarker feature vectors from liquid biopsy biomarker data; specifically: First, the standardized clinical laboratory data (such as CA19-9, ALP, TBil, ALT, AST, platelet count, age, gender coding, etc., a total of M features) are concatenated into a one-dimensional feature vector. , ; The standardized liquid biopsy biomarker data (such as miR-21 expression level, FGFR2 fusion state encoding, etc., a total of N features) are concatenated into a one-dimensional feature vector. , ; Next, a multilayer perceptron (MLP) network consisting of two fully connected hidden layers and one output layer is used to process the one-dimensional feature vector. and one-dimensional characteristic vector Input the first hidden layer, and the output is: ; ; In the formula, This represents the first weight matrix of the first fully connected hidden layer in the clinical laboratory data path. , This represents the bias vector of the first fully connected hidden layer in the clinical laboratory data path. , This represents the output vector of the clinical trial data after passing through the first fully connected hidden layer. This represents the output vector of the liquid biopsy biomarker data after passing through the first fully connected hidden layer. This represents the first weight matrix of the first fully connected hidden layer in the liquid biopsy biomarker data path. This represents the bias vector of the first fully connected hidden layer in the liquid biopsy biomarker data path, as described in this embodiment. =64; Will , Input the second hidden layer, and the output is: ; ; In the formula, This represents the second weight matrix of the second fully connected hidden layer in the clinical laboratory data path. , This represents the bias vector of the second fully connected hidden layer in the clinical laboratory data path. , This represents the second weight matrix of the second fully connected hidden layer in the liquid biopsy biomarker data path. This represents the bias vector of the second fully connected hidden layer in the liquid biopsy biomarker data path, as described in this embodiment. 2 = 32; Finally, the output layer outputs the following: Clinical feature vector , ; Biomarker Feature Vector , ; In the formula, This represents the third weight matrix of the output layer in the clinical laboratory data path. , This represents the bias vector of the output layer in the clinical laboratory data path. ; This represents the third weight matrix of the output layer in the liquid biopsy biomarker data path. , This represents the bias vector of the output layer in the liquid biopsy biomarker data path.

[0031] The feature fusion module employs a self-attention mechanism to weightedly fuse high-level semantic feature vectors, clinical feature vectors, and biomarker feature vectors to generate a global fused feature vector. This allows for in-depth mining of the intrinsic relationships among the high-level semantic feature vectors, clinical feature vectors, and biomarker feature vectors, achieving information complementarity. Specifically, the feature fusion module performs the following operations: Step A1, divide the dimensions as follows: , , High-level semantic feature vector Clinical feature vector and biomarker feature vectors Through a learnable weight matrix and bias Linear projection to a unified dimension The expression is: , ; Step A2, for each projected feature vector Generate query vectors Key vector Sum value vector The expressions are as follows: ; ; ; In the formula, , , It is a learnable matrix; Step A3, calculate the modes With all modes Attention weights between The expression is: ; Step A4, for the value vector Perform a weighted summation to obtain the modalities. context vector The expression is: ; Step A5: Perform average pooling on the context vectors of all modalities to obtain the global fused feature vector. : ; The obtained global fusion feature vector It includes in-depth correlation information across three modalities: imaging, clinical, and biomarkers.

[0032] The shared module is used to transform the global fusion feature vector into a hidden representation vector through a shared fully connected layer; the fully connected layer is linked to the cholangiocarcinoma identification module, the tumor staging module, and the lymph node metastasis prediction module, respectively. The fully connected layer is used to globally fuse feature vectors. Transform it into a more compact and discriminative hidden representation vector. The expression is: ; In the formula, This represents the learnable weight matrix. , This represents a learnable bias vector. This represents the modified linear unit activation function, introducing nonlinearity, and is expressed as: R ; This means randomly setting the output of the hidden layer neurons to zero with a preset probability; This represents the hidden representation vector. ; Represents the hidden representation vector The dimension; The bile duct cancer identification module takes the hidden representation vector as input, calculates the bile duct cancer identification probability, and outputs a positive or negative identification result based on the bile duct cancer identification probability. In this embodiment, the bile duct cancer identification module employs a pre-trained bile duct cancer identification classifier, which has an independent linear output layer for processing the hidden representation vector. This is mapped to the probability of bile duct cancer identification; the bile duct cancer identification module specifically performs the following operations: Step B1: A pre-trained bile duct cancer identification classifier receives the hidden representation vector output by the fully connected layer. Output the probability of bile duct cancer identification. The expression is: ; ; In the formula, , representing the first weight vector specific to the bile duct cancer identification task, the first weight vector Used for hidden representation vectors The characteristics of the lesion were weighted according to morphological, clinical biochemical, and molecular evidence. , indicating the first bias specific to the bile duct cancer identification task; This indicates that the dot product operation is used on the hidden representation vector. The overall score is obtained by weighted summation of the characteristics of lesion morphology, clinical biochemical abnormalities, and molecular abnormalities. This represents a nonlinear probability mapping function. In this embodiment, the nonlinear probability mapping function is the Sigmoid activation function, which maps logistic regression values... Compress to the (0,1) interval; Step B2, based on the preset probability threshold Probability of cholangiocarcinoma identification The recognition result is represented as: .

[0033] In the embodiments of this application When the diagnosis result is positive, that is If the result is positive, it indicates a confirmed diagnosis of cholangiocarcinoma. The tumor staging module will then proceed. If the diagnosis result is negative, the diagnosis result will be output directly and the process will terminate.

[0034] Furthermore, the core task of the cholangiocarcinoma identification classifier in this application embodiment is to determine whether cholangiocarcinoma exists. To achieve this goal, the system design of this classifier must be able to comprehensively evaluate three key types of evidence: morphological evidence of lesions (M), clinical biochemical abnormalities (C), and molecular abnormalities (B). Therefore, during the pre-training phase, the cholangiocarcinoma identification classifier is configured with... The binary cross-entropy loss (BCE loss) is used to optimize the parameters of the classifier. ,make The numerical distribution can specifically target the feature dimensions (denoted as the dimension set) in the hidden vector hh that encode the evidence M, C, and B respectively. This assigns a larger absolute value (positive or negative) weight. For example, the dimension encoding "arterial phase enhancement" might be weighted +0.8, while the dimension encoding "CA19-9 normal" might be weighted -0.6. Therefore, the calculation... The process essentially involves a weighted synthesis of these key pieces of evidence, identifying evidence of lesion morphology, biochemical abnormalities, and molecular abnormalities, and interpreting the training of a specific linear classifier. This allows its weight distribution to amplify the features corresponding to these pieces of evidence, and through dot product operations. Real-time execution of evidence synthesis and judgment.

[0035] The formula for binary cross-entropy loss is: ; In the formula, Indicates the number of training samples; , indicating the first The true label of each sample (0 represents non-cholangiocarcinoma, 1 represents cholangiocarcinoma); Indicates the first The predicted probability of bile duct cancer for each sample.

[0036] The tumor staging module operates when the identification result is positive. The tumor staging module uses a tumor staging classification stage, inputting the hidden representation vector into a pre-trained tumor staging classifier and outputting the staging result. The tumor staging classifier in this embodiment has an independent task output layer, which inputs the hidden representation vector... Mapped to the probability of bile duct carcinoma staging, the tumor staging module specifically performs the following operations: Step C1: When a positive diagnosis of cholangiocarcinoma is made, the pre-trained tumor staging classifier receives the hidden representation vector output by the fully connected layer. Output the probability of bile duct cancer staging. The expression is: ; ; In the formula, , representing the second weight vector specific to the tumor staging task, the second weight vector Used for hidden representation vectors The tumor burden and invasion characteristics, biliary obstruction characteristics, and progressive biological characteristics were weighted together. This indicates a second bias specific to the tumor staging task; This indicates that the dot product operation is used on the hidden representation vector. The total score is obtained by weighted summation of features related to tumor burden and invasion, biliary obstruction, and progressive biological characteristics. This represents a nonlinear probability mapping function; the nonlinear probability mapping function in this embodiment uses the Sigmoid activation function. Step C2, based on the preset probability threshold Probability of staging of bile duct cancer The phased results are represented as follows: .

[0037] In this embodiment, the tumor staging classifier is tasked with assessing the degree of tumor invasion and progression. Its decisions depend on tumor burden and invasion characteristics (T), biliary obstruction characteristics (O), and progressive biological characteristics (P). During the pre-training phase, the tumor staging classifier is configured... ,train This enables it to work with hidden representation vectors The dimensions of the encoded features T, O, and P are given high weights. Because... Focusing on progression characteristics, the hidden representation vector of an advanced tumor sample. In the middle, features T, O, and P have high intensity and corresponding weights. The value is also large, therefore The consequences will be significant. Using the same binary cross-entropy loss, the formula is: ; In the formula, This indicates the number of cholangiocarcinoma samples in a training batch. , indicating the first The true stage label of each cholangiocarcinoma sample (0 represents early stage, 1 represents late stage). Indicates the first The predicted staging probability of a single cholangiocarcinoma sample; Loss is calculated through gradient backpropagation. Will be updated and the parameters of the shared fully connected layer This allows the hidden representation vector to be hidden. Learn to encode information that is useful for phased tasks; After training, its high-weight dimension and hidden representation vector The coding of tumor burden is associated with dimensions of invasion features (such as size and vascular invasion), biliary obstruction features (such as bile duct dilatation), and progressive biological features. Therefore, It reflects the overall invasiveness of the tumor.

[0038] If the diagnosis is "early stage", a final diagnostic report containing "early stage bile duct cancer positive" will be output and the process will terminate; if the diagnosis is "late stage", the lymph node metastasis prediction module will proceed to the next risk assessment.

[0039] The lymph node metastasis prediction module operates when the staging result is late stage. The module employs a pre-trained lymph node metastasis prediction classifier, using a hidden representation vector as input, and outputs the predicted metastasis probability of lymph nodes. Specifically, the lymph node metastasis prediction module performs the following operations: Step D1: When the staging result is late stage, a pre-trained lymph node metastasis prediction classifier is used to receive the hidden representation vector output by the fully connected layer. Output lymph node metastasis probability The expression is: ; ; In the formula, , representing the third weight vector specific to the lymph node metastasis prediction task, the third weight vector It is configured as a sparse vector, with a large number of non-zero weights concentrated on the hidden representation vector. Feature dimensions of coded lymph node status characteristics and metastatic molecular characteristics; , representing a learnable bias scalar specific to the lymph node metastasis prediction task; This indicates that the dot product operation is used on the hidden representation vector. The comprehensive score is obtained by weighted summation of the feature dimensions of the encoded lymph node status features and metastatic molecular features; the nonlinear probability mapping function of this application embodiment. Use the Sigmoid activation function; Step D2, based on the preset probability threshold and the probability of lymph node metastasis The output transition prediction result is represented as follows: .

[0040] The task of the lymph node metastasis prediction classifier is to assess the risk of lymphatic dissemination based on lymph node status characteristics (L), metastatic molecular characteristics (Met), and systemic response characteristics (S). During the training phase, the lymph node metastasis prediction classifier is configured... Because the training samples (late-stage cases) are relatively few, a regularization term was introduced into the loss function when training the classifier; under the regularization constraint, training... This makes it sparsely and powerfully focused on features L and Met; therefore, the third weight vector It is optimized into a sparse vector, meaning that its non-zero weights are largely concentrated in the hidden representation vector. The model focuses on a few feature dimensions that are most directly related to lymph node status (imaging) and metastatic molecular features, thereby enhancing its generalization ability and robustness in small samples.

[0041] To enhance clinical trust, an interpretability module is also included, which has dual interpretable outputs and is configured as follows: Image interpretability: Based on Gradient-Weighted Class Activation Map (Grad-CAM) technology. Gradient information from the cholangiocarcinoma identification classifier is backpropagated to the feature map of the last convolutional layer of ResNet18, generating a heatmap. This heatmap is overlaid on the original CT image, visually highlighting the lesion area upon which the model's "cholangiocarcinoma" diagnosis is based. Fig. 2 The image shown is a CT image from the portal venous phase. The thermal map shows that the red highlighted areas are concentrated in the thickened bile duct wall (coordinates [100,150] to [200,250]), indicating that the model focuses on enhancement characteristics. If there are no lesions, the thermal map is evenly distributed without obvious peaks.

[0042] Data interpretability: Based on SHAP (SHapley Additive exPlanations) value analysis technique. On the test set, the probability of each clinical and biomarker feature (such as CA19-9, miR-21) for the model's final prediction is calculated. The contribution (SHAP value) is calculated, and a bar chart ranking the feature importance is generated. For example... Fig. 3 As shown, in a positive prediction case, the SHAP plot showed that CA19-9 contributed +0.35 (positive, driving the positive probability), miR-21 contributed +0.25, and ALT contributed -0.10 (negative); the ordination plot was displayed as bars, with the top three features being CA19-9 (importance 0.35), total bilirubin (0.28), and FGFR2 mutation (0.22), helping clinicians understand the basis for decision-making.

[0043] Experimental verification: To verify the effectiveness of this system, experiments were conducted on a multi-center dataset.

[0044] Dataset: Data from 500 patients across three hospitals (300 for training, 100 for validation, and 100 for testing) included CT / MRCP images, clinical indicators (e.g., abnormal CA19-9 levels >37 U / mL), and liquid biopsies (e.g., miR-21 expression >2-fold as high). Labels were based on the gold standard (pathological confirmation): 200 cases of positive cholangiocarcinoma (100 early-stage and 100 late-stage, including 50 metastatic cases).

[0045] Training details: PyTorch framework was used, with the Adam optimizer (learning rate 0.001, decay 0.0001); batch size 32; early stopping mechanism (patient 10 epochs); loss function was binary cross-entropy with regularization (weight decay 0.01). Training was conducted on an NVIDIA RTX 3090 GPU for 50 epochs.

[0046] Performance metrics: Cholangiocarcinoma identification: accuracy 92%, AUC 0.95, sensitivity 90%, specificity 94%.

[0047] Tumor staging: accuracy 85%, F1 score 0.87 (early vs. late).

[0048] Lymph node metastasis prediction: accuracy 88%, recall 85% (for late-stage cases).

[0049] Comparative Experiment: Compared with benchmark models (Benchmark 1: Simple feature concatenation + MLP; Benchmark 2: ResNet without self-attention fusion). As shown in Table 1: Table 1: Comparison of Experiments Results analysis: The accuracy of this system decreased by only 3% under noisy data (with 10% random perturbation added), which is better than the baseline of 10%. Interpretable outputs (such as heatmaps) helped clinical validation, with an average trust score improvement of 20% (based on feedback from 10 physicians). These results demonstrate the superiority and practicality of this invention.

[0050] In the description of the embodiments of this application, it should be noted that the terms "inner" and "outer" and other terms indicating direction or positional relationship are based on the direction or positional relationship shown in the drawings. This is only for the convenience of description and does not indicate or imply that the device or component must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation of this application.

[0051] In the description of this application, the references to terms such as "an embodiment," "some embodiments," "in this embodiment," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in a suitable manner in any one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0052] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A step-by-step multimodal data fusion system for the early diagnosis and staging of cholangiocarcinoma, characterized in that, include: The data preprocessing module is used to acquire and preprocess medical imaging data, clinical laboratory data, and liquid biopsy biomarker data; The feature extraction module is used to extract high-level semantic feature vectors from medical imaging data, clinical feature vectors from clinical laboratory data, and biomarker feature vectors from liquid biopsy biomarker data. The feature fusion module uses a self-attention mechanism to weightedly fuse high-level semantic feature vectors, clinical feature vectors, and biomarker feature vectors to generate a global fused feature vector. The shared module converts the globally fused feature vector into a hidden representation vector through a fully connected layer; The bile duct cancer identification module takes the hidden representation vector as input, calculates the bile duct cancer identification probability through dot product operation and nonlinear probability mapping, and outputs the identification result of whether the bile duct cancer identification probability is positive. The tumor staging module is connected to the bile duct cancer identification module. When the identification result is positive, it takes the hidden representation vector as input, calculates the probability of late-stage bile duct cancer through dot product operation and nonlinear probability mapping, and outputs the tumor staging result based on the probability of late-stage bile duct cancer. The lymph node metastasis prediction module is connected to the tumor staging module. When the staging result is late stage, it takes the hidden representation vector as input, calculates the lymph node metastasis probability through dot product operation and nonlinear probability, and outputs the lymph node metastasis prediction result based on the metastasis probability.

2. The step-by-step multimodal data fusion system for early diagnosis and staging of cholangiocarcinoma according to claim 1, characterized in that, The medical imaging data includes computed tomography images and / or magnetic resonance cholangiopancreatography images in the plain scan, arterial phase, portal venous phase, and delayed phase; the clinical laboratory data includes at least one of carbohydrate antigen 19-9, alkaline phosphatase, total bilirubin, ALT, AST, and platelet count; the liquid biopsy biomarker data includes the expression level of exosomal microRNA and / or the mutation status of circulating free deoxyribonucleic acid.

3. The step-by-step multimodal data fusion system for early diagnosis and staging of cholangiocarcinoma according to claim 1, characterized in that, The preprocessing of the data preprocessing module specifically includes: Medical imaging data were resampled and intensity normalized, and window width and window level were used to adjust the medical imaging data of the liver or bile duct region. Numerical variables in clinical laboratory data and liquid biopsy biomarker data are normalized, and categorical variables in clinical laboratory data and liquid biopsy biomarker data are converted into numerical vectors using one-hot encoding.

4. The step-by-step multimodal data fusion system for early diagnosis and staging of cholangiocarcinoma according to claim 1, characterized in that, The feature extraction module includes an image feature extraction unit, a clinical feature extraction unit, and a biomarker feature extraction unit; The image feature extraction unit is used to extract high-level semantic feature vectors from medical image data to characterize the morphological characteristics, enhancement characteristics, and relationships with surrounding tissues of lesions. The image feature extraction unit uses a ResNet18 pre-trained on a natural image dataset as its backbone network, removes the classification layer at the end of the ResNet18, and applies global average pooling to the last convolutional layer to form a convolutional neural network. The standardized medical image data is input into the convolutional neural network, and the output is a high-level semantic feature vector. ; The clinical feature extraction unit is used to extract clinical feature vectors from clinical laboratory data to characterize the patient's overall health status and the pathophysiological state of the biliary system. ; The biomarker feature extraction unit is used to extract biomarker feature vectors from liquid biopsy biomarker data to characterize the integration patterns and functional states of key molecular markers. ; Both the clinical feature extraction unit and the biomarker feature extraction unit employ a multilayer perceptron network to output clinical feature vectors from the standardized clinical laboratory data and liquid biopsy biomarker data, respectively. and biomarker feature vectors .

5. The step-by-step multimodal data fusion system for early diagnosis and staging of cholangiocarcinoma according to claim 4, characterized in that, The feature fusion module specifically performs the following operations: Step A1, divide the dimensions as follows: , , High-level semantic feature vector Clinical feature vector and biomarker feature vectors Through a learnable weight matrix and bias Linear projection to a unified dimension The expression is: , ; Step A2, for each projected feature vector Generate query vectors Key vector Sum value vector The expressions are as follows: ; ; ; In the formula, , , It is a learnable matrix; Step A3, calculate the modes With all modes Attention weights between The expression is: ; Step A4, for the value vector Perform a weighted summation to obtain the modalities. context vector The expression is: ; Step A5: Perform average pooling on the context vectors of all modalities to obtain the global fused feature vector. : 。 6. The step-by-step multimodal data fusion system for early diagnosis and staging of cholangiocarcinoma according to claim 5, characterized in that, The shared module will globally fuse feature vectors. The hidden representation vector is obtained through a shared fully connected layer. The expression is: ; In the formula, This represents the learnable weight matrix. , This represents a learnable bias vector. This indicates a modified linear unit activation function, introducing nonlinearity, expressed as: ; This represents the hidden representation vector. ; Represents the hidden representation vector Dimensions.

7. The step-by-step multimodal data fusion system for early diagnosis and staging of cholangiocarcinoma according to claim 6, characterized in that, The bile duct cancer identification module performs the following operations: Step B1: A pre-trained bile duct cancer identification classifier receives the hidden representation vector output by the fully connected layer. Output the probability of bile duct cancer identification. The expression is: ; ; In the formula, , representing the first weight vector specific to the bile duct cancer identification task, the first weight vector Used for hidden representation vectors The characteristics of the lesion were weighted according to morphological, clinical biochemical, and molecular evidence. , indicating the first bias specific to the bile duct cancer identification task; This indicates that the dot product operation is used on the hidden representation vector. The overall score is obtained by weighted summation of the characteristics of lesion morphology, clinical biochemical abnormalities, and molecular abnormalities. Represents a nonlinear probability mapping function; Step B2, based on the preset probability threshold Probability of cholangiocarcinoma identification The recognition result is represented as: 。 8. The step-by-step multimodal data fusion system for early diagnosis and staging of cholangiocarcinoma according to claim 6, characterized in that, The tumor staging module performs the following operations: Step C1: When the identification result is positive, a pre-trained tumor staging classifier is used to receive the hidden representation vector output by the fully connected layer. Output the phase probability The expression is: ; ; In the formula, , representing the second weight vector specific to the tumor staging task, the second weight vector Used for hidden representation vectors The tumor burden and invasion characteristics, biliary obstruction characteristics, and progressive biological characteristics were weighted together. This indicates a second bias specific to the tumor staging task; This indicates that the dot product operation is used on the hidden representation vector. The total score is obtained by weighted summation of features related to tumor burden and invasion, biliary obstruction, and progressive biological characteristics. Represents a nonlinear probability mapping function; Step C2, based on the preset probability threshold and the probability of phases The phased results are represented as follows: 。 9. The step-by-step multimodal data fusion system for early diagnosis and staging of cholangiocarcinoma according to claim 6, characterized in that, The lymph node metastasis prediction module performs the following operations: Step D1: When the staging result is late stage, a pre-trained lymph node metastasis prediction classifier is used to receive the hidden representation vector output by the fully connected layer. Output lymph node metastasis probability The expression is: ; ; In the formula, , representing the third weight vector specific to the lymph node metastasis prediction task, the third weight vector It is configured as a sparse vector, with a large number of non-zero weights concentrated on the hidden representation vector. Feature dimensions of coded lymph node status characteristics and metastatic molecular characteristics; , representing a learnable bias scalar specific to the lymph node metastasis prediction task; This indicates that the dot product operation is used on the hidden representation vector. The comprehensive score is obtained by weighted summation of the feature dimensions of the encoded lymph node status characteristics and metastatic molecular characteristics. Step D2, based on the preset probability threshold and transition probability The output transition prediction result is represented as follows: 。 10. The step-by-step multimodal data fusion system for early diagnosis and staging of cholangiocarcinoma according to claim 1, characterized in that, It also includes an interpretive module, which is configured to: Based on the bile duct cancer identification probability output by the bile duct cancer identification module, a gradient-weighted activation map-based technique is used to generate a heat map identifying the lesion area on medical imaging data. Based on SHAP value analysis technology, the contribution of each feature in the clinical laboratory data and liquid biopsy biomarker data to the identification results, staging results and / or lymph node metastasis prediction results is analyzed, and a feature importance ranking chart is generated.