Structural entropy guided esophageal cancer multi-modal data discriminative representation learning method
By constructing a structural entropy-guided discriminative representation learning method for multimodal data, using a three-layer coding tree and dynamic weight adjustment, the hierarchical structure and complex association problems in esophageal cancer data representation learning were solved, and efficient analysis and accurate diagnosis of multimodal data were achieved.
Patent Information
- Application Number
- CN202510768483.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-10-17
AI Technical Summary
Existing methods have difficulty capturing the hierarchical structure and complex associations of data in esophageal cancer data representation learning, and representation collapse is prone to occur in unsupervised or semi-supervised scenarios, affecting the semantic space alignment and hierarchical structure preservation during multimodal data fusion.
A discriminative representation learning method for multimodal data guided by structural entropy is adopted. By constructing a regularized framework of a three-layer hierarchical coding tree, the structural entropy of the intermediate layer nodes is maximized. Combined with dynamic weight adjustment and multimodal data alignment, soft label and pseudo label strategies are introduced to achieve hierarchical structure capture and semantic space alignment of multimodal data.
It improves the accuracy and robustness of esophageal cancer data analysis, solves the problem of insufficient utilization of data structure information in traditional methods, effectively reflects the disease pathological mechanism, and improves the semantic space alignment and hierarchical structure preservation capabilities of multimodal data fusion.
Smart Images

Figure CN120809140A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of medical image engineering and the field of semantic understanding, and in particular to a structural entropy guided esophageal cancer multi-modal data discriminative representation learning method. BACKGROUND
[0002] Under the background of the deep integration of medical informatization and artificial intelligence, the precise diagnosis and treatment of esophageal cancer puts forward urgent needs for efficient analysis of multi-modal data. With the explosive growth of esophageal cancer related data such as electronic medical records, endoscopic images, pathological sections, and gene sequencing, how to extract discriminative semantic features from complex heterogeneous data to achieve precise semantic segmentation, semantic recognition, efficacy prediction, and treatment decision support has become a core challenge in the current esophageal cancer AI field.
[0003] However, existing methods face two major difficulties in esophageal cancer data representation learning: 1) Traditional models are difficult to capture the hierarchical structure and complex associations of data. Esophageal cancer diagnosis terms follow the ICD-10 tree system (such as C15.0 esophageal upper segment squamous carcinoma, C15.9 cancer of unspecified site), and gene expression data involve PI3K-AKT, NOTCH, and other multi-pathway cooperative networks. Mainstream CNN / RNN can only extract single-dimensional features, ignoring the structured association between features, resulting in the inability to reflect the pathological evolution rule of esophageal cancer from intraepithelial neoplasia to infiltration and metastasis, or the molecular mechanism of immune therapy response. 2) In unsupervised / semi-supervised scenarios, "representation collapse" is prone to occur. Taking esophageal cancer endoscopic image clustering as an example, if the contrast learning method ignores the structural semantics such as lesion edge shape and blood vessel distribution, it will cause confusion in the feature embedding of ulcerative and protruding squamous cell carcinoma, affecting subtype classification. In the multi-modal fusion scenario, cross-modal semantic analysis alignment and hierarchical structure preservation are still challenges. SUMMARY
[0004] To solve the above problems, the present application provides a structural entropy guided esophageal cancer multi-modal data discriminative representation learning method. The present application aims to solve the problems of insufficient utilization of data structure information in existing esophageal cancer data representation learning, "representation collapse" in unsupervised or semi-supervised scenarios, and semantic space alignment and hierarchical structure preservation in multi-modal data fusion, to provide a more interpretable and robust general representation learning framework for esophageal cancer artificial intelligence diagnosis and treatment, and to improve the performance of esophageal cancer precise diagnosis, efficacy prediction, and treatment decision support.
[0005] To achieve the above-mentioned technology, the steps are as follows: S1, collecting esophageal cancer multi-modal data including text data, image data, and gene data; Text data is obtained from electronic medical records and medication records, and can be obtained from hospital systems and desensitized; Image data is acquired by CT and endoscopic images, along with shooting parameters from the image system; Gene data includes physiological signals and test data (such as electrocardiogram, blood indicators), respectively from monitoring equipment and laboratory systems.
[0006] S2, perform preprocessing operation on the collected esophageal cancer multi-modal data; Medical data preprocessing is a key step to improve data quality and enhance data usability, providing high-quality training samples for machine learning models. Due to the wide range of esophageal cancer medical data sources and complex structure, there are problems such as text errors and redundant records in electronic medical records, different resolutions and formats of endoscopic images, and missing values and outliers in gene test data. Therefore, data preprocessing is a necessary prerequisite to ensure the effectiveness of esophageal cancer data analysis and model training; Medical data preprocessing realizes the transformation of raw esophageal cancer data from chaos to standardization through systematic data governance process. The preprocessed data can significantly improve the modeling ability of machine learning models for esophageal cancer pathological mechanism, improve the training efficiency and prediction accuracy of the model, and lay a solid data foundation for artificial intelligence applications such as esophageal cancer intelligent auxiliary diagnosis and treatment decision support; Preprocessing operations include: S2.1, data cleaning operation; data cleaning is the first step of preprocessing, aiming to identify and correct errors, duplicates and invalid information in the data; Data cleaning operations include: For spelling errors of diagnosis names and repeated medical records in esophageal cancer electronic medical record system, data deduplication algorithm and medical terminology verification mechanism are used for correction and elimination; For missing values in gene expression data and blood test data, use mean filling method or time series interpolation method for completion; For lesion region annotation abnormalities and temperature exceeding physiological limits in endoscopic images, combine esophageal cancer clinical professional knowledge for labeling or deletion processing; S2.2, data format unification operation; data format unification is the core step to realize data standardization; Data format unification includes: For text data processing, unstructured diagnosis description (such as "esophageal upper segment visible ulcerative mass") needs to be mapped to the International Classification of Diseases (ICD-10) coding system to eliminate differences in terminology; For medical image data (such as esophageal endoscopic images, CT images), normalization processing is required, including uniform image resolution (such as adjusting to 224x224 pixels), grayscale standardization and format conversion; For gene data, it is necessary to map the detection results of different experimental platforms to a unified numerical interval through a standardization transformation to ensure the consistency and comparability of the data; S2.3, multi-modal data alignment operation; the multi-modal data alignment focuses on aligning the heterogeneous data generated at different time nodes of the same esophageal cancer patient; The mode of the multi-modal data alignment operation is: taking the time when the patient is diagnosed with esophageal cancer as the reference, the endoscopic image, the pathological section image, the gene test result and the multi-source data of the electronic medical record text are spatio-temporally aligned to establish a correlation relationship network; for example, the endoscopic image features of a patient at the time of diagnosis are associated with the contemporaneous gene expression data and the pathological diagnosis text, which can effectively capture the potential correlation between the image features, molecular features and clinical phenotypes in the pathological evolution process from intraepithelial neoplasia to infiltration and metastasis of esophageal cancer, and provide structured data support for subsequent analysis.
[0007] S3, a regularization framework based on structural entropy is constructed to capture the hierarchical structure and complex dependency relationship in medical data; In order to solve the problem of insufficient utilization of structural information of medical data by traditional methods, the present application designs a structural entropy regularization framework of a three-layer hierarchical coding tree, maximizes the structural entropy of the intermediate layer nodes, constrains the distribution of latent variables to promote their separation, thereby enhancing the expression ability of the latent representation, enabling the model to learn a feature representation with clear hierarchical division and better reflecting the pathological mechanism of the disease; The input of the structural entropy regularization framework is the esophageal cancer multi-modal data features after the pre-processing operation, the adjacency matrix and the allocation matrix are constructed after the data embedding representation, the regularization structural entropy loss calculation is performed through the adjacency matrix and the allocation matrix and the three-layer hierarchical coding tree, and the dynamic weight adjustment is introduced to improve the robustness and interpretability of the structural entropy regularization framework; The three-layer hierarchical coding tree includes: a root node, an intermediate layer and a leaf node; the data is input into the intermediate layer after passing through the root node, the hierarchical separation operation of the multi-modal data features is performed by maximizing the structural entropy of the intermediate layer nodes, and the individual data samples are output through the leaf nodes; The structural entropy of the intermediate layer nodes of the three-layer coding tree As follows: In the formula, represents the number of categories, represents the index of the number of categories; represents the constructed medical data graph; denotes the set of intermediate nodes in the coding tree, for each , denotes the total weight of edges between the sub-trees and its complement ; denotes the volume of ; denotes the structural entropy of the intermediate nodes that is maximized; The construction step includes: S3.1, encoding the pre-processed multi-source data into an embedding representation for unifying the multi-source data into a computable low-dimensional embedding and preserving semantic information; S3.2, constructing an adjacency matrix and an assignment matrix based on the embedding representation ; The adjacency matrix is denoted as: wherein, denotes a sigmoid activation function that ensures the positive values in A; the adjacency matrix can be obtained from the embedding representation ; The medical data graph can be obtained from the embedding representation , which reflects the similarity of latent features; wherein, denotes the number of leaf nodes, denotes the number of intermediate nodes (i.e., clusters or semantic groups); each element in the binary assignment matrix , if denotes that the th leaf node is assigned to the th category, which means the hard partition of nodes; such discrete assignment structure makes it possible to explicitly model the hierarchical relationship and supports the construction of a multi-scale semantic tree that captures the local and global organization in the latent space; such adaptive adjacency can capture high-order dependency relationships between samples, thereby improving the representation quality; S3.3, obtaining the structural entropy of the intermediate nodes of the three-layer coding tree based on the adjacency matrix and the assignment matrix in the form of a regularization loss, and introducing a dynamic weight factor; The regularization loss form of the structural entropy of the intermediate nodes of the three-layer coding tree is as follows: wherein, denotes an all-1 matrix with a shape of ; denotes the total volume of the graph , i.e., , the graph The sum of all edge weights reflects the total correlation strength of the entire medical data graph. The volume of the category The total weight of the edges between the parts of the all-1 matrix, the adjacency matrix and the assignment matrix of the graph , that is . The volume of the category , that is . The dynamic weight factor expression is as follows: In the formula, denotes the assignment entropy of the i-th node at the t-th time point, and the expression is as follows: In the formula, denotes the number index of the leaf node; The loss function expression after introducing the dynamic weight adjustment is as follows: . The dynamic weight adjustment mechanism adjusts the regularization strength of each node, enhances regularization to guide the structure learning of medical data (such as identifying key pathways or disease characteristics), and reduces constraints to optimize detailed features (such as specific gene or image detail representation), thereby improving the robustness and interpretability of the model in medical tasks. Through the above regularization framework based on structural entropy, the hierarchical structure and complex dependency relationship in medical data are effectively captured.
[0008] S4, input the preprocessed esophageal cancer multi-modal data into the constructed regularization framework based on structural entropy, and perform single-modal task classification clustering based on structural entropy; The single modal includes: classification task (text), regression task (gene data) and clustering task (image data); The single-modal task classification includes: hard assignment and soft assignment; wherein the soft assignment mode includes: introducing soft label and pseudo label strategies in the three-layer coding tree in the regularization framework based on structural entropy; The clustering method includes: S4.1, a discrete structure with probability coding is used to perform hard classification prediction on the classification task; Probability coding unifies probability representation learning and task prediction in a single module. In order to realize gradient-based optimization, the reparameterization technique is applied, which allows sampling from the distribution without introducing gradient bias; Unlike existing methods, our method uses an encoder-only architecture that removes the decoder and directly samples and infers labels from the learned latent distribution. Our method integrates probability coding and task prediction into a single encoder module through the network and applies non-parametric operations to obtain the predicted output. The overall loss of the Structured Entropy Probabilistic Coding (SEPC) for discrete structures of probabilistic coding for classification tasks is as follows: Where, represents the task loss of probabilistic encoding; Represents a hyperparameter that controls the regularization loss based on structural entropy The weights of , force the model to learn hierarchical discriminative diagnostic features (such as key pathological indicators for distinguishing pneumonia from influenza); S4.2. We introduce a soft label strategy into the three-layer coding tree in the regularized framework based on structural entropy to construct a coding tree for regression tasks and perform soft assignment prediction for regression tasks. Discretized coding trees constructed using true labels are inherently designed for classification tasks. However, this hard partitioning scheme is not directly applicable to other discriminative representation learning scenarios, such as regression and unsupervised clustering. To address this limitation and expand the applicability of coding trees, this paper proposes two soft assignment strategies that enable data points to be flexibly and continuously assigned to the hierarchy. For regression tasks, a soft labeling method is introduced, in which each data point is assigned to a probability distribution over all possible categories. Based on this idea, this paper proposes a probabilistic coding tree, which extends the structural entropy theory to the soft classification scenario. Unlike traditional coding trees that rely on hard assignment, this method relaxes the constraint that each child node is linked to only one parent node. Instead, each child node is probabilistically connected to all higher-level nodes to capture subtle structural dependencies. The construction methods include: The structural entropy framework is extended to the regression scenario by first discretizing the continuous regression label space into intermediate nodes, i.e. discrete boxes; each box consists of a central point It means that for each data point, its distance to all bin centers is calculated to estimate the category affinity: Where, represents the true regression label vector; Represents the absolute distance matrix between the sample label and the center of each box, ; To obtain soft assignment, apply a softmax operation to negative distance values so that closer boxes have higher assignment probabilities: ,in, denotes the soft assignment probability matrix (the probability of each sample belonging to each bin); This soft assignment allows each data point to contribute to multiple semantic regions in the latent space, defining a volume and cut edge weights under continuous conditions, and shows that their form is consistent with the discrete setting, where each intermediate node has a soft volume computed as: where denotes the total number of samples; denotes the samples, and denotes the multi-source data; denotes the soft assignment of the th node to the th class; denotes the degree of the sample (the sum of the adjacency matrix row); In the soft assignment setting, quantifies the expected cut edge weight between and its complement and is computed as the weighted sum of edge connections, where the weight is given by the product of the probabilities of each pair of vertices lying on either side of the partition; for each pair of nodes , the probability that one node belongs to and the other belongs to its complement is: where denotes the element in the soft assignment matrix ; The total expected cut weight is given by: where, let , the final regularization loss formula based on structural entropy is the same as , the key difference is that the assignment matrix now encodes soft probabilities rather than hard class membership; this formula enables structural entropy to be effectively used for regression tasks by modeling through probability labels and soft hierarchical encoding; S4.3, by introducing a pseudo-label strategy to the three-layer encoding tree in the structural entropy-based regularization framework, an encoding tree for clustering tasks is constructed, and soft assignment prediction is performed on the clustering task; The construction method includes: first, using the K-Means algorithm to obtain pseudo centers, wherein corresponds to the number of clusters; in order to convert hard assignment to probability (soft) assignment, apply the softmax operation to the similarity score calculated using the Gaussian kernel, for a given data point , let be its embedding, be the th cluster center, The distance between and is calculated as: The obtained distance is used to define soft assignment probabilities using a temperature scaled softmax function, which is expressed as: where denotes the th pseudo-center; denotes the cluster center of the th pseudo-center; is a temperature hyper-parameter that controls the sharpness of the distribution, and the temperature parameter controls the sharpness of the soft assignment, i.e., low values produce deterministic (more sharp) distributions, while high values encourage softer, more uniform distributions, such that .
[0009] S5, based on the structure entropy-based regularization framework, a structure entropy-guided unsupervised multi-modal clustering framework (SEUMC) is constructed, and the preprocessed esophageal cancer multi-modal data is input into the structure entropy-guided unsupervised multi-modal clustering framework (SEUMC) to perform multi-modal data clustering; the multi-modal task is to fuse multiple data types (such as text + image + gene) data; The clustering method is: S5.1, input multi-modal data to obtain multi-modal data modalities; According to the characteristics of medical data, a modal-specific encoder is used to extract features, a fine-tuned pre-trained BERT model is used, and a [CLS] token embedding is used as a sentence-level feature, which is then projected to a low-dimensional space through linear projection; For videos, Swin Transformer is used to obtain frame-level features, and WavLM is used to extract audio features after waveform processing; After feature extraction, modal-specific representations are obtained, which correspond to the modalities of multi-modal data text , multi-modal data video and multi-modal data audio respectively; specifically, denotes the dimension of the text embedding; denotes the video length, is the feature dimension; and respectively represent the feature dimension of audio length and audio modality; such as the medical record text + endoscopic video + inquiry recording data of esophageal cancer patients; S5.2, model the inter-modal association through a graph structure, and introduce a gating mechanism to dynamically adjust the importance of intra-modal features, so as to realize cross-modal semantic alignment; The method is: taking the text feature as a query node, and taking the feature vector of the visual / auditory modality as a key-value node, to construct a cross-modal bipartite graph , wherein: contains a text node and all feature nodes of modality M, represents the edge between the text node and each feature node of modality , and the weight is determined by semantic association; The attention weight of the text to the modality is calculated through a multi-head graph attention mechanism: In the formula, = , which represents the set of multi-modal data video and multi-modal data audio ; is the number of attention heads, represents the index of the attention head number; , , is the projection matrix of the query (Query), key (Key) and value (Value) of each head; represents a LeakyReLU activation function, which enhances the non-linear modeling capability; represents a scaling factor, which is used to control the numerical range of dot product (Dot-Product) to prevent gradient vanishing or explosion; by calculating the weighted sum of each feature vector of modality M, the “soft selection” of cross-modal semantic information is realized; S5.3, after cross-modal semantic alignment, introduce a gated self-attention mechanism to optimize the text and aligned cross-modal features within the modality, and retain the modality-specific structure (such as the spatial texture of medical images and the timing rhythm of speech signals), and the expression is as follows: In the formula, represents a gated self-attention mechanism; respectively represent the multi-modal data text after intra-modal optimization and the aligned cross-modal features; The steps include: S5.3.1, perform a semantic fusion operation on the text features and the cross-modal alignment features to obtain spliced features, and the expression is as follows: wherein, represents the concatenation feature; S5.3.2, linearly projecting the concatenation feature to generate a query of the concatenation feature , key , value ; S5.3.3, querying the concatenation feature , key , value Introducing a gating vector to dynamically adjust the attention score; The expression of the gating vector is as follows: wherein, represents a Sigmoid activation function, used to compress the input to the interval (0, 1); represents a learnable weight matrix, used to control the generation direction of the gating after feature fusion; The expression of dynamically adjusting the attention score is as follows: wherein, represents element-wise multiplication, and the gating vector can suppress irrelevant features (such as noise regions in images and background noise in speech); S5.3.4, outputting a training result that preserves the original feature information through residual connection and layer normalization operation, the expression is as follows: ; S5.4, introducing a cross-modal mutual information maximization target loss function to ensure that the visual and acoustic modalities are aligned to a unified text semantic space; The expression of the cross-modal mutual information maximization target loss function is as follows: wherein, represents the text-anchored visual feature (such as the representation of the endoscope image aligned with the text); represents the text-anchored acoustic feature (such as the representation of the inquiry recording aligned with the text); represents the conditional probability distribution; represents the expectation; The way to ensure that the visual and acoustic modalities are aligned to a unified text semantic space is: by minimizing the loss function, forcing the text-anchored visual feature and the acoustic feature to share high-order semantic information in the latent space (such as the image texture corresponding to “cough” and the speech keyword).
[0010] S5.5, introducing enhanced view contrast learning strategy to enhance cross-modal discriminative representation learning; The specific of enhanced view contrast learning strategy is: constructing multiple semantic-aligned but structurally diverse representations and forcing their consistency. Firstly, the optimized intra-modal features are fused to obtain joint representations: In the formula, represents the optimized text features (after gated self-attention); represents the optimized text-anchored visual features (text + image aligned representation); represents the optimized text-anchored acoustic features (text + speech aligned representation); represents the feature concatenation along the channel dimension; is a projection layer for mapping to a shared embedding space; three views and are enhanced perspectives of the same latent semantic concept; Subsequently, instance-level contrast learning is performed between the three views to promote cross-modal consistency; for each positive sample pair in the generated enhanced sample , the contrast loss is defined as: In the formula, represents the cosine similarity; represents the temperature scaling factor; and represent the multi-modal embeddings (from three views and ) of the positive sample and the sample , respectively; represents the projection head; represents the temperature coefficient, which controls the sharpness of the distribution (the smaller the contrast is more strict); represents other samples, covering all non-matching samples (such as features of different patients), pushing the model to distinguish the semantic differences between different instances; The final contrast objective of the entire batch is the average of all valid positive sample pairs, expressed as follows: In the formula, represents the positive sample index set (multi-modal view) that belongs to the same category as the sample , and is the enhanced sample index set that shares the same category as ; represents the contrast loss for a single positive sample; represents the average contrastive loss of all valid positive samples in the whole batch; this contrastive objective encourages the fused representation to be close to its modality-specific counterpart representation while pushing away mismatched views from other instances, thus improving the alignment and clusterability of the multi-modal embeddings; S5.6, guiding multi-modal data clustering through a structural entropy-based regularization framework; The way of guiding multi-modal data clustering through a structural entropy-based regularization framework is: After the multi-modal contrastive pre-training phase, semantic representations that fuse information from all modalities are obtained These representations constitute the unified feature space for the downstream clustering task Further division into meaningful medical semantic clusters (such as diabetes subtypes, postoperative complication types) is required, and the way to divide is: using the K-Means++ algorithm to divide the samples into semantic groups; Introducing a contrastive loss Utilizing the sample to capture the high-level similarity relationship between the pair of samples; Introducing a structural entropy-based regularization loss formula Relieving representation collapse (i.e., the problem that sample embeddings become difficult to distinguish and class boundaries become blurred), thus improving the accuracy and robustness of clustering; The overall training objective combines instance-level contrastive alignment and global structural regularization, and the total loss formula is: In the formula, represents a balance coefficient for controlling the contribution of the structural entropy term.
[0011] Advantages of the present application The present application introduces the structural entropy theory to construct a three-layer coding tree to simulate the "gene→pathway→tissue→clinical phenotype" hierarchy (with esophageal cancer as the root node and molecular subtypes as the intermediate layer), maximize the structural entropy of the intermediate layer, and force the model to learn hierarchical feature representations. For the tree-like hierarchical structure of esophageal cancer electronic medical records and the multi-pathway collaborative network of gene expression data, traditional CNN / RNN can only extract single variable features, ignoring the structural correlation between features, resulting in difficulty in reflecting the pathological mechanism of esophageal cancer. The present application solves the problem of insufficient utilization of structural information of esophageal cancer data by traditional methods.
[0012] The present application designs different strategies for different tasks. In the scene where esophageal cancer rare subtypes are insufficient, the existing algorithm lacks semantic guidance, which easily leads to confusion of feature embedding. The present application designs, such as classification task: adopt probability coding discrete structure; regression task: introduce probability coding tree, divide survival time into discrete box, calculate soft assignment probability based on gene expression distance, avoid hard binning defects; clustering task: combine K-Means pseudo center and Gaussian kernel soft assignment, generate probability distribution through temperature scaling softmax, relieve the representation collapse of rare subtypes, the present application solves the problem of "representation collapse" in unsupervised or semi-supervised scene.
[0013] The present application realizes semantic space alignment and hierarchical structure preservation in multi-modal medical data fusion by fusing multi-modal data such as esophageal endoscopic images. BRIEF DESCRIPTION OF DRAWINGS
[0014] Figure 1 The step flowchart of the present application. DETAILED DESCRIPTION The present application will be further described in detail below in combination with specific embodiments.
[0015] A structure entropy guided esophageal cancer multi-modal data discriminative representation learning method, comprising the following steps: S1, collecting esophageal cancer multi-modal data including text data, image data and gene data; Text data is obtained through diagnosis information and medication records in electronic medical records, which can be obtained from hospital systems and desensitized; Image data is obtained through CT and endoscopic images, and is collected from the image system together with shooting parameters; Gene data includes physiological signals and test data (such as electrocardiogram, blood indicators), respectively from monitoring equipment and laboratory system.
[0016] S2, performing pretreatment operation on the collected esophageal cancer multi-modal data; Medical data preprocessing is a key step to provide high-quality training samples for machine learning models by systematically organizing and optimizing raw esophageal cancer medical data to improve data quality and enhance data usability; Because esophageal cancer medical data is widely sourced and has complex structure, there are problems such as text errors and redundant records in electronic medical records, endoscopic image resolution and format are different, gene test data contains missing values and outliers, etc. Therefore, data preprocessing is a necessary prerequisite to ensure the effectiveness of esophageal cancer data analysis and model training; The medical data preprocessing realizes the transformation of the original esophageal cancer data from disorder to standardization through a systematic data governance process. The preprocessed data can significantly improve the modeling ability of machine learning models for the pathological mechanism of esophageal cancer, improve the training efficiency and prediction accuracy of the models, and lay a solid data foundation for artificial intelligence applications such as intelligent auxiliary diagnosis and treatment decision support of esophageal cancer. The preprocessing operation includes: S2.1, data cleaning operation; data cleaning is the first step of preprocessing, aiming to identify and correct errors, duplicates and invalid information in the data; The data cleaning operation includes: For the spelling errors of diagnosis names and repeated medical records in the esophageal cancer electronic medical record system, the data deduplication algorithm and medical terminology verification mechanism are used for correction and elimination; For missing values in gene expression data and blood test data, the same disease mean filling method or time series interpolation method is used for completion; For the abnormal lesion region annotation in endoscopic images and the obvious outliers of body temperature exceeding the physiological limit, combined with the clinical professional knowledge of esophageal cancer, the outliers are marked or deleted; S2.2, data format unification operation; data format unification is the core step to realize data standardization; Data format unification includes: For text data processing, non-structured diagnosis description (such as "esophageal upper segment visible ulcerative mass") needs to be mapped to the International Classification of Diseases (ICD-10) coding system to eliminate differences in terminology; For medical image data (esophageal endoscopic images, CT images), normalization processing is required, including unifying image resolution (adjusting to 224x224 pixels), gray scale standardization and format conversion; For gene data, standardized transformation is required to map the detection results of different experimental platforms to a unified numerical interval to ensure data consistency and comparability; S2.3, multi-modal data alignment operation; multi-modal data alignment focuses on aligning heterogeneous data generated at different time nodes for the same esophageal cancer patient; The multi-modal data alignment operation is as follows: taking the time of diagnosis of esophageal cancer as the reference, the endoscopic image, pathological section image, gene test result and electronic medical record text multi-source data are spatio-temporally aligned to establish a correlation network; for example, the endoscopic image features of a patient at the time of diagnosis are associated with the same period gene expression data and pathological diagnosis text, which can effectively capture the potential association between image features, molecular features and clinical phenotypes in the pathological evolution process from intraepithelial neoplasia to invasive metastasis of esophageal cancer, and provide structured data support for subsequent analysis.
[0017] S3, construct a regularization framework based on structural entropy to capture the hierarchical structure and complex dependency in medical data; To solve the problem of insufficient utilization of medical data structure information in traditional methods, the present application designs a structural entropy regularization framework of a three-layer hierarchical coding tree, maximizes the structural entropy of the intermediate layer nodes, constrains the distribution of latent variables to promote their separation, thereby enhancing the expression capacity of the latent representation, enabling the model to learn a feature representation with clear hierarchical division and better reflecting the pathological mechanism of the disease; The input of the structural entropy regularization framework is the esophageal cancer multi-modal data features after pre-processing operation, and the adjacency matrix and allocation matrix are constructed after data embedding representation. The regularization structural entropy loss is calculated through the adjacency matrix and allocation matrix and the three-layer hierarchical coding tree, and the dynamic weight adjustment is introduced to improve the robustness and interpretability of the structural entropy regularization framework. The three-layer hierarchical coding tree includes: root node, intermediate layer and leaf node; the data is input into the intermediate layer after passing through the root node, the hierarchical separation operation of multi-modal data features is performed by maximizing the structural entropy of the intermediate layer nodes, and the individual data samples are output through the leaf nodes; The structural entropy of the intermediate layer nodes of the three-layer coding tree As follows: In the formula, represents the number of categories, represents the index of the number of categories; represents the constructed medical data graph; represents the set of intermediate layer nodes in the coding tree, for each , represents the total weight of the edges between the subtree and its complement ; represents the volume of ; represents the structural entropy of the intermediate layer nodes; The construction steps include: S3.1, encode the input pre-processed multi-source data into embedding representation , which is used to unify the multi-source data into a computable low-dimensional embedding and retain semantic information; S3.2, construct adjacency matrix and allocation matrix based on embedding representation ; The adjacency matrix is represented as: , wherein, Represents the sigmoid activation function that ensures that the elements in A are positive; by calculating the adjacency matrix It can be expressed from the embedding Get medical data map ,picture It reflects the similarity of underlying characteristics; Define the binary allocation matrix as: ,in, represents the number of leaf nodes, represents the number of intermediate nodes (i.e., clusters or semantic groups); each element in the binary assignment matrix ,like Indicates the The leaf nodes are assigned to categories, which means a hard partitioning of nodes; this discrete assignment structure makes it possible to explicitly model hierarchical relationships and supports the construction of multi-scale semantic trees that capture local and global organization in the latent space; this adaptive adjacency can capture high-order dependencies between samples, thereby improving representation quality; S3.3. Based on the adjacency matrix and the allocation matrix, the structural entropy of the middle layer nodes of the three-layer coding tree is obtained. Regularized loss form and introduction of dynamic weight factors; Structural entropy of the middle layer nodes of the three-layer coding tree The regularization loss is as follows: Where, Indicates the shape All-1 matrix; Representation diagram The total volume of ,picture The sum of all edge weights in reflects the total association strength of the entire medical data graph; Representation category With Figure The total weight of the edges between the parts except the all-1 matrix, the adjacency matrix and the assignment matrix, that is ; Representation category The volume of ; The dynamic weight factor expression is as follows: Where, Indicates the Nodes in The distribution entropy at the moment is expressed as follows: Where, a number index representing a leaf node; The loss function expression after introducing the dynamic weight adjustment is as follows: ; The dynamic weight adjustment mechanism enhances regularization to guide the structural learning of medical data (such as identifying key pathways or disease characteristics), and reduces constraints to optimize detailed features (such as specific gene or image detail representation), thereby improving the robustness and interpretability of the model in medical tasks; Through the above regularization framework based on structural entropy, the hierarchical structure and complex dependency relationship in medical data are effectively captured.
[0018] S4, input the preprocessed esophageal cancer multi-modal data into the constructed regularization framework based on structural entropy, and perform single-modal task classification and clustering based on structural entropy; The single modal includes: classification task (text), regression task (gene data) and clustering task (image data); The single-modal task classification includes: hard assignment and soft assignment; wherein the soft assignment mode includes: introducing soft label and pseudo label strategy in the three-layer coding tree of the regularization framework based on structural entropy; The clustering method includes: S4.1, a discrete structure with probability coding is used for hard classification prediction of the classification task; The probability coding unifies the probability representation learning and the task prediction in a single module. In order to realize gradient-based optimization, the reparameterization technique is applied, which allows sampling from the distribution without introducing gradient bias; Unlike existing methods, the present application uses an encoder-only architecture, which removes the decoder, directly samples and infers labels from the learned latent distribution; the probability coding of the present application integrates the probability coding and task prediction into a single encoder module through the network, and applies a non-parametric operation to obtain the prediction output; The overall loss of the structural entropy probability coding (SEPC) of the discrete structure with probability coding for the classification task is as follows: In the formula, represents the task loss of the probability coding; represents a hyperparameter, which controls the weight of the regularization loss based on structural entropy , forcing the model to learn diagnostic features with hierarchical discriminability (such as key pathological indicators to distinguish pneumonia from influenza); S4.2, an encoding tree for regression task is constructed by introducing a soft label strategy in the three-layer coding tree of the regularization framework based on structural entropy, and a soft assignment prediction is performed on the regression task; Discretized coding trees constructed with real labels are essentially designed for classification tasks; however, for other discriminative representation learning scenarios (e.g., regression and unsupervised clustering), this hard partitioning scheme does not directly apply; to address this limitation and extend the applicability of coding trees, two soft assignment strategies are proposed, enabling data points to be flexibly and continuously assigned to the hierarchical structure; For regression tasks, a soft label approach is introduced, where each data point is assigned to a probability distribution over all possible classes; based on this idea, a probabilistic coding tree is proposed, extending the structural entropy theory to the soft classification scenario; unlike traditional coding trees that rely on hard assignments, this approach relaxes the constraint that each child node is linked to only one parent node, and instead probabilistically connects each child node to all higher-level nodes to capture subtle structural dependencies; The way to build a discretized coding tree includes: By first discretizing the continuous label space, the structural entropy framework is extended to the regression scenario; this process first divides the continuous regression label space into intermediate nodes, i.e., discrete bins; each bin is represented by a center point For each data point, the distance to all bin centers is calculated to estimate the class affinity: where, represents the real regression label vector; represents the absolute distance matrix between the sample label and the bin centers, To obtain soft assignment, the softmax operation is applied to the negative distance values, allowing closer bins to have higher assignment probabilities: where, represents the soft assignment probability matrix (the probability of each sample belonging to each bin); This soft assignment allows each data point to contribute to multiple semantic regions in the latent space, defining the volume and cut edge weights under continuous conditions, and showing that their forms are consistent with the discrete setting, where each intermediate node has a soft volume calculated as: where, represents the total number of samples; represents the sample, and represents multi-source data; represents the soft assignment of the th node to the th class; represents the degree of the sample (the sum of the adjacency matrix row); In the soft assignment setting, quantified and its complement The expected cut edge weight between , one node belongs to and the other belongs to its complement, is computed as the weighted sum of edge connections, where the weight is given by the product of the probability that each pair of vertices lies on opposite sides of the partition; for each pair of nodes , the probability that one node belongs to and the other belongs to its complement is: where denotes the element in the soft assignment matrix The total expected cut weight is expressed as: where we let The final regularizer based on structural entropy is identical to , and denotes the element of the adjacency matrix for the sample, and denotes the strength of the association between and The key difference is that the assignment matrix now encodes soft probabilities rather than hard class membership, and this formulation enables structural entropy to be effectively used for regression tasks by modeling probabilities and soft hierarchical encoding; S4.3, by introducing a pseudo-label strategy to the three-layer encoding tree in the structural entropy-based regularization framework, an encoding tree for clustering tasks is constructed, and soft assignment prediction is performed for clustering tasks; The construction method includes: first, using the K-Means algorithm to obtain pseudo centers, wherein corresponds to the number of clusters; in order to convert hard assignment to probability (soft) assignment, the softmax operation is applied to the similarity score calculated using the Gaussian kernel, and for a given data point , let be its embedding, be the distance between and be calculated as: ; The obtained distance uses a temperature scaled softmax function to define the soft assignment probability, and the expression is as follows: wherein denotes the first pseudo center; denotes the cluster center of the pseudo center; is a temperature hyperparameter that controls the sharpness of the distribution, and the temperature parameter Controlling the sharpness of soft assignment, i.e. low values produce a more determined (sharper) distribution, while high values encourage a softer, more uniform distribution, leads to The regularization loss formula based on structural entropy is: .
[0019] S5. Based on the structural entropy-based regularization framework, a structural entropy-guided unsupervised multi-modal clustering framework (SEUMC) is constructed, and the preprocessed esophageal cancer multi-modal data is input into the structural entropy-guided unsupervised multi-modal clustering framework (SEUMC) for multi-modal data clustering; the multi-modal task is to fuse multiple data types (such as text + image + gene) data; The clustering method is: S5.1, input multi-modal data to obtain multi-modal data modal; According to the characteristics of medical data, modal-specific encoders are used to extract features, a pre-trained BERT model is fine-tuned, and [CLS] token embeddings are used as sentence-level features, which are then projected to a low-dimensional space through linear projection; For videos, Swin Transformer is used to obtain frame-level features, and WavLM is used to extract audio features after waveform processing; After feature extraction, modal-specific representations are obtained, which correspond to the modal of multi-modal data text , multi-modal data video and multi-modal data audio respectively; specifically, represents the dimension of the text embedding; represents the video length, is the feature dimension; and represent the audio length and the feature dimension of the audio modal respectively; for example, the medical record text + endoscopic video + audio recording data of an esophageal cancer patient; S5.2, model the inter-modal association through a graph structure and introduce a gating mechanism to dynamically adjust the importance of intra-modal features to achieve cross-modal semantic alignment; The method is: the text features are taken as query nodes, and the feature vectors of visual / acoustic modalities are taken as key-value nodes to construct a cross-modal bipartite graph , wherein: contains text nodes and all feature nodes of modal M, represents the establishment of edges between the text node and each feature node of modal , and the weight is determined by semantic association; The attention weight of the text to the modal is calculated through a multi-head graph attention mechanism: where, represents a collection of multi-modal data video and multi-modal data audio ; is the number of attention heads, represents the index of the attention head; is the projection matrix of Query, Key and Value of each head; represents the LeakyReLU activation function, which enhances the non-linear modeling ability; represents the scaling factor, which is used to control the numerical range of Dot-Product, preventing gradient vanishing or explosion; by calculating the weighted sum of each feature vector of the modality M, the "soft selection" of cross-modal semantic information is realized; S5.3, after the alignment of cross-modal semantics, the gated self-attention mechanism is introduced, and the modal intra-optimization is performed on the text and the aligned cross-modal features, and the modal specificity structure (such as the spatial texture of medical images and the timing rhythm of speech signals) is preserved, and the expression is as follows: where, represents the gated self-attention mechanism; respectively represent the multi-modal data text after modal intra-optimization and the aligned cross-modal features; The steps include: S5.3.1, perform a semantic fusion operation on the text features and the cross-modal alignment features to obtain spliced features, and the expression is as follows: where, represents the spliced features; S5.3.2, linearly project the spliced features to generate the query , key , and value of the spliced features; S5.3.3, introduce a gating vector to the query , key , and value of the spliced features to dynamically adjust the attention score; The expression of the gating vector is as follows: where, represents the Sigmoid activation function, which is used to compress the input to the interval (0, 1); denotes the learnable weight matrix for controlling the direction of the generated gate after feature fusion; The expression for dynamically adjusting the attention score is as follows: In the formula, denotes element-wise multiplication, and the gate vector Irrelevant features (such as noise regions in images and background noise in speech) can be suppressed; S5.3.4, through residual connection and layer normalization operation, output the training result which preserves the original feature information, the expression is as follows: ; S5.4, introduce cross-modal mutual information maximization target loss function, to ensure that the visual and acoustic modalities are aligned to the unified text semantic space; The expression of the cross-modal mutual information maximization target loss function is as follows: In the formula, denotes the text-anchored visual feature (such as the representation of the endoscopic image aligned with the text); denotes the text-anchored acoustic feature (such as the representation of the inquiry recording aligned with the text); denotes the conditional probability distribution; denotes expectation; The way to ensure that the visual and acoustic modalities are aligned to the unified text semantic space is to force the text-anchored visual feature and the acoustic feature share high-order semantic information (such as image texture and speech keywords corresponding to "cough") in the latent space.
[0020] S5.5, introduce enhanced view contrast learning strategy to enhance cross-modal discriminative representation learning; The specific of the enhanced view contrast learning strategy is to construct multiple semantically aligned but structurally diverse representations and force their consistency. First, fuse the optimized intra-modal features to obtain a joint representation: In the formula, denotes the optimized text feature (after gated self-attention); denotes the optimized text-anchored visual feature (text + image aligned representation); denotes the optimized text-anchored acoustic feature (text + speech aligned representation); denotes feature concatenation along the channel dimension; is a projection layer for mapping to a shared embedding space; three views and enhanced perspectives of the same potential semantic concept; Subsequently, instance-level contrastive learning is performed among the three views to promote cross-modal consistency; for each positive sample pair in the generated enhanced samples The contrastive loss is defined as: wherein, denotes the cosine similarity; denotes the temperature scaling factor; and denote the multi-modal embeddings (from the three views and ) of the positive sample and the sample ; denotes the projection head; denotes the temperature coefficient, which controls the sharpness of the distribution (the smaller, the stricter the contrast); denotes other samples, covering all non-matching samples (such as features of different patients), to promote the model to distinguish the semantic differences between different instances; The final contrastive objective for the entire batch is the average of all valid positive sample pairs, expressed as follows: wherein, denotes the set of positive sample indices (multi-modal views) belonging to the same category as the sample , is the set of enhanced sample indices sharing the same category as ; denotes the contrastive loss for a single positive sample; denotes the average contrastive loss for all valid positive samples in the entire batch; this contrastive objective encourages the fused representation to be close to its modality-specific counterpart, while pushing away mismatched views from other instances, thereby improving the alignment and clusterability of the multi-modal embeddings; S5.6, guide multi-modal data clustering through a structural entropy-based regularization framework; The way to guide multi-modal data clustering through a structural entropy-based regularization framework is: After the multi-modal contrastive pre-training phase, semantic representations that fuse all modality information are obtained These representations constitute a unified feature space for downstream clustering tasks which need to be further divided into meaningful medical semantic clusters (such as diabetes subtypes, postoperative complication types); the way to divide is to use the K-Means++ algorithm to divide the samples into semantic groups; Introduce the contrastive loss , utilize the sample to capture the high-level similarity relationship between the sample pairs; A regularization loss formula based on structural entropy is introduced , alleviate the representation collapse (i.e., the problem that sample embeddings become indistinguishable and class boundaries are blurred), thereby improving the accuracy and robustness of clustering; The overall training target integrates instance-level contrast alignment and global structure regularization, and the total loss formula is represented as: In the formula, Balancing coefficient, used to control the contribution of the structural entropy term.
[0021] The structural entropy theory provides a new path for breaking the situation. Entropy, as a measure of system uncertainty, can quantify the hierarchical dependence of esophageal cancer data. The essence of the occurrence and development of esophageal cancer is the dynamic evolution of multi-level biological networks, and the feature space should present a tree-like hierarchy of “driver gene variation → signal pathway abnormality → cell malignant transformation → tissue invasion → organ metastasis”. By maximizing the structural entropy, the model can learn a hierarchical feature representation that captures the complex mapping of phenotype and genotype (such as EGFR amplification status).
[0022] Compared with traditional methods, the structural entropy framework explicitly models the hierarchical correlation between features, forcing the model to capture cross-scale semantics from micro-molecules to macro-clinical phenotypes in esophageal cancer data. In the unsupervised scenario, the soft assignment strategy based on structural entropy can guide the separation of feature embeddings and alleviate the representation collapse caused by the lack of rare subtype samples; in multi-modal fusion, through text-anchored attention and structural entropy regularization clustering, semantic alignment and hierarchical preservation of endoscopic images, pathology reports, and genetic data can be achieved, improving the accuracy of prognosis prediction.
[0023] This innovation provides a more interpretable and robust general framework for esophageal cancer AI diagnosis and treatment, and is expected to promote the transition of esophageal cancer from “empirical diagnosis and treatment” to “data-driven precision stratified treatment”, especially in the application of rare subtype identification and treatment response prediction.
[0024] In summary, the present application solves the problems of hierarchical modeling, unsupervised separation, and multi-modal fusion of esophageal cancer data through the structural entropy theory, and constructs a cross-scale semantic framework from molecules to clinics, providing a new paradigm for precision diagnosis and treatment of esophageal cancer and promoting the application of AI in subtype identification and efficacy prediction.
[0025] Although embodiments of the present application have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made without departing from the principles and spirit of the present application, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A discriminative representation learning method for esophageal cancer multimodal data guided by structural entropy, characterized by: The following steps are involved: S1. Collect multimodal data of esophageal cancer, including text data, image data, and gene data; S2. performing preprocessing operations on the collected esophageal cancer multimodal data; S3. Construct a regularization framework based on structural entropy; The structural entropy-based regularization framework includes: a three-layer coding tree that maximizes the structural entropy of intermediate layer nodes; the three-layer coding tree includes: a root node, an intermediate layer, and leaf nodes; data passes through the root node and is input into the intermediate layer, and a hierarchical separation operation of esophageal cancer multimodal data features is performed by maximizing the structural entropy of the intermediate layer nodes, and individual data samples are output through the leaf nodes; S4. The pre-processed esophageal cancer multimodal data are input into the constructed regularization framework based on structural entropy, and single-modal task classification and clustering are performed based on structural entropy; The single-modal tasks include: classification tasks, regression tasks and clustering tasks; The unimodal task classification includes: hard assignment and soft assignment; The soft assignment method includes: introducing a strategy of soft labels and pseudo labels to the three-layer coding tree in a regularization framework based on structural entropy; S5. Based on the regularization framework based on structural entropy, an unsupervised multimodal clustering framework guided by structural entropy is constructed. The preprocessed esophageal cancer multimodal data are input into the unsupervised multimodal clustering framework guided by structural entropy for multimodal data clustering, and the discriminative representation learning of esophageal cancer multimodal data is completed.
2. The method for learning discriminative representation of esophageal cancer multimodal data guided by structural entropy according to claim 1, characterized in that: The steps of constructing a regularization framework based on structural entropy include: S3.
1. Encode the pre-processed multi-source data into an embedded representation , used to unify multi-source data into a computable low-dimensional embedding and preserve semantic information; S3.
2. Embedding-based representation Construct adjacency matrix and assignment matrix; The adjacency matrix is represented as: ,in, Represents the sigmoid activation function that ensures that the elements in A are positive; Define the binary distribution matrix as: ,in, represents the number of leaf nodes, Indicates the number of intermediate nodes; S3.
3. Based on the adjacency matrix and the assignment matrix, the regularized loss form of the structural entropy of the middle layer nodes of the three-layer coding tree is obtained, and a dynamic weight factor is introduced; The regularized loss form of the structural entropy of the middle layer nodes of the three-layer encoding tree is as follows: Where, Indicates the shape All-1 matrix; Representing medical data graphs The total volume of , medical data graph The sum of all edge weights in reflects the total association strength of the entire medical data graph; Representation category Graph with medical data The total weight of the edges between the parts except the all-one matrix, the adjacency matrix and the assignment matrix; Representation category volume; The dynamic weight factor expression is as follows: Where, Indicates the Nodes in The distribution entropy at each moment, and the loss function expression after introducing dynamic weight adjustment are as follows: 。 3. The method for learning discriminative representation of esophageal cancer multimodal data guided by structural entropy according to claim 1, characterized in that: The steps of inputting the pre-processed esophageal cancer multimodal data into the constructed regularization framework based on structural entropy and performing single-modal task classification and clustering based on structural entropy include: S4.1, using a probabilistically coded discrete structure to perform hard classification predictions for classification tasks; S4.
2. We introduce a soft label strategy into the three-layer coding tree in the regularized framework based on structural entropy to construct a coding tree for regression tasks and perform soft assignment prediction for regression tasks. Build methods include: Divide the continuous regression label space into intermediate nodes, i.e. discrete boxes; each box consists of a central point It means that for each data point, its distance to all bin centers is calculated to estimate the category affinity: Where, represents the true regression label vector; Represents the absolute distance matrix between the sample label and the center of each box, ; To obtain a soft assignment, apply a softmax operation to the negative distance values: ,in, represents the soft assignment probability matrix; When the distribution matrix When , the regularized loss formula based on the structural entropy in the encoding tree for regression tasks is the same as same. S4.
3. We construct a coding tree for clustering tasks by introducing a pseudo-label strategy into the three-layer coding tree in a regularized framework based on structural entropy, and perform soft assignment prediction for clustering tasks. The construction methods include: First, use the K-Means algorithm to obtain pseudo-centers, among which Corresponding to the number of clusters; To convert hard assignments into soft assignments, a softmax operation is applied to the similarity scores calculated using the Gaussian kernel. For a given data point ,make For its embedding, For the cluster centers, and The distance between is calculated as: ; The obtained distance uses the temperature-scaled softmax function to define the soft assignment probability, which is expressed as follows: Where, Indicates the pseudo center number; Indicates the The cluster centers of pseudo centers; is a temperature hyperparameter that controls the sharpness of the distribution; When the distribution matrix , the regularized loss formula of the structural entropy of the encoding tree for clustering tasks is the same as same.
4. The method for learning discriminative representation of esophageal cancer multimodal data guided by structural entropy according to claim 1, characterized in that: The steps of constructing an unsupervised multimodal clustering framework guided by structural entropy based on a regularization framework based on structural entropy, inputting the preprocessed esophageal cancer multimodal data into the unsupervised multimodal clustering framework guided by structural entropy to perform multimodal data clustering, and completing the discriminative representation learning of esophageal cancer multimodal data include: S5.
1. Input multimodal data to obtain multimodal data modalities; Based on the characteristics of medical data, a modality-specific encoder is used to extract features. A pre-trained BERT model is fine-tuned, and [CLS] token embedding is used as sentence-level features. The features are then linearly projected into a low-dimensional space. For video, Swin Transformer is used to obtain frame-level features, and audio features are extracted using WavLM after waveform processing; After feature extraction, modality-specific representations are obtained , corresponding to the multimodal data text , multimodal data video and multimodal data audio The modality; S5.
2. Modeling inter-modal relationships through graph structures and introducing a gating mechanism to dynamically adjust the importance of intra-modal features to achieve cross-modal semantic alignment. The alignment method is to use text features as query nodes and feature vectors of visual / acoustic modalities as key nodes to construct a cross-modal bipartite graph. ,in: Contains text nodes and all feature nodes of modality M, Represents text nodes and modals Establish edges between each feature node, and the weight is determined by the semantic relevance; S5.
3. After cross-modal semantic alignment, a gated self-attention mechanism is introduced to perform intra-modal optimization on the text and aligned cross-modal features, preserving modality-specific structures (such as the spatial texture of medical images and the temporal rhythm of speech signals). The expression is as follows: Where, represents the gated self-attention mechanism; Represent the optimized multimodal data text within the modality and the aligned cross-modal features respectively; S5.
4. Introduce a cross-modal mutual information maximization objective loss function to ensure that the visual and acoustic modalities are aligned to a unified textual semantic space; The objective loss function for maximizing cross-modal mutual information is expressed as follows: Where, Visual features representing text anchors (e.g., alignment of the endoscopic image with the text); Acoustic features that represent text anchors (e.g., representations of medical interview recordings aligned with text); represents the conditional probability distribution; express expectations; S5.
5. Introduce enhanced view contrast learning strategy to enhance cross-modal discriminative representation learning; First, the optimized intra-modality features are fused to obtain a joint representation: Where, Represents the optimized text features; Visual features representing optimized text anchors; The acoustic features representing the optimized text anchors; represents the concatenation of features along the channel dimension; is the projection layer mapped to the shared embedding space; three views and are considered as enhanced perspectives of the same underlying semantic concept; Subsequently, instance-level contrastive learning is performed between the three views to promote cross-modal consistency; for each positive sample pair in the generated enhanced samples , the contrastive loss is defined as: Where, represents cosine similarity; represents the temperature scaling factor; and Represents positive samples and samples Multimodal embedding of Indicates projection head; represents the temperature coefficient, which controls the sharpness of the distribution; Indicates other samples, Cover all non-matching samples and push the model to distinguish the semantic differences between different instances; The final comparison target for the entire batch is the average of all valid positive sample pairs, expressed as follows: Where, Representation and Sample The index set of positive samples belonging to the same category, For Sharing the enhanced sample index set of the same category; Represents the contrast loss for a single positive sample; Represents the average contrast loss of all valid positive samples in the entire batch; S5.
6. Guiding multimodal data clustering via a regularization framework based on structural entropy. The way to guide multimodal data clustering through the regularization framework based on structural entropy is: After the multimodal contrast pre-training phase, a semantic representation that integrates all modal information is obtained. , these representations constitute a unified feature space for downstream clustering tasks ,need to be further divided into meaningful medical semantic clusters, and the division method is as follows:,using the K-Means++ algorithm to divide the samples into semantic groups; Introducing contrast loss ,Use samples to capture high-level similarity relations between pairs of samples; Introducing a regularization loss formula based on structural entropy , alleviating representation collapse, thereby improving clustering accuracy and robustness; The overall training objective combines instance-level contrast alignment and global structure regularization, and the total loss formula is expressed as: Where, represents the balance coefficient, which is used to control the contribution of the structural entropy term.
5. The method for learning discriminative representation of esophageal cancer multimodal data guided by structural entropy according to claim 4, characterized in that: The steps of introducing a gated self-attention mechanism after cross-modal semantic alignment, performing intra-modal optimization on the text and aligned cross-modal features, and preserving the modality-specific structure include: S5.3.
1. Perform semantic fusion operation on text features and cross-modal alignment features to obtain concatenated features, which are expressed as follows: Where, Represents splicing features; S5.3.
2. Generate a query for stitching features by linearly projecting the stitching features ,key ,value ; S5.3.
3. Query of splicing features ,key ,value Introducing a gate vector to dynamically adjust the attention score; S5.3.
4. Through residual connection and layer normalization operations, the training results that retain the original feature information are output.
Citation Information
Cited By
Heart and cerebral vessel scientific research multi-modal data semantic alignment method and medium
CN121191795A