Multi-illumination esophageal cancer early screening and labeling system based on probability embedding

By adopting a multi-light esophageal cancer early screening and labeling system based on probability embedding in the detection of esophageal cancer, the problem of insufficient detection accuracy in traditional technology under different lighting conditions is solved, and higher detection accuracy and labeling accuracy are achieved, improving the robustness of image analysis.

CN119722597BActive Publication Date: 2025-06-06BEIHANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411768792.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-04
Publication Date
2025-06-06
Estimated Expiration
2044-12-04

AI Technical Summary

Technical Problem

Traditional medical imaging technology shows low adaptability and robustness when dealing with different lighting conditions, individual patient differences and diversity of imaging equipment, resulting in limited accuracy of early detection of esophageal cancer.

Method used

A multi-light esophageal cancer early screening and labeling system based on probability embedding is adopted. The system includes an encoder module, a multi-light fusion module, a classification and screening module, a structural entropy regularization module and an uncertainty-based labeling guidance module. A unified and robust image representation is generated through probability embedding and multi-light fusion technology, combining structural entropy regularization and uncertainty annotation guidance to improve detection accuracy and labeling accuracy.

Benefits of technology

It significantly improves the accuracy of early detection of esophageal cancer and the accuracy of labeling, enhances the robustness and adaptability of image analysis, and provides more reliable support for clinical diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119722597B_ABST
    Figure CN119722597B_ABST
Patent Text Reader

Abstract

The present application discloses a multi-illumination esophageal cancer early screening and labeling system based on probability embedding, which aims to provide more effective technical support for the early screening of esophageal cancer by improving the accuracy of early detection, ensuring the robustness of image analysis and optimizing the labeling accuracy of the labeling process. The system is used to perform the following steps: using the encoder module to perform the generation of probability embedding; using the multi-illumination fusion module to perform the generation of fusion embedding; using the classification and screening module to generate probability classification; using the structural entropy regularization module to perform structural entropy regularization; using an end-to-end training method and using the Adam optimizer to optimize the model comprehensive loss through the overall loss function to perform model training; based on the uncertainty-based labeling guidance module to perform labeling guidance, identify and determine high uncertainty areas.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of image data processing technology, and in particular to a multi-illumination esophageal cancer early screening and labeling system based on probability embedding. Background Art

[0002] With the rapid development of medical imaging technology, the accuracy and efficiency of cancer screening have become the focus of clinical medical research. Especially in the early detection of esophageal cancer, timely identification of lesions is crucial to improving patient survival. At present, mainstream medical imaging technologies such as endoscopy, ultrasound imaging and computed tomography (CT) have improved the diagnostic ability of cancer to a certain extent, but there are still many shortcomings.

[0003] Traditional medical imaging technology mostly relies on deterministic representation, that is, analyzing images through fixed image processing algorithms. This method often shows low adaptability and robustness when faced with physiological differences between patients, the diversity of imaging equipment, and the imaging quality under different lighting conditions. For example, the tissue characteristics of different patients, the morphological changes of lesions, and the differences in imaging environments can lead to significant differences in imaging results, which makes the lesion identification and annotation process based on traditional methods complicated and error-prone.

[0004] To solve this problem, researchers began to explore multi-illumination imaging technology. This technology can more comprehensively reveal the diversity and characteristics of tissues by collecting images under different illumination conditions. However, existing multi-illumination imaging methods are still insufficient in information integration and lack effective strategies to deal with the variability and uncertainty in multi-source medical imaging. These methods are usually unable to fully utilize the information obtained from different illumination conditions, resulting in limited accuracy in early cancer detection.

[0005] In the above context, this paper proposes a multi-illumination imaging technology based on probabilistic embedding. This technology can not only effectively fuse image information from different illumination conditions, but also produce a unified and robust representation through a probabilistic model. This representation method can fully consider the uncertainty in the imaging process, thereby enhancing the accuracy of image analysis. The introduction of this technology means that the recognition accuracy and annotation accuracy can be significantly improved when dealing with early detection of esophageal cancer, providing more reliable support for clinical diagnosis. Summary of the invention

[0006] In order to overcome the above technical defects, the purpose is to effectively integrate image information taken under various lighting conditions to improve the early screening and labeling of esophageal cancer. The embodiment of the present application provides a multi-lighting esophageal cancer early screening and labeling system based on probability embedding, including an encoder module, a multi-lighting fusion module, a classification and screening module, a structural entropy regularization module, and an uncertainty-based labeling guidance module. The system realizes multi-lighting esophageal cancer early screening and labeling, and performs the following operations:

[0007] S2. Perform probabilistic embedding generation using the encoder module:

[0008] A convolutional neural network based on ResNet-50 is used as the feature extractor for esophageal images. Probabilistic embedding is generated through two fully connected layers, where the first layer generates the mean of the probability distribution and the second layer generates the covariance of the probability distribution, and the outputs of the two fully connected layers are ensured to be positive.

[0009] S3. Generate fusion embedding using multi-illumination fusion module:

[0010] For at least two probabilistic embeddings from the encoder, the multi-head attention mechanism in the Transformer encoder architecture is used, with each head processing a different subset of embeddings to generate weighted embeddings, which are then integrated into the final fused embedding through a linear layer;

[0011] S4. Generate probabilistic classification using classification and screening modules:

[0012] The fusion embedding is input into the LLaVA large model after LoRA fine-tuning for classification processing, and the Softmax activation function is used to output the category probability distribution;

[0013] S5. Use the structural entropy regularization module to perform structural entropy regularization:

[0014] By constructing an encoding tree of latent representations and using a hierarchical clustering algorithm to group embeddings, the distribution of latent representations is optimized through structural entropy regularization to enhance the ability to distinguish different categories and maximize the entropy difference between categories, thereby capturing the structural information between latent variables;

[0015] S6. Model training:

[0016] Adopt end-to-end training mode and use Adam optimizer to optimize the comprehensive loss of the model through the overall loss function for model training. The overall loss function includes classification cross entropy loss and structural entropy regularization loss.

[0017] S7. Annotation guidance based on uncertainty annotation guidance module:

[0018] By evaluating the covariance matrix of each embedding, regions of high uncertainty are identified and marked.

[0019] Further optionally, in S2, ResNet-50 is a deep network constructed by stacking multiple residual blocks, each residual block includes two convolutional layers and a shortcut connection;

[0020] The ReLU activation function is used in the first fully connected layer to ensure that the output is positive;

[0021] The Softplus function is used in the second fully connected layer to ensure that the output is positive, thereby ensuring the positive definiteness of the covariance matrix.

[0022] Further optionally, in S3, using the multi-illumination fusion module to generate the fusion embedding specifically includes:

[0023] S3.1. Input embedding: Receive at least two probabilistic embeddings from the encoder, each embedding corresponding to an image under a specific lighting condition;

[0024] S3.2. Generate query, key and value: For each input embedding, generate query Q, key K and value vector V, calculated as follows: Q = zW Q , K = zW K , V = zW V , where W Q , W K and W V is the learned weight matrix;

[0025] S3.3. Calculate attention weight: For each attention head, calculate the attention weight A:

[0026] where d k is the dimension of the key vector;

[0027] S3.4, weighted output: Use the calculated attention weight A to weight the value vector V and output the result of each attention head: Attention(Q, K, V) = AV;

[0028] S3.5. Integrate multi-head outputs: Concatenate the outputs of all attention heads and transform them into the final fused embedding z through a linear layer f :z f =Concat(head 1 , head 2 ,...,head h )W O , where W O is the output weight matrix.

[0029] Further optionally, in S4, generating a probability classification using a classification and screening module specifically includes:

[0030] S4.1, perform LoRA fine-tuning on the LLaVA model;

[0031] S4.2. After completing the LoRA fine-tuning, the fine-tuned LLaVA-1.5 model is used to screen the esophageal images processed by the multi-illumination fusion module. The specific screening process is as follows:

[0032] S4.2.1. Feature extraction: The LLaVA-1.5 model is used to process the input fused embedding through its deep neural network structure, analyze the fused embedding information, identify the difference between normal tissue and abnormal tissue, and locate potential lesion areas;

[0033] S4.2.2, Class probability distribution output: After feature extraction is completed, based on the extracted features, it is determined whether each embedding has a lesion and is classified as normal or abnormal, generating a class prediction probability p for each region: p(y=1|z f )=σ(Wz f +b), where σ is the Sigmoid function, W is the weight matrix, and b is the bias term.

[0034] Further optionally, in S5, using the structural entropy regularization module to perform structural entropy regularization specifically includes:

[0035] S5.1. Construct a three-layer encoding tree, where the middle-layer nodes represent the categories of the classification task;

[0036] S5.2. Calculate the structural entropy of the coding tree: The structural entropy of the middle layer node is defined as:

[0037] where r is the number of categories, Is an intermediate node Its complement The sum of the weights of the edges between yes Volume;

[0038] Using the adjacency matrix And the allocation matrix C is used to calculate the structural entropy regularization loss, which is calculated as follows:

[0039] Structural entropy regularization loss L SE It is defined as constraining the distribution of the latent variable by maximizing the entropy difference between classes.

[0040] Further optional, in S6, classification cross entropy loss: used to measure the classification accuracy of the model for each image, denoted by L PC =L(θ), the expression is:

[0041] Among them, L(θ) is the loss function, N is the number of samples, y i is the true label, is the probability predicted by the model;

[0042] Structural entropy regularization loss: used to enhance the distribution of potential representation and maximize the entropy difference between categories. Its expression is: Where r is the number of categories, C is the assignment matrix, is the adjacency matrix;

[0043] Thus, the overall loss function is: L SEPC =L PC -γL SE , where γ is a hyperparameter used to control the weight of the structural entropy regularization loss.

[0044] Further optionally, in S7, the specific steps of marking guidance include:

[0045] S7.1. Uncertainty assessment: The variance of each embedding is assessed by analyzing the size of the diagonal elements.

[0046] S7.2. Marking of uncertainty areas: Based on the variance of the uncertainty assessment, areas with higher variance are marked as high uncertainty areas. When marking high uncertainty areas, the marking strategy adopted is: for areas that have been clearly marked as cancerous, use strong marking; for areas that are marked as high uncertainty but have not yet been confirmed as cancerous, use soft marking.

[0047] Further optionally, the esophageal image is acquired by a high-resolution medical imaging device under different lighting conditions, and is input into the system after image cleaning, image standardization and structuring, and data augmentation technology processing.

[0048] The embodiments of the present application adopt the above technical solution to achieve the following technical effects:

[0049] 1. Improve the accuracy of early detection: The primary goal of the system is to improve the accuracy of early detection of esophageal cancer. By converting each image under different lighting conditions into a probabilistic embedding, the system can effectively capture the variability and uncertainty in the image. This process ensures that the embeddings of normal and abnormal tissues are clearly clustered in the latent space, thereby improving the system's ability to identify early lesions of esophageal cancer. To achieve this goal, the image screening part of this technology uses the LLaVA-1.5 large model, which uses its powerful image understanding capabilities to efficiently screen input images to ensure that the most relevant images are provided to subsequent analysis links. Traditional methods often rely on fixed algorithms and are prone to missing subtle changes in early lesions, while models based on probabilistic embeddings allow the system to maintain consistency under different imaging conditions, improving sensitivity to early lesions.

[0050] 2. Ensure the robustness of image analysis: the type of imaging equipment and changes in lighting conditions. Traditional analysis methods often show low robustness when dealing with these changes, resulting in inconsistency in diagnostic results. By using a multi-illumination fusion module, this technology is able to integrate the probabilistic embeddings from different illumination conditions into a unified and robust representation. This fusion mechanism not only improves the adaptability of the system, but also ensures the consistency of the analysis results, and can effectively cope with different imaging environments, thereby enhancing the overall image analysis capabilities. In addition, the annotation part uses a convolutional neural network (CNN), which takes advantage of its advantages in image processing to further improve the ability to recognize tissue features and ensure accurate annotation of abnormal tissues.

[0051] 3. Optimize the labeling process and results: The system is also committed to optimizing the labeling process and improving the accuracy of labeling. By introducing structural entropy regularization technology, the system separates the embedding of normal and abnormal tissues in the latent space, increasing the sensitivity of detection and the accuracy of labeling. At the same time, the uncertainty-based labeling guidance mechanism allows labelers to focus on high-uncertainty areas, thereby simplifying the labeling process and reducing human errors. This strategy not only improves the efficiency of labeling, but also provides clinicians with a more reliable basis to ensure faster identification of potential cancerous areas in early screening.

[0052] In summary, this early screening and annotation system aims to provide more effective technical support for early screening of esophageal cancer by improving early detection accuracy, ensuring the robustness of image analysis, and optimizing the annotation process. This not only helps to improve the survival rate of patients, but also promotes the innovation and development of medical imaging technology. By effectively integrating and analyzing image information under multiple lighting conditions, the system provides a more reliable auxiliary diagnostic tool for clinicians, ultimately promoting the advancement of early screening of esophageal cancer. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] The accompanying drawings exemplarily illustrate the embodiments and constitute a part of the specification, and together with the text description of the specification, are used to explain the exemplary implementation of the embodiments. The embodiments shown are for illustrative purposes only and do not limit the scope of the claims. In all drawings, the same reference numerals refer to similar but not necessarily identical elements.

[0054] Figure 1 A schematic diagram of the structure of the multi-illumination esophageal cancer early screening and labeling system based on probability embedding in this application;

[0055] Figure 2 This is a workflow diagram of the multi-illumination esophageal cancer early screening and labeling system based on probability embedding in this application. DETAILED DESCRIPTION

[0056] In order to make the purpose, technical solutions and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of the present application.

[0057] It should be noted that the descriptions involving "first", "second", etc. in the embodiments of the present application are only for descriptive purposes and cannot be understood as indicating or implying their relative importance or implicitly indicating the number of technical features indicated. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features. In addition, the technical solutions between the various embodiments can be combined with each other, but they must be based on the ability of ordinary technicians in the field to implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be deemed that such combination of technical solutions does not exist and is not within the scope of protection required by this application.

[0058] In the description of the present application, it should be understood that the numerical labels before the steps do not indicate the order in which the steps are executed, but are only used to facilitate the description of the present application and to distinguish each step, and therefore should not be understood as a limitation on the present application.

[0059] This application proposes a multi-illumination esophageal cancer early screening and labeling system based on probability embedding, which aims to effectively integrate image information taken under various lighting conditions to improve the early screening and labeling of esophageal cancer. With the increasing incidence of esophageal cancer, early detection is crucial to improving patient survival rates. However, traditional medical imaging technology often shows limitations when dealing with different lighting conditions, individual differences among patients, and the diversity of imaging equipment. By introducing a probabilistic embedding model, the present invention can generate a probabilistic representation for each image and handle inherent variability, thereby significantly improving detection performance and labeling accuracy. The design of this system not only enhances the accuracy and robustness of image analysis, but also provides clinicians with a more reliable auxiliary diagnostic tool, ultimately promoting early screening and treatment of patients.

[0060] First, the following terms are explained:

[0061] 1. Probabilistic Embedding

[0062] Probabilistic embedding is a technique for mapping data to a probability distribution space. It is commonly used in fields such as machine learning, natural language processing, and recommendation systems. Its main purpose is to capture the structure and relationship of data through representation in a low-dimensional space, thereby improving the efficiency and performance of the algorithm. Probabilistic embedding maps high-dimensional data to a low-dimensional space so that similar data points maintain similar distances in the embedding space. Usually, this embedding takes into account the distribution characteristics of the data and uses a probability model to define the relationship between data points. For example, a Gaussian distribution or other probability distribution is used to describe the characteristics of data points.

[0063] In the field of natural language processing, word embedding (such as Word2Vec or GloVe) maps words into vectors to capture the semantic relationship between words. Probabilistic models can help understand the frequency and co-occurrence patterns of vocabulary in context. In graph data analysis, nodes can be embedded into a probability distribution to capture the connection patterns and properties between nodes. Probabilistic embedding is a powerful tool that can effectively capture the potential structure and relationship of data by mapping data into a probability space, and is widely used in many fields.

[0064] 2. LLaVA

[0065] The LLaVA (Large Language and Vision Assistant) model shows important application potential in the diagnosis of esophageal cancer. By combining vision and language processing capabilities, the model can analyze esophageal images and identify cancer-related features and lesion areas. With its powerful feature extraction capabilities, LLaVA can extract subtle tissue structure information from endoscopic images or CT scans to help doctors detect early lesions. In the diagnosis of esophageal cancer, LLaVA's multimodal characteristics enable it to parse complex information in images. This process not only improves doctors' understanding of images, but also helps them make more accurate judgments when diagnosing. For example, LLaVA can generate descriptions for specific images, point out possible abnormal areas, and provide data support for clinical decision-making.

[0066] The LLaVA-1.5 model, which has undergone data augmentation and LoRA fine-tuning, can effectively adapt to the task of esophageal cancer screening. By screening a large number of processed images, it can quickly identify potential cancerous areas. The evaluation results generated by the model provide an important reference for clinicians, helping them to give priority to high-risk patients. This intelligent auxiliary screening system not only improves screening efficiency, but also improves the accuracy of diagnosis, which can ultimately improve the treatment effect and prognosis of patients and promote the development of early screening technology for esophageal cancer.

[0067] 3. Structural entropy

[0068] Structural entropy is a concept used to measure the complexity and information content of a system or structure, and is often applied in fields such as network science and information theory. Intuitively, structural entropy methods encode tree structures by characterizing the uncertainty of hierarchical topology. The structural entropy of a graph G is defined as the minimum total number of bits required to determine the codewords of nodes in G. Structural entropy has achieved success in fields such as information retrieval, traffic prediction, and reinforcement learning. By minimizing the structural entropy of a given graph G, the hierarchical clustering results of the vertices in G are retained by the associated coding trees.

[0069] The definition of structural entropy under graph G and coding tree T is as follows:

[0070] H T (G) = ∑ α∈T,a≠λ H T (G; α), where Here, g α Is connected to T a The sum of the weights of the edges between internal points and external points (i.e., T a and its complement The weight of the cut edge between them), the volume of the graph G vol(G) is the sum of the degrees of all data points X, that is: vol(G) = Σ x∈X d x ,and, It is T a The volume, α - is the parent node of α.

[0071] 4. Coding Tree

[0072] Given a graph G = {X, E, W}, where X is the set of input data points and E is the set of edges, is the set of edge weights. For each point x∈X, its degree d x is defined as the sum of the weights of the edges associated with it. The coding tree T is a multi-subtree with the following properties:

[0073] Correspondence between nodes and data points: Each tree node α corresponds to a subset of data points In particular, for the root node λ of the tree, we define its associated point set as Tλ = X. For a leaf node α, Tα is a set containing only a single data point x∈X.

[0074] Node hierarchical relationship: For each non-leaf node α, its i-th direct child node is α , whose parent node is represented by α - .

[0075] Subset relationship: For each non-leaf node α, there is:

[0076]

[0077] Where N α is the number of child nodes of node α.

[0078] With these properties, the depth of a node in the encoding tree depicts a partition of the set of data points X, with nodes with lower depths representing more coarse-grained partitions.

[0079] Secondly, in order to facilitate those skilled in the art to understand the technical solutions provided in the embodiments of the present application, the relevant technologies are described below:

[0080] In the embodiment of the present application, the multi-illumination esophageal cancer early screening and labeling system based on probability embedding is as follows: Figure 1 As shown in Figure 1, it includes: encoder module, multi-illumination fusion module, classification and screening module, structural entropy regularization module. The system workflow is as follows: Figure 2 As shown, perform the following steps:

[0081] S1. Image acquisition and preprocessing:

[0082] The system receives multiple esophageal images captured under different lighting conditions to ensure that normal tissues and various pathological states are covered. These images are collected by high-resolution medical imaging devices (such as endoscopes, ultrasound imagers, etc.). The image input is standardized and resized to 224x224 pixels, and data augmentation techniques such as random rotation, flipping, and brightness adjustment are applied to improve the generalization ability of the model.

[0083] S1.1 Data Collection

[0084] Data acquisition is the basis of the system, which aims to collect multiple esophageal images captured under different lighting conditions. The specific steps are as follows:

[0085] Imaging equipment selection: Use high-resolution medical imaging equipment, such as endoscopes, ultrasound imagers, and other related imaging technologies to ensure clear and detailed esophageal images. These devices have the ability to take pictures under a variety of lighting conditions to adapt to changes in clinical environments.

[0086] Image acquisition under multiple lighting conditions: The system has designed a standardized imaging process to ensure image acquisition under different lighting conditions. This includes natural light, artificial light, and settings of different light intensities to capture the performance of esophageal tissue under various lighting conditions. In this way, the system can obtain more comprehensive tissue features, covering images of normal tissues and various pathological states.

[0087] Data storage and quality control: All collected images will be stored in a secure database to ensure the privacy of the participants. At the same time, in order to ensure the high quality of the data, the system will perform image quality assessment to check the clarity, contrast and integrity of the image to ensure that each image meets the analysis requirements.

[0088] Through a rigorous acquisition process, the system can effectively collect diverse esophageal image data, laying a solid foundation for subsequent large-model image screening, probabilistic embedding generation, and labeling work.

[0089] S1.2 Data preprocessing

[0090] Data preprocessing is a key step to ensure the quality and applicability of image data. Through a series of processing operations, the original image will be converted into a format suitable for subsequent analysis and model training. The specific steps are as follows:

[0091] Image cleaning: Input the original multi-illumination esophageal images and remove duplicate images, blurred images, and low-quality images. The system will automatically evaluate the clarity and contrast of the image based on the set threshold.

[0092] The output is a set of cleaned, high-quality images.

[0093] Image standardization and structuring: Input a cleaned high-quality image set, resize and standardize the color of each image so that all images have the same resolution and color space (such as RGB). This step can reduce the impact of image size and color differences. Structure the processed data into a format that is easy to use for machine learning models to ensure effective organization of the data.

[0094] The output is a standardized and structured image set, where all images have the same size, standardized colors, and are converted into a format that is easy for the model to recognize.

[0095] Data augmentation technology: In order to improve the diversity of the data set and the generalization ability of the model, data augmentation technology is used to perform the following rotation transformation on the processed esophageal images:

[0096] Rotating the image 90 degrees clockwise helps the model learn tissue structures in different directions; rotating the image 180 degrees clockwise enhances the model’s robustness to upside-down situations; rotating the image 270 degrees clockwise further increases the diversity of training samples.

[0097] Through these rotation operations, the number of images in the original dataset can be expanded to four times. For example, if there are 1,000 images in the original dataset, after data augmentation, the dataset will increase to 4,000 images. This process improves the utilization of data and effectively avoids model overfitting.

[0098] Probabilistic embedding screening and labeling system

[0099] The multi-illumination esophageal cancer early screening and labeling system based on probability embedding of the present invention uses convolutional neural network (CNN) and structured entropy regularization (SEPC) technology to improve the detection accuracy and labeling efficiency of early esophageal cancer. The system captures images under different illumination conditions and emphasizes different tissue characteristics, thereby improving the detection rate of early esophageal cancer. Images under each illumination condition provide unique perspectives and information. However, traditional medical imaging methods often rely on deterministic representations, which leads to certain deficiencies in dealing with differences between patients, the diversity of illumination conditions, and changes in imaging equipment.

[0100] The core of this system is to use probabilistic embedding technology to transform each image into a Gaussian distribution to effectively capture the variability and uncertainty of image data. Subsequently, the system integrates the probabilistic embeddings from different lighting conditions into a unified representation through a multi-illumination fusion module, and at the same time enhances the separation of different categories in the latent space through structural entropy regularization technology, thereby improving the sensitivity of detection and the accuracy of annotation.

[0101] S2. Perform probabilistic embedding generation using the encoder module:

[0102] The main function of the encoder module is to convert the input medical image into a probabilistic embedding, so that each image is represented not only as a fixed point, but as a Gaussian distribution. This can effectively capture the variability and uncertainty in the image and provide rich feature information for subsequent processing. Each image is processed by the encoder, and a convolutional neural network based on ResNet-50 is used as a feature extractor. The probabilistic embedding is generated through two fully connected layers, where the first layer generates the mean μ of the probability distribution, and the second layer generates the covariance ∑ of the probability distribution, and ensures that ∑ is positive definite.

[0103] The structure of the S2.1 encoder consists of the following main components:

[0104] Feature Extractor: Use ResNet-50 as the base network. ResNet-50 is a deep residual network that can effectively capture complex features in images. Its residual connection design makes the network more stable when training deep models and reduces the problem of gradient disappearance.

[0105] Specifically, ResNet-50 builds a deep network by stacking multiple residual blocks. Each residual block contains two convolutional layers and a shortcut connection. Its structure is as follows:

[0106] Convolutional layer: Each convolutional layer uses a 3x3 convolution kernel with appropriate padding to maintain the size of the feature map. The convolutional layer is followed by a ReLU activation function to introduce nonlinear characteristics.

[0107] Quick connection: In each residual block, the input is directly added to the output through a quick connection, thus forming residual learning. This design helps improve the performance of the model in deep learning, allowing the network to learn residuals instead of directly learning mappings.

[0108] Fully connected layers: After ResNet-50 extracts features, the output feature maps are flattened and passed to two fully connected layers. These two fully connected layers are used to generate mean and covariance parameters respectively.

[0109] After ResNet-50, the feature maps are flattened and fed into two fully connected layers:

[0110] The first fully connected layer: generates the mean μ parameter and uses the ReLU activation function to ensure that the output is non-negative.

[0111] The second fully connected layer: generates the covariance ∑ parameter using the Softplus function, which is defined as:

[0112] Softplus(x)=log(1+e x ), ensuring that the output is positive, thus guaranteeing the positive definiteness of the covariance matrix.

[0113] Generation of probabilistic embeddings:

[0114] After the fully connected layer, the probability embedding of the encoder output is expressed as:

[0115] z~N(z;μ,∑), where μ is the mean vector generated by the first fully connected layer, and ∑ is the covariance matrix generated by the second fully connected layer, which is processed by the Softplus function to ensure its positive definiteness.

[0116] This generation method allows the encoder to capture image variability caused by different lighting conditions, providing rich information for the subsequent multi-illumination fusion module.

[0117] S3. Generate fusion embedding using multi-illumination fusion module:

[0118] The multi-illumination fusion module integrates these embeddings using a self-attention mechanism, specifically a multi-head attention mechanism in the Transformer encoder architecture. Each head processes a different subset of embeddings, generates a weighted embedding, and integrates it into the final fused embedding through a linear layer.

[0119] 1. Self-Attention Mechanism

[0120] The self-attention mechanism allows the model to dynamically assign weights to different parts when processing different input embeddings. This mechanism can capture information dependencies under different lighting conditions.

[0121] S3.1 Input Embedding: The module receives multiple probabilistic embeddings z from the encoder 1 , z 2 , ..., z n , each embedding corresponds to an image under a specific lighting condition. Here, multiple means at least two or more, that is, 2 or more.

[0122] S3.2 Generate query, key, value: For each input embedding, generate query Q, key K and value vector V, calculated as follows: Q = zW Q , K = zW K , V = zW V , where W Q , W K and W V is the learned weight matrix.

[0123] 2. Multi-head attention mechanism

[0124] The multi-head attention mechanism allows the model to extract information from different subspaces by using multiple attention heads in parallel.

[0125] S3.3. Calculate attention weight: For each attention head, calculate the attention weight A:

[0126] where d k is the dimension of the key vector.

[0127] S3.4. Weighted output: Use the calculated attention weights to perform weighted summation on the value vector and output the result of each attention head: Attention(Q, K, V) = AV.

[0128] S3.5. Integrate multi-head outputs: Concatenate the outputs of all attention heads and transform them into the final fused embedding z through a linear layer f :

[0129] z f =Concat(head 1 , head 2 , ..., head h )W O , where W O is the output weight matrix.

[0130] 3. Generation of the final fused embedding

[0131] Through the above process, the final embedding z generated by the fusion module f It contains important feature information from different lighting conditions. This fused embedding will be used as the input of subsequent processing modules to ensure that the model fully utilizes the information under multiple lighting conditions when performing tasks.

[0132] S4. Generate probabilistic classification using classification and screening modules:

[0133] The fused embedding is input into the LLaVA model for classification processing, and the Softmax activation function is used to output the category probability distribution. At the same time, structural entropy regularization is applied to maximize the entropy difference between categories by constructing a coding tree of the potential representation to ensure good separation of categories in the latent space.

[0134] The LLaVA large model plays a key role in medical image screening, especially in the analysis of esophageal images. The model uses its powerful multimodal understanding capabilities to effectively extract key features in images and identify potential lesion areas. After data augmentation and LoRA fine-tuning, the LLaVA-1.5 model is able to adapt to specific esophageal cancer data classification and screening tasks. The detailed steps are as follows:

[0135] S4.1 LoRA fine-tuning

[0136] Set training parameters: During fine-tuning, set the training parameters to 15 epochs and a batch size of 32. During training, the entire dataset is traversed 15 times, processing 32 images each time. Choosing an appropriate number of epochs can ensure that the model is continuously optimized during training.

[0137] Training process: In each epoch, the model will traverse the entire training dataset. The model performance is evaluated by calculating the loss function, and the parameters of the low-rank adaptation layer are updated according to the gradient descent algorithm, thereby gradually optimizing the performance of the model.

[0138] Monitoring and evaluation: Regularly evaluate model performance during training, including training loss and validation accuracy, and adjust training strategies in a timely manner to ensure that the model achieves good results on the validation set.

[0139] The loss function can be expressed as:

[0140] Among them, L(θ) is the loss function, N is the number of samples, and y i is the true label, is the probability predicted by the model.

[0141] After training is completed, save the fine-tuned model for subsequent application and evaluation.

[0142] S4.2 Model screening process

[0143] After completing LoRA fine-tuning, the fine-tuned LLaVA-1.5 model will be used to screen the esophageal images processed by the multi-illumination fusion module.

[0144] Input fusion embedding: Input the fusion probability embedding output by the multi-illumination fusion module into the fine-tuned LLaVA-1.5 model. Make sure the input embedding format meets the requirements of the model for effective feature extraction and classification.

[0145] Feature extraction: The LLaVA-1.5 model processes the input fusion embedding through its deep neural network structure. Since the model has learned rich visual features in the pre-training stage, it can effectively identify the key information in the embedding.

[0146] Lesion region identification: Analyze the fused embedded information, identify the difference between normal tissue and abnormal tissue, and locate potential lesion areas, such as tumors, inflammation or other abnormal phenomena. The model generates a class prediction probability p for each region: p(y=1|z f )=σ(Wz f +b, where σ is the Sigmoid function, W is the weight matrix, and b is the bias term.

[0147] Class probability distribution output: After feature extraction is completed, the model uses these features for classification tasks to generate a class probability distribution for each image. Based on the extracted features, the model determines whether each embedding has a lesion and classifies it as normal or abnormal.

[0148] S5. Use the structural entropy regularization module to perform structural entropy regularization:

[0149] In the multi-illumination esophageal cancer early screening and labeling system based on probability embedding, the goal of the structural entropy regularization module is to optimize the distribution of latent representations to enhance the ability to distinguish different categories (such as healthy tissue and cancerous tissue). This module captures the structural information between latent variables by constructing an encoding tree of latent representations and using a hierarchical clustering algorithm to group the embeddings to maximize the entropy difference between categories.

[0150] S5.1. Construction of coding tree

[0151] First, the input data X is encoded into a probabilistic embedding Z. Next, the graph G is constructed, ensuring that the adjacency matrix The value of is positive:

[0152] Among them, H Z is the representation of the probability embedding Z, and σ is the sigmoid activation function.

[0153] Then, a three-layer coding tree is constructed, and the nodes in the middle layer represent the categories of the classification task. Each leaf node (i.e., input data X) is assigned to the corresponding middle node according to its label. Define the assignment matrix C∈{0,1} n×r , where n is the number of leaf nodes and r is the number of intermediate nodes. If C ij =1, it means that the i-th leaf node belongs to the,th category.

[0154] S5.2. Calculation of structural entropy

[0155] In order to enhance the ability of latent representation, the module proposes to maximize the structural entropy of the intermediate layer nodes and constrain the probability distribution of the latent variables to ensure the separation between categories. The structural entropy of the intermediate layer nodes is defined as:

[0156] where r is the number of categories, Is an intermediate node Its complement The sum of the weights of the edges between yes volume.

[0157] Using the adjacency matrix and the allocation matrix C, the expression of the structural entropy regularization loss is as follows:

[0158]

[0159] Structural entropy regularization loss L SE It is defined as constraining the distribution of latent variables by maximizing the entropy difference between categories. By optimizing this regularization term, the model is able to learn the differences between categories, thereby improving classification capabilities, especially in medical images under multiple illumination conditions.

[0160] S6. Use end-to-end training and Adam optimizer for model training:

[0161] The system adopts an end-to-end training method to optimize the performance of the multi-illumination esophageal cancer early screening and labeling system based on probability embedding. The optimization target is the weighted sum of the classification cross entropy loss and the structural entropy regularization term to improve the model's detection accuracy and labeling efficiency for esophageal cancer. The Adam optimizer is used with a learning rate of (1x10 -4 ), the batch size is 32, and the training is done for 100 epochs.

[0162] 1. Training objectives

[0163] The training loss function consists of two parts:

[0164] Categorical cross entropy loss L PC : It is used to measure the classification accuracy of the model for each image. Its calculation formula is: L PC =L(θ),

[0165] Among them, L(θ) is the loss function, N is the number of samples, and y i is the true label, is the probability predicted by the model.

[0166] Structural entropy regularization loss L SE : Used to enhance the distribution of potential representations and maximize the entropy difference between categories. The loss term is calculated as:

[0167] Where r is the number of categories, C is the assignment matrix, is the adjacency matrix.

[0168] 2. Comprehensive loss function

[0169] To sum up, the overall loss function is: L SEPC =L PC -γL SE , where γ is a hyperparameter that controls the weight of the structural entropy regularization loss. By optimizing this comprehensive loss, the model is able to simultaneously consider the classification accuracy and the entropy difference between categories, thereby improving the overall performance, especially when dealing with cancer detection under different lighting conditions.

[0170] 3. Training process

[0171] The model training uses the Adam optimizer, which is a widely used optimization algorithm that can effectively accelerate convergence due to its adaptive learning rate. The following are the specific settings of the training process:

[0172] Learning rate: Set the initial learning rate to , which has been experimentally verified to maintain good convergence during training.

[0173] Batch size: Select a batch size of 32 to process an appropriate amount of data in each iteration. This batch size can effectively balance memory usage and training efficiency, ensuring that the model can update parameters stably during training, while also ensuring the efficiency of the training process.

[0174] Number of training rounds: The training process is set to 100 epochs. In each epoch, the model will traverse the entire training data set to ensure that each sample has the opportunity to participate in the learning of the model, thereby improving the learning effect.

[0175] 4. Monitoring and evaluation

[0176] During the training process, the model performance will be evaluated regularly, including training loss and validation accuracy. By monitoring these indicators, the training strategy can be adjusted in time to ensure that the model achieves good results on the validation set. The loss function changes and accuracy indicators during the training process will be recorded for subsequent analysis and model tuning.

[0177] S7. Annotation guidance based on uncertainty annotation guidance module:

[0178] The main goal of the uncertainty-based annotation guidance module is to identify and mark high-uncertainty regions that may contain critical lesion information in esophageal cancer detection. By evaluating the covariance matrix of each embedding, high-uncertainty regions are marked. These regions are prioritized for annotation, thereby improving the efficiency and accuracy of annotation and ensuring that the model focuses on potential lesion information.

[0179] 1. Uncertainty Assessment

[0180] Uncertainty estimation is performed by computing the diagonal elements of the covariance matrix of each embedding. For each input probability embedding Z, its covariance matrix Γ is computed:

[0181] Γ=Cov(Z)=E[(Z-μ)(Z-μ) T ], where μ is the embedded mean vector and E represents the expected value.

[0182] Extract the diagonal elements from the covariance matrix, which represent the variance of the individual features:

[0183] Var(Z i )=Γ ii , where Γ ii is the variance of the ith feature of the covariance matrix.

[0184] Evaluate uncertainty: Evaluate the variance of each embedding by analyzing the size of the diagonal elements. Regions with higher variance are marked as regions of high uncertainty. This process can be achieved by setting a threshold τ, τ = k Var max , where k is a coefficient less than 1, and this paper chooses k to be 0.8:

[0185] 2. Marking of uncertainty areas

[0186] Based on the results of uncertainty assessment, the model marks high variance areas as high uncertainty areas. These areas usually represent potential difficulties for the model in esophageal cancer detection and may contain more complex lesion features or ambiguous information caused by factors such as illumination changes and imaging noise.

[0187] Regions of high uncertainty may indicate the presence of early cancer or atypical lesions that may not be obvious in imaging, resulting in a reduced ability of the model to identify their features. For example, tiny tumors or irregular tissue structures may present high variance in images, making the model more uncertain when classifying these regions.

[0188] Under multiple lighting conditions, different regions of an image may show different characteristics due to uneven lighting. High uncertainty regions may include those where features are difficult to clearly identify due to light reflections, shadows, or blur. In this case, the model's classification confidence in that region will be reduced, which is reflected as a higher variance.

[0189] Different labeling strategies: When labeling areas of high uncertainty, the model will use different labeling strategies to reflect the certainty and different probabilities of cancer:

[0190] Identify cancerous regions: For regions that have been clearly marked as cancerous, the model should give higher weights and use strong labels (such as "cancer"), even if the characteristics of these regions show uncertainty in some cases. This is because the existence of these regions itself indicates the presence of lesions, although the model's understanding of their characteristics may be limited.

[0191] Uncertainty area: For areas marked as high uncertainty but not yet confirmed as cancerous, the model will perform soft labeling. For example, a probabilistic labeling method is used to indicate the correlation between the area and cancer, such as a probability of 70%. This labeling method allows the model to remain flexible in clinical applications while providing a basis for further verification.

[0192] Prioritization strategy: By guiding the model to perform more training and annotation in high-uncertainty areas, the system's ability to identify esophageal cancer can be improved. The model can perform detailed feature learning in these areas and gradually improve its ability to identify complex lesions. This will help ensure accuracy in clinical applications and reduce missed diagnosis rates.

[0193] Finally, in order to verify the performance of the model, a 5-fold cross-validation method was used to evaluate the performance of the model. The data set was divided into five subsets so that each subset could be tested in different training and validation combinations to ensure the reliability of the evaluation results. The evaluation indicators include accuracy (indicating the proportion of samples correctly classified by the model), recall (indicating the proportion of correctly identified samples in actual positive samples), F1 score (the harmonic mean of accuracy and recall), and the area under the ROC curve (AUC, which evaluates the overall performance of the model at different thresholds. The closer the AUC value is to 1, the better the model performance). Together, these indicators help users fully understand the performance of the model in the task of esophageal cancer detection.

[0194] In order to comprehensively evaluate the performance of the model, the following indicators are selected:

[0195] Accuracy: It indicates the proportion of correctly classified samples to the total samples. The formula is:

[0196]

[0197] Among them, TP is true positive, FN is true negative, FP is false positive, and FN is false negative.

[0198] Recall rate: also known as sensitivity, it indicates the proportion of actual positive examples that are correctly identified. The formula is:

[0199]

[0200] F1 score: It takes into account the harmonic mean of precision and recall, and is suitable for evaluating unbalanced data. The formula is:

[0201]

[0202] Among them, Precision is the accuracy rate, which indicates the proportion of samples identified as positive examples that are actually positive examples.

[0203] Area under the ROC curve (AUC): The ROC curve depicts the relationship between the true positive rate and the false positive rate. The AUC value represents the overall performance of the model at different thresholds. The closer the AUC value is to 1, the better the model performance.

[0204] This paper proposes a multi-illumination esophageal cancer early screening and annotation system based on probabilistic embedding, aiming to improve the detection accuracy of esophageal cancer under different illumination conditions. By using convolutional neural network (CNN) to extract medical image features and converting these features into probabilistic embedding, the system can effectively capture the variability and uncertainty of image data.

[0205] The multi-illumination fusion module in the system uses the self-attention mechanism to dynamically integrate embeddings from different illumination conditions to ensure effective weighting of key information. This is particularly important in medical image processing, as illumination changes often have a significant impact on image quality and feature extraction. This module allows the model to automatically identify and enhance features under different illumination when processing diverse inputs, thereby improving the accuracy and robustness of classification results.

[0206] The introduction of structural entropy regularization is a major innovation of this system. By constructing a coding tree of latent representations and maximizing the entropy difference between categories, the model's ability to distinguish between different categories (such as healthy tissue and cancerous tissue) is significantly enhanced. This regularization method helps ensure that the categories in the latent space are well separated, reducing the confusion of the model in classification tasks, thereby improving overall performance.

[0207] In summary, the present invention not only technically realizes effective support for early screening of esophageal cancer, but also provides clinicians with a powerful tool to quickly identify potential lesion areas in complex medical images. By improving the accuracy and efficiency of detection, this system provides important technical support for the early diagnosis and treatment of esophageal cancer, and has broad application prospects and significant clinical value.

[0208] Obviously, those skilled in the art should understand that the modules or steps of the above-mentioned embodiments of the present application can be implemented by general-purpose computer devices, they can be concentrated on a single computer device, or distributed on a network composed of multiple computer devices, optionally, they can be implemented by executable program codes of computer devices, so that they can be stored in a storage device and executed by the computer device, and in some cases, the steps shown or described can be executed in a different order from that herein, or they can be made into individual integrated circuit modules, or multiple modules or steps therein can be made into a single integrated circuit module for implementation. In this way, the embodiments of the present application are not limited to any specific combination of hardware and software.

[0209] It should be noted that the above are only preferred embodiments of the present application, and the patent protection scope of the present application is not limited thereto. Any equivalent structure or equivalent process transformation made using the contents of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.

Claims

1. A multi-illumination esophageal cancer early screening and labeling system based on probability embedding, comprising an encoder module, a multi-illumination fusion module, a classification and screening module, a structural entropy regularization module, and an uncertainty-based labeling guidance module, characterized in that: The system implements multi-light esophageal cancer early screening and labeling, and performs the following operations: S1, image acquisition and preprocessing; S2. Perform probabilistic embedding generation using the encoder module: A convolutional neural network based on ResNet-50 is used as the feature extractor for esophageal images. Probabilistic embedding is generated through two fully connected layers, where the first layer generates the mean of the probability distribution and the second layer generates the covariance of the probability distribution, and the outputs of the two fully connected layers are ensured to be positive. S3. Generate fusion embedding using multi-illumination fusion module: For at least two probabilistic embeddings from the encoder, the multi-head attention mechanism in the Transformer encoder architecture is used, with each head processing a different subset of embeddings to generate weighted embeddings, which are then integrated into the final fused embedding through a linear layer; S4. Generate probabilistic classification using classification and screening modules: The fusion embedding is input into the LLaVA large model after LoRA fine-tuning for classification processing, and the Softmax activation function is used to output the category probability distribution; In S4, the classification and screening module is used to generate probabilistic classification, including: S4.1, perform LoRA fine-tuning on the LLaVA model; S4.

2. After completing the LoRA fine-tuning, the fine-tuned LLaVA-1.5 model is used to screen the esophageal images processed by the multi-illumination fusion module. The specific screening process is as follows: S4.2.

1. Feature extraction: The LLaVA-1.5 model is used to process the input fused embedding through its deep neural network structure, analyze the fused embedding information, identify the difference between normal tissue and abnormal tissue, and locate potential lesion areas; S4.2.2, Class probability distribution output: After feature extraction is completed, based on the extracted features, it is determined whether each embedding has a lesion and is classified as normal or abnormal, generating a class prediction probability p for each region: p(y=1|z f )=σ(Wz f +b), where σ is the Sigmoid function, W is the weight matrix, and b is the bias term; S5. Use the structural entropy regularization module to perform structural entropy regularization: By constructing an encoding tree of latent representations and using a hierarchical clustering algorithm to group embeddings, the distribution of latent representations is optimized through structural entropy regularization to enhance the ability to distinguish different categories and maximize the entropy difference between categories, thereby capturing the structural information between latent variables; In S5, the structural entropy regularization module is used to perform structural entropy regularization, specifically including: S5.

1. Construct a three-layer encoding tree, where the middle-layer nodes represent the categories of the classification task; S5.

2. Calculate the structural entropy of the coding tree: The structural entropy of the middle layer node is defined as: where r is the number of categories, Is an intermediate node Its complement The sum of the weights of the edges between yes Volume; Using the adjacency matrix And the allocation matrix C is used to calculate the structural entropy regularization loss, which is calculated as follows: Structural entropy regularization loss L SE It is defined as constraining the distribution of latent variables by maximizing the entropy difference between categories; S6. Model training: Adopt end-to-end training mode and use Adam optimizer to optimize the comprehensive loss of the model through the overall loss function for model training. The overall loss function includes classification cross entropy loss and structural entropy regularization loss. S7. Annotation guidance based on uncertainty annotation guidance module: By evaluating the covariance matrix of each embedding, regions of high uncertainty are identified and marked.

2. The system according to claim 1, characterized in that In S2, ResNet-50 is a deep network built by stacking multiple residual blocks, each of which consists of two convolutional layers and a shortcut connection; The ReLU activation function is used in the first fully connected layer to ensure that the output is positive; The Softplus function is used in the second fully connected layer to ensure that the output is positive, thereby ensuring the positive definiteness of the covariance matrix.

3. The system according to claim 1, characterized in that In S3, the generation of fused embedding using the multi-illumination fusion module specifically includes: S3.

1. Input embedding: Receive at least two probabilistic embeddings from the encoder, each embedding corresponding to an image under a specific lighting condition; S3.

2. Generate query, key and value: For each input embedding, generate query Q, key K and value vector V, calculated as follows: Q = zW Q , K = zW K , V = zW V , where W Q , W K and W V is the learned weight matrix; S3.

3. Calculate attention weight: For each attention head, calculate the attention weight A: where d k is the dimension of the key vector; S3.4, weighted output: Use the calculated attention weight A to weight the value vector V and output the result of each attention head: Attention(Q, K, V) = AV; S3.

5. Integrate multi-head outputs: Concatenate the outputs of all attention heads and transform them into the final fused embedding z through a linear layer f :z f =Concat(head1, head2,..., head h )W O , where W O is the output weight matrix.

4. The system according to claim 1, characterized in that In S6, classification cross entropy loss: used to measure the classification accuracy of the model for each image, denoted by L PC =L(θ), the expression is: Among them, L(θ) is the loss function, N is the number of samples, and y i is the true label, is the probability predicted by the model; Structural entropy regularization loss: used to enhance the distribution of potential representation and maximize the entropy difference between categories. Its expression is: Where r is the number of categories, C is the assignment matrix, is the adjacency matrix; Thus, the overall loss function is: L SEPC =L PC -γL SE , where γ is a hyperparameter used to control the weight of the structural entropy regularization loss.

5. The system according to claim 1, characterized in that In S7, the specific steps of annotation guidance include: S7.

1. Uncertainty assessment: The variance of each embedding is assessed by analyzing the size of the diagonal elements. S7.

2. Marking of uncertainty areas: Based on the variance of the uncertainty assessment, areas with higher variance are marked as high uncertainty areas. When marking high uncertainty areas, the marking strategy adopted is: for areas that have been clearly marked as cancerous, use strong marking; for areas that are marked as high uncertainty but have not yet been confirmed as cancerous, use soft marking.

6. The system according to claim 1, characterized in that The esophageal images are collected by high-resolution medical imaging equipment under different lighting conditions, and are input into the system after image cleaning, image standardization and structuring, and data augmentation technology processing.

Citation Information

Patent Citations

  • Probability embedding combination retrieval method based on CLIP

    CN116578734A

  • Multi-source domain adaptive EEG emotional state classification method based on knowledge distillation

    CN116821764A