Endoscopic esophageal-gastric junction adenocarcinoma staging diagnosis method and system
By combining convolutional neural networks and local and global feature extractors of visual big models, the problem of precise staging diagnosis of early EGJA lesions is solved, and a highly accurate endoscopic endoscopic staging diagnosis of adenocarcinoma in the adenocarcinoma in the esophageal gastric junction is achieved.
Patent Information
- Application Number
- CN202510266444.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-07-29
AI Technical Summary
Existing deep learning techniques are difficult to use for the precise staging diagnosis of endoscopic esophageal gastric junction adenocarcinoma (EGJA), especially for early lesions identification, and existing CNNs and visual mockups do not perform well in medical imaging tasks.
A local feature extractor based on convolutional neural networks and a global feature extractor based on visual big model are combined to combine local features and global features through a hybrid expert model, and a classifier is used to predict and diagnose.
It significantly improves the accuracy of endoscopic EGJA staging diagnosis, captures subtle changes in the lesion area, reduces the inference cost and improves the staging diagnostic performance of the model.
Smart Images

Figure CN120388710A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical fields of computer vision and medical image processing, and particularly relates to a method and system for endoscopic staging diagnosis of adenocarcinoma of the esophagogastric junction. Background Art
[0002] Endoscopic examination has become an important and effective diagnostic tool for digestive tract diseases, especially some gastrointestinal lesions and tumors. However, compared with other gastrointestinal diseases, the lesion area of esophageal-gastric junction adenocarcinoma (EGJA) under endoscopy is more atypical, and its lesion characteristics are more difficult to identify. Therefore, the conventional manual endoscopic diagnosis method will have a relatively high missed diagnosis rate for EGJA. However, the early and accurate diagnosis of EGJA plays a crucial role in improving the prognosis of patients with this disease. This makes it particularly necessary to develop a computer-aided technology that can achieve automatic staging diagnosis of EGJA, which not only helps to significantly improve the survival rate of EGJA patients, but also provides strong support for clinical decision-making.
[0003] In recent years, deep learning technology has developed rapidly. In medical image analysis, technologies and methods represented by Convolutional Neural Networks (CNN) have become the mainstream. With the sparse connection and weight sharing characteristics of CNN, it can automatically extract local features of images, thus efficiently and accurately completing visual tasks related to disease diagnosis. Currently, some classic CNN architectures such as ResNet, DenseNet, etc. have shown excellent results in tasks such as disease classification and lesion segmentation. In the endoscopic field, the wide application of CNN has also significantly improved the diagnostic accuracy of many common gastrointestinal diseases. However, the deep learning technology specifically for EGJA staging diagnosis is still in a blank state. On the other hand, directly applying existing CNN architectures to EGJA automatic staging diagnosis usually causes many problems. Especially in the diagnosis of early EGJA, its accuracy is often difficult to reach an ideal level.
[0004] Meanwhile, vision large models based on Transformer have gradually received increasing attention in the field of deep learning. These models are mostly trained on large-scale natural image datasets. At the same time, with the help of the self-attention mechanism of Transformer, it can effectively extract the global information of the generalized input image, thus showing strong generality and adaptability in many downstream tasks. However, when directly migrating such models to medical image tasks, significant performance degradation usually occurs. This is mainly because there are essential differences between medical imaging modalities and natural images, and coupled with the specific privacy protection requirements of medical data, it becomes extremely difficult to obtain a sufficiently large-scale high-quality medical dataset for model fine-tuning. For EGJA, its unique pathological features such as the esophagogastric junction being difficult to fully expose, early lesions lacking obvious symptoms, and subtle mucosal changes further increase the complexity of constructing an effective dataset. In such a situation, even using the most advanced vision large models, their performance may be inferior to that of optimized CNNs.
[0005] In summary, although some existing technologies have proposed endoscopic disease diagnosis methods based on CNN or vision large models, they are generally not suitable for the targeted diagnosis of EGJA. In particular, these methods usually fail to fully consider the inherent limitations of CNN or vision large models and propose effective improvement measures, which further limits their staging diagnosis effect for EGJA. Summary of the Invention
[0006] This application provides a method for staging diagnosis of adenocarcinoma of the esophagogastric junction under endoscopy to solve the problem in the prior art that existing deep learning technologies are difficult to be directly used for the accurate staging diagnosis of EGJA.
[0007] Correspondingly, this application also provides a staging diagnosis system for adenocarcinoma of the esophagogastric junction under endoscopy, an electronic device, and a computer-readable storage medium to ensure the implementation and application of the above method.
[0008] To solve the above technical problems, this application discloses a method for staging diagnosis of adenocarcinoma of the esophagogastric junction under endoscopy, and the method includes:
[0009] Extracting local features in the endoscopic image using a local feature extractor based on a convolutional neural network;
[0010] Extracting global features in the endoscopic image using a global feature extractor based on a vision large model;
[0011] Fusing the local features and the global features using a mixture of experts model to obtain deeply fused features;
[0012] Classifying the deeply fused features using a classifier to obtain a predicted diagnosis result.
[0013] Preferably, the local feature extractor adopts the ResNet50 network;
[0014] Extract the local features in the endoscopic image by using a local feature extractor based on a convolutional neural network, including:
[0015] Perform convolution and max pooling operations on the endoscopic image to obtain low-dimensional local features;
[0016] Extract features from the low-dimensional local features through convolutional modules in four stages to obtain local features;
[0017] Among them, the convolutional modules in the four stages all contain residual blocks, and the output of each residual block is added to the input through an identity mapping to form a residual connection.
[0018] Preferably, the global feature extractor adopts the DINOv2 model;
[0019] Extract the global features in the endoscopic image by using a global feature extractor based on a large vision model, including:
[0020] Divide the endoscopic image into multiple non-overlapping image patches;
[0021] Flatten and merge multiple non-overlapping image patches, and linearly project them into a high-dimensional feature space to generate feature vectors;
[0022] Merge the category labels for each feature vector and add position information to obtain an embedded vector;
[0023] Extract features from the embedded vector through stacked ViT encoders to obtain global features.
[0024] Preferably, the ViT encoder includes a multi-head attention mechanism and a feed-forward neural network.
[0025] Preferably, use a mixture of experts model to fuse the local features and global features to obtain deep fusion features, including:
[0026] Assign corresponding local weights and global weights to the local features and global features through a gating network;
[0027] Based on the local weights and global weights, fuse the local features and global features at the element level to obtain deep fusion features.
[0028] Preferably, the gating network includes three convolutional layers;
[0029] Assign corresponding local weights and global weights to the local features and global features through a gating network, including:
[0030] After aligning the local features and global features, merge them in the first dimension to obtain a merged feature;
[0031] The combined features are processed sequentially by three convolutional layers to obtain intermediate features;
[0032] The intermediate features are processed by the softmax function to obtain final features;
[0033] The final features are separated into local weights and global weights.
[0034] Preferably, the classifier is constructed based on a multi-layer perceptron structure and includes three fully connected layers;
[0035] The classifier is used to classify the depth fusion features to obtain a predicted diagnosis result, including:
[0036] The depth fusion features are sequentially input into the first two fully connected layers, and the first classification features are obtained as the output;
[0037] After the first classification features are normalized, activated, and then randomly inactivated, the second classification features are obtained;
[0038] The second classification features are input into the third fully connected layer, and the third classification features are obtained as the output;
[0039] The third classification features are input into the softmax layer to obtain the probabilities of the corresponding disease categories, and the predicted diagnosis result is output according to the probabilities.
[0040] This application also discloses an endoscopic esophagogastric junction adenocarcinoma staging diagnosis system, which includes:
[0041] A local feature extraction module, which is used to extract local features in the endoscopic image by using a local feature extractor based on a convolutional neural network;
[0042] A global feature extraction module, which is used to extract global features in the endoscopic image by using a global feature extractor based on a vision large model;
[0043] A feature fusion module, which is used to fuse the local features and the global features by using a mixture of experts model to obtain depth fusion features;
[0044] A classification prediction module, which is used to classify the depth fusion features by using a classifier to obtain a predicted diagnosis result.
[0045] This application also discloses an electronic device, which includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements one or more of the methods in this application.
[0046] The present application also discloses a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the methods described in one or more of the present application are implemented.
[0047] The present application has at least the following beneficial effects:
[0048] 1. The present application uses a local feature extractor based on a convolutional neural network and a global feature extractor based on a vision large model to extract local features and global features in endoscopic images respectively, and fuses the local features and global features to obtain deeply fused features. Finally, a classifier is used to classify the fused features to obtain a predicted diagnosis result. The present application realizes the first application of the vision large model in the field of EGJA diagnosis, and makes full use of the complementary advantages of the convolutional neural network and the vision large model, significantly improving the accuracy of endoscopic EGJA staging diagnosis.
[0049] 2. The local feature extractor of the present application can effectively extract local features of the input endoscopic image by stacking multiple convolutional modules with residual connections, enabling the model to capture subtle changes in the EGJA lesion area. These local features are crucial for distinguishing between advanced and early lesion areas of EGJA, thus effectively improving the staging diagnosis performance of the model.
[0050] 3. The global feature extractor of the present application based on the DINOv2 model uses a vision large model after knowledge distillation, which can efficiently compress the general knowledge extracted by the vision large model into a smaller model, with only a minimal loss of accuracy, and can significantly reduce the inference cost. Therefore, the designed global feature extractor not only provides rich global information for EGJA staging diagnosis, but also does not significantly increase the computational burden of the overall model. At the same time, the combination with the local feature extractor also significantly reduces the dependence of the vision large model encoder on the amount of training data.
[0051] Additional aspects and advantages of the present application will be given in the following description section, which will become apparent from the following description, or can be understood through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] The above and / or additional aspects and advantages of the present application will become apparent and be readily understood from the following description of the embodiments in conjunction with the drawings, where:
[0053] Figure 1 is a flowchart of the method for staging diagnosis of esophageal-gastric junction adenocarcinoma under endoscopy provided by an embodiment of the present application;
[0054] Figure 2 is a schematic diagram of the overall structure of the EGJA staging diagnosis model provided by an embodiment of the present application;
[0055] Figure 3 Schematic diagram of the local feature extractor provided by the embodiment of the present application;
[0056] Figure 4 Schematic diagram of the global feature extractor provided by the embodiment of the present application;
[0057] Figure 5 Schematic diagram of the endoscopic esophagogastric junction adenocarcinoma staging diagnosis system provided by the embodiment of the present application;
[0058] Figure 6 Schematic diagram of the electronic device provided by the embodiment of the present application. Detailed implementation manners
[0059] The embodiments of the present application will be described in detail below. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are only used to explain the present application, and cannot be construed as a limitation to the present application.
[0060] Those skilled in the art of the present technology can understand that unless specifically stated, the singular forms "a", "an", "the" and "said" used herein may also include the plural forms. It should be further understood that the term "including" used in the specification of the present application means the presence of features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or their combinations. It should be understood that when we say that an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The phrase "and / or" used herein includes all or any unit and all combinations of one or more related listed items.
[0061] Those skilled in the art can understand that unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meaning as the general understanding of those of ordinary skill in the art to which the present invention belongs. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted in an idealized or overly formal sense unless specifically defined as here.
[0062] The solution provided by the embodiments of this application can be executed by any electronic device. For example, it can be a terminal device or a server. Among them, the server can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. The terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto. The terminal and the server can be directly or indirectly connected through wired or wireless communication methods, and this application does not make any restrictions here. Regarding the technical problems existing in the prior art, the endoscopic esophageal-gastric junction adenocarcinoma staging diagnosis method and system provided by this application aim to solve at least one of the technical problems in the prior art.
[0063] The following uses specific embodiments to elaborate in detail on the technical solution of this application and how the technical solution of this application solves the above technical problems. These several specific embodiments below can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The following will describe the embodiments of this application with reference to the accompanying drawings.
[0064] The embodiments of this application provide a possible implementation. The embodiments of this application propose an EGJA staging diagnosis model, including a local feature extractor, a global feature extractor, a mixture of experts model, and a classifier. Based on this model, an endoscopic esophageal-gastric junction adenocarcinoma staging diagnosis method is provided. As Figure 1 shown in the figure, this method may include the following steps:
[0065] Step 101, use a local feature extractor based on a convolutional neural network to extract local features in the endoscopic image.
[0066] The local feature extractor adopts multiple stacked convolutional modules, which can effectively extract local features in the input endoscopic image, so as to accurately capture the subtle features of EGJA lesions in a smaller area.
[0067] Step 102, use a global feature extractor based on a large vision model to extract global features in the endoscopic image.
[0068] The global feature extractor in the embodiments of this application is based on the VisionTransformer (ViT) structure pre-trained on a large scale of unlabeled images, and has a powerful global feature capture ability.
[0069] Step 103, use a mixture of experts model to fuse the local features and the global features to obtain deeply fused features.
[0070] Compared with using only a single feature extractor, fusing the features extracted by two feature extractors using a mixture of experts model can obtain more comprehensive and integrated deep fusion features. Using this deep fusion feature, the model in the embodiments of the present application can more easily identify and distinguish the image features of different staging stages of EGJA, which is of great significance for improving the accuracy of the EGJA staging diagnosis of the model.
[0071] Step 104, classify the deep fusion feature using a classifier to obtain a predicted diagnosis result.
[0072] In the embodiments of the present application, a local feature extractor based on a convolutional neural network and a global feature extractor based on a vision large model are used to extract local features and global features in the endoscopic image respectively, and the local features and global features are fused to obtain a deep fusion feature. Finally, a classifier is used to classify the fused feature to obtain a predicted diagnosis result. The embodiments of the present application realize the first application of the vision large model in the field of EGJA diagnosis, and make full use of the complementary advantages of the convolutional neural network and the vision large model, significantly improving the accuracy of the endoscopic EGJA staging diagnosis.
[0073] In an optional embodiment, the local feature extractor adopts a ResNet50 network;
[0074] Using a local feature extractor based on a convolutional neural network to extract local features in the endoscopic image includes:
[0075] Perform convolution and max pooling operations on the endoscopic image to obtain low-dimensional local features;
[0076] Extract features from the low-dimensional local features through convolutional modules in four stages to obtain local features;
[0077] Among them, the convolutional modules in the four stages all contain residual blocks, and the output of each residual block is added to the input through an identity mapping to form a residual connection.
[0078] The local feature extractor in the embodiments of the present application is based on the ResNet50 network. The ResNet50 network is a classic deep residual network. By introducing convolutional modules with residual connections, the problems of gradient disappearance and gradient explosion in deep neural networks can be effectively solved, enabling the network to perform deep learning. ResNet50 mainly has a 50-layer neural network structure. In the local feature extractor used in the embodiments of the present application, the original fully connected layer located at the end of the network in ResNet50 is removed, and the rest of the structure remains unchanged. As Figure 2 and Figure 3As shown in the figure, starting from the input layer, the endoscopic image with a size of 488×488×3 passes through a 7×7 convolutional layer (including 7×7 convolution, normalization, and activation operations) and a 3×3 max-pooling operation, extracting a low-dimensional local feature with a size of 112×112×64, while significantly reducing the dimension of the input image, thereby reducing the computational cost of subsequent modules.
[0079] Subsequently, the image features are further extracted through four stages of convolutional modules; these four stages respectively consist of 3, 4, 6, and 3 residual blocks (Base Block). These residual blocks are the core components of ResNet50, and each residual block contains three convolutional layers of 1×1, 3×3, and 1×1. Among them, the 1×1 convolutional layer is used to reduce the number of channels and computational complexity; the 3×3 convolutional layer is used to extract local image features and capture texture and edge information; the last 1×1 convolutional layer is used to integrate the number of channels to ensure alignment with the input feature channels. The output of each residual block is added to the input through an identity mapping to form a residual connection, so that during backpropagation, the network can more effectively propagate gradient information, thus supporting a deeper network structure. ResNet50 effectively extracts the local feature information of the input image through its unique residual block design. In addition, the design of the residual connection enables the network to better retain low-level features and avoid information loss in the deep network. After the low-dimensional local feature with a size of 112×112×64 undergoes feature extraction in stage 1, stage 2, stage 3, and stage 4, features with sizes of 112×112×256, 56×56×512, 28×28×1024, and 14×14×2048 are obtained in sequence. Finally, the ResNet encoded features output by the last residual block are used as the output of the entire local feature extractor, denoted as H1, W1, and C1 respectively represent the C length, width, and number of channels of F.
[0080] In the embodiments of the present application, the local feature extractor can effectively extract the local features of the input endoscopic image by stacking multiple convolutional modules with residual connections, enabling the model to capture the subtle changes in the EGJA lesion area. These local features are crucial for distinguishing the advanced and early lesion areas of EGJA, thereby effectively improving the staging diagnosis performance of the model.
[0081] In an alternative embodiment, the global feature extractor adopts the DINOv2 model;
[0082] Extracting the global features in the endoscopic image using the global feature extractor based on the large vision model, including:
[0083] Dividing the endoscopic image into multiple non-overlapping image patches;
[0084] Flatten and merge multiple non-overlapping image patches, and linearly project them into a high-dimensional feature space to generate feature vectors;
[0085] Merge the class labels with the feature vectors and add position information to obtain embedding vectors;
[0086] Extract features from the embedding vectors through a stacked ViT encoder to obtain global features.
[0087] In an optional embodiment, the ViT encoder includes a multi-head attention mechanism and a feed-forward neural network.
[0088] In the embodiments of the present application, the DINOv2 model is used to extract preliminary global information. In order to further enhance the extracted global information, the remaining features output by the DINOv2 model except for the class token are reused, and the dimensions of the remaining features are aligned with the class token through operations such as average pooling and convolution, and the information of the two is merged to obtain the final global information.
[0089] DINOv2 is a large vision model based on ViT. The pre-training process of this model uses a discriminative self-supervised method and a teacher-student network architecture based on knowledge distillation. By comparing the corresponding features extracted from the student and teacher networks through global and local contrast loss functions, the ability to capture robust and generalizable knowledge is finally achieved. The global feature extractor adopted in the embodiments of the present application is based on the ViT-S / 14 model. As Figure 2 and Figure 4 shown, in the global feature extractor, an endoscopic image with an input size of 448×448×3 is divided into 14×14 non-overlapping image patches after Patch Embedding, and the size of the image patches is (32×32)×384. These image patches are then flattened and merged, and linearly projected into a high-dimensional feature space. After the generated feature vectors are merged with the class labels (with a size of 1×384), features with a size of 1025×384 are obtained, and these features are input into the stacked ViT encoder after Position Embedding. For ViT-S / 14, it has 12 ViT encoders. Each encoder has the same structure, and mainly contains a multi-head self-attention (MHSA) module and a feed-forward network (FFN) inside. Among them, the multi-head self-attention mechanism can capture the feature dependencies between different regions of the image, and the feed-forward network further enhances the expression ability of the extracted features. Layer normalization is performed before both the multi-head self-attention mechanism and the feed-forward network. Finally, the global feature extractor will generate two feature outputs and H2, W2 and C2 respectively represent FG The length, width, and number of channels. Among them, F cls is the class token separated from the corresponding position of the output of the last ViT encoder, and F G is the feature after dimensional transformation of the remaining output after the separation of the class token. To further enhance the global information represented by the output features of the global feature extractor, F G is fused with F cls . F G is compressed to the same size of 1×C2 as F cls through global pooling operations on H2 and W2; subsequently, the two are added together to obtain the fusion, resulting in DINOv2 encoded features of size 1025×384 and it is used as the final output of the DINOv2 feature extractor, thereby obtaining the final global feature.
[0090] In the embodiments of the present application, the global feature extractor based on the DINOv2 model is based on the pre-trained ViT-S / 14 and has a powerful global feature capture ability. The pre-trained ViT-S / 14 used is obtained by model distillation from the pre-trained ViT-G / 14 with 1B parameters. This method can efficiently compress the general knowledge of large vision models into smaller models, and only incur a minimal loss in accuracy while significantly reducing the inference cost. Therefore, the designed global feature extractor not only provides rich global information for EGJA staging diagnosis, but also does not significantly increase the computational burden of the overall model. At the same time, the combination with the local feature extractor also significantly reduces the dependence of the large vision model encoder on the amount of training data.
[0091] In an optional embodiment, a mixture of experts model is used to fuse local features and global features to obtain deep fusion features, including:
[0092] Assign corresponding local weights and global weights to local features and global features through a gating network;
[0093] Based on the local weights and global weights, fuse local features and global features at the element level to obtain deep fusion features.
[0094] In the embodiments of the present application, a mixture of experts model is adopted to achieve deep fusion of local features and global features extracted by the local feature extractor and the global feature extractor, so as to achieve complex tasks that are difficult to complete by a single feature extractor. Optionally, the mixture of experts model adopts an element-wise feature fusion mechanism. Through a gating network, learnable fusion weights are provided for each corresponding element of the local features and global features to be fused, and the two are fused at the element level, thereby generating a more comprehensive and integrated deep fusion feature compared to only using a single feature extractor.
[0095] In an alternative embodiment, the gating network includes three convolutional layers;
[0096] Assigning corresponding local weights and global weights to the local features and global features through the gating network includes:
[0097] After aligning the local features and global features, merge them in the first dimension to obtain merged features;
[0098] Process the merged features sequentially using three convolutional layers to obtain intermediate features;
[0099] Process the intermediate features through the softmax function to obtain final features;
[0100] Separate the final features into local weights and global weights.
[0101] In the embodiments of the present application, as Figure 2 shown, the purpose of the gating network is to achieve the deep fusion of the local features F C extracted by the local feature extractor and the global features F DINO output by the global feature extractor. Using the global average pooling operation, the local features F C output by the local feature extractor are compressed from H1×W1×C1 to 1×C1, and then the number of channels of the local features F C is transformed from C1 to C2 through a linear layer (fully connected layer) to obtain features DINO aligned with the global features F Merge F Res and F DINO in the first dimension to generate a merged feature of size 1×C2×2 as the input to the gating network. This network is used to fuse the C2 pairs of elements in F DINO and F Res The network mainly includes three one-dimensional convolutional layers (kernel size is 1, and the number of filters is 1024, 1024, 2 respectively). The outputs of the first two convolutional layers will be normalized and activated, and then passed through a dropout layer to randomly discard some nodes to improve the generalization of the model and prevent overfitting of the model. The output of the last convolutional layer will pass through a softmax layer to generate an output of 1×C2×2 with the same size as the input to obtain the final features. The final features are further separated into local weights and global weights These are two learnable weight vectors used for element-wise fusion of F Res and F DINO to generate the deep fusion feature
[0102] F Fus= W Res ⊙F Res + W DINO ⊙F DINO
[0103] Wherein, ⊙ represents the dot product operation of the corresponding positions of two vectors.
[0104] In the embodiments of the present application, the mixture of experts model realizes element - level feature fusion by designing a gating network. This network can assign a learnable fusion weight to each corresponding element in the local and global features to be fused. Compared with the traditional vector - level fusion method, the design of the embodiments of the present application can make the information fusion in the features more sufficient, thereby generating a more accurate fused feature representation. Using these fused features, the classifier of the model can output a more accurate prediction and diagnosis result.
[0105] In an alternative embodiment, the classifier is constructed based on a multi - layer perceptron structure and includes three fully - connected layers;
[0106] Using the classifier to classify the deep - fused features to obtain a prediction and diagnosis result, including:
[0107] Input the deep - fused features into the first two fully - connected layers in sequence, and output to obtain the first classification feature;
[0108] After normalizing and activating the first classification feature, perform dropout to obtain the second classification feature;
[0109] Input the second classification feature into the third fully - connected layer, and output to obtain the third classification feature;
[0110] Input the third classification feature into the softmax layer to obtain the probability of the corresponding disease category, and output the prediction and diagnosis result according to the probability.
[0111] The classifier in the embodiments of the present application aims to classify the input image based on the deep - fused features, and output the prediction probability distribution of the image on all possible categories through a multi - layer perceptron; finally, the model will output the most likely prediction and diagnosis result.
[0112] Among them, the classifier based on the multi - layer perceptron structure includes three fully - connected layers. The outputs of the first two fully - connected layers will be randomly discarded by some nodes through a dropout layer after normalization and activation. The output of the last fully - connected layer will pass through a softmax layer to generate the EGJA prediction and diagnosis probability distribution [p1, p2, p3] of the endoscopic image input to the model, where p i represents the probability that the input endoscopic image is diagnosed by the model as the corresponding EGJA stage or normal. For example Figure 2As shown, p1 represents the probability that the input endoscopic image is diagnosed as advanced EGJA by the model, p2 represents the probability that the input endoscopic image is diagnosed as early EGJA by the model, and p3 represents the probability that the input endoscopic image is diagnosed as normal by the model. Finally, the model will output the most likely predicted diagnosis result for the input image.
[0113] In the embodiment of the present application, a dataset of 9,191 white light and narrow band imaging endoscopic images was collected, and experiments were carried out using the model proposed in the embodiment of the present application (hereinafter simply referred to as the model of the present application), and compared with some existing classic CNNs and large vision models. As shown in Table 1, the model proposed in the embodiment of the present application performs excellently in evaluation indicators such as precision, recall, F1 score, and area under the receiver operating characteristic curve (AUC), exceeding some existing classic CNNs and large vision models.
[0114] Table 1 Performance comparison with other deep learning models in the overall diagnosis of EGJA
[0115] Model Precision Recall F1Score AUC Time (s) ResNet50 0.9134 0.9125 0.9128 0.9790 0.0092 Inception-ResNet-V2 0.8959 0.8906 0.8918 0.9732 0.0270 EfficientNetV2 0.8784 0.8807 0.8789 0.9716 0.0181 ViT-B / 16 0.8789 0.8796 0.8783 0.9718 0.0088 CLIP-RN50 0.8713 0.8676 0.8687 0.9577 0.0096 The model of this application 0.9254 0.9256 0.9254 0.9818 0.0158
[0116] In addition, a human-machine comparison experiment was also carried out in the embodiment of the present application. 15 endoscopic physicians were selected and divided into three groups: junior, intermediate, and senior according to their experience levels, and compared and analyzed with the model of the present application. The experimental data are shown in Tables 2 and 3. Among them, the experimental results in Table 2 show that the model of the present application has significant advantages over all endoscopic physician groups in terms of precision, recall, F1 score, and diagnostic time.
[0117] Table 2 Performance comparison with endoscopic physicians with different experiences in the overall diagnosis of EGJA
[0118] Endoscopist / Model Precision Recall F1Score Time (s) Junior physician group 0.7692 0.7190 0.7146 4.146 Intermediate physician group 0.8074 0.7404 0.7442 5.626 Senior physician group 0.8271 0.8144 0.8157 2.712 The model of this application 0.9254 0.9256 0.9254 0.0158
[0119] The experimental results in Table 3 show that the model of the present application is significantly better than the senior physician group with the most experience in all classification evaluation indicators for the early diagnosis of EGJA. The above results further verify the superiority of the model of the present application in terms of the overall diagnosis accuracy and diagnosis efficiency of EGJA, especially in the early diagnosis of EGJA. This shows that the model of the present application has high clinical application value and can provide reliable and valuable diagnostic suggestions for endoscopic physicians, so as to assist them in making more accurate final decisions.
[0120] Table 3 Performance comparison with endoscopic physicians with different experiences in the early diagnosis of EGJA
[0121] Endoscopist / Model Precision Recall F1Score Specificity NPV Junior physician group 0.5069 0.5963 0.5234 0.8336 0.8836 Intermediate physician group 0.5083 0.7474 0.5756 0.7662 0.9207 Senior physician group 0.6752 0.7243 0.6974 0.8978 0.9183 The model of this application 0.8502 0.8462 0.8482 0.9561 0.9547
[0122] Therefore, the method in this application performs excellently in the early diagnosis of EGJA and can achieve significantly better accuracy than existing manual diagnosis methods and some representative deep learning diagnosis methods.
[0123] Based on the same principle as the method provided in the embodiments of this application, the embodiments of this application also provide an endoscopic esophagogastric junction adenocarcinoma staging diagnosis system, as Figure 5 shown, the system includes:
[0124] A local feature extraction module 501, configured to extract local features in an endoscopic image by using a local feature extractor based on a convolutional neural network;
[0125] A global feature extraction module 502, configured to extract global features in an endoscopic image by using a global feature extractor based on a vision large model;
[0126] A feature fusion module 503, configured to fuse local features and global features by using a mixture of experts model to obtain deeply fused features;
[0127] A classification prediction module 504, configured to classify the deeply fused features by using a classifier to obtain a predicted diagnosis result.
[0128] In the embodiments of this application, a local feature extractor based on a convolutional neural network and a global feature extractor based on a vision large model are used to extract local features and global features in an endoscopic image respectively, and the local features and global features are fused to obtain deeply fused features. Finally, a classifier is used to classify the fused features to obtain a predicted diagnosis result. The embodiments of this application realize the first application of the vision large model in the field of EGJA diagnosis, and make full use of the complementary advantages of the convolutional neural network and the vision large model, significantly improving the accuracy of endoscopic EGJA staging diagnosis.
[0129] The endoscopic esophagogastric junction adenocarcinoma staging diagnosis system provided in the embodiments of this application can implement Figures 1 to 4 each process implemented in the method embodiments. To avoid repetition, it will not be elaborated here.
[0130] The endoscopic esophagogastric junction adenocarcinoma staging diagnosis system of the embodiments of this application can execute the endoscopic esophagogastric junction adenocarcinoma staging diagnosis method provided in the embodiments of this application, and its implementation principle is similar. The actions performed by each module and unit in the endoscopic esophagogastric junction adenocarcinoma staging diagnosis system in the embodiments of this application correspond to the steps in the endoscopic esophagogastric junction adenocarcinoma staging diagnosis method in the embodiments of this application. For the detailed function descriptions of each module of the endoscopic esophagogastric junction adenocarcinoma staging diagnosis system, reference can specifically be made to the descriptions in the corresponding endoscopic esophagogastric junction adenocarcinoma staging diagnosis method shown above, and it will not be elaborated here.
[0131] Based on the same principle as the method shown in the embodiments of the present application, the embodiments of the present application also provide an electronic device, which may include but is not limited to: a processor and a memory; the memory is used to store a computer program; the processor is used to execute the endoscopic esophagogastric junction adenocarcinoma staging diagnosis method shown in any optional embodiment of the present application by calling the computer program. Compared with the prior art, the endoscopic esophagogastric junction adenocarcinoma staging diagnosis method provided by the present application realizes the first application of the vision large model in the field of EGJA diagnosis, and makes full use of the complementary advantages of the convolutional neural network and the vision large model, significantly improving the accuracy of endoscopic EGJA staging diagnosis.
[0132] In an optional embodiment, an electronic device is also provided, as Figure 6 shown Figure 6 The electronic device 600 shown can be a server, including: a processor 601 and a memory 603. Among them, the processor 601 and the memory 603 are connected, such as connected through a bus 602. Optionally, the electronic device 600 may further include a transceiver 604. It should be noted that in actual applications, the transceiver 604 is not limited to one, and the structure of the electronic device 600 does not constitute a limitation on the embodiments of the present application.
[0133] The processor 601 may be a CPU (Central Processing Unit, central processor), a general-purpose processor, a DSP (Digital Signal Processor, data signal processor), an ASIC (Application Specific Integrated Circuit, application-specific integrated circuit), an FPGA (Field Programmable Gate Array, field programmable gate array) or other programmable logic devices, transistor logic devices, hardware components or any combination thereof. It can implement or execute various exemplary logic blocks, modules and circuits described in connection with the disclosure of the present application. The processor 601 may also be a combination that implements a computing function, such as a combination including one or more microprocessors, a combination of a DSP and a microprocessor, etc.
[0134] The bus 602 may include a path for transmitting information between the above components. The bus 602 may be a PCI (Peripheral Component Interconnect, peripheral component interconnect standard) bus or an EISA (Extended Industry Standard Architecture, extended industry standard architecture) bus, etc. The bus 602 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 6It is represented by only one thick line, but it does not mean that there is only one bus or one type of bus.
[0135] The memory 603 can be a ROM (Read Only Memory), or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory), or other types of dynamic storage devices that can store information and instructions. It can also be an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.
[0136] The memory 603 is used to store the application program code for implementing the solution of this application, and is controlled by the processor 601 for execution. The processor 601 is used to execute the application program code stored in the memory 603 to implement the content shown in the foregoing method embodiments.
[0137] Among them, the electronic device includes but is not limited to: mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 6 The shown electronic device is only an example and should not bring any restrictions to the functions and usage scopes of the embodiments of this application.
[0138] The server provided by this application can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, as well as big data and artificial intelligence platforms. The terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto. The terminal and the server can be directly or indirectly connected through wired or wireless communication methods, and this application does not make any restrictions here.
[0139] An embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program runs on a computer, the computer can execute the corresponding content in the foregoing method embodiment.
[0140] It should be understood that although the steps in the flowchart of the accompanying drawings are shown in sequence according to the indication of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps is not strictly restricted by order, and they can be executed in other orders. Moreover, at least some of the steps in the flowchart of the accompanying drawings may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or sub-steps or stages of other steps.
[0141] It should be noted that the computer-readable storage medium in the present application may also be a computer-readable signal medium or a combination of a computer-readable storage medium and a computer-readable storage medium. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable storage medium may be any tangible medium that contains or stores a program, and this program can be used by or combined with an instruction execution system, apparatus, or device. In the present application, the computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which the computer-readable program code is carried. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium may also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable signal medium can send, propagate, or transmit a program for use by or combined with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0142] The above computer-readable medium may be included in the above electronic device; or it may exist separately without being assembled into the electronic device.
[0143] The above computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to execute the methods shown in the above embodiments.
[0144] According to one aspect of the present application, there is provided a computer program product or a computer program. The computer program product or the computer program includes computer instructions, which are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the endoscopic esophagogastric junction adenocarcinoma staging diagnosis method and system provided in the above various alternative implementation manners.
[0145] Computer program code for performing the operations of the present application may be written in one or more programming languages or combinations thereof. The above programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (for example, by connecting through the Internet using an Internet service provider).
[0146] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0147] The modules involved in the embodiments of the present application can be implemented in software or in hardware. Among them, the name of the module does not constitute a limitation on the module itself in some cases. For example, the local feature extraction module can also be described as "a local feature extraction module for extracting local features in endoscopic images using a local feature extractor based on a convolutional neural network".
[0148] The above description is only a preferred embodiment of the present application and an explanation of the technical principles applied. Those skilled in the art should understand that the scope of disclosure involved in the present application is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) having similar functions disclosed in the present application.
Claims
1. An endoscopic staging diagnostic method for adenocarcinoma of the esophagogastric junction, characterized in that, The method includes: Extracting local features in the endoscopic image using a local feature extractor based on a convolutional neural network; Extracting global features in the endoscopic image using a global feature extractor based on a large vision model; Fusing the local features and the global features using a mixture of experts model to obtain deep fusion features; Classifying the deep fusion features using a classifier to obtain a predicted diagnosis result.
2. The endoscopic staging diagnostic method for adenocarcinoma of the esophagogastric junction according to claim 1, wherein The local feature extractor uses a ResNet50 network; The extracting local features in the endoscopic image using a local feature extractor based on a convolutional neural network includes: Performing convolution and max pooling operations on the endoscopic image to obtain low-dimensional local features; Performing feature extraction on the low-dimensional local features through convolutional modules in four stages to obtain the local features; Among them, the convolutional modules in the four stages all contain residual blocks, and the output of each residual block is added to the input through an identity mapping to form a residual connection.
3. The endoscopic staging diagnostic method for adenocarcinoma of the esophagogastric junction according to claim 1, characterized in that, The global feature extractor uses a DINOv2 model; The extracting global features in the endoscopic image using a global feature extractor based on a large vision model includes: Dividing the endoscopic image into multiple non-overlapping image patches; Flattening and merging multiple non-overlapping image patches and linearly projecting them into a high-dimensional feature space to generate feature vectors; Merging the category labels of the feature vectors and adding position information to obtain embedding vectors; Performing feature extraction on the embedding vectors through stacked ViT encoders to obtain the global features.
4. The endoscopic staging diagnostic method for adenocarcinoma of the esophagogastric junction according to claim 3, wherein The ViT encoder includes a multi-head attention mechanism and a feed-forward neural network.
5. The endoscopic staging diagnosis method for adenocarcinoma of the esophagogastric junction according to claim 1, characterized in that, The fusing the local features and the global features using a mixture of experts model to obtain deep fusion features includes: Assigning corresponding local weights and global weights to the local features and the global features through a gating network; Based on the local weights and the global weights, fusing the local features and the global features at the element level to obtain the deep fusion features.
6. The endoscopic staging diagnosis method for adenocarcinoma of the esophagogastric junction according to claim 5, wherein The gating network includes three convolutional layers; The assigning corresponding local weights and global weights to the local features and the global features through a gating network includes: After aligning the local features and the global features, merging them in the first dimension to obtain merged features; Successively processing the merged features using three convolutional layers to obtain intermediate features; Processing the intermediate features through a softmax function to obtain final features; Separating the final features into the local weights and the global weights.
7. The endoscopic staging diagnosis method for adenocarcinoma of the esophagogastric junction according to claim 1, characterized in that, The classifier is constructed based on a multi-layer perceptron structure and includes three fully connected layers; The classifying the deep fusion features using a classifier to obtain a predicted diagnosis result includes: Successively inputting the deep fusion features into the first two fully connected layers, and outputting to obtain a first classification feature; Normalizing and activating the first classification feature and then performing dropout to obtain a second classification feature; Inputting the second classification feature into the third fully connected layer, and outputting to obtain a third classification feature; Inputting the third classification feature into a softmax layer to obtain the probabilities of corresponding disease categories, and outputting the predicted diagnosis result according to the probabilities.
8. An endoscopic staging diagnosis system for adenocarcinoma of the esophagogastric junction, characterized in that, The system includes: A local feature extraction module, configured to extract local features in the endoscopic image by using a local feature extractor based on a convolutional neural network; A global feature extraction module, configured to extract global features in the endoscopic image by using a global feature extractor based on a large vision model; A feature fusion module, configured to fuse the local features and the global features by using a mixture of experts model to obtain deep fusion features; A classification and prediction module, configured to classify the deep fusion features by using a classifier to obtain a predicted diagnosis result.
9. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the method described in any one of claims 1-7 is implemented.
10. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium. When the computer program is executed by a processor, the method described in any one of claims 1-7 is implemented.