Enhanced interpretable multi-mode esophageal cancer target detection method and system

Through the labeling and processing of energy spectrum CT images and multimodal classification, a deep learning fusion model was constructed to generate an esophageal tumor detection report, which solved the problem of inconsistent lymph node dissection standards after esophageal cancer, and achieved efficient and accurate automatic detection of esophageal tumors, improving diagnostic accuracy and interpretability.

CN120259215APending Publication Date: 2025-07-04SUN YAT SEN UNIVERSITY CANCER CENTER (CANCER HOSPITAL AFFILIATED TO SUN YAT SEN UNIVERSITY CANCER RESEARCH INSTITUTE OF SUN YAT SEN UNIVERSITY)
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510318557.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

In the prior art, the standards for lymph node dissection after esophageal cancer are not unified, the tumor area is not accurate, the classification effect is not ideal, the high-resolution specific category activation map interpretation disorders and the focus of medical reports is unclear, which affects the accuracy of diagnosis.

Method used

By collecting energy spectrum CT images for labeling, multimodal classification, building a fusion model based on multimodal deep learning, generating an esophageal tumor detection report, establishing an intelligent detection cloud platform, and realizing image data transmission, storage management and automatic detection.

Benefits of technology

It improves the accuracy and robustness of image segmentation of esophageal tumors, achieves efficient and accurate automatic detection, assists clinicians to formulate reasonable surgical paths, and reduces postoperative recurrence rates and complications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259215A_ABST
    Figure CN120259215A_ABST
Patent Text Reader

Abstract

The invention discloses an enhanced interpretable multi-mode esophageal cancer target detection method and system, and the method comprises the steps: collecting a medical image of an energy spectrum CT detection esophageal tumor, and carrying out the marking processing of the obtained CT image; according to the marked CT image, carrying out multi-modal classification on the medical image and the clinical data information; constructing a fusion model based on multi-modal deep learning to perform fusion processing on the multi-modal features obtained by classification; according to fusion features obtained by fusion, generating an esophageal tumor detection report based on weak supervision contrast learning; constructing an intelligent detection cloud platform by taking the optimized AI diagnosis model as a kernel; the intelligent detection cloud platform is used for transmission, storage management, image data preprocessing, automatic detection and query statistics of energy spectrum CT images. According to the embodiment of the invention, the accuracy and robustness of image segmentation can be improved, efficient and accurate automatic detection of esophageal tumors is realized, and the method can be widely applied to the technical field of computers.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular, to an enhanced interpretable multi-modal esophageal cancer target detection method and system. Background Art

[0002] Esophageal cancer is one of the most common malignant tumors worldwide, with its incidence and mortality ranking 7th and 6th among all malignant tumors respectively. The lymph node metastasis of esophageal cancer has the characteristic of "skipping", and the residual tiny lymph node metastasis foci after esophageal cancer surgery are the root causes of tumor recurrence and metastasis. Thorough lymph node dissection can effectively reduce the recurrence after esophageal cancer surgery, but the corresponding surgical risks and the incidence of postoperative complications increase accordingly. Therefore, there is still no consensus on the standard of thoracic lymph node dissection for esophageal cancer internationally.

[0003] Artificial intelligence deep learning has currently achieved great success in many fields, such as image recognition, speech recognition, and natural language processing. The rapid development of deep learning technology has provided new ideas for medical image analysis and greatly improved the accuracy of target recognition in medical images. In addition, in response to problems faced in the field of medical images such as different sizes of lesions, the models and algorithms of deep learning are also continuously optimized to ensure performance improvement. Therefore, the deep combination of artificial intelligence and medical images will, to a great extent, relieve the work pressure of clinicians and improve work efficiency. The combination of spectral CT and artificial intelligence deep learning not only has the potential to further improve the diagnostic accuracy of regional lymph node metastasis before esophageal cancer surgery; but also by establishing a multi-modal intelligent detection system of spectral CT and new tumor markers through artificial intelligence, the diagnostic accuracy of esophageal cancer can be further improved.

[0004] The following problems exist in the related prior art: inaccurate annotation of the lesion location in the tumor region, unsatisfactory classification effect, obstacles in interpreting high-resolution activation maps of specific categories, and unclear focus in medical reports. Summary of the Invention

[0005] The main purpose of the embodiments of the present invention is to propose an enhanced interpretable multi-modal esophageal cancer target detection method and system, which can improve the accuracy and robustness of image segmentation and achieve efficient and accurate automatic detection of esophageal tumors.

[0006] To achieve the above object, on the one hand, an embodiment of the present invention proposes an enhanced interpretable multi-modal esophageal cancer target detection method, including the following steps:

[0007] Collect medical images of spectral CT for detecting esophageal tumors, and perform annotation processing on the obtained CT images;

[0008] According to the annotated CT images, perform multi-modal classification on the medical images and clinical data information;

[0009] Construct a fusion model based on multimodal deep learning to fuse the classified multimodal features;

[0010] Generate an esophageal tumor detection report based on the fused features obtained by weak-supervised contrastive learning;

[0011] Build an intelligent detection cloud platform with the optimized AI diagnosis model as the core; the intelligent detection cloud platform is used for the transmission, storage management, image data preprocessing, automatic detection and query statistics of spectral CT images.

[0012] Another aspect of the embodiments of the present invention also provides an enhanced interpretable multimodal esophageal cancer target detection system, including:

[0013] The first module is used to collect medical images of spectral CT detecting esophageal tumors and perform annotation processing on the obtained CT images;

[0014] The second module is used to perform multimodal classification on medical images and clinical data information according to the annotated CT images;

[0015] The third module is used to construct a fusion model based on multimodal deep learning to fuse the classified multimodal features;

[0016] The fourth module is used to generate an esophageal tumor detection report based on the fused features obtained by weak-supervised contrastive learning;

[0017] The fifth module is used to build an intelligent detection cloud platform with the optimized AI diagnosis model as the core; the intelligent detection cloud platform is used for the transmission, storage management, image data preprocessing, automatic detection and query statistics of spectral CT images.

[0018] To achieve the above object, another aspect of the embodiments of the present invention proposes an electronic device, the electronic device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the above-mentioned method is implemented.

[0019] To achieve the above object, another aspect of the embodiments of the present invention proposes a computer-readable storage medium, the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned method is implemented.

[0020] The embodiments of the present invention also disclose a computer program product or a computer program, the computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device can read the computer instructions from the computer-readable storage medium, and when the processor executes the computer instructions, the computer device executes the above-mentioned method.

[0021] The embodiments of the present invention at least include the following beneficial effects: The present invention provides an enhanced interpretable multi-modal esophageal cancer target detection method and system. This solution collects medical images of esophageal tumors detected by spectral CT, and performs annotation processing on the obtained CT images; according to the annotated CT images, multi-modal classification is performed on the medical images and clinical data information; a fusion model based on multi-modal deep learning is constructed to fuse the classified multi-modal features; according to the fused features obtained by fusion, an esophageal tumor detection report is generated based on weakly supervised contrast learning; with the optimized AI diagnosis model as the core, an intelligent detection cloud platform is constructed; the intelligent detection cloud platform is used for the transmission, storage management, image data preprocessing, automatic detection, and query statistics of spectral CT images. The embodiments of the present invention can improve the accuracy and robustness of image segmentation, and achieve efficient and accurate automatic detection of esophageal tumors. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 FIG. is a schematic diagram of an implementation environment provided by an embodiment of the present invention;

[0023] Figure 2 FIG. is a flowchart of the overall steps provided by an embodiment of the present invention;

[0024] Figure 3 FIG. is a schematic diagram of the overall idea provided by an embodiment of the present invention;

[0025] Figure 4 FIG. is the main process of the esophageal tumor annotation method provided by an embodiment of the present invention;

[0026] Figure 5 FIG. is the main framework of the Deeplabv3+ network with 4 channels provided by an embodiment of the present invention;

[0027] Figure 6 FIG. is a diagram illustrating the working principle of TandemNet provided by an embodiment of the present invention;

[0028] Figure 7 FIG. is the workflow of the end-to-end model (gCAM-CCL) for automatic classification and interpretation of multi-modal data fusion provided by an embodiment of the present invention;

[0029] Figure 8 FIG. is a weakly supervised contrast learning framework provided by an embodiment of the present invention;

[0030] Figure 9 FIG. is a physical topology diagram of the interpretable multi-modal tumor intelligent detection system provided by an embodiment of the present invention;

[0031] Figure 10 FIG. is an architecture diagram of the interpretable multi-modal tumor intelligent detection system provided by an embodiment of the present invention;

[0032] Figure 11 It is a visualization diagram of the segmentation results of esophageal cancer, where (a) is the result of U-NetPlus; (b) is the result of U-Net; (c) is the result of LinkNet; (d) is the result of SegNet;

[0033] Figure 12 It is the flowchart of Channel-attention U-Net;

[0034] Figure 13 It is the visualization of the segmentation result;

[0035] Figure 14 It is the flowchart of Eso-Net;

[0036] Figure 15 It is the visualization diagram of the 2.5D segmentation result;

[0037] Figure 16 It is the CT image of the lymph nodes in the main pulmonary artery window. Among them, the iodine-based substance density is 2.2 mg / ml, and the effective atomic number is 8.77. On the same layer, the iodine-based substance density of the thoracic aorta is 5.7 mg / ml, and the effective atomic number is 10.05;

[0038] Figure 17 It is the ROC curve graph of the combined diagnosis result. Specific implementation manners

[0039] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the embodiments of the present invention. They are only examples of devices and methods consistent with some aspects of the embodiments of the present invention as detailed in the appended claims.

[0040] The enhanced interpretable multi-modal esophageal cancer target detection method and system provided by the embodiments of the present invention relate to the field of computer technology. The enhanced interpretable multi-modal esophageal cancer target detection method provided by the embodiments of the present invention can be applied to a terminal, a server, or software running on a terminal or a server. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, a vehicle-mounted terminal, etc., but is not limited thereto; the server side can be configured as an independent physical server, or as a server cluster or distributed system composed of multiple physical servers, or as a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The server can also be a node server in a blockchain network; the software can be an application implementing the enhanced interpretable multi-modal esophageal cancer target detection method, etc., but is not limited to the above forms.

[0041] The present invention can be used in numerous general-purpose or special-purpose computer system environments or configurations. For example: personal computers, server computers, hand-held or portable devices, tablet-type devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. The present invention can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present invention can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0042] As Figure 1 shown, it is a schematic diagram of an implementation environment provided by the embodiments of the present invention. Referring to Figure 1 , this implementation environment includes at least one terminal 102 and a server 101. The terminal 102 and the server 101 can be network-connected by wireless or wired means to complete data transmission and exchange.

[0043] The server 101 can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.

[0044] In addition, the server 101 can also be a node server in a blockchain network. Among them, the blockchain is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithms.

[0045] The terminal 102 can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart speaker, a smart watch, etc. Among them, the terminal 102 can also be an in-vehicle terminal of various device types exemplified above, but is not limited thereto. The terminal 102 and the server 101 can be directly or indirectly connected through wired or wireless communication methods, and the embodiments of the present invention do not limit this here.

[0046] Exemplarily based on Figure 1 the described implementation environment, the embodiments of the present invention provide an enhanced interpretable multi-modal esophageal cancer target detection method. Taking the application of this enhanced interpretable multi-modal esophageal cancer target detection method in the server 101 as an example for description, it can be understood that this method can also be applied to the terminal 102.

[0047] Referring to Figure 2 Figure 2 is a flowchart of the enhanced interpretable multi-modal esophageal cancer target detection method applied to the server provided by the embodiments of the present invention. The execution subject of this method can be any of the aforementioned computer devices (including servers or terminals). Referring to Figure 2 this figure, this method may include the following steps:

[0048] Collect medical images of esophageal tumors detected by spectral CT, and perform annotation processing on the obtained CT images;

[0049] According to the annotated CT images, perform multi-modal classification on the medical images and clinical data information;

[0050] Construct a fusion model based on multi-modal deep learning to fuse the classified multi-modal features;

[0051] According to the fused features obtained by fusion, generate an esophageal tumor detection report based on weakly supervised contrast learning;

[0052] ​Construct an intelligent detection cloud platform with the optimized AI diagnosis model as the core; the intelligent detection cloud platform is used for the transmission, storage management, image data preprocessing, automatic detection, and query statistics of spectral CT images.

[0053] In some embodiments, collect the medical images of spectral CT for esophageal tumor detection, and perform annotation processing on the obtained CT images, including the following steps:

[0054] Delete the areas that do not contribute to the annotation, and adjust the remaining area of the spectral CT image to 384px × 384px through bilinear interpolation to complete the preprocessing of the image;

[0055] Use HRDEN to establish the depth map of the spectral CT image, where HRDEN consists of 4 modules: an encoder, a decoder, a multi-scale feature fusion module, and a refinement module;

[0056] Adopt the DeepLabV3+ model to predict the EEC area, support 4-channel input including the depth channel C5 by modifying the first-layer convolution kernel, and maintain the feature extraction performance of the deep CNN module at the same time; use the ASPP module to fuse multi-scale features, simplify the low-level features and concatenate them with the high-level features in the decoding stage, and process the fused features to output a multi-scale information map;

[0057] Optimize the prediction result through post-processing, fill the small holes and apply morphological operations until the binary image C8 area is stable as C9 and the edge is smooth; then map C9 to the original RGB image C1 to generate the final annotation image C10, thus completing the clinically friendly automatic annotation process.

[0058] In some embodiments, perform multi-modal classification on the medical images and clinical data information according to the annotated CT images, including the following steps:

[0059] Extract the visual features and semantic features of the input image and text through the visual model and the language model;

[0060] Construct the interaction between the visual features and semantic features through the attention mechanism to generate the context vector;

[0061] Introduce a random modality transfer function to approximately generate semantic features in the absence of text, combine the refined visual features with the context vector, and input them into the fully connected neural network for classification; jointly optimize through the least squares method and the cross-entropy loss.

[0062] In some embodiments, the extraction of the visual features and semantic features of the input image and text through the visual model and the language model includes the following steps:

[0063] Find the optimal interaction calculation unit I between the visual feature V and the semantic feature S such that the following formula holds:

[0064]

[0065] In the formula, p is the label possibility;

[0066] wherein, the semantic feature S comes from the text encoding S when the text is available r , or from the simulated test encoding S when the text is not available f .

[0067] In some embodiments, constructing the interaction between the visual feature and the semantic feature through the attention mechanism to generate a context vector includes the following steps:

[0068] Construct an R-radix interaction module, which includes an attention module, a visual-semantic interaction module, and a prediction module. Use the attention module to refine visual knowledge and consider the multi-modal segmentation relationship to construct the interaction between the image feature and the text feature; wherein, the input of the attention module consists of a query vector Q, a set of vectors called keywords in matrix K, and a set of vectors called values in matrix Q. The defined expression of the attention module is: where d q is the dimension of the query Q; the output is a context vector, which is the weighted average of Q The function outputs a weight vector, which is used to characterize the importance of each value in Q;

[0069] The text feature includes N hidden states Hidden state is used to query the visual feature V and serves as the target text embedding, which is simulated by the proposed channel transfer function;

[0070] The defined expression of the visual-semantic interaction module is: V = P(S T W q , V T W K , V T W Q ) * W i , where V represents the output of the visual-semantic interaction module; P represents the parallel multi-head attention module; S T represents the text feature; W q represents the query weight matrix; W K represents the key weight matrix; V T represents the visual feature; W Q represents the value weight matrix; W i represents the output projection matrix;

[0071] Define the expression for constructing the joint context vector: where \(v\in\mathbb{R}\) C is the average context vector containing the extracted visual features; \(R\) is a cardinality factor selected empirically;

[0072] The defined expression of the prediction module is: The prediction module is used to concatenate the visual feature \(V\) with \(v\) to form a two-layer fully connected neural network with LeakyReLU as the activation function.

[0073] In some embodiments, the introduction of the random modality transfer function approximately generates semantic features in the absence of text, combines the refined visual features with the context vector, and inputs them into the fully connected neural network for classification; jointly optimized by the least squares method and the cross-entropy loss, including the following steps:

[0074] Construct a modality transfer function for learning the visual-to-semantic approximation. The expression of the modality transfer function is: \(h\) T =T(\(\Delta(V)\); \(\theta\) T ), where \(h\) T \(\in\mathbb{R}\) D*1 is used to represent the simulated test encoding \(S\) F \(\in\mathbb{R}\) D*1 , \(T\) represents a two-layer fully connected network, and \(\theta\) T represents the trainable parameters;

[0075] The training expression of the trainable parameters is: \(\min\|\Delta(S r ) - T(\(\Delta(V)\); \(\theta\) T )\|^2\). The training loss of this training process is used to penalize the image model to approach the semantic features.

[0076] In some embodiments, the construction of the fusion model based on multimodal deep learning performs fusion processing on the classified multimodal features, including the following steps:

[0077] Collect and preprocess multimodal data, including feature vectors and esophageal images;

[0078] Learn features from the feature data using one-dimensional convolution through gCAM-CCL and learn features from the imaging data using two-dimensional convolution; where the outputs of the two convolutional neural networks are flattened and then fused in the loss function in the collaboration layer;

[0079] Select two intermediate layers, use gradient-based weights to combine the feature maps therefrom, and correspondingly generate the Grad-CAM activation map for a specific class;

[0080] A fine-grained activation map is calculated by using the guiding BP to project the gradient from the collaborative layer back to the input layer, and the obtained activation map indicates the contribution of pixels to the decision of interest.

[0081] In some embodiments, based on the fused features obtained by fusion, generating an esophageal tumor detection report based on weakly supervised contrastive learning includes the following steps:

[0082] Extract visual features from esophageal tumor detection images, and then use a memory-driven transformer model as the backbone model to generate report text;

[0083] Use a fine-tuned BERT model to embed and cluster the report to identify semantic similarities;

[0084] Maximize the similarity between the image and the report through weakly supervised contrastive learning, while minimizing the similarity between negative samples;

[0085] Improve the accuracy and diversity of the model when generating long text descriptions by mixing and optimizing the cross-entropy loss and the contrastive loss.

[0086] Next, taking a specific application scenario as an example, the implementation process of the embodiments of the present invention will be described in detail:

[0087] Aiming at the problems existing in the prior art, the embodiments of the present invention study an enhanced interpretable multi-modal tumor intelligent detection system and its report generation method, covering automatic annotation, multi-modal fusion classification, interpretable model optimization, and weakly supervised contrastive learning report generation:

[0088] First, apply deep learning technology to develop an automatic annotation and segmentation tool to realize the automatic annotation of early esophageal cancer lesions in spectral CT images and reduce manual intervention. Then, establish a multi-modal medical image semantic fusion classification model, use advanced algorithms to improve the classification accuracy, achieve more accurate diagnosis, and promote the feature extraction and classification process of the project. This method improves the model structure, establishes an interpretable medical image analysis model, visualizes the inference mechanism of the model, so that the model has interpretability, and combines with clinical applications to make the prediction results closer to pathological diagnosis. Finally, this method fuses semantic and multi-modal image features, and proposes a weakly supervised contrastive learning framework for generating esophageal tumor detection reports to improve the generation performance of medical image reports. The method of the present invention can be applied to technical fields such as deep learning, multi-modal image processing, interpretable AI, and weakly supervised learning.

[0089] As Figure 3 shown, this embodiment provides a method for an interpretable multi-modal tumor intelligent detection system and report generation. It includes the following steps:

[0090] S1: Perform annotation through an automatic annotation technology for esophageal tumor images based on deep learning;

[0091] Collect medical images of esophageal tumors detected by spectral CT, remove the black background area in the spectral CT images, and equalize the size of the medical images;

[0092] The depth map of the spectral CT image is initially extracted by a DL network.

[0093] The additional depth information is fused with the original spectral CT image and then sent to another DL network to obtain an accurate annotation of the esophageal cancerous area.

[0094] S2: Perform multi-modal classification on medical images and clinical data information;

[0095] Combine the spectral CT image and clinical detection data to find the optimal interaction calculation unit between visual features and semantic features;

[0096] Control the coupling degree between visual and semantic information, so that TandemNet considers the multi-modal segmentation relationship when constructing the interaction between images and texts;

[0097] S3: An interpretable fusion model based on multi-modal deep learning;

[0098] Use Grad-CAM-guided convolutional collaborative learning (gCAM-CCL) to achieve the interpretability problem by combining intermediate feature maps with gradient-based weights;

[0099] Use the gCAM-CCL model to generate interpretable activation maps to quantify the pixel-level contributions of input features.

[0100] S4: Weakly supervised contrastive learning for esophageal tumor detection report generation;

[0101] Introduce weakly supervised contrastive loss, assign more weights to reports with semantic proximity to the target, and assign more weights to the reports close to the target during training, that is, pay more attention.

[0102] S5: Build an intelligent detection cloud platform software with an optimized AI diagnosis model as the core, including the transmission and storage management of spectral CT images, image data preprocessing, automatic detection, and query statistics modules, to improve the efficient and accurate automatic detection of esophageal tumors and provide a platform for multi-center verification and system optimization.

[0103] In the specific implementation process, first, medical images of esophageal tumors detected by spectral CT are collected, and the black background area in the spectral CT images is removed and the size of the medical images is equalized. An HRDEN is used to establish the depth map of the spectral CT images. The HRDEN consists of four modules: an encoder, a decoder, a multi-scale feature fusion (MFF) module, and a refinement module. The decoder and the refinement module respectively achieve the preliminary feature extraction of the high-resolution map and the basic feature extraction of the high-depth map. The DeepLabV3+ model is used to predict the EEC region. By modifying the first-layer convolutional kernel to support 4-channel input including the depth channel C5, while maintaining the feature extraction performance of the deep CNN module. The ASPP module is used to fuse multi-scale features, and the low-level features are simplified and concatenated with the high-level features in the decoding stage. The fused features are processed to output a multi-scale information map. Finally, the prediction results are optimized through post-processing, small holes are filled, and morphological operations are applied until the binary image C8 region is stabilized as C9 with smooth edges. Then C9 is mapped to the original RGB image C1 to generate the final annotation image C10, thus completing the clinically friendly automatic annotation process.

[0104] Then, the visual features (V) and semantic features (S) of the input image and text are extracted through the visual model and the language model; the interaction between the visual features and the semantic features is constructed through the attention mechanism to generate the context vector; a random modality transfer function (T) is introduced to approximately generate the semantic features in the absence of text. The refined visual features are combined with the context vector and input into the fully connected neural network for classification; jointly optimized through the least squares method and the cross-entropy loss to ensure that the model maintains high accuracy in asynchronous training and testing;

[0105] Then, multi-modal data, including feature vectors and esophageal images, are collected and preprocessed. The gCAM-CCL is used to learn features from the feature data using one-dimensional convolution and learn features from the imaging data using two-dimensional convolution. The outputs of the two convolutional neural networks are flattened and then fused in the loss function in the collaboration layer. Two intermediate layers are selected, from which the feature maps can be combined using gradient-based weights and the Grad-CAM activation maps of specific classes are generated accordingly. At the same time, the fine-grained activation maps are calculated by projecting the gradients from the collaboration layer back to the input layer using guided BP, and the obtained activation maps indicate the contribution of the pixels to the decision of interest.

[0106] Then, the visual features are extracted from the esophageal tumor detection images, and then the memory-driven transformer model is used as the backbone model to generate the report text; on this basis, the fine-tuned BERT model is used to embed and cluster the reports to identify semantic similarities; the similarity between the image and the report is maximized through weakly supervised contrastive learning, while the similarity between the negative samples is minimized; the accuracy and diversity of the model in generating long text descriptions are improved by mixing and optimizing the cross-entropy loss and the contrastive loss;

[0107] Finally, with the optimized AI diagnosis model as the core, an intelligent detection cloud platform software is constructed, which includes modules for the transmission and storage management of spectral CT images, image data preprocessing, automatic detection, and query statistics, to improve the efficient and accurate automatic detection of esophageal tumors and provide a platform for multi-center verification and system optimization. It provides a non-invasive artificial intelligence-assisted detection method for clinical esophageal tumors, assisting in formulating a reasonable surgical path, determining the scope of intraoperative lymph node dissection, and reducing the postoperative recurrence rate and complications.

[0108] This method uses Deeplabv3+ to achieve accurate annotation of esophageal tumor lesions in spectral CT images. By combining spectral CT images and clinical detection data, the optimal interaction calculation unit between visual features and semantic features is found, and the coupling degree between visual and semantic information is controlled, enabling TandemNet to consider the multi-modal segmentation relationship when constructing the interaction between images and texts, thereby improving scalability to adapt to the visual tasks of esophageal cancer classification. The Grad-CAM-guided convolutional collaborative learning (gCAM-CCL), which is achieved by combining the intermediate feature map with gradient-based weights, can generate interpretable activation maps to quantify the pixel-level contributions of input features. In addition, the estimated activation maps are class-specific, thus facilitating the identification of biomarkers for different groups. Finally, a weakly supervised contrastive loss is introduced to assign more weights to reports that are semantically close to the target. By exposing the model to these "hard" negative examples during training, it can learn to capture a robust representation of the essence of medical images and generate high-quality reports, thereby improving the clinical correctness of unseen images.

[0109] In step S1, annotation is performed through an automatic annotation technology for esophageal tumor images based on deep learning, and the steps include:

[0110] S1.1: First, delete the areas that do not contribute to the annotation, and adjust the remaining area of the spectral CT image to 384px × 384px through bilinear interpolation to complete the preprocessing of the image.

[0111] S1.2: Use HRDEN to establish a depth map of the spectral CT image, where HRDEN consists of 4 modules: an encoder, a decoder, a multi-scale feature fusion (MFF) module, and a refinement module. The decoder and the refinement module respectively achieve the preliminary feature extraction of the high-resolution map and the basic feature extraction of the high-depth map.

[0112] S1.3: Use the DeepLabV3+ model to predict the EEC region. Modify the first-layer convolutional kernel to support 4-channel input including the depth channel C5, while maintaining the feature extraction performance of the deep CNN module. Use the ASPP module to fuse multi-scale features, simplify the low-level features in the decoding stage and concatenate them with the high-level features. Process the fused features to output a multi-scale information map.

[0113] S1.4: Optimize the prediction results through post-processing, fill small holes and apply morphological operations until the binary image C8 region is stable as C9 with smooth edges. Then map C9 to the original RGB image C1 to generate the final annotation image C10, thus completing the clinically friendly automatic annotation process.

[0114] As Figure 4 (Main process of esophageal tumor annotation method) shown, the automatic annotation consists of four steps. The first step is preprocessing, removing the black background area and equalizing the size of the spectral CT image. In the second step, the preprocessed image is sent to HRDEN to achieve the prediction of the depth map. Then, in step 3, the RGB image and the corresponding depth map are sent to the 4-channel DeepLabV3+ network to obtain the prediction of the EEC region. Finally, post-processing is performed on the EEC prediction results to complete the final annotation.

[0115] In the step S1.1, first preprocess the spectral CT image. The black background area of the spectral CT image records some auxiliary text information such as time, gender, and age. These areas do not contribute to the annotation, so they are first deleted by a fixed window. Then the remaining area of the spectral CT image is adjusted to 384px × 384px by bilinear interpolation.

[0116] Then, based on the DL network, the prediction of the depth map is realized. This network consists of 4 modules: an encoder, a decoder, a multi-scale feature fusion (MFF) module, and a refinement module, as Figure 4 shown. The two modules C2 and C4 respectively achieve the preliminary feature extraction of the high-resolution map and the basic feature extraction of the high-depth map. For ease of description, the network in step S2 is called the "high-resolution depth estimation network" (HRDEN).

[0117] Compared with the traditional depth estimation network, HRDEN has made two improvements. The first is the extraction and fusion of multi-scale features, reducing the loss of the spatial resolution of the depth estimation map; the second is the improved loss function, which further improves the reconstruction accuracy.

[0118] Through the above improvements, HRDEN has achieved state-of-the-art performance in depth prediction. Therefore, here HRDEN is used to calculate the depth map of the spectral CT image. By utilizing and refining multi-scale features, HRDEN effectively reduces the loss of spatial resolution.

[0119] To simplify the decoding of high-level features, this DL network does not upscale the feature map to the size of the original input. Therefore, the estimated size of the depth map is less than 1 / 2 of the input image. Similarly, in this work, the size of C5 is only 152px × 114px. Therefore, this method performs bilinear interpolation on C4 to increase the size of the depth map back to 384px × 384px, denoted as C5. The final depth map C5 is used as the input for the fourth channel in the subsequent DL network.

[0120] In the step S1.2, the main framework after modifying DeepLabV3+ is as Figure 5 shown.

[0121] The color image is initially processed by the depth CNN module in the DeepLabV3+ encoder. However, the original version of this module only supports 3 input channels (i.e., the red, green, and blue channels). To utilize the additional depth channel C5, the number of channels of the convolutional kernel in the first layer is modified from 3 to 4. The subsequent layers of this deep CNN module remain unchanged to maintain its remarkable feature extraction performance.

[0122] Next, ASPP is performed. The features extracted by the 4-channel depth CNN are sent to parallel layers, the main content of which is dilated convolutions with different dilation rates. After ASPP, the features of different scales are fused through 1×1 convolutions. The feature map C6 that records multi-scale information is obtained.

[0123] In the decoding stage, the low-level features extracted by the 4-channel depth CNN are simplified through a 1×1 convolutional layer to obtain the low-level feature map C7. Then, skip connections are performed. The low-level feature C7 and the high-level multi-scale features upsampled from C6 are concatenated.

[0124] Finally, the fused features are processed through convolutional, upsampling, and activation layers. It should be noted that in the original Deeplabv3+, the function used in the last activation layer is Softmax because the original DeepLabV3+ is designed for the segmentation of 21 different types of objects in natural scenes.

[0125] However, this automatic annotation process only needs to distinguish two objects, namely the EEC region and the non-EEC region. Therefore, this method changes the activation function from Softmax to Sigmoid to make the network more competent for the binary classification task of EEC annotation. After the activation layer performs classification at the pixel level, a binary image C8 is obtained, where the white region (i.e., the "1" region) records the location of EEC damage.

[0126] To make the annotation results more acceptable to clinicians, this method proposes a post - processing step to smooth the serrated and perforated prediction results. The specific description is as follows. For the binary image C8, if the number of pixels in the small holes is respectively lower than the thresholds n1 and n2, first fill the small holes with "0" and "1". After that, perform morphological "closing" and "opening" operations in sequence. For the "closing" operation, morphological "erosion" and "dilation" are achieved through a disk - shaped structuring element with a radius of r1. Similarly, morphological "opening" is achieved through a disk - shaped element with a radius of r2. Repeat the above hole - filling and morphological operations until the area of "1" in C8 becomes constant. The constant binary image is represented by C9. As Figure 5 shown, the edge of C9 becomes smooth, solving the problem of low visual accuracy. Finally, in this embodiment, the final annotation image C10 is obtained by mapping C9 to the original RGB image C1. The complete process of the entire automatic annotation has been completed.

[0127] In step S2, multi - modal classification is performed on medical images and clinical data information. The steps include:

[0128] S2.1: Extract the visual feature (V) and semantic feature (S) of the input image and text through a visual model and a language model;

[0129] S2.2: Construct the interaction between the visual feature and the semantic feature through an attention mechanism to generate a context vector;

[0130] S2.3: Introduce a random modality transfer function (T) to approximately generate semantic features in the absence of text. Combine the refined visual features with the context vector and input them into a fully - connected neural network for classification; jointly optimize through the least - squares method and cross - entropy loss to ensure that the model maintains high accuracy in asynchronous training and testing;

[0131] Step S2.1 includes:

[0132] It is necessary to find the optimal interaction calculation unit I between the visual feature V and the semantic feature S such that the following formula holds.

[0133]

[0134] In the formula, p is the label probability.

[0135] This method proposes a technique for controlling the coupling degree between visual and semantic information (i.e., avoiding dedicated embeddings), Figure 6The method of the present invention is outlined. In addition to the visual-semantic interaction module, this embodiment also proposes a technique for controlling the coupling degree between visual and semantic information (i.e., avoiding dedicated embeddings), the basic idea of which is to generate "simulated" text encodings when the text is not given to achieve asynchronous training or testing behavior. Therefore, the semantic feature S either comes from the text encoding S when the text is available r (i.e., w / text), or from the simulated test encoding S when the text is not available f (i.e., w / o text).

[0136] The step S2.2 includes:

[0137] Two key issues followed by the R-radix interaction module (TandemNet): (1) Using the attention mechanism to refine visual knowledge; (2) Considering the multimodal segmentation relationship when constructing the interaction between images and texts.

[0138] Attention mechanism. The basic attention unit is inspired by the recent large-scale dot-product attention for sequence transduction. The input of this module consists of a query vector Q, a set of vectors called keys in matrix K, and a set of vectors called values in matrix Q. The definition of the attention module is shown in the following formula.

[0139]

[0140] where, d q is the dimension of the query Q. The output is a context vector, which is the weighted average of Q The function outputs a weight vector, which determines the importance of each value in Q.

[0141] The text feature S includes N hidden states Finally, it represents N individual concepts. Finally, it provides two functions: (1) It is used to query the visual feature V; (2) It is the target text embedding, which is simulated by the proposed channel transfer function.

[0142] The basic visual-semantic interaction module of this method is defined as shown in the following formula.

[0143] V = P(S T W q , V T W K , V T W Q ) * W i

[0144] The input before the attention module is first processed by the learnable matrix W q ∈R D*D′ , W k , WQ ∈R C*D′ and W v ∈R D′*C Input. Output V ∈ R N*C is regarded as R C a set of context vectors in. The embedding dimension inside the D' unit is configured to be less than D and C, aiming to reduce the computational amount. W v maps the output back to the original dimension. This embodiment also applies a batch normalization layer after V. In addition, the cardinality of this module can be increased by increasing the number of P in parallel, so that the definition of the joint context vector is as shown in the following formula.

[0145]

[0146] where, v ∈ R C is the average context vector containing the extracted visual features. This embodiment also applies a small spatial dropout to V before the input of each module, and its goal is to reduce the correlation from multiple interaction modules v i s. R is an empirically selected cardinality factor.

[0147] Prediction module. The prediction model is defined as concatenates the visual feature V with v to form a two-layer fully connected neural network with LeakyReLU as the activation function.

[0148] The step S2.3 includes:

[0149] To support asynchronous training and testing behaviors, make the model indistinguishable without text and make equally accurate image classifications, this method proposes a modality transfer function to explicitly learn the visual-to-semantic approximation in order to switch from S r to S f . Given the global (average) visual feature Δ(V) of the image, T approximates a simulated text encoding that can implicitly describe the image content. The definition of the transfer function is as shown in the following formula.

[0150] h T = T(Δ(V); θ T )

[0151] where h T ∈R D*1 is used to represent S f ∈R D*1 (note that the number of text encodings in S r is N, while in S f it is 1). T is implemented as a two-layer fully connected (FC) network (defined as FC → LeakyReLU → FC → Tanh) with trainable parameters θ T, which is trained using the mean least squares loss, and the formula is as follows.

[0152] min‖Δ(S r ) - T(Δ(V); θ T )‖2

[0153] This loss only penalizes the image model to be close to the semantic features. When jointly trained with the image classification (cross-entropy) loss of the complete model, this embodiment encourages the CNN to generate discriminative and semantic-preserving representations.

[0154] In the step S3, the interpretable multi-modal deep learning-based fusion model includes the following steps:

[0155] S3.1: Collect and preprocess multi-modal data, including feature vectors and esophageal images;

[0156] S3.2: gCAM-CCL uses one-dimensional convolution to learn features from the feature data and two-dimensional convolution to learn features from the imaging data. The outputs of the two convolutional neural networks are flattened and then fused in the loss function in the collaboration layer;

[0157] S3.3: Select two intermediate layers from which the feature maps can be combined using gradient-based weights and the Grad-CAM activation maps for specific classes are generated accordingly;

[0158] S3.4: At the same time, the fine-grained activation maps are calculated by projecting the gradients from the collaboration layer back to the input layer using guided BP, and the obtained activation maps indicate the contribution of the pixels to the decision of interest;

[0159] Currently, a large number of model interpretation methods have been proposed, aiming to interpret the decisions of deep neural networks by providing human-understandable interpretations. However, most methods cannot be fully integrated into medical image processing. This method intends to establish a new interpretable method for classification and result interpretation. The Grad-CAM-guided convolutional collaboration learning (gCAM-CCL) model generates activation maps to show the pixel contributions of the input images and genetic vectors. Specifically, it calculates the gradients of each feature map, combines the gradients using global average pooling to combine the feature maps. In addition, the activation maps are specific to classes, which further facilitates the analysis of class differences and the discovery of potential biological mechanisms.

[0160] Compared with the DCL model, gCAM-CCL adopts a new structure and a new loss function to incorporate Grad-CAM. Since calculating the activation maps requires one layer of feature maps, gCAM-CCL replaces the multi-layer perceptron (MLP) network with two convolutional neural networks, so that multi-channel feature maps can be obtained, which is also beneficial to model training because the convolutional neural network greatly reduces the number of parameters by forcing the sharing of kernel weights.

[0161] In addition, both DCCA and DCL include the parameter of the sample size in their loss functions, resulting in an issue of batch size adjustment. In other words, due to the existence of the relevant term U1′Z1′Z2U2, the loss function depends on the batch size, so a large batch size is required in network training. In this work, this embodiment proposes a new loss function to alleviate the problem of batch size dependence, as expressed in the formula. The formula is as follows, and the population-level relevant term is replaced by the sum of the sample-level losses.

[0162] In addition, the relevant term is replaced by the regression loss, that is, the cross-entropy loss, because it has been proven that the optimization of the relevant term is equivalent to the optimization of the regression loss.

[0163]

[0164] where, are the outputs of two convolutional neural networks, as Figure 7 shown.

[0165] This batch-independent loss function is easier to extend to multi-class and multi-view scenarios, and the extended loss function is as follows.

[0166]

[0167] where m represents the number of views and C represents the number of classes.

[0168] gCAM-CCL uses one-dimensional convolution to learn features from the feature data and two-dimensional convolution to learn features from the imaging data. The outputs of the two convolutional neural networks are flattened and then fused in the collaboration layer with the loss function in the formula, which takes into account both cross-modal associations and their fit with the phenotype / label y. Then, two intermediate layers are selected from which the feature maps can be combined using gradient-based weights, and the class-specific Grad-CAM activation maps are generated accordingly. At the same time, the fine-grained activation maps are calculated by projecting the gradients from the collaboration layer back to the input layer using guided BP, and the obtained activation maps indicate the contribution of pixels to the decision of interest, such as predicting important biomarkers.

[0169] In addition, the ideal class-specific activation map should only emphasize the features related to the corresponding class, but the features related to other classes may have a negative contribution to the prediction, resulting in noise features in the activation map. To remove the noise or irrelevant features, this embodiment applies the ReLU function to the gradients, as shown in the formula. The ReLU function ensures the positive effect, so the pixels with negative contributions can be filtered out.

[0170]

[0171] Among them, it represents the predicted score for class C.

[0172] In step S4, for the weakly supervised contrastive learning for esophageal tumor detection report generation, the steps include:

[0173] S4.1: Extract visual features from esophageal tumor detection images, and then use a memory-driven transformer model as the backbone model to generate report text;

[0174] S4.2: On this basis, use a fine-tuned BERT model to embed and cluster the report to identify semantic similarities;

[0175] S4.3: Maximize the similarity between the image and the report through weakly supervised contrastive learning, while minimizing the similarity between negative samples;

[0176] S4.4: Improve the accuracy and diversity of the model when generating long text descriptions by mixing and optimizing the cross-entropy loss and the contrastive loss;

[0177] Medical report generation aims to automatically generate descriptions of esophageal tumor spectral CT images, which has attracted the attention of the machine learning and medical communities. Many methods have been proposed to solve this problem. Recently, learning visual semantic embeddings has been proposed for cross-modal retrieval in a contrastive environment and has achieved improvements in identifying abnormal results. However, this is difficult to scale or generalize because this embodiment requires constructing template abnormal statements for new datasets. And using a memory-augmented transformer model to improve the ability to generate long and coherent text, but does not specifically solve the problem of generating explicit normal results. The work of this embodiment proposes incorporating contrastive learning into the training of a generation-based model, which benefits from the contrastive loss to encourage diversity and is easy to scale compared to retrieval-based methods.

[0178] This embodiment uses the recently proposed memory-driven transformer as the backbone model and uses a memory module to record key information when generating long text. Given an esophageal tumor detection image I, its visual feature X is extracted by a pre-trained convolutional neural network. Then this embodiment uses a standard encoder in the transformer to obtain the hidden visual feature H X . The decoding process at each time step t can be formalized as the following formula.

[0179]

[0180] This embodiment uses the cross-entropy (CE) loss to maximize the conditional log-likelihood value log P θ (Y|X), for a given set of N observations (X (i) , Y (i) i=1 )N , the formula is as follows.

[0181]

[0182] As Figure 8 shown, in this embodiment, the embedded content of each report is first extracted from ChexBERT, which is a BERT model pre-trained with biomedical text and fine-tuned for esophageal tumor report tagging. This embodiment uses the [CLS] embedding of BERT to represent the report-level features. Then, this embodiment applies K-Means to divide the reports into K groups. After clustering, each report Y is assigned a corresponding cluster label l, and reports in the same cluster are considered to be semantically close to each other.

[0183] To standardize the training process, this embodiment proposes a weakly supervised contrastive loss (WCL). This embodiment first projects the hidden representations of the image and the target sequence into the latent space, and the formula is as follows.

[0184]

[0185] In the above formula, H X and H Y are the average pools of the hidden states H X and H Y from the transformer, and φ x and φ y are two fully connected layers with ReLU activation. Then, this embodiment maximizes the similarity between the source image pair and the target sequence while minimizing the similarity between the negative image pairs, and the formula is as follows:

[0186]

[0187] In the above formula, sim is the cosine similarity between two vectors, τ is the temperature parameter, and α is a hyperparameter used to measure the importance of negative samples that are semantically close to the target sequence, that is, the negative samples with the same cluster label l i = l j in the formula. Empirically, this embodiment finds that these samples are "hard" negative samples, and by assigning more weights to distinguish these samples, the performance of the model will be better.

[0188] Generally speaking, the model is optimized by a mixture of cross-entropy loss and weakly supervised contrastive loss, and the formula is as follows.

[0189] L loss = (1 - λ)L CE + λL WCL

[0190] Among them, λ is a hyperparameter that measures two losses.

[0191] In the step S5, regarding the improvement of the scalability and adaptability of the system, the steps include:

[0192] Based on design concepts such as domain-oriented design, hexagonal architecture, separation of query and command, etc., an improved program architecture that meets the integration technology requirements of microservices, distributed systems, and Docker is implemented. The application programs are independent in a modular way. Modules such as infrastructure, RESTFull HTTP interfaces, MQ consumption interfaces, and RPC interfaces are used to implement the interaction between the extended application and external systems as needed through various adapters, with high scalability and adaptability.

[0193] The method includes: automatic annotation of esophageal tumor images based on deep learning, multimodal classification based on medical images and clinical detection data, an interpretable fusion model based on multimodal deep learning, and weakly supervised contrast learning technology for generating esophageal tumor detection reports.

[0194] The system developed by this method mainly collects data by docking with the hospital's PACS or directly with the device. The data is pushed to the algorithm server on the PACS system or the device, and the algorithm runs. The doctor client accesses the server through the hospital's internal local area network to view the results of the algorithm operation. The physical topology diagram involved is as Figure 9 shown:

[0195] 1), Hospital PACS or device: DICOM image server, responsible for pushing image data that needs auxiliary diagnosis.

[0196] 2), Energy spectrum CT algorithm application server: This server is placed locally in the medical institution, and the doctor workstation can access it through the hospital's internal local area network.

[0197] Figure 10 It is a system structure diagram, where

[0198] 1), Front end: The front part of the website, a web page that runs on browsers such as the PC side and is presented to users for browsing.

[0199] 2), Microservices: That is, the microservices architecture, a method of developing a single application as a set of small services. Each application runs in its own process and communicates with lightweight mechanisms (usually HTTP resource APIs).

[0200] 3), User service: Responsible for implementing functions such as account login and logout.

[0201] 4), Image service: Responsible for implementing functions such as image reception, image parsing, and image storage in the database.

[0202] 5) Esophageal lesion automatic detection service: Responsible for realizing functions such as lesion detection, one-key reporting, delineation and annotation.

[0203] 6) PACS service: Responsible for functions such as image list, query on the film viewing page, etc.

[0204] The experimental results of the method of the embodiment of the present invention will be described in detail below in conjunction with the accompanying drawings of the specification:

[0205] As Figures 11 - 17 shown, specifically including:

[0206] 1) Preliminary research on semantic segmentation of esophagus and esophageal cancer;

[0207] Preliminary research on preoperative regional lymph nodes of esophageal cancer by machine learning (U-Net Plus): In this embodiment, U-Net Plus is used to segment the esophagus and esophageal cancer from two-dimensional CT slices. In the new network architecture, two blocks are adopted to enhance the feature extraction performance of complex abstract information, which can effectively solve irregular and fuzzy boundaries. The block is a skip connection operation similar to convolution. The architecture was trained on a dataset of 1924 slices from 10 CT scans and tested on 295 slices from 6 CT scans. The training and test datasets were expanded tenfold to simulate the segmentation of 3-D CT images. Using the new framework, this embodiment reports a random value of 0.79±0.20 and a Hausdorff distance of 5.87±9.91. Then a semi-automatic scheme is designed for 3-D segmentation of the esophagus or esophageal cancer. Implementing 3D rendering of the esophagus or esophageal cancer can help diagnose esophageal cancer.

[0208] Table 1 Comparison between U-Net plus and several popular networks

[0209] U-Net Plus U-Net LinkNet SegNet DV 0.79±0.20 0.72±0.19 0.66±0.29 0.61±0.24 HD (mm) 5.87±9.91 7.04±7.45 11.12±19.74 14.45±27.00

[0210] Figure 11 For visualization of the segmentation results of esophageal cancer. Among them: (a) Results of U-Net Plus; (b) Results of U-Net; (c) Results of LinkNet; (d) Results of SegNet. Red: Overlapping area (true positive) between the delineated area and the segmented area; Blue: Areas in the delineated area that are not in the segmented area (false negative); Green: Areas in the segmented area that are not in the delineated area (false positive).

[0211] This embodiment proposes a new U-Net—Channel-attention U-Net for segmenting the esophagus and esophageal cancer from tomographic images. This network combines a channel attention module and a cross-layer feature fusion module (CFFM). The former differentiates the esophagus from surrounding tissues by emphasizing and suppressing channel features, and the latter enhances the network's generalization ability by using high-level features to weight low-level features. Since high-level features represent specific tissue information and low-level features represent features such as edges and contours, the network can learn specific detailed features of specific tissues. In addition, to better locate the esophageal region, a three-dimensional semi-automatic esophageal and esophageal cancer segmentation method is proposed. The proposed network is trained using 46,400 CT images as the training set and 11,600 CT images are segmented from the dataset at a ratio of 0.2 as the validation set. Finally, 7,250 CT images are used as the test set to test the performance of the network. The experimental results show that the IoU value of this network can reach 0.625, the Dice value is 0.732, and the Hausdorff distance is 3.193.

[0212] Among them, the flowchart of Channel-attention U-Net is as Figure 12 shown.

[0213] This model is based on U-Net and is widely used in medical image segmentation tasks and pixel-level classification due to its superior performance. On this basis, this embodiment embeds a cross-level feature fusion module (CFFM) to form the Channel-attention U-Net of this embodiment. Channel-attention U-Net is an end-to-end network architecture composed of two main parts. The first component is the cross-layer feature fusion module (CFFM), which consists of several channel attention modules (CAM). CFFM can fuse the features of the top-level feature map layer by layer into the bottom-level feature map and use computer-aided manufacturing to achieve the purpose of guiding the feature selection of the bottom-level feature map. The second component is U-Net, which is an encoder-decoder structure. The first half is used for feature extraction, and the second half is upsampled. In particular, there is a skip connection structure between the encoder and the decoder. The skip connection links the corresponding downsampled and upsampled feature maps. This process can solve the problem of information loss caused by downsampling and the relationship with surrounding tissues, and improve the attention to convolutional feature extraction. The segmentation effect is as Figure 13 shown, and the experimental result indicators are shown in the following table.

[0214] Table 2 Experimental Result Indicators

[0215]

[0216] As can be seen in this embodiment, the proposed method has achieved better results than the traditional method. From Table 2, the method of this embodiment performs better than several variants of U-Net, FCN, and SegNet. Channel-attention U-Net achieved the best segmentation performance with the highest IoU (0.625), DV (0.732), and the lowest HD (3.193 mm).

[0217] To solve the problem that two-dimensional tomography of esophageal cancer lacks three-dimensional structural information, a new 2.5D segmentation network, called Eso-Net, based on an encoder-decoder architecture for cancerous esophagus was proposed. A 3D enhancement filter called multi-structure response filter (MSRF) was designed to extract 3D structural information as prior knowledge. In the case of 3D structural prior, a priority attention module (PAM) was incorporated into the network to facilitate the transmission of relevant spatial information. The experiment was conducted on a dataset of 30 esophageal cancer patients. Finally, this embodiment obtained a Dice similarity coefficient of 84.839%, a precision of 85.955%, a sensitivity of 83.752%, and a Hausdorff distance of 2.583, which indicates that the proposed method is superior to other existing segmentation networks in this task. Therefore, this method can be extended to this topic and solve the problem that two-dimensional spectral CT multimodal images lack three-dimensional structural information.

[0218] Among them, the flowchart of Eso-Net is as Figure 14 shown.

[0219] In the first stage, the input CT slices are preprocessed. First, the window level (WL) and window width (WW) of the CT slices are set to 40 and 200 respectively to increase the contrast between the esophagus and other tissues and organs. The intensity values between 160 and 240 are linearly normalized to [0, 1] to accelerate model convergence during training. Intensities less than 160 are set to 0, and intensities greater than 240 are set to 1. Then, the multi-structural response filter (MSRF) is used to enhance tissues and organs with specific geometric structures at multiple scales, which can provide additional prior knowledge for the segmentation network. Finally, the pixel regions at the centers of the original image and the enhanced image are cropped to obtain a region of interest (ROI) large enough to contain the entire esophagus. In addition, the cropping operation also reduces the interference of irrelevant tissues and organs and speeds up model training and inference. In the second stage, channel-wise 2.5D segmentation is performed by Eso-Net. The CT image to be segmented is concatenated with two adjacent images as one of the network inputs, which helps to effectively utilize the z-axis information without significantly increasing the number of parameters. If the image to be segmented is the first or the last image of a CT scan, the unavailable adjacent images are replaced by the image to be segmented itself. In addition, another input is the corresponding enhanced image concatenated in the same way. Finally, the network outputs the segmentation map of the CT image in the middle channel. The segmentation effect is as Figure 15 shown.

[0220] In Figure 15 , the size of the CT image is 144×144. True positives, false positives, and false negatives are represented by red, green, and blue respectively. Each row shows the same CT image, and it can be observed in this embodiment that the size and shape of the esophagus are different. Some segmentation networks, such as FCN-8s, SegNet, LinkNet, and U-Net plus, cannot obtain satisfactory segmentation results. For these networks, some similar tissues and organs are mispredicted as the cancerous esophagus. Since the esophageal cancer region is a key diagnostic basis in CT images, inaccurate segmentation results will bring serious consequences. In contrast, this method can obtain high-precision segmentation results. The last row in the figure shows a challenging segmentation case, where the esophagus with a large tumor has a blurred boundary and low contrast with the surrounding tissues. In this case, the proposed method achieves better segmentation performance than other networks. These visualization examples show that the proposed method can provide reliable segmentation results for the diagnosis and treatment of esophageal cancer. The experimental result metrics are shown in the following table:

[0221] Table 3 Experimental Result Metrics

[0222] Method DSC (%) PRE (%) SEN (%) FCN-8s[8] 75.703 74.479 76.968 U-Net

[11] 78.432 80.330 76.622 SegNet

[12] 72.167 66.181 79.344 LinkNet

[13] 66.073 68.182 64.091 3D U-Net

[14] 83.693 84.274 83.120 U-Net plus

[15] 61.033 62.713 59.441 U-Net++

[30] 79.025 74.992 83.516 Attention U-Net

[34] 82.425 82.494 82.356 Ours 84.839 85.955 83.752

[0223] From the table, LinkNet and SegNet have poor performance in terms of DSC and HD. One of the reasons is that LinkNet and SegNet were proposed for scene segmentation and are not suitable for segmenting objects with blurred boundaries. In addition, except for U-Net plus, U-Net and its variants achieved better segmentation results. The improvement in performance is attributed to the correct use of skip connections, which helps to recover spatial information. Among them, 3D U-Net achieved the best performance with a DSC of 83.693%, PRE of 84.274%, SEN of 83.120%, and HD of 2.591 mm, demonstrating the superiority of 3D convolutional networks in medical image segmentation. Two 2D convolutional networks also achieved competitive performance. U-Net++ integrated U-Nets with different depths and obtained a DSC of 79.025%, PRE of 74.992%, SEN of 83.516%, and HD of 2.875 mm. The attention U-Net that uses a spatial attention mechanism in the skip connection obtained a higher DSC (82.425%), PRE (82.494%), and a lower HD (2.751 mm). However, neither of them exceeded the 3D network. The experimental results show that none of the existing segmentation networks can achieve the best performance in all evaluation metrics because they do not take into account the particularity of esophageal carcinogenesis. To complete this task, the Eso-Net of this embodiment performs channel-based 2.5D segmentation and uses dilated convolution and residual connections for multi-scale feature extraction and fusion. In addition, PAMs are designed to enhance the effectiveness of skip connections. From the above figure, it obtained the highest DSC (84.839%), PRE (85.955%), SEN (83.752%), and the lowest HD (2.583 mm), indicating that it is superior to other existing segmentation networks in this task.

[0224] 2) Preliminary study on differential diagnosis of preoperative regional lymph node metastasis of esophageal cancer by multi-modal quantitative analysis of spectral CT

[0225] In the preliminary study of the method of this embodiment, this embodiment collected 58 patients with esophageal cancer who were confirmed by esophagoscopy or gastroscopy and were scheduled for surgical treatment. None of them received radiotherapy or chemotherapy before surgery. Spectral CT examinations were performed within 2 weeks before surgery. The spectral CT images before surgery were compared, and the lymph nodes were accurately located and marked during the operation, and postoperative pathological evaluation was carried out. The results showed that iodine-based material density, standardized iodine-based material density, spectral curve slope, and standardized spectral curve slope were statistically significant for the differential diagnosis of benign and malignant preoperative regional lymph nodes of esophageal cancer. The results are as follows:

[0226] As Figure 16As shown, the iodine-based material density of the lymph nodes in the main pulmonary artery window is 2.2 mg / ml, and the effective atomic number is 8.77. The iodine-based material density of the thoracic aorta at the same level is 5.7 mg / ml, and the effective atomic number is 10.05.

[0227] As shown in Table 4, the iodine-based material density, standardized iodine-based material density, energy spectrum curve slope, and standardized energy spectrum curve slope are all statistically significant for the preoperative differential diagnosis of benign and malignant regional lymph nodes in esophageal cancer, with p < 0.05.

[0228] Table 4 Iodine-based material density, standardized iodine-based material density, energy spectrum curve slope, standardized energy spectrum curve slope

[0229]

[0230]

[0231] As Figure 17 shown, the ROC curve analysis of the combined diagnostic results indicates that when the cut-off point is 0.558, the maximum AUC is 0.943. At this time, the sensitivity of spectral CT in differentiating preoperative regional lymph node metastasis in esophageal cancer is 0.882, the specificity is 0.932, the Youden index is 0.814, and the accuracy is 0.905. The results of this study won an oral presentation at the annual meeting of the Radiological Society of North America (RSNA) held in Chicago in December 2019.

[0232] In summary, this embodiment proposes a method for semantic segmentation of the esophagus and esophageal cancer and multi-modal quantitative analysis of spectral CT to differentiate preoperative regional lymph node metastasis in esophageal cancer. By demonstrating the research process and research effects, it explains how the network improves the segmentation performance through cross-level feature fusion and channel attention mechanism; demonstrates the advantages of the network in segmentation effect, and verifies its high-precision segmentation results through visualization examples; demonstrates the subsequent spectral CT images for analysis and pathological evaluation results. In the previous research on semantic segmentation of the esophagus and esophageal cancer, this method uses the U-Net Plus network to segment the esophagus and esophageal cancer from two-dimensional CT tomograms, and enhances the segmentation performance by introducing a new network architecture and feature extraction blocks. Then, a semi-automatic scheme is designed for three-dimensional segmentation of the esophagus or esophageal cancer, and the effectiveness of this scheme is verified through experiments. In addition, the Channel-attention U-Net and Eso-Net networks are respectively used to segment the esophagus and esophageal cancer from tomographic images, and to solve the problem of lack of three-dimensional structural information in two-dimensional tomograms. These networks improve the accuracy and robustness of segmentation by combining techniques such as channel attention modules, cross-layer feature fusion modules, and multi-structure response filters.

[0233] In the study of differentiating preoperative regional lymph node metastasis of esophageal cancer by multi-modal quantitative analysis of spectral CT, by collecting spectral CT images of esophageal cancer patients and postoperative pathological evaluation results, through comparative analysis, it was found that indicators such as iodine-based material density, standardized iodine-based material density, spectral curve slope, and standardized spectral curve slope were statistically significant for differentiating benign and malignant preoperative regional lymph nodes of esophageal cancer. In addition, through ROC curve analysis, the optimal cut-off point for combined diagnosis was determined, and indicators such as sensitivity, specificity, Youden index, and accuracy were calculated to verify the effectiveness of this method.

[0234] The research in the above examples provides an important foundation and guidance for subsequent methods. In the previous research on semantic segmentation of the esophagus and esophageal cancer, researchers successfully achieved accurate segmentation from two-dimensional CT tomograms to three-dimensional esophagus and esophageal cancer through innovations in network architectures such as U-Net Plus, Channel-attention U-Net, and Eso-Net. These studies not only improved the accuracy and robustness of segmentation but also provided key image processing and segmentation technologies for subsequent multi-modal tumor intelligent detection systems. Especially in the research of Eso-Net, by introducing technologies such as multi-structure response filters and priority attention modules, the problem of lack of three-dimensional structure information in two-dimensional tomograms of esophageal cancer was effectively solved, providing richer image features for subsequent multi-modal analysis.

[0235] In the previous research on differentiating preoperative regional lymph node metastasis of esophageal cancer by multi-modal quantitative analysis of spectral CT, researchers collected and analyzed spectral CT images of esophageal cancer patients and postoperative pathological evaluation results, and found that indicators such as iodine-based material density, standardized iodine-based material density, spectral curve slope, and standardized spectral curve slope were of great significance for differentiating benign and malignant preoperative regional lymph nodes of esophageal cancer. These research results provide important clinical indicators and basis for feature selection for subsequent multi-modal tumor intelligent detection systems.

[0236] Based on the foundation and guidance of the previous research, the subsequent method further integrates technologies such as deep learning, multi-modal classification, interpretable models, and weakly supervised contrast learning to construct a method for an interpretable multi-modal tumor intelligent detection system and report generation. This system not only realizes automatic annotation and accurate segmentation of esophageal tumors but also improves the accuracy and credibility of diagnosis through multi-modal classification and interpretable models. At the same time, the introduction of weakly supervised contrast learning also provides new ideas and methods for the optimization and performance improvement of the system. Finally, by constructing an intelligent detection cloud platform software, efficient and accurate automatic detection of esophageal tumors is achieved, providing platform support for conducting multi-center verification and system optimization.

[0237] Another aspect of the embodiments of the present invention also provides an enhanced interpretable multi-modal esophageal cancer target detection system, including:

[0238] The first module is used to collect medical images of esophageal tumors detected by spectral CT and perform annotation processing on the obtained CT images;

[0239] The second module is used to perform multi-modal classification on the medical images and clinical data information according to the annotated CT images;

[0240] The third module is used to construct a fusion model based on multi-modal deep learning to fuse the classified multi-modal features;

[0241] The fourth module is used to generate an esophageal tumor detection report based on weakly supervised contrast learning according to the fused features obtained by fusion;

[0242] The fifth module is used to construct an intelligent detection cloud platform with the optimized AI diagnosis model as the core; the intelligent detection cloud platform is used for the transmission, storage management, image data preprocessing, automatic detection and query statistics of spectral CT images.

[0243] It can be understood that the content in the above method embodiments is applicable to the system embodiments of the present invention. The functions specifically implemented by the system embodiments of the present invention are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.

[0244] The embodiments of the present invention also provide an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the above enhanced interpretable multi-modal esophageal cancer target detection method is implemented. The electronic device can be any intelligent terminal including a tablet computer, an in-vehicle computer, etc.

[0245] It can be understood that the content in the above method embodiments is applicable to the device embodiments of the present invention. The functions specifically implemented by the device embodiments of the present invention are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.

[0246] The embodiments of the present invention also provide a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, the above enhanced interpretable multi-modal esophageal cancer target detection method is implemented.

[0247] It can be understood that the content in the above method embodiments is applicable to the storage medium embodiments of the present invention. The functions specifically implemented by the storage medium embodiments of the present invention are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.

[0248] The memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include memories remotely located relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0249] It should be noted that in each specific embodiment of the present invention, when it comes to relevant processing based on data related to the user's identity or characteristics, such as user information, user behavior data, user historical data, and user location information, the user's permission or consent will be obtained first. Moreover, the collection, use, and processing of these data will comply with relevant laws, regulations, and standards. In addition, when the embodiments of the present invention need to obtain sensitive personal information of the user, the user's separate permission or separate consent will be obtained through methods such as pop-up windows or redirecting to a confirmation page. After clearly obtaining the user's separate permission or separate consent, the necessary user-related data for the normal operation of the embodiments of the present invention will be obtained.

[0250] The embodiments described in the embodiments of the present invention are to more clearly illustrate the technical solutions of the embodiments of the present invention, and do not constitute a limitation on the technical solutions provided by the embodiments of the present invention. Those skilled in the art know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present invention are equally applicable to similar technical problems.

Claims

1. An enhanced interpretable multi-modal esophageal cancer object detection method, characterized in that, Including the following steps: Collect medical images of esophageal tumors detected by spectral CT, and perform annotation processing on the obtained CT images; Based on the annotated CT images, perform multi-modal classification on medical images and clinical data information; Construct a fusion model based on multi-modal deep learning to fuse the multi-modal features obtained by classification; Based on the fused features obtained by fusion, generate an esophageal tumor detection report based on weakly supervised contrast learning; Taking the optimized AI diagnosis model as the core, construct an intelligent detection cloud platform; the intelligent detection cloud platform is used for the transmission, storage management, image data preprocessing, automatic detection and query statistics of spectral CT images.

2. An enhanced interpretable multi-modal esophageal cancer target detection method according to claim 1, characterized in that, The step of collecting medical images of esophageal tumors detected by spectral CT and performing annotation processing on the obtained CT images includes the following steps: Delete the areas that do not contribute to the annotation, and adjust the remaining area of the spectral CT image to 384px×384px by bilinear interpolation to complete the preprocessing of the image; Use HRDEN to establish a depth map of the spectral CT image, where HRDEN consists of 4 modules: an encoder, a decoder, a multi-scale feature fusion module, and a refinement module; Use the DeepLabV3+ model to predict the EEC area, support 4-channel input including the depth channel C5 by modifying the first-layer convolution kernel, and at the same time maintain the feature extraction performance of the deep CNN module; use the ASPP module to fuse multi-scale features, simplify the low-level features in the decoding stage and concatenate them with the high-level features, and process the fused features to output a multi-scale information map; Optimize the prediction result through post-processing, fill the small holes and apply morphological operations until the binary image C8 area is stable as C9 and the edge is smooth; then map C9 to the original RGB image C1 to generate the final annotation image C10, thus completing the clinically friendly automatic annotation process.

3. An enhanced interpretable multi-modal esophageal cancer target detection method according to claim 1, characterized in that, The step of performing multi-modal classification on medical images and clinical data information based on the annotated CT images includes the following steps: Extract the visual features and semantic features of the input image and text through a visual model and a language model; Construct the interaction between the visual features and semantic features through an attention mechanism to generate a context vector; Introduce a random modality transfer function to approximately generate semantic features in the absence of text, combine the refined visual features with the context vector, and input them into a fully connected neural network for classification; optimize jointly through the least squares method and cross-entropy loss.

4. An enhanced interpretable multi-modal esophageal cancer target detection method according to claim 3, characterized in that, The step of extracting the visual features and semantic features of the input image and text through a visual model and a language model includes the following steps: Find the optimal interaction calculation unit I between the visual feature V and the semantic feature S, so that the following formula holds: In the formula, p is the label possibility; Among them, the semantic feature S comes from the text encoding S when the text is available r , or comes from the simulation test encoding S when the text is not available f .

5. An enhanced interpretable multi-modal esophageal cancer target detection method according to claim 3, characterized in that The step of constructing the interaction between the visual features and semantic features through an attention mechanism to generate a context vector includes the following steps: Construct an R-radix interaction module, where the R-radix interaction module includes an attention module, a vision-semantic interaction module, and a prediction module. Use the attention module to refine visual knowledge and consider multi-modal segmentation relationships to build the interaction between image features and text features. Among them, the input of the attention module consists of a query vector Q, a set of vectors called keys in matrix K, and a set of vectors called values in matrix Q. The defined expression of the attention module is: where d q is the dimension of query Q; the output is a context vector, which is the weighted average of Q The function outputs a weight vector, which is used to characterize the importance of each value in Q; The text features include N hidden states Hidden state used to query the visual feature V and serve as the target text embedding, which is simulated by the proposed channel transfer function; The defined expression of the visual-semantic interaction module is: V = P(S T W q ,V T W K ,V T W Q ) * W i , where V represents the output of the visual-semantic interaction module; P represents the parallel multi-head attention module; S T represents the text feature; W q represents the query weight matrix; W K represents the key weight matrix; V T represents the visual feature; W Q represents the value weight matrix; W i represents the output projection matrix; Define an expression for constructing a joint context vector: where v ∈ R C is the average context vector containing the extracted visual features; R is a cardinality factor selected empirically; The defined expression of the prediction module is as follows: The prediction module is used to concatenate the visual feature V with v to form a two-layer fully connected neural network with LeakyReLU as the activation function.

6. An enhanced interpretable multi-modal esophageal cancer target detection method according to claim 3, characterized in that The step of introducing a random modality transfer function to approximately generate semantic features in the absence of text, combining the refined visual features with the context vector, inputting them into a fully connected neural network for classification, and jointly optimizing through the least squares method and cross-entropy loss includes the following steps: Construct a modal transfer function for learning the approximation from vision to semantics. The expression of the modal transfer function is: h T = T(Δ(V); θ T ), where h T ∈ R D*1 is used to represent the simulated test encoding S F ∈ R D*1 , T represents a two-layer fully connected network, and θ T represents the trainable parameters; The training expression of the trainable parameters is: min‖Δ(S r ) - T(Δ(V); θ T )‖2, and the training loss of this training process is used to penalize the image model to approximate the semantic features.

7. An enhanced interpretable multi-modal esophageal cancer target detection method according to claim 1, characterized in that The constructed fusion model based on multimodal deep learning performs fusion processing on the classified multimodal features, including the following steps: Collect and preprocess multimodal data, including feature vectors and esophageal images; Use one-dimensional convolution to learn features from the feature data and two-dimensional convolution to learn features from the imaging data through gCAM-CCL; wherein, the outputs of the two convolutional neural networks are flattened and then fused in the loss function in the collaboration layer; Select two intermediate layers, use gradient-based weights to combine the feature maps therefrom, and correspondingly generate class-specific Grad-CAM activation maps; Calculate fine-grained activation maps by projecting the gradients from the collaboration layer back to the input layer using guided BP, and the obtained activation maps indicate the contribution of pixels to the decision of interest.

8. An enhanced interpretable multi-modal esophageal cancer target detection method according to claim 1, characterized in that, Based on the fused features obtained by fusion, generate an esophageal tumor detection report based on weakly supervised contrast learning, including the following steps: Extract visual features from the esophageal tumor detection images, and then use the memory-driven transformer model as the backbone model to generate report text; Use the fine-tuned BERT model to embed and cluster the reports and identify semantic similarities; Maximize the similarity between the images and the reports through weakly supervised contrast learning while minimizing the similarity between negative samples; Improve the accuracy and diversity of the model when generating long text descriptions by mixing and optimizing the cross-entropy loss and the contrast loss.

9. An enhanced interpretable multimodal esophageal cancer target detection system, characterized in that, Including: The first module is used to collect medical images of esophageal tumors detected by spectral CT and perform annotation processing on the obtained CT images; The second module is used to perform multimodal classification on the medical images and clinical data information according to the annotated CT images; The third module is used to construct a fusion model based on multimodal deep learning to perform fusion processing on the classified multimodal features; The fourth module is used to generate an esophageal tumor detection report based on weakly supervised contrast learning according to the fused features obtained by fusion; The fifth module is used to construct an intelligent detection cloud platform with the optimized AI diagnosis model as the core; the intelligent detection cloud platform is used for the transmission, storage management, image data preprocessing, automatic detection and query statistics of spectral CT images.

10. An electronic device, characterized in that, Including a processor and a memory; The memory is used to store programs; The processor executes the program to implement the method according to any one of claims 1 to 8.

Citation Information

Cited By

  • Asset information intelligent completion method and system fused with multi-modal large model

    CN120747981A

  • Prediction method and device based on multi-modal data, equipment and medium

    CN121122671A

  • Brain tumor imaging diagnosis large model pre-training method, diagnosis method and system

    CN121306517A