A deep learning-based system for early and accurate detection of endoscopic colorectal cancer
Patent Information
- Application Number
- DE202025104839
- Authority / Receiving Office
- DE · DE
- Patent Type
- Utility models
- Current Assignee / Owner
- Filing Date
- 2025-08-19
- Publication Date
- 2025-10-23
- Estimated Expiration
- 2035-08-31
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
AREA OF INVENTION
[0001] The present disclosure relates to a deep learning-based system for the early and accurate detection of endoscopic colorectal cancer. More precisely, the present invention relates to a deep learning-based system for the detection of colorectal cancer using endoscopic images. The system is configured to combine the results of several CNN models to achieve higher accuracy and predictive capability. BACKGROUND OF THE INVENTION
[0002] Colorectal cancer (CRC) represents a serious global health challenge and is among the leading causes of cancer-related deaths worldwide. Traditional diagnostic procedures primarily rely on the histopathological analysis of tissue samples using hematoxylin and eosin (H&E) staining, followed by manual microscopic examination by pathologists. However, this conventional method has inherent limitations, including subjective interpretation, labor-intensive processes, and significant variability among findings. Diagnostic agreement can vary by as much as 25% between pathologists.
[0003] With the introduction of computer-aided diagnostic (CAD) systems, computer-aided solutions were introduced to address these diagnostic challenges. However, existing systems are significantly limited in their ability to accurately detect colorectal cancer from histopathological images due to the complexity and diversity of tissue patterns. While individual convolutional neural network (CNN) models such as ResNet50, DenseNet121, and EfficientNetV2-S have demonstrated effectiveness in medical image analysis, single-model approaches often fail to capture the full spectrum of pathological features required for reliable cancer detection.
[0004] Current deep learning systems for medical image analysis typically require large, labeled datasets for optimal performance, which are difficult to obtain in the medical field due to data privacy restrictions and expert annotation requirements. While transfer learning techniques have partially addressed this challenge by utilizing pre-trained models from large image libraries, the complexity of histopathological patterns in colorectal cancer tissue necessitates more sophisticated approaches.
[0005] Therefore, there is an urgent need in the field for an automated system that combines multiple CNN architectures through ensemble learning techniques to improve diagnostic accuracy, reduce observer variability, and enable consistent, reliable detection of colorectal cancer from endoscopic and histopathological images, while overcoming the limitations of single-model approaches. Summary of the invention
[0006] The present disclosure relates to a deep learning-based system for the early and accurate detection of endoscopic colorectal cancer. More specifically, the present invention provides a deep learning-based system for the early and accurate detection of endoscopic colorectal cancer using ensemble convolutional neural networks with transfer learning. The system processes endoscopic images using multiple CNN models and combines their results using ensemble techniques to achieve improved diagnostic accuracy in the classification of gastrointestinal diseases.
[0007] To provide a deep learning-based system for the early and accurate detection of endoscopic colorectal cancer. The system comprises: an input storage module for storing the endoscopic images to be used as input for the convolutional model for colorectal cancer detection, the images being sourced from a Kvasir dataset containing endoscopic images from eight categories of gastrointestinal diseases; a preprocessing module associated with the input storage module for preprocessing the images, the preprocessing module being configured to resize images, apply data extensions, and normalize the images;A feature extraction module connected to the preprocessing module for receiving the preprocessed endoscopic images and extracting meaningful features from the images using pretrained convolutional neural network (CNN) models, wherein the feature extraction module is configured to use a variety of CNN models; an ensemble module connected to the feature extraction module and configured to combine the outputs of the CNN models implemented by the feature extraction module for feature extraction from preprocessed endoscopic images, wherein the ensemble module is configured to average the predicted probabilities from the CNN models; a classification output module connected to the ensemble module and configured to generate final classification results for endoscopic colorectal cancer diagnosis;and a user interface connected to the classification output module and configured to display the classified results.
[0008] One objective of the present disclosure is to provide a deep learning-based system for the early and accurate detection of colon cancer through endoscopy.
[0009] Another objective of the present disclosure is the development of an automated deep learning system that combines several pre-trained CNN models through ensemble learning techniques to improve the accuracy and reliability of endoscopic colorectal cancer detection compared to single-model approaches.
[0010] Another objective of this disclosure is to implement transfer learning methods that utilize pre-trained weights from ImageNet and optimize multiple CNN architectures on medical image datasets to overcome the limitations of small labeled medical datasets.
[0011] Another objective of the present disclosure is to provide a diagnostic system that can accurately predict colon cancer using endoscopic images.
[0012] To further clarify the advantages and features of the present disclosure, the invention is explained in more detail with reference to specific embodiments illustrated in the accompanying drawings. These drawings merely show typical embodiments of the invention and are therefore not to be understood as limiting its scope. The invention is described and explained more precisely and in greater detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE FIGURES
[0013] These and other features, aspects, and advantages of the present disclosure will be better understood if the following detailed description is read with reference to the accompanying drawings, in which identical symbols consistently represent identical parts. The following applies: Fig. Figure 1 shows a block diagram of a deep learning-based system for the early and accurate detection of endoscopic colorectal cancer according to an embodiment of the present disclosure. Fig. Figure 2 shows a block diagram of the proposed ensemble architecture of the system according to an embodiment of the present disclosure.
[0014] Experts will also recognize that the elements in the drawings are presented for the sake of simplicity and are not necessarily to scale. For example, the flowcharts illustrate the process by highlighting the main steps to enhance understanding of the aspects of this disclosure. Furthermore, with regard to the design of the device, one or more components of the device may be represented in the drawings by conventional symbols, and the drawings may show only the specific details relevant to understanding the embodiments of this disclosure, so as not to clutter the drawings with details that are readily apparent to those skilled in the art after reading this description. DETAILED DESCRIPTION:
[0015] For a better understanding of the inventive principles, reference is made below to the embodiment shown in the drawings, which is described in specific language. However, this does not limit the scope of the invention. Changes and further modifications of the illustrated system, as well as further applications of the inventive principles, are possible, as would normally occur to a person skilled in the art in this field.
[0016] It is clear to the person skilled in the art that the preceding general description and the following detailed description are exemplary and explanatory of the invention and are not intended as a limitation of it.
[0017] References in this specification to “an aspect”, “another aspect”, or similar expressions mean that a particular feature, structure, or property described in connection with the embodiment is included in at least one embodiment of the present disclosure. Therefore, occurrences of the expressions “in one embodiment”, “in another embodiment”, and similar expressions in this specification may all refer to the same embodiment, but need not.
[0018] The terms "includes," "include," or other variations thereof are intended to cover non-exclusive inclusion, such that a process or method that includes a list of steps may not only contain those steps but may also include other steps not expressly listed or inherent in such process or method. Likewise, the statement "includes..." in the case of one or more devices, subsystems, elements, structures, or components does not, without further limitations, preclude the existence of other devices, subsystems, elements, structures, components, or additional devices, subsystems, elements, structures, or components.
[0019] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as understood by a person skilled in the art in the field of the invention. The system, methods, and examples provided here serve only for illustration and are not to be construed as a limitation.
[0020] Embodiments of the present disclosure are described in detail below with reference to the accompanying drawings.
[0021] The functional units described in this specification are referred to as devices. A device may be implemented in programmable hardware such as processors, digital signal processors, central processing units, field-programmable gate arrays, programmable array logic systems, programmable logic devices, cloud processing systems, or the like. Devices may also be implemented in software for execution by various processor types. An identified device may contain executable code and consist, for example, of one or more physical or logical blocks of computer instructions, which may be organized, for example, as an object, procedure, function, or other construct.However, the executable file of an identified device does not need to be physically stored in the same location, but can consist of different commands stored in different locations which, logically linked together, form the device and fulfill its purpose.
[0022] The executable code of a device or module can consist of one or more instructions and may even be distributed across multiple code segments, different applications, and multiple storage devices. Similarly, operational data within the device can be identified and represented, and organized in any form and data structure. This operational data can be captured as a single data record or distributed across different locations, including various storage devices, and may exist, at least partially, as electronic signals within a system or network.
[0023] References in this description to “a selected embodiment”, “an embodiment”, or “an embodiment” mean that a particular feature, structure, or property described in connection with the embodiment is included in at least one embodiment of the disclosed subject matter. Therefore, the expressions “a selected embodiment”, “in an embodiment”, or “in an embodiment” appearing at different points in this description do not necessarily refer to the same embodiment.
[0024] Furthermore, the described features, structures, or properties can be combined in any way in one or more embodiments. The following description contains numerous specific details to enable a comprehensive understanding of the embodiments of the disclosed subject matter. However, those skilled in the art will recognize that the disclosed subject matter can also be implemented without these specific details or with other methods, components, materials, etc. In other cases, known structures, materials, or processes are not presented or described in detail so as not to obscure aspects of the disclosed subject matter.
[0025] According to the exemplary embodiments, the disclosed computer programs or modules can be executed in a variety of ways, for example, as an application in the memory of a device or as a hosted application running on a server and communicating with the device application or browser via various standard protocols such as TCP / IP, HTTP, XML, SOAP, REST, JSON, and other suitable protocols. The disclosed computer programs can be written in exemplary programming languages that are executed from the device's memory or from a hosted server, such as BASIC, COBOL, C, C++, Java, Pascal, or scripting languages such as JavaScript, Python, Ruby, PHP, Perl, or other suitable programming languages.
[0026] Some of the disclosed embodiments involve or otherwise involve data transmission over a network, for example, the transmission of various inputs or files over the network. The network may include, for example, the Internet, wide area networks (WANs), local area networks (LANs), analog or digital wired and wireless telephone networks (e.g., PSTN, Integrated Services Digital Network (ISDN), mobile networks, and Digital Subscriber Line (xDSL)), radio, television, cable, satellite, and / or other transmission or tunneling mechanisms for data transmission. The network may comprise multiple networks or subnetworks, each containing, for example, a wired or wireless data path. The network may include a circuit-switched voice network, a packet-switched data network, or another network for transmitting electronic communications.For example, the network can include networks based on the Internet Protocol (IP) or the Asynchronous Transmission Mode (ATM) and supporting voice communication via VoIP, Voice over ATM, or other comparable protocols. In one implementation, the network includes a cellular network configured for exchanging text or SMS messages.
[0027] Examples of networks include a Personal Area Network (PAN), a Storage Area Network (SAN), a Home Area Network (HAN), a Campus Area Network (CAN), a Local Area Network (LAN), a Wide Area Network (WAN), a Metropolitan Area Network (MAN), a Virtual Private Network (VPN), an Enterprise Private Network (EPN), the Internet, a Global Area Network (GAN), etc.
[0028] Fig. Figure 1 shows a block diagram of a deep learning-based system (100) for the early and accurate detection of endoscopic colorectal cancer according to an embodiment of the present disclosure.
[0029] Referring to Fig. 1 The system (100) comprises: an input storage module (102) configured to store the endoscopic images to be used as input for the convolutional model for colorectal cancer detection, wherein the images are taken from a Kvasir dataset containing endoscopic images covering eight categories of gastrointestinal diseases; a preprocessing module (104) connected to the input storage module (102) and configured to preprocess the images, wherein the preprocessing module (104) is configured to resize the images, apply data expansion, and normalize the images;a feature extraction module (106) connected to the preprocessing module (104) and configured to receive the preprocessed endoscopic images and extract meaningful features from the images using pretrained convolutional neural network (CNN) models, wherein the feature extraction module (106) is configured to use a variety of CNN models; an ensemble module (108) connected to the feature extraction module (106) and configured to combine the outputs of the CNN models implemented by the feature extraction module (106) for feature extraction from preprocessed endoscopic images, wherein the ensemble module (108) is configured to average the predicted probabilities from the CNN models;a classification output module (110) connected to the ensemble module (108) and configured to generate final classification results for endoscopic colorectal cancer diagnosis; and a user interface (112) connected to the classification output module (110) and configured to display the classified results.
[0030] In one embodiment, the input storage module (102) comprises a Kvasir dataset with 4,000 endoscopic images evenly distributed across eight different classes of gastrointestinal diseases, wherein the system (100) is configured to split the dataset into training data, which constitute 80% of the total dataset, and test data, which constitute 20% of the total dataset.
[0031] In one embodiment, the image preprocessing module (104) is also configured to use a Torchvision transformation module to implement data extension techniques and image normalization.
[0032] In one embodiment, the preprocessing module (104) is also configured to receive endoscopic images from a dataset containing several categories of gastrointestinal diseases, to resize the endoscopic images to a standardized resolution of 224 x 224 pixels, to apply data extension techniques that include random horizontal mirroring, rotations of up to 10 degrees and color adjustments including brightness, contrast, saturation and hue changes, and to normalize the processed images using predefined means and standard deviations.
[0033] In one embodiment, the multiple CNN models implemented by the feature extraction module (106) are initialized with ImageNet weights and fine-tuned to the data set used for training, i.e., the Kvasir data set.
[0034] In one embodiment, the multiple convolutional neural network models implemented by the feature extraction module (106) are fine-tuned by transfer learning on the preprocessed endoscopic images. This includes replacing the original classification layer with a user-defined classification layer comprising a fully connected layer with 512 units and ReLU activation, dropout layers with rates of 0.4 and 0.3, and a final output layer corresponding to the number of gastrointestinal disease classes.
[0035] In one embodiment, the system (100) further comprises an optimization module (114) configured to implement an optimizer on the CNN models with a learning rate of 0.0001 and a low weight decay to reduce overfitting, and wherein the optimization module (114) is further configured to implement a learning rate planner to lower the learning rate if the validation loss has not improved after several epochs.
[0036] In one embodiment, the system (100) further comprises a cross-validation module (116) configured to perform 5-fold cross-validation by splitting the data set into training subsets containing 80% of the data and validation subsets containing 20% of the data, wherein the cross-validation module (116) is configured to repeat the validation process multiple times with different data splits.
[0037] In one embodiment, the ensemble processing module (108) is configured to implement a weighted average ensemble that combines prediction probabilities from the multitude of CNN models used by the feature extraction module (106), wherein the ensemble processing module (108) is further configured to generate ensemble predictions by averaging the weighted prediction probabilities of all models.
[0038] In one embodiment, each of the plurality of convolutional neural network models is configured to be trained for 25 epochs, and the cross-validation module (116) is configured to store the highest-performing model from each convolution based on the validation accuracy.
[0039] The present invention addresses the urgent need for automated and accurate detection of colorectal cancer using endoscopic images by implementing a sophisticated deep learning system that combines multiple convolutional neural network architectures through ensemble learning techniques. The system represents a significant advance over conventional manual diagnostic systems and single-model deep learning approaches by leveraging the combined strengths of multiple CNN models to achieve superior diagnostic performance.
[0040] The system architecture comprises several interconnected modules that process endoscopic images and generate reliable classification results. The input storage module manages the Kvasir dataset, which contains 4,000 endoscopic images assigned to eight categories of gastrointestinal diseases. The system is configured to split this dataset into training and test subsets in an 80:20 ratio. This dataset forms the basis for training and evaluating the various CNN models used by the system.
[0041] The preprocessing module plays a crucial role in preparing endoscopic raw images for analysis by implementing standardized image processing techniques. This module scales all input images to a uniform resolution of 224×224 pixels and applies comprehensive data enhancement techniques, including random horizontal mirroring, rotations of up to 10 degrees, and color adjustments for brightness, contrast, saturation, and hue. The module also normalizes the processed images using predefined means and standard deviations, leveraging the Torchvision transformation module for consistent implementation of these preprocessing operations.
[0042] The feature extraction module is the central computational component of the system. It implements several pre-trained convolutional neural network models, initialized with ImageNet weights and subsequently adapted to the Kvasir dataset using transfer learning techniques. The transfer learning process involves replacing the original classification levels of these pre-trained models with user-defined classification architectures. These include fully connected levels with 512 units and ReLU activation, dropout levels with rates of 0.4 and 0.3 for regularization, and final output levels corresponding to the number of gastrointestinal disease classes in the dataset.
[0043] The system includes an optimization module that implements advanced training strategies to ensure optimal model performance while preventing overfitting. This module uses an AdamW optimizer with a learning rate of 0.0001 and applies weight-loss parameters to reduce overfitting tendencies. Additionally, the optimization module implements a learning rate planner that dynamically reduces the learning rate if the validation loss does not improve over several epochs, thus ensuring stable convergence throughout the training process.
[0044] To ensure statistical robustness and reliable performance evaluation, the system includes a cross-validation module that implements a five-fold cross-validation method. This module divides the dataset into training subsets comprising 80% of the data and validation subsets comprising 20% of the data. This validation process is repeated multiple times with different data divisions to enable comprehensive performance evaluation across various data configurations.
[0045] The ensemble module represents a key innovation of the system. It combines the results of multiple CNN models using a weighted ensemble approach. This module obtains prediction probabilities from all CNN models implemented by the feature extraction module and generates final ensemble predictions by calculating the weighted averages of these probabilities. This ensemble approach significantly improves classification accuracy and reliability compared to individual model predictions by leveraging the complementary strengths of different CNN architectures.
[0046] The classification output module processes the ensemble predictions to generate final diagnostic results for endoscopic colorectal cancer detection, while the user interface module provides medical professionals with an intuitive display of the classified results. The system is configured so that each CNN model is trained for 25 epochs, with the cross-validation module storing the highest-performing model from each fold based on validation accuracy metrics.
[0047] Fig. Figure 2 shows a block diagram of the proposed ensemble architecture of the system according to an embodiment of the present disclosure.
[0048] According to Fig. The system is configured for colorectal cancer detection and uses the Kvasir V1 dataset, which contains 4000 endoscopic images evenly distributed across eight different classes. The proposed system employs several convolutional neural network (CNN) models, namely ResNet50, EfficientNetV2-S, and a modified version of DenseNet121. The DenseNet121 CNN model is extended by adding additional fully connected layers to improve its learning and feature extraction capabilities. The CNN models are trained, validated, and tested on the Kvasir dataset using a transfer learning approach.
[0049] As in Fig. As shown in Figure 2, the deep learning-based system for the early and accurate detection of endoscopic colorectal cancer comprises the following components: The input storage module is configured to use the publicly available Kvasir dataset, which comprises 4,000 high-resolution endoscopic images from eight categories of gastrointestinal diseases. The images are annotated by medical experts to ensure clinical relevance and reliability. The input storage module is also configured to divide the dataset into a training subset containing 80% of the images and a test subset containing the remaining 20%. Additionally, the dataset is stratified using 5-fold cross-validation within the training subset to improve statistical robustness and prevent overfitting during model training and evaluation.
[0050] The preprocessing module, connected to the input storage module, prepares the endoscopic images for feature extraction. It scales all images to a standardized resolution of 224 x 224 pixels. It also applies various data enhancement techniques, including random horizontal mirroring, small rotations of up to 10 degrees, and color adjustments that modify brightness, contrast, saturation, and hue. After enhancement, the preprocessing module normalizes the images using predefined means and standard deviations. These transformations are implemented using the Torchvision transformation module.
[0051] The feature extraction module associated with the preprocessing module is configured to extract discriminatory features from the preprocessed images using multiple convolutional neural network (CNN) models. Specifically, the feature extraction module integrates DenseNet121, ResNet50, and EfficientNetV2S, each initialized with ImageNet weights and tuned to the Kvasir dataset. For transfer learning, the original classification levels of these models are replaced by a custom classification architecture consisting of a fully connected 512-unit level with ReLU activation, followed by dropout levels with dropout rates of 0.4 and 0.3, and a final output level corresponding to the number of gastrointestinal disease classes.
[0052] The optimization module is configured to train the CNN models using the AdamW optimizer, with a learning rate of 0.0001 and a small weight decay to prevent overfitting. The system implements cross-entropy loss to handle the multi-class classification problem. The optimization module is also configured with a learning rate scheduler (ReduceLROnPlateau) that dynamically reduces the learning rate if the validation loss does not improve over several epochs. Each CNN model is trained for 25 epochs, and the highest-performing model from each cross-validation fold is stored based on validation accuracy.
[0053] The cross-validation module, integrated into the optimization and input storage modules, is configured to perform five-fold cross-validation. This involves splitting the training data into five different configurations, comprising training subsets (80%) and validation subsets (20%). The cross-validation module ensures that the model is not biased by a single data split and calculates the average performance across all splits to aid in model selection.
[0054] The ensemble module, which is linked to the feature extraction module, is configured to implement and evaluate several ensemble techniques, including Weighted Average Ensemble, Max Confidence Voting, and Majority Voting with Confidence. The Weighted Average Ensemble method was chosen because it yields better results by averaging the prediction probabilities of DenseNet121, ResNet50, and EfficientNetV2S.
[0055] The classification output module associated with the ensemble module is configured to generate the final diagnostic results for the classification of colorectal cancer. The user interface associated with the classification output module is configured to display the predicted results and the corresponding diagnostic metrics.
[0056] The evaluation unit integrated into the classification output module is also configured to calculate key performance indicators, including accuracy, precision, recall, F1 score, and a confusion matrix. The evaluation confirms that the ensemble-based approach delivers a more balanced and superior classification performance compared to individual CNN models.
[0057] The input storage module of the deep-learning-based system is configured to store the Kvasir dataset, which was specifically curated for the development and evaluation of computer-aided diagnostic (CAD) systems in gastrointestinal endoscopy. The dataset comprises 4,000 annotated endoscopic images evenly distributed across eight clinically and anatomically relevant classes: polyps, esophagitis, ulcerative colitis, Z-line, pylorus, appendix, stained and elevated polyps, and stained resection margins. These images were acquired at the Vestre Hospital. The images are from the Viken Health Trust in Norway and were validated by experienced endoscopists to ensure diagnostic accuracy and clinical relevance.The image resolution varies between 720 × 576 and 1920 × 1072 pixels and is organized in class-specific folders, enabling structured access for systematic data preprocessing, model training, and evaluation. The input storage module is also configured to support the organized loading and labeling of these class-specific data subsets, thus facilitating seamless integration with downstream processing modules.
[0058] The preprocessing module associated with the input storage module is configured to standardize and improve the quality of the input images before they are fed into the feature extraction pipeline. First, the preprocessing module resizes all endoscopic images to a fixed resolution of 224 x 224 pixels to maintain dimensional consistency across the entire dataset. This is followed by color conversion to ensure the images are in the correct format required by the downstream convolutional neural networks (CNNs). To improve the generalizability of the classification models and prevent overfitting, the preprocessing module is configured to apply a series of data augmentation techniques.These techniques include random horizontal mirroring, rotation within a range of ±10 degrees, and color jitter, which introduces controlled variations in brightness, contrast, saturation, and hue. These enhancements increase the diversity of training examples without altering the core visual features required for classification. Following the enhancements, the preprocessing module normalizes all images using means and standard deviations from the ImageNet dataset. This adjusts the pixel intensity distribution to match the distribution expected by the pretrained CNN models of the feature extraction module. This normalization step improves feature extraction efficiency and enables faster and more stable convergence during model training.Through combined operations of resizing, expanding, color matching, and normalization, the preprocessing module ensures that all inputs for the feature extraction module are clean, diverse, and standardized. This contributes to improved performance, generalizability, and robustness of the deep learning-based classification system.
[0059] The feature extraction module of the deep learning-based system is configured to implement a variety of convolutional neural network models, including ResNet50, EfficientNetV2-S and DenseNet121, each finely tuned and adapted for multi-class classification of endoscopic images from the Kvasir dataset.
[0060] The ResNet50 model implemented in the feature extraction module is initialized with pretrained ImageNet weights and adapted for classification by replacing the last fully connected (fc) layer with a task-specific linear layer. The modified layer outputs a vector corresponding to the number of classes in the Kvasir dataset, as follows: model.fc=nn.Linear(model.fc.in_features,num_classes)
[0061] The ResNet50 architecture comprises residual learning blocks that utilize identity joins to facilitate gradient flow. Each residual block is configured to learn a transformation F(x) such that the output becomes y = F(x) + x. This mitigates vanishing gradient problems and improves convergence during training. Structurally, ResNet50 employs bottleneck residual blocks with 1×1, 3×3, and 1×1 convolutions arranged in a [3, 4, 6, 3] configuration across four levels. The result is a 50-layer architecture optimized for feature extraction from medical images.
[0062] The EfficientNetV2-S model, also included in the feature extraction module, is initialized with ImageNet weights to utilize common visual representations. The model's final classification level is replaced with a task-specific linear level to correspond to the number of disease classes in the Kvasir dataset. model.classifier=nn.Linear(model.classifier[1].in_features,num_classes)
[0063] The classification result can be formally described as follows: y=Softmax(W⋅z+b)
[0064] Where: • y is the predicted class probability distribution, • z is the feature vector generated by the model's backbone, • W and b are the learnable weights and bias parameters of the classifier layer.
[0065] EfficientNetV2-S uses a combination of Fused MBConv and MBConv blocks to efficiently capture spatial and channel-wise features. The transformation applied by a representative block can be expressed as follows: y_output=x+SE(Conv3×3(BN(Conv1×1(x))))
[0066] Where SE stands for Squeeze-and-Excitation Attention, BN for Batch Normalization, and Conv for Convolution Operations. This structure provides residual connections for improved gradient flow, and the SE mechanism enhances feature recalibration, thus enabling robust and efficient classification of endoscopic images.
[0067] The DenseNet121 model, additionally embedded in the feature extraction module, is loaded with pre-trained ImageNet weights and adapted with a custom multi-layer classification head for improved performance on the Kvasir dataset. DenseNet121 is characterized by its dense connectivity structure, where each layer receives the outputs of all preceding layers. This connectivity promotes feature reuse, mitigates vanishing gradients, and enables efficient training on relatively small medical datasets. The standard classification layer has been replaced with a task-specific deep head consisting of two fully connected layers with ReLU activation, dropout regularization, and a final linear output layer corresponding to the number of classes. model.classifier=nn.Sequential(Linear→ReLU→Dropout→Linear→ReLU→Dropout→Linear)
[0068] This classification header performs the following transformation: y=Softmax(W3⋅Drop0.3(ReLU(W2⋅Drop0.4(ReLU(W1⋅z+b1))+b2))+ b3)
[0069] Where: • z is the output feature vector from the DenseNet backbone, • W1, W2, W3 and b1, b2, b3 are the weights and biases of the fully connected layers, • Drop refers to dropout layers with respective dropout probabilities of 0.4 and 0.3, • ReLU is the activation function used for the nonlinearity between layers.
[0070] The last head consists of: Linear(1024→512)→ReLU→Dropout(0,4) Linear(1024→512)→ReLU→Dropout(0,4) Linear(256→Number_of_classes)
[0071] This hierarchical classifier architecture enables DenseNet121 to learn discriminating features more effectively, especially for subtle and complex visual patterns such as polyps in colonoscopy images. The dense connectivity and additional classifier head reduce overfitting and improve generalization, making DenseNet121 particularly effective for small annotated medical datasets like Kvasir.
[0072] The system's ensemble module is configured to improve classification performance by aggregating outputs from multiple CNN models. It supports three ensemble strategies: Max Confidence Voting, Majority Voting with Confidence, and Weighted Average Ensemble. • In Max Confidence Voting, the final prediction is selected based on the class with the highest confidence value of all models. • In majority voting with confidence, the class predicted by the majority among the models is determined, and the prediction with the highest confidence is selected from this majority group. • The most effective strategy, the Weighted Average Ensemble, calculates the final prediction probabilities by assigning performance-based weights to the outputs of the individual models: P−Ensemble=∑(wi×pi) Where pi is the softmax probability from the i-th model, wi is the corresponding weight and ∑ wi = 1.
[0073] The weights are derived from the validation performance, allowing more accurate models to have a greater impact on the final prediction. This ensemble method demonstrated superior performance and robustness, particularly in differentiating complex gastrointestinal disease classes, and is used within the system as the standard ensemble approach.
[0074] The system fine-tunes and adapts the RestNet50, EfficiencyNetV2S, and DenseNet121 models for colorectal cancer detection. To optimize the models, the system uses transfer learning and adjusts the models' internal layers, improving their ability to extract key features from histopathological images. The system is configured to use a new classifier head developed for DenseNet121 by adding dense, dropout, and activation layers, thereby improving the model's feature learning capability and reducing overfitting. The system is configured to implement three ensemble techniques: majority voting with confidence, weighted average, and maximum confidence voting. These techniques facilitate the combination of predictions from all models, resulting in a more reliable and consistent final outcome.The performance of each model and ensemble technique is carefully evaluated using multiple metrics, including accuracy, sensitivity, specificity, and AUC-ROC. Confusion matrices and learning curves are employed to assess model performance. Among all tested techniques, the weighted average ensemble achieved the highest accuracy of 99.25%, outperforming any single model. The results demonstrate that combining deep learning models with ensemble techniques can significantly improve both the accuracy and reliability of colorectal cancer detection.
[0075] In one implementation, the experimental design and evaluation of the deep learning-based colon cancer detection system were performed on a machine running Microsoft Windows 10 Pro, equipped with an Intel(R) Core(TM) i5-6300U CPU at 2.40 GHz, Intel HD Graphics 520, two cores, four logical processors, and 8 GB of RAM. Regardless of the system requirements, all experiments were conducted using Google Colab, which provided the necessary computational resources for training and evaluating the deep learning model. The implementation was carried out using Python and a number of widely used libraries and frameworks, including Pandas, NumPy, Matplotlib.pyplot, Seaborn, Scikit-Learn, Torch, TorchVision, and Google Colab utilities. These tools were used for data manipulation, visualization, model evaluation, and training.
[0076] In one embodiment, several carefully selected hyperparameter settings were used to optimize the learning process for the system's model training. Hyperparameters such as learning rate, batch size, number of epochs, dropout rate, optimizer type, and activation functions were predefined and configured to control the training behavior. Specifically, the Adam optimizer was used with a learning rate planner to dynamically adjust the learning rate to compensate for validation losses. Dropout and weight decay techniques were implemented to prevent overfitting and improve model generalization. These measures ensured efficient and stable model training across multiple epochs.
[0077] In one implementation, several performance metrics were used to assess the classification models. These included accuracy, precision, recall, F1 score, confusion matrix, ROC curve, and AUC (area under curve). The confusion matrix is a tabular representation containing true positives (TP), true negatives (TN), false positives (FP), and false negatives (FN), which are essential for calculating other performance indicators. Accuracy is defined as the ratio of correctly predicted instances (TP and TN) to the total number of predictions. Precision quantifies the proportion of true positive predictions among all positive predictions, while recall measures the model's ability to correctly identify actually positive cases. The F1 score, calculated as the harmonic mean of precision and recall, provides a balanced performance metric that is particularly useful in cases of class imbalance.In addition to these measures, ROC curves were created to evaluate the classifier's performance by illustrating the trade-off between the true positive rate (TPR) and the false positive rate (FPR) across various classification thresholds. The area under the ROC curve (AUC) provides a single scalar value that summarizes the model's ability to distinguish between classes. A higher AUC value indicates better classification performance. Values close to 1 indicate near-perfect class separation, while values close to 0 indicate poor or even reversed classification results. K-fold cross-validation was used to ensure the reliability of the model's performance across different data splits. Five-fold cross-validation (k=5) was applied, in which the dataset was divided into five equal parts.In each partition, four parts were used for training and one for testing. This process was repeated five times, so that each subset served as a test set once. The average performance metric across the five partitions was then calculated to represent the final evaluation result. This method ensured that the performance results were statistically robust and not dependent on a single partition of the data, thereby improving the reliability of the reported accuracy and other metrics.
[0078] In one embodiment, a deep learning-based system is configured for the early and accurate detection of colorectal cancer using colonoscopy images from the KvasirV1 dataset. This system utilizes a weighted average ensemble framework instead of relying solely on individual convolutional transfer learning models. The system includes a data preprocessing module that performs image resizing, normalization using ImageNet statistics, and augmentation techniques such as rotation, mirroring, and contrast adjustments to improve feature richness and model robustness. A cross-validation engine implements a five-fold cross-validation strategy to promote generalization and minimize overfitting.
[0079] A feature extraction module runs three pre-trained CNN architectures—DenseNet121, EfficientNetV2S, and ResNet50—in parallel to extract diverse and representative deep features from the pre-processed input images. The outputs of these models are processed by an ensemble learning module that supports three ensemble strategies: Max Confidence Voting, Majority Voting with Confidence, and a Weighted Average Ensemble, the latter assigning weights based on model-specific validation performance. The Weighted Average Ensemble proved to be the optimal approach, achieving a maximum classification accuracy of 99.25% while maintaining high precision, hit rate, and a high F1 score, as confirmed by confusion matrix analysis.The system also offers improvements at the classifier head level, integrates dropout and batch normalization levels, and supports fine-tuning of deeper CNN levels to reduce overfitting while improving clinical reliability. Empirical results demonstrate the system's superior diagnostic capability, computational power, and practical applicability for automated colorectal cancer diagnosis. The integrated framework significantly contributes to the advancement of ensemble learning techniques in medical imaging and provides a reliable foundation for intelligent, pathology-based diagnostic support systems.
[0080] The deep learning-based system for endoscopic colorectal cancer detection using medical images utilizes lightweight convolutional neural network (CNN) architectures. Unlike previous systems, which often rely on complex and computationally intensive models, the present invention emphasizes both high classification accuracy and computational efficiency, making it better suited for real-time clinical applications. While deeper CNNs are known for extracting more detailed features, the system uses three models: EfficientNetV2S, ResNet50, and DenseNet121. The results showed that each of these models demonstrated strong individual performance. EfficientNetV2S achieved an accuracy of 92.00%, ResNet50 achieved 92.63%, and DenseNet121 performed best with an accuracy of 93.20%.These results surpass several previous complex architectures and demonstrate that well-optimized transfer learning models can deliver outstanding performance when applied to specific medical imaging tasks. The present invention provides an approach that significantly improves recognition accuracy. The system is configured to integrate the CNN models into a weighted average ensemble. The ensemble approach resulted in an impressive accuracy of 99.25%, with corresponding precision, recall, and F1 score values of 98.74%, 98.73%, and 98.74%, respectively. As shown in Table 5, this ensemble outperformed all other models, highlighting the advantage of combining efficient models to improve classification accuracy.
[0081] In contrast to previous studies based on manual feature extraction or dimensionality reduction techniques such as PCA, DWT, or FHWT, the proposed invention operates fully end-to-end. This approach preserves the semantic richness of the data, reduces the need for manual intervention, and simplifies the model pipeline. Furthermore, the chosen base models outperform previously used lightweight models such as ShuffleNet and MobileNet, demonstrating that deeper, yet computationally efficient architectures can deliver superior results.
[0082] To overcome limitations due to the size of the training dataset, the system is configured to implement comprehensive data expansion strategies that significantly improve the model's generalizability. Unlike studies that rely heavily on synthetic data or limited datasets, our model was trained using real colorectal cancer images and validated through five-fold cross-validation, ensuring robust and unbiased analysis.
[0083] Overall, the results showed that the proposed system exhibits high accuracy while remaining computationally intensive. It outperforms many resource-intensive systems and is therefore suitable for use in real-world medical diagnostics.
[0084] The drawings and the preceding description show examples of embodiments. Those skilled in the art will recognize that one or more of the described elements can be combined to form a single functional element. Alternatively, certain elements can be divided into several functional elements. Elements of one embodiment can be added to another embodiment. For example, the sequence of the processes described here can be changed and is not limited to the manner described herein. Furthermore, the actions of a flowchart need not be implemented in the sequence shown; nor does it necessarily have to be performed by all actions. Actions that are not dependent on other actions can also be performed in parallel with the other actions. The scope of the embodiments is in no way limited by these specific examples.Numerous variations are possible, whether explicitly stated in the specification or not, such as differences in structure, dimensions, and material use. The range of embodiments is at least as broad as specified in the following claims.
[0085] Advantages, further benefits, and problem solutions have been described above with reference to specific embodiments. However, the advantages, benefits, problem solutions, and all components that can lead to an advantage, benefit, or solution occurring or becoming more apparent are not to be construed as critical, necessary, or essential features or components of individual or all claims. REFERENCES 100 A Deep Learning Based System for the Early and Accurate Detection of Endoscopic Colon Cancer. 102 Input memory module 104 Preprocessing module 106 Module for Feature Extraction 108 Ensemble Module 110 - Classification output module 112 User interface 114 Optimization module 116 Cross-validation module 202 Input of an Endoscopic Medical Image 204 Image preprocessing 206 Efficientnetv2s CNN model 208 Weighted average ensemble 210 Classification and Performance Measurement 212 Resnet50 CNN model 214 Densenet121 CNN model
Claims
[1] A deep learning-based system for the early and accurate detection of endoscopic colorectal cancer, consisting of: an input storage module configured to store the endoscopic images to be used as input for the folding model for colorectal cancer detection, the images being taken from a Kvasir dataset comprising endoscopic images covering eight categories of gastrointestinal diseases; a preprocessing module connected to the input storage module, which is configured to preprocess the images, wherein the preprocessing module is configured to resize the images, apply data extension, and normalize the images; a preprocessing module connected to a feature extraction module configured to receive the preprocessed endoscopic images and extract meaningful features from the images using pre-trained Convolutional Neural Networks (CNN) models, wherein the feature extraction module is configured to use a variety of CNN models; an ensemble module connected to the feature extraction module and configured to combine the outputs of the CNN models implemented by the feature extraction module for feature extraction from preprocessed endoscopic images, wherein the ensemble module is configured to average the probabilities predicted by the CNN models; a classification output module connected to the ensemble module, configured to generate final classification results for endoscopic colorectal cancer diagnosis; and A user interface connected to the classification output module, configured to display the classified results. [2] System according to claim 1, wherein the input storage module comprises a Kvasir data set with 4,000 endoscopic images evenly distributed across eight different classes of gastrointestinal diseases, wherein the system is configured to split the data set into training data comprising 80% of the total data set and test data comprising 20% of the total data set. [3] System according to claim 1, wherein the image preprocessing module is further configured to use a Torchvision transformation module for implementing the data extension techniques and image normalization. [4] System according to claim 1, wherein the preprocessing module is further configured to: receive endoscopic images from a dataset comprising several categories of gastrointestinal diseases; scale the endoscopic images to a standardized resolution of 224 x 224 pixels; apply data enhancement techniques comprising random horizontal mirroring, rotations of up to 10 degrees, and color adjustments including brightness, contrast, saturation, and hue changes; and normalize the processed images using predetermined means and standard deviations. [5] System according to claim 1, wherein the plurality of CNN models implemented by the feature extraction module are initialized with ImageNet weights and fine-tuned to the data set used for training, i.e. the Kvasir data set. [6] System according to claim 1, wherein the plurality of convolutional neural network models implemented by the feature extraction module are subjected to fine-tuning by transfer learning on the preprocessed endoscopic images, comprising replacing the original classification level with a user-defined classification level consisting of a fully connected level with 512 units and ReLU activation, dropout levels with rates of 0.4 and 0.3, and a final output level corresponding to the number of gastrointestinal disease classes. [7] System according to claim 1, further comprising an optimization module configured to implement an optimizer on the CNN models with a learning rate of 0.0001 and a low weight decay to reduce overfitting, and wherein the optimization module is further configured to implement a learning rate planner to lower the learning rate if the validation loss has not improved after several epochs. [8] System according to claim 1, further comprising a cross-validation module configured to implement 5-fold cross-validation by dividing the data set into training subsets containing 80% of the data and validation subsets containing 20% of the data, wherein the cross-validation module is configured to repeat the validation process multiple times with different data divisions; [9] System according to claim 1, wherein the ensemble processing module is configured to implement a weighted average ensemble that combines prediction probabilities from the plurality of CNN models used by the feature extraction module, wherein the ensemble processing module is further configured to generate ensemble predictions by averaging the weighted prediction probabilities of all models. [10] System according to claims 5 and 6, wherein each of the plurality of convolutional neural network models is configured to be trained for 25 epochs, and the cross-validation module is configured to store the model with the best performance from each convolution based on the validation accuracy.