Classification method for cross-modal renal cancer pathology upgrade based on double-input fusion network

Through the feature extraction and fusion of the dual-input fusion network, the time-consuming error and high computational complexity of renal cancer pathological grading in traditional methods are solved, and the automatic and accurate classification of renal cancer pathological upgrades are achieved.

CN120298776APending Publication Date: 2025-07-11DALIAN NEUSOFT UNIV OF INFORMATION
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510359160.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing image-based renal cancer pathological grading method relies on manual feature extraction, which is time-consuming and error-prone. Traditional deep learning models lack fine-grained feature capture capabilities when processing complex images, have high computational complexity, and have overfitting, making it difficult to achieve stable diagnosis of renal cancer pathological upgrades.

Method used

Using a dual-input fusion network method, the CT images are enhanced through the CLAHE algorithm, and the original and enhanced CT images are extracted respectively. The feature fusion modules of the Bottom branch and the Upper branch are used to combine the feature fusion module and the full connection layer to realize the classification of renal cancer pathological upgrade.

Benefits of technology

It realizes automated and efficient classification of renal cancer CT images, improves the capture ability of fine-grained features, reduces calculation costs, improves the accuracy and stability of classification results, and is suitable for pathological upgrading diagnosis of complex images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298776A_ABST
    Figure CN120298776A_ABST
Patent Text Reader

Abstract

The invention discloses a classification method for cross-modal renal cancer pathology upgrading based on a double-input fusion network, and the method comprises the steps: building the double-input fusion network which comprises a Bottom branch, a Upper branch and a feature fusion module, and obtaining the nonlinear enhanced high-level semantic features of an original renal cancer CT image through the original renal cancer CT image and the Upper branch; meanwhile, through the kidney cancer CT image subjected to image enhancement and Bottom branches, nonlinear enhancement high-level semantic features of the enhanced kidney cancer CT image are obtained, and finally classification of the kidney cancer CT image is completed through a feature fusion module. According to the double-input fusion network, cross-modal classification of the images can be automatically and efficiently achieved, when complex images are processed, the capturing capacity of fine-grained features is high, the calculation cost is low, the accuracy of the classification result of the images is high, and the double-input fusion network has great significance in achieving kidney cancer pathology upgrading diagnosis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of cross-modal renal cancer pathological grade classification, and particularly relates to a classification method for cross-modal renal cancer pathological grade upgrade based on a dual-input fusion network. Background Art

[0002] Renal cell carcinoma (RCC) is the most common malignant tumor of the kidney, accounting for about 90% of all kidney tumors. Early diagnosis is of crucial significance for improving the survival rate and treatment effect of patients. Although preoperative enhanced CT has become the preferred imaging examination method for preoperative evaluation of renal cancer, the image details are limited and it cannot fully reveal all the characteristics of the tumor. During the diagnosis and treatment process, the pathological grade of renal cancer may change, especially during the pathological evaluation from biopsy specimens to surgical resection specimens, and the phenomenon of common pathological grade upgrade occurs. This pathological grade upgrade not only affects the tumor grade, but also has an important impact on prognosis evaluation and the selection of treatment strategies.

[0003] With the continuous development of computed tomography (CT) technology, using the texture features of CT images for non-invasive grade diagnosis has become an important trend. However, most of the existing image-based grading methods rely on manual feature extraction, which is not only time-consuming but also prone to errors. In addition, since the preoperative imaging judgment of renal tumor staging is not always completely accurate, renal biopsy is usually required for pathological detection, but renal biopsy is often accompanied by complications such as bleeding and infection.

[0004] Deep learning models analyze the lesion features in a large number of medical images and update the connection weights between the original neurons. Traditional deep learning models and their improved models have made significant progress in the classification of renal cancer pathological grade upgrade, such as the AlexNet network, SqueezeNet network, GooleNe network, ResNet network, and the recently proposed TL-MobileNet network, NDAG-Net network, R-AlexNet network, DenseNet-TF network model.

[0005] However, traditional machine learning classification methods require a large sample size and manual setting of image features and probability values. Therefore, the results are unstable, the false positive rate is high, and the feasibility in practical applications is low. For the most advanced network model algorithms, the R-AlexNet network extracts features from low-level textures to high-level semantics through a multi-layer network structure and has certain advantages when processing large-scale image data. However, when processing complex images, its ability to capture fine-grained features is weak, and the large number of parameters is also prone to overfitting. The DenseNet-TF network realizes feature reuse and efficient gradient propagation through a dense connection mechanism and can extract richer features. However, this mechanism also increases the computational complexity, especially in the case of high-resolution feature maps, significantly increasing the computational cost and requiring a large amount of video memory to store intermediate feature maps. The NDAG-Net network introduces a directed acyclic graph structure and can integrate feature information from multiple levels and different sources. However, its excessive dependence on regional importance assessment may ignore non-critical regions, thus limiting the global feature extraction ability. TL-MobileNet, as a lightweight network, uses depthwise separable convolutions to reduce the computational amount and is suitable for resource-constrained environments. However, its lightweight design may sacrifice some feature expression ability, thus affecting the accuracy of complex image classification tasks. Summary of the Invention

[0006] The present invention discloses a classification method for cross-modal renal cancer pathological upgrading based on a dual-input fusion network to overcome the above technical problems.

[0007] To achieve the above object, the technical solution of the present invention is as follows:

[0008] A classification method for cross-modal renal cancer pathological upgrading based on a dual-input fusion network includes the following steps:

[0009] S1: Obtain the original renal cancer CT image, and obtain the enhanced renal cancer CT image according to the CLAHE algorithm;

[0010] S2: Establish a dual-input fusion network, where the dual-input fusion network includes a Bottom branch, an Upper branch, and a feature fusion module;

[0011] S3: According to the original renal cancer CT image and the Upper branch, obtain the non-linearly enhanced high-level semantic features of the original renal cancer CT image; obtain the feature map of the original renal cancer CT image for extracting the non-linearly enhanced high-level semantic features;

[0012] S4: According to the enhanced renal cancer CT image after image enhancement and the Bottom branch, obtain the non-linearly enhanced high-level semantic features of the enhanced renal cancer CT image; obtain the feature map of the enhanced renal cancer CT image for extracting the non-linearly enhanced high-level semantic features;

[0013] S5: Based on the original renal cancer CT image feature map for extracting non - linear enhanced high - level semantic features, the enhanced renal cancer CT image feature map for extracting non - linear enhanced high - level semantic features, and the feature fusion module, obtain the probability that the original renal cancer CT image belongs to a high - grade renal cancer CT image and the probability that it belongs to a low - grade renal cancer CT image, so as to complete the classification of renal cancer CT images.

[0014] Further, the structures of the Bottom branch and the Upper branch are the same;

[0015] The Upper branch includes a first feature extraction module, a second feature extraction module, and a third feature extraction module connected in sequence from start to end;

[0016] The first feature extraction module is used to obtain the dimensionality - reduced underlying significant features of the original renal cancer CT image according to the original renal cancer CT image, so as to obtain the original renal cancer CT image feature map for extracting dimensionality - reduced underlying significant features;

[0017] The second feature extraction module is used to obtain the dimensionality - reduced middle - layer significant features of the original renal cancer CT image according to the original renal cancer CT image feature map for extracting dimensionality - reduced underlying significant features;

[0018] The third feature extraction module is used to obtain the non - linear enhanced high - level semantic features of the original renal cancer CT image according to the original renal cancer CT image feature map for extracting dimensionality - reduced middle - layer significant features, so as to obtain the original renal cancer CT image feature map for extracting non - linear enhanced high - level semantic features.

[0019] Further, the first feature extraction module includes a first convolutional layer, a first activation layer, and a first max - pooling layer;

[0020] The first convolutional layer is used to obtain the underlying local features according to the original renal cancer CT image, so as to obtain the original renal cancer CT image feature map for extracting underlying local features;

[0021] The first activation layer is used to obtain non - linear enhanced underlying local features according to the original renal cancer CT image feature map for extracting underlying local features, so as to obtain the original renal cancer CT image feature map for extracting non - linear enhanced underlying local features;

[0022] The first max - pooling layer is used to obtain the dimensionality - reduced underlying significant features of the original renal cancer CT image according to the original renal cancer CT image feature map for extracting non - linear enhanced underlying local features, so as to obtain the original renal cancer CT image feature map for extracting dimensionality - reduced underlying significant features.

[0023] Further, the second feature extraction module includes a second convolutional layer, a second activation layer, and a second max - pooling layer;

[0024] The second convolutional layer is used to obtain middle-level features based on the original renal cancer CT image feature map that extracts the dimension-reduced underlying significant features, so as to obtain the original renal cancer CT image feature map that extracts the middle-level features;

[0025] The second activation layer is used to obtain non-linearly enhanced middle-level features based on the original renal cancer CT image feature map that extracts the middle-level features, so as to obtain the original renal cancer CT image feature map that extracts the non-linearly enhanced middle-level features;

[0026] The second max pooling layer is used to obtain the dimension-reduced middle-level significant features of the original renal cancer CT image based on the original renal cancer CT image feature map that extracts the non-linearly enhanced middle-level features, so as to obtain the original renal cancer CT image feature map that extracts the dimension-reduced middle-level significant features.

[0027] Further, the third feature extraction module includes a third convolutional layer and a third activation layer;

[0028] The third convolutional layer is used to obtain the high-level semantic features of the original renal cancer CT image based on the original renal cancer CT image feature map that extracts the dimension-reduced middle-level significant features, so as to obtain the original renal cancer CT image feature map that extracts the high-level semantic features;

[0029] The third activation layer is used to obtain non-linearly enhanced high-level semantic features based on the original renal cancer CT image feature map that extracts the high-level semantic features, so as to obtain the original renal cancer CT image feature map that extracts the non-linearly enhanced high-level semantic features.

[0030] Further, the relationship between the first convolutional layer, the second convolutional layer, and the third convolutional layer is:

[0031] The number of convolution kernels of the third convolutional layer = 2 * the number of convolution kernels of the second convolutional layer = 4 * the number of convolution kernels of the first convolutional layer.

[0032] Further, the feature fusion module includes an element-wise addition module, a fourth feature extraction module, a fifth feature extraction module, and a fully connected layer;

[0033] The element-wise addition module is used to obtain the feature-fused renal cancer CT image feature map based on the original renal cancer CT image feature map that extracts the non-linearly enhanced high-level semantic features and the enhanced renal cancer CT image feature map that extracts the non-linearly enhanced high-level semantic features;

[0034] The fourth feature extraction module includes a fourth convolutional layer and a fourth activation layer;

[0035] The fourth convolutional layer is used to obtain the fused low-level features based on the feature-fused renal cancer CT image feature map, so as to obtain the renal cancer CT image feature map that extracts the fused low-level features;

[0036] The fourth activation layer is used to obtain non-linearly enhanced fused low-level features based on the renal cancer CT image feature map that extracts and fuses low-level features, so as to obtain a renal cancer CT image feature map that extracts non-linearly enhanced fused low-level features;

[0037] The fifth feature extraction module includes a fifth convolutional layer and a fifth activation layer;

[0038] The fifth convolutional layer is used to obtain fused high-level features based on the renal cancer CT image feature map that extracts non-linearly enhanced fused low-level features, so as to obtain a renal cancer CT image feature map that extracts fused high-level features;

[0039] The fifth activation layer is used to obtain non-linearly enhanced fused high-level features based on the renal cancer CT image that extracts fused high-level features, so as to obtain a renal cancer CT image feature map that extracts non-linearly enhanced fused high-level features;

[0040] The fully connected layer is used to obtain the probability of the renal cancer category based on the renal cancer CT image feature map that extracts non-linearly enhanced fused high-level features, so as to complete the classification of the pathological upgrade of renal cancer.

[0041] Beneficial effects: A classification method for cross-modal pathological upgrade of renal cancer based on a dual-input fusion network of the present invention establishes a dual-input fusion network including a Bottom branch, an Upper branch, and a feature fusion module. Through the original renal cancer CT image and the Upper branch, non-linearly enhanced high-level semantic features of the original renal cancer CT image are obtained; at the same time, through the enhanced renal cancer CT image and the Bottom branch, non-linearly enhanced high-level semantic features of the enhanced renal cancer CT image are obtained. Finally, through the feature fusion module, the classification of the renal cancer CT image is completed. The dual-input fusion network of the present invention can automatically and efficiently achieve cross-modal classification of images. When processing complex images, it has strong ability to capture fine-grained features, low computational cost, high accuracy of the classification results of images, and is of great significance for realizing the diagnosis of pathological upgrade of renal cancer. Description of the Drawings

[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0043] Figure 1 It is a flowchart of the classification method for cross-modal pathological upgrade of renal cancer based on a dual-input fusion network of the present invention;

[0044] Figure 2 Schematic diagram of the overall algorithm flow in the embodiment of the present invention;

[0045] Figure 3a Original renal cancer CT image and corresponding grayscale image in the embodiment of the present invention;

[0046] Figure 3b Enhanced renal cancer CT image and corresponding grayscale image in the embodiment of the present invention;

[0047] Figure 4 Schematic diagram of the dual-input fusion network structure in the embodiment of the present invention;

[0048] Figure 5 Schematic diagram of feature fusion and classification in the embodiment of the present invention;

[0049] Figure 6 Feature fusion visualization heat map in the embodiment of the present invention;

[0050] Figure 7 Schematic diagram of the data set in the embodiment of the present invention;

[0051] Figure 8 Schematic diagram of the algorithm experimental results in the embodiment of the present invention. Detailed implementation manners

[0052] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0053] This embodiment introduces a classification method for cross-modal renal cancer pathological upgrading based on a dual-input fusion network (Dual-Input Fusion Network, DIF-Net), as Figure 1 and Figure 2 shown, which includes the following steps:

[0054] S1: Obtain the original renal cancer CT image, and obtain the enhanced renal cancer CT image according to the CLAHE algorithm;

[0055] S2: Establish a dual-input fusion network, and the dual-input fusion network includes a Bottom branch, an Upper branch and a feature fusion module;

[0056] S3: According to the original renal cancer CT image and the Upper branch, input the original renal cancer CT image into the Upper branch to obtain the non-linearly enhanced high-level semantic features of the original renal cancer CT image; to obtain the feature map of the original renal cancer CT image for extracting the non-linearly enhanced high-level semantic features;

[0057] S4: According to the image-enhanced renal cancer CT image and the Bottom branch, input the image-enhanced renal cancer CT image into the Bottom branch to obtain the non-linearly enhanced high-level semantic features of the image-enhanced renal cancer CT image; to obtain the feature map of the image-enhanced renal cancer CT image for extracting the non-linearly enhanced high-level semantic features;

[0058] S5: According to the feature map of the original renal cancer CT image for extracting the non-linearly enhanced high-level semantic features, the feature map of the image-enhanced renal cancer CT image for extracting the non-linearly enhanced high-level semantic features, and the feature fusion module, input the feature map of the original renal cancer CT image for extracting the non-linearly enhanced high-level semantic features and the feature map of the image-enhanced renal cancer CT image for extracting the non-linearly enhanced high-level semantic features into the feature fusion module to obtain the probability that the original renal cancer CT image belongs to a high-grade renal cancer CT image and the probability that it belongs to a low-grade renal cancer CT image, so as to complete the classification of the renal cancer CT image.

[0059] Preferably, the structures of the Bottom branch and the Upper branch are the same;

[0060] The Upper branch includes a first feature extraction module Block1, a second feature extraction module Block2, and a third feature extraction module Block3 that are connected in sequence from beginning to end;

[0061] The first feature extraction module is used to obtain the dimension-reduced underlying significant features of the original renal cancer CT image according to the original renal cancer CT image to obtain the feature map of the original renal cancer CT image for extracting the dimension-reduced underlying significant features;

[0062] The first feature extraction module includes a first convolutional layer, a first activation layer, and a first max pooling layer;

[0063] The first convolutional layer is used to obtain the underlying local features according to the original renal cancer CT image to obtain the feature map of the original renal cancer CT image for extracting the underlying local features;

[0064] The first activation layer is used to obtain the non-linearly enhanced underlying local features according to the feature map of the original renal cancer CT image for extracting the underlying local features to obtain the feature map of the original renal cancer CT image for extracting the non-linearly enhanced underlying local features;

[0065] The first maximum pooling layer is used to obtain the reduced-dimensional underlying significant features of the original renal cancer CT image based on the original renal cancer CT image feature map that extracts locally enhanced non-linear underlying features, so as to obtain the original renal cancer CT image feature map that extracts the reduced-dimensional underlying significant features;

[0066] The second feature extraction module is used to obtain the reduced-dimensional middle-level significant features of the original renal cancer CT image based on the original renal cancer CT image feature map that extracts the reduced-dimensional underlying significant features

[0067] The second feature extraction module includes a second convolutional layer, a second activation layer, and a second maximum pooling layer;

[0068] The second convolutional layer is used to obtain middle-level features based on the original renal cancer CT image feature map that extracts the reduced-dimensional underlying significant features, so as to obtain the original renal cancer CT image feature map that extracts the middle-level features;

[0069] The second activation layer is used to obtain non-linearly enhanced middle-level features based on the original renal cancer CT image feature map that extracts the middle-level features, so as to obtain the original renal cancer CT image feature map that extracts the non-linearly enhanced middle-level features;

[0070] The second maximum pooling layer is used to obtain the reduced-dimensional middle-level significant features of the original renal cancer CT image based on the original renal cancer CT image feature map that extracts the non-linearly enhanced middle-level features, so as to obtain the original renal cancer CT image feature map that extracts the reduced-dimensional middle-level significant features;

[0071] The third feature extraction module is used to obtain the non-linearly enhanced high-level semantic features of the original renal cancer CT image based on the original renal cancer CT image feature map that extracts the reduced-dimensional middle-level significant features, so as to obtain the original renal cancer CT image feature map that extracts the non-linearly enhanced high-level semantic features;

[0072] The third feature extraction module includes a third convolutional layer and a third activation layer;

[0073] The third convolutional layer is used to obtain the high-level semantic features of the original renal cancer CT image based on the original renal cancer CT image feature map that extracts the reduced-dimensional middle-level significant features, so as to obtain the original renal cancer CT image feature map that extracts the high-level semantic features;

[0074] The third activation layer is used to obtain non-linearly enhanced high-level semantic features based on the original renal cancer CT image feature map that extracts the high-level semantic features, so as to obtain the original renal cancer CT image feature map that extracts the non-linearly enhanced high-level semantic features;

[0075] The relationship among the first convolutional layer, the second convolutional layer, and the third convolutional layer is:

[0076] The number of convolutional kernels in the third convolutional layer = 2 * the number of convolutional kernels in the second convolutional layer = 4 * the number of convolutional kernels in the first convolutional layer;

[0077] Preferably, the feature fusion module includes an element-wise addition module, a fourth feature extraction module, and a fifth feature extraction module;

[0078] The element-wise addition module is used to obtain a feature-fused renal cancer CT image feature map based on the original renal cancer CT image feature map for extracting non-linearly enhanced high-level semantic features and the enhanced renal cancer CT image feature map for extracting non-linearly enhanced high-level semantic features;

[0079] The fourth feature extraction module includes a fourth convolutional layer and a fourth activation layer;

[0080] The fourth convolutional layer is used to obtain fused low-level features based on the feature-fused renal cancer CT image feature map, so as to obtain a renal cancer CT image feature map for extracting fused low-level features;

[0081] The fourth activation layer is used to obtain non-linearly enhanced fused low-level features based on the renal cancer CT image feature map for extracting fused low-level features, so as to obtain a renal cancer CT image feature map for extracting non-linearly enhanced fused low-level features;

[0082] The fifth feature extraction module includes a fifth convolutional layer and a fifth activation layer;

[0083] The fifth convolutional layer is used to obtain fused high-level features based on the renal cancer CT image feature map for extracting non-linearly enhanced fused low-level features, so as to obtain a renal cancer CT image feature map for extracting fused high-level features;

[0084] The fifth activation layer is used to obtain non-linearly enhanced fused high-level features based on the renal cancer CT image for extracting fused high-level features, so as to obtain a renal cancer CT image feature map for extracting non-linearly enhanced fused high-level features;

[0085] The fully connected layer is used to obtain the probability of the renal cancer category based on the renal cancer CT image feature map for extracting non-linearly enhanced fused high-level features, so as to complete the classification of the pathological upgrading of renal cancer.

[0086] Specifically, the network in the classification method for cross-modal renal cancer pathological upgrading based on the Dual-Input Fusion Network (DIF-Net) in this embodiment consists of two branches, namely the Upper branch and the Bottom branch, which respectively process two sets of images: the original renal cancer CT images and the enhanced renal cancer CT images. Each set of images contains two types of renal cancer CT images: high-grade and low-grade. These two sets of images are converted into input features and respectively input into the dual-input fusion network. The dual-input fusion network contains two branches, and each branch has three feature extraction modules for extracting multi-level features of the images. Then, the features extracted by these two branches are fused, and the fused features are further extracted through two feature extraction modules, and finally, classification is completed through a classifier.

[0087] Specifically, high-grade and low-grade are conventional terms in pathology. The pathological grading of renal cell carcinoma is mainly evaluated by histological examination according to the morphological characteristics of tumor cells. The most commonly used grading system currently is the WHO / ISUP grading system. The grading criteria of this system include:

[0088] · Grade 1: The tumor cells are well-differentiated, with small cell nuclei, and the nucleoli are invisible or very small.

[0089] · Grade 2: The tumor cells have moderate differentiation, with larger cell nuclei, visible nucleoli, and irregular cell arrangement.

[0090] · Grade 3: The tumor cells are poorly differentiated, with significantly enlarged cell nuclei, obvious nucleoli, and more disordered cell arrangement.

[0091] · Grade 4: The tumor cells are highly undifferentiated, with large nuclei and large nucleoli, and the cell arrangement is chaotic, often accompanied by highly malignant features such as necrosis and vascular invasion.

[0092] The grading system of the WHO / ISUP grading system can be simplified into a two-level classification: Grade 1 and Grade 2 are classified as low-grade, and Grade 3 and Grade 4 are classified as high-grade.

[0093] The input of the dual-input fusion network is two sets of images from different sources, including high-grade renal cancer CT images and low-grade renal cancer CT images:

[0094] The Bottom branch inputs the CT image of renal cell carcinoma after image enhancement using the CLAHE algorithm. It mainly improves the contrast of the image through histogram equalization in the local area and avoids over-enhancement through contrast limitation. By applying the CLAHE algorithm to the image, the local details in the image can be made more obvious, thus helping the model better identify the features that may be ignored in the low-contrast areas of the original image. This channel focuses on enhancing the local contrast of the image, providing more detailed information, and helping to improve the model's learning ability in terms of details.

[0095] The Upper branch inputs the original CT image of renal cell carcinoma to preserve the original information of the image and avoid losing useful features that may be lost during the preprocessing process. The original image contains the overall information of the image. Although the image enhanced by the CLAHE algorithm may be clearer in details, some original global features may change after being processed by the CLAHE algorithm. Therefore, retaining the original image input can ensure that the model does not lose these important information.

[0096] Then, the features extracted from the two branches are fused, and feature extraction is performed on the fused features. Finally, high-grade and low-grade classification of the CT image of renal cell carcinoma is achieved. Such a model design can simultaneously utilize the advantages of the CLAHE-enhanced image and the original image, thereby improving the recognition ability for different types of features, making the performance of the model more stable, especially when facing different image qualities, it can adapt to more changes.

[0097] The system of the classification method for cross-modal pathological upgrading of renal cell carcinoma based on a dual-input fusion network in this embodiment has three modules: the renal cell carcinoma CT image enhancement module based on the CLAHE algorithm, the dual-input fusion network model module, and the feature fusion and loss function calculation module.

[0098] Renal cell carcinoma CT image enhancement module based on the CLAHE algorithm:

[0099] Specifically, CLAHE is an image enhancement technique that improves contrast by performing local histogram equalization on images, and is particularly suitable for images with uneven illumination or low contrast. The algorithm first divides the image into multiple non-overlapping sub-blocks of equal size. For each small block, the algorithm calculates its grayscale histogram separately. This means that each small block has its own grayscale distribution statistics, rather than calculating a global histogram for the entire image as in traditional histogram equalization. After obtaining the histogram of each small block, CLAHE equalizes the histogram of each small block. It redistributes the grayscale values of the pixels within each small block, making the histogram more uniform across the entire grayscale range. To avoid the over-enhancement problem of traditional adaptive histogram equalization, CLAHE sets a contrast limit threshold, redistributes the pixels exceeding the threshold to other gray levels, and then equalizes the corrected histogram. To eliminate possible boundary artifacts between sub-blocks, CLAHE performs bilinear interpolation on the results of adjacent sub-blocks to make the transition smoother, and finally combines all the processed sub-blocks into a complete enhanced image. In this way, it can effectively enhance the contrast of renal cell carcinoma CT images, while retaining the local detail information of the images, reducing the problems of noise amplification and over-enhancement, and thus obtaining an enhanced image with better deep features of the image.

[0100] Figure 3a and Figure 3b Shown are the histogram changes of the original renal cancer CT image and the renal cancer CT image after image enhancement by CLAHE. Here, the histogram is a normalized histogram. To more clearly see the change in the pixel proportion of the kidney and its surrounding area, in this embodiment, the original histogram is truncated, and the maximum value of the ordinate is truncated to 0.01.

[0101] The grayscale histogram of the original image shows typical grayscale distribution characteristics. There is an obvious peak in the left area of low intensity in the background area, and the soft tissue and the kidney are mainly concentrated in the middle area of the histogram. This is because the grayscale values of these tissues are similar in the CT image, resulting in multiple overlapping peaks in the histogram. In addition, the right area of the histogram corresponds to high-density structures, with relatively low pixel values, but reflects the distribution of denser tissues.

[0102] After being processed by CLAHE, the histogram distribution of the enhanced image significantly tends to be more uniform. Although the peak in the left area still exists, its distribution is more uniform due to the enhanced contrast; in the middle area, the peak range of the histogram becomes wider and the distribution is smoother; for the right area, its peak characteristics are more prominent and easier to distinguish.

[0103] In summary, the CLAHE algorithm can effectively enhance the contrast between soft tissues and the kidney, enabling tissues with similar grayscale values in the original image to be more clearly distinguished in the enhanced image. Meanwhile, the fine structures within the kidney and the potentially diseased areas exhibit more obvious differences in the enhanced histogram. This image enhancement algorithm provides stronger support for the pathological grading and classification of renal cancer.

[0104] Dual-input fusion network model module:

[0105] Specifically, to address the problems of insufficient ability to capture detailed features, lack of global feature representation, and excessive computational overhead in current algorithms, this embodiment constructs a dual-input fusion network, as Figure 4 shown. The dual-input fusion network model sets the input image as where H is the image height, W is the width, and C is the number of channels, denoting the set of real numbers; and the input image size is set to 227×227×3. Each group of images extracts features through their respective independent convolutional neural network branches. To fully extract the detailed features of renal cancer CT images at different grades, three feature extraction modules containing combinations of convolution and pooling, namely Block1, Block2, and Block3 in the figure, are designed in each branch, which are responsible for feature extraction at different levels. The specific steps are as follows:

[0106] Step 1: Low-level local feature extraction

[0107] The input image I of the model first passes through Block1, whose main function is to extract the low-level local features of the image I through the first convolutional layer, and then downsample the features through the pooling layer to reduce the computational amount and parameter complexity. The specific calculation of this process is as follows.

[0108] First, the input image I passes through the first convolutional layer, and its calculation formula is formula (1).

[0109]

[0110] In the formula: both i and j are parameters for the positions of pixel points in the feature map, C is the number of input channels, and c is the channel index; N is the size of the convolutional kernel, and both m and n are index variables for traversing the elements of the convolutional kernel, both being indices that vary within the range [0, N - 1], where m is used to traverse each row of elements of the convolutional kernel, and n is used to traverse each column of elements of the convolutional kernel; is used to traverse each value in the convolutional kernel, l is the output channel index, and assuming the number of convolutional kernels is k, then 1 ≤ l ≤ k. represents the output of the feature map after the first layer of convolution at the position (i, j) in channel l, and I i+m,j+n,c represents the pixel value of the input image I at the position (i + m, j + n) in channel c, Represents the convolution kernel of the first layer, which represents the bias term of the l-th channel in the first layer. In this layer, the algorithm sets the convolution kernel size to 3×3, the number of convolution kernels to 8, and the stride to 1.

[0111] Then, after the input image I undergoes the first-layer convolution operation, features are obtained. At this time, the ReLU activation function is connected for non-linear transformation. This calculation is shown in formula (2).

[0112]

[0113] where max represents the operation of taking the maximum value, represents the output of the activated feature map at the l-th channel and position (i, j). This step aims to introduce non-linear factors into the network model, enabling the neural network to fit more complex functions and thus improving the network's ability to represent data.

[0114] After that, in this embodiment, the activated output features are input into the max-pooling layer for downsampling operation, and this operation is shown in formula (3).

[0115]

[0116] where f is a parameter representing the size of the pooling window, s represents the stride, u and v are index variables for traversing the pooling window forever, which are indices varying within the range [0, f - 1] and are used to traverse each element within the pooling window, represents the output of the feature map after max-pooling at the l-th channel and position (i, j). is the input feature map in the corresponding local window elements of the output feature map. We set the window size of the max-pooling layer to 2×2 and the stride to 2, so that the size of the feature map after pooling is halved, reducing the computational cost.

[0117] Step 2: Intermediate feature extraction

[0118] After passing through Block1, the model can extract low-level local features of the renal cancer CT image. At this time, in this embodiment, the features are input into Block2 to extract higher-level intermediate features, and at the same time, continue to downsample through the max-pooling layer to further reduce the computational overhead.

[0119] Block2 still consists of a convolutional layer, a ReLU activation function, and a pooling layer. The convolution kernel size is still 3×3. In order to extract more features, the model increases the number of convolution kernels to 16. In addition, the feature map obtained after the second-layer convolution is defined as The features after passing through the activation function are defined as The features after passing through the second max - pooling layer are defined as The calculation processes of these three features are the same as those in formula (1), formula (2), and formula (3).

[0120] Step 3: Advanced feature extraction

[0121] After that, the features obtained from Block2 are input into Block3 to extract the high - level features of the image and capture more abstract patterns and semantic information. This layer consists of only one convolutional layer and one ReLU activation function, and the number of convolutional kernels is increased to 32. In addition, the features obtained after passing through the third convolutional layer are defined as The features after passing through the activation function are defined as Its calculation process is the same as that of formula (1) and formula (2).

[0122] After the feature extraction in the above three steps, the network can extract the deep - level detailed features of the input image, laying a foundation for subsequent classification.

[0123] Module 3: Feature fusion and loss calculation

[0124] To better solve the problem of insufficient global feature representation, the present invention introduces a feature fusion mechanism in the dual - input fusion network, aiming to enhance the model's comprehensive processing ability for local details and original information of the image. Specifically, the dual - input fusion network contains two branches. One branch focuses on enhancing the local details of the image, while the other branch retains the original information of the image. The output features of the two branches are fused by element - wise addition, thus achieving true feature fusion. The fused features are further subjected to feature extraction and finally classified through a fully - connected layer.

[0125] This design makes full use of the complementarity between the original image and the pre - processed image, enabling the network to comprehensively learn image features from multiple perspectives. In this way, not only the problem of insufficient global feature representation is solved, but also the performance of the classification task is significantly improved. The specific processes of feature fusion and classification are as Figure 4 shown.

[0126] First: Element - wise addition feature fusion

[0127] The model names the features of the original CT image of renal cancer extracted by the Upper branch as The features of the enhanced CT image of renal cancer extracted by the Bottom branch as At this time, the sizes of the feature maps obtained by the two branches are both 56×56×32, and then their feature fusion is performed, as shown in formula (4).

[0128]

[0129] Among them, represents the fused feature at position (i, j) in channel l. In Figure 5 and Figure 6 , in this embodiment, the elements at the red positions in the first channel of the feature maps obtained from the two branches are fused, and the value after fusion is the element corresponding to the red position in the first channel of the fused feature. The size of the feature map after fusion is still 56×56×32.

[0130] Secondly: Feature extraction of the fused feature

[0131] After feature fusion, the model further extracts the features of the fused feature through two feature extraction modules (Block4 and Block5). Both of these two feature extraction modules consist of a convolutional layer and a ReLU activation function. In Block4, the number of convolution kernels is 64, and the feature obtained after passing through the convolutional layer in Block4 is defined as The feature after passing through the activation function is defined as In Block5, the number of convolution kernels is 128, and the feature obtained after passing through the convolutional layer in Block5 is defined as The feature after passing through the activation function is defined as This is also the final feature of the CT image extracted by the entire network. In this embodiment, it is renamed as F end . The calculation processes of the convolutional layer and the activation function in these two modules are the same as those in formula (1) and formula (2).

[0132] After the feature extraction through these two modules, the network can extract more detailed features after the fusion of the CT image features from different sources. After feature extraction, in this embodiment, F end is input into the fully connected layer for the final classification operation. The calculation of this process is as shown in formula (5).

[0133] F fc = W fc ·F end + b fc (5)

[0134] Among them, F fc represents the final classification output, W fc is the weight matrix of the fully connected layer, and b fc is the bias vector of the fully connected layer.

[0135] Finally, the difference between the prediction result and the true label is calculated through the loss function to optimize the network parameters.

[0136] To visually demonstrate the effectiveness of feature fusion, in this embodiment, comparative analysis is carried out through heatmaps. As Figure 5 shown, the features extracted by the Upper branch are mainly concentrated in the local areas of the image, especially in and around the center. The values in these areas of the heatmap are higher, indicating that this branch is more sensitive to the features in these areas. However, overall, the feature extraction ability of the Upper branch still has certain limitations. In contrast, the heatmap of the Bottom branch shows that its features are mainly concentrated in the internal areas of the image (such as key parts like the kidneys), and the heat values are relatively high. This indicates that the Bottom branch pays more attention to the features of internal tissues and can effectively capture the parts that the Upper branch does not focus on, verifying the effectiveness of enhancing image details through the CLAHE algorithm in this embodiment.

[0137] The fused heatmap combines the advantages of the two branches. Its overall characteristics are closer to those of the Bottom branch, and at the same time, it shows relatively high heat values in the center, edges, and some internal areas of the image. It should be noted that in some areas, the feature values of the fused heatmap are higher than those of a single branch, indicating that feature fusion is not just a simple superposition, but also generates a synergistic effect in some areas, significantly enhancing the response to key features. This fused feature provides more valuable information for subsequent analysis, further verifying the effectiveness of feature fusion.

[0138] Again: Loss function calculation

[0139] In this network model, the calculation of the loss function is carried out after the feature fusion of the two branches. The loss function used in this algorithm is the Cross Entropy Loss, as shown in formula (6). After the loss function calculation, the model will adjust the network weights according to the difference between the prediction result and the true label to minimize the error.

[0140]

[0141] Among them, h s represents the true label, represents the predicted probability of the model, S represents the number of classes, and s represents the class index.

[0142] Specifically, in this embodiment, an actual dataset from the Second Affiliated Hospital of Dalian Medical University is constructed. Doctors mainly collect the images of patients with inconsistent imaging and pathological staging and construct an actual hospital image dataset based on the clinically graded results based on pathology. This dataset selects the renal tumor images of the maximum section of the arterial phase, venous phase, and delayed phase of enhanced CT for each patient. This dataset is divided into high-grade images and low-grade images. There are 176 high-grade images and 293 low-grade images in the dataset. Example images of the dataset are as Figure 7 shown.

[0143] In this embodiment, the dataset is divided into a training set, a validation set, and a test set in a ratio of 7:2:1. The initial learning rate is set to 0.0001 and is dynamically adjusted during training by the Adam optimizer. The training batch size is set to 16, and the number of training epochs for the model is set to 50. Finally, the model with the highest accuracy in the validation set is saved. The experimental results of this embodiment are as Figure 8 shown.

[0144] To evaluate the performance advantages of this embodiment in the classification of renal cancer pathological grades, this embodiment is comprehensively compared and verified with traditional models and the state-of-the-art (SOTA) models, as shown in Table 1. The experimental results show that the method proposed in this embodiment is superior to other comparison models in all evaluation indicators, showing obvious performance advantages, and maintaining low overhead in terms of running time.

[0145] Table 1 Comparison experimental results with other models

[0146]

[0147] In summary, this embodiment has the following beneficial effects:

[0148] (1) This embodiment constructs an actual dataset for the pathological upgrade of renal cancer: To ensure the authenticity, clinical relevance, and non-invasive diagnosis of the data, this embodiment collects the images of patients with inconsistent imaging and pathological staging and trains the model using the actual hospital image dataset constructed based on the pathological clinical grading results. The output results of the model can assist clinicians in achieving more accurate diagnoses and promoting the progress of early detection and personalized treatment of renal cell carcinoma.

[0149] (2) This embodiment adopts an image enhancement algorithm: The Contrast Limited Adaptive Histogram Equalization (CLAHE) algorithm is used for image enhancement, effectively improving the visibility of deep lesion features in the renal cell carcinoma region and providing richer visual information for subsequent analysis.

[0150] (3) This embodiment constructs a dual-input fusion network, which fuses the deep features extracted by the CLAHE algorithm and the basic features to construct a deep learning model with good universality and adaptability to improve the automated diagnosis performance of renal cell carcinoma.

[0151] Therefore, the classification method of this embodiment can achieve automated and efficient cross-modal image classification, which is of great significance for realizing the diagnosis of renal cancer pathological upgrade.

[0152] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A classification method for cross-modal renal cancer pathological upgrading based on a dual-input fusion network, characterized in that, It includes the following steps: S1: Obtain the original renal cancer CT image, and according to the CLAHE algorithm, obtain the enhanced renal cancer CT image; S2: Establish a dual-input fusion network, which includes a Bottom branch, an Upper branch, and a feature fusion module; S3: According to the original renal cancer CT image and the Upper branch, obtain the non-linearly enhanced high-level semantic features of the original renal cancer CT image; to obtain the feature map of the original renal cancer CT image for extracting the non-linearly enhanced high-level semantic features; S4: According to the enhanced renal cancer CT image after image enhancement and the Bottom branch, obtain the non-linearly enhanced high-level semantic features of the enhanced renal cancer CT image; To obtain the feature map of the enhanced renal cancer CT image for extracting the non-linearly enhanced high-level semantic features; S5: According to the feature map of the original renal cancer CT image for extracting the non-linearly enhanced high-level semantic features, the feature map of the enhanced renal cancer CT image for extracting the non-linearly enhanced high-level semantic features, and the feature fusion module, obtain the probability that the original renal cancer CT image belongs to a high-grade renal cancer CT image and the probability that it belongs to a low-grade renal cancer CT image, so as to complete the classification of the renal cancer CT image.

2. The classification method for cross-modal renal cancer pathological upgrading based on a dual-input fusion network according to claim 1, wherein The structures of the Bottom branch and the Upper branch are the same; The Upper branch includes a first feature extraction module, a second feature extraction module, and a third feature extraction module connected in sequence at the head and tail; The first feature extraction module is used to obtain the dimension-reduced low-level significant features of the original renal cancer CT image according to the original renal cancer CT image, so as to obtain the feature map of the original renal cancer CT image for extracting the dimension-reduced low-level significant features; The second feature extraction module is used to obtain the dimension-reduced middle-level significant features of the original renal cancer CT image according to the feature map of the original renal cancer CT image for extracting the dimension-reduced low-level significant features; The third feature extraction module is used to obtain the non-linearly enhanced high-level semantic features of the original renal cancer CT image according to the feature map of the original renal cancer CT image for extracting the dimension-reduced middle-level significant features, so as to obtain the feature map of the original renal cancer CT image for extracting the non-linearly enhanced high-level semantic features.

3. The classification method for cross-modal renal cancer pathological upgrade based on a dual-input fusion network according to claim 2, wherein, The first feature extraction module includes a first convolutional layer, a first activation layer, and a first max pooling layer; The first convolutional layer is used to obtain the low-level local features according to the original renal cancer CT image, so as to obtain the feature map of the original renal cancer CT image for extracting the low-level local features; The first activation layer is used to obtain the non-linearly enhanced low-level local features according to the feature map of the original renal cancer CT image for extracting the low-level local features, so as to obtain the feature map of the original renal cancer CT image for extracting the non-linearly enhanced low-level local features; The first max pooling layer is used to obtain the dimension-reduced low-level significant features of the original renal cancer CT image according to the feature map of the original renal cancer CT image for extracting the non-linearly enhanced low-level local features, so as to obtain the feature map of the original renal cancer CT image for extracting the dimension-reduced low-level significant features.

4. A classification method for cross-modal renal cancer pathological upgrading based on a dual-input fusion network according to claim 3, characterized in that The second feature extraction module includes a second convolutional layer, a second activation layer, and a second max pooling layer; The second convolutional layer is used to obtain middle-level features based on the original renal cancer CT image feature map that extracts the dimension-reduced low-level significant features, so as to obtain the original renal cancer CT image feature map that extracts the middle-level features; The second activation layer is used to obtain non-linearly enhanced middle-level features based on the original renal cancer CT image feature map that extracts the middle-level features, so as to obtain the original renal cancer CT image feature map that extracts the non-linearly enhanced middle-level features; The second max pooling layer is used to obtain the dimension-reduced middle-level significant features of the original renal cancer CT image based on the original renal cancer CT image feature map that extracts the non-linearly enhanced middle-level features, so as to obtain the original renal cancer CT image feature map that extracts the dimension-reduced middle-level significant features.

5. A classification method for cross-modal renal cancer pathological upgrading based on a dual-input fusion network according to claim 4, characterized in that, The third feature extraction module includes a third convolutional layer and a third activation layer; The third convolutional layer is used to obtain the high-level semantic features of the original renal cancer CT image based on the original renal cancer CT image feature map that extracts the dimension-reduced middle-level significant features, so as to obtain the original renal cancer CT image feature map that extracts the high-level semantic features; The third activation layer is used to obtain non-linearly enhanced high-level semantic features based on the original renal cancer CT image feature map that extracts the high-level semantic features, so as to obtain the original renal cancer CT image feature map that extracts the non-linearly enhanced high-level semantic features.

6. The classification method for cross-modal renal cancer pathological upgrading based on a dual-input fusion network according to claim 5, wherein The relationship between the first convolutional layer, the second convolutional layer, and the third convolutional layer is: The number of convolutional kernels of the third convolutional layer = 2 * the number of convolutional kernels of the second convolutional layer = 4 * the number of convolutional kernels of the first convolutional layer.

7. A classification method for cross-modal renal cancer pathological upgrading based on a dual-input fusion network according to claim 5, characterized in that, The feature fusion module includes an element-wise addition module, a fourth feature extraction module, a fifth feature extraction module, and a fully connected layer; The element-wise addition module is used to obtain the feature-fused renal cancer CT image feature map based on the original renal cancer CT image feature map that extracts the non-linearly enhanced high-level semantic features and the enhanced renal cancer CT image feature map that extracts the non-linearly enhanced high-level semantic features; The fourth feature extraction module includes a fourth convolutional layer and a fourth activation layer; The fourth convolutional layer is used to obtain fused low-level features based on the feature-fused renal cancer CT image feature map, so as to obtain the renal cancer CT image feature map that extracts the fused low-level features; The fourth activation layer is used to obtain non-linearly enhanced fused low-level features based on the renal cancer CT image feature map that extracts the fused low-level features, so as to obtain the renal cancer CT image feature map that extracts the non-linearly enhanced fused low-level features; The fifth feature extraction module includes a fifth convolutional layer and a fifth activation layer; The fifth convolutional layer is used to obtain fused high-level features based on the renal cancer CT image feature map that extracts the non-linearly enhanced fused low-level features, so as to obtain the renal cancer CT image feature map that extracts the fused high-level features; The fifth activation layer is used to obtain non-linearly enhanced fused high-level features based on the renal cancer CT image that extracts the fused high-level features, so as to obtain the renal cancer CT image feature map that extracts the non-linearly enhanced fused high-level features; The fully connected layer is used to obtain the probability of the renal cancer category based on the renal cancer CT image feature map that extracts the non-linearly enhanced fused high-level features, so as to complete the classification of the pathological upgrade of renal cancer.