Bile duct cancer detection method
The method enhances bile duct cancer detection using Raman spectroscopy and machine learning models, addressing inefficiencies and costs of current methods by improving model training and feature extraction for accurate bile duct cancer diagnosis.
Patent Information
- Application Number
- CN202510429238.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-04-07
AI Technical Summary
Current methods for detecting bile duct cancer (CCA) are inefficient, costly, and lack sensitivity and specificity, with ERCP having low sensitivity and gene sequencing being time-consuming and expensive.
A bile duct cancer detection method using machine learning techniques on Raman spectroscopy data, employing a KPCA-LDA-SVM model and a CNN-RF model with an improved fish swarm algorithm for feature extraction and fusion, and data augmentation via GANs to enhance model training.
Improves the efficiency and accuracy of bile duct cancer detection while reducing costs, leveraging machine learning models to analyze Raman spectra for rapid and precise diagnosis.
Smart Images

Figure CN120275364A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of cancer auxiliary detection, and particularly to a method for detecting cholangiocarcinoma. Background Art
[0002] Cholangiocarcinoma (CCA) is a malignant tumor originating from bile duct epithelial cells. It has aggressive biological behavior and a relatively high cancer-related mortality rate. Surgical resection for therapeutic purposes remains the best option for achieving long-term survival. Since surgical operations usually involve resection of peripheral organs and require complex digestive tract reconstruction, misdiagnosis can lead to a high incidence of perioperative complications. Therefore, accurate detection of CCA is crucial for subsequent timely and appropriate treatment.
[0003] In the prior art, the detection methods for cholangiocarcinoma are as follows: (1) Endoscopic retrograde cholangiopancreatography (ERCP) tissue / cytology biopsy, but its sensitivity is relatively low (usually less than 50%), resulting in delayed diagnosis of the disease. (2) Bile directly contacts the bile duct lesions and is rich in metabolites secreted by the bile duct system. Therefore, bile is a potential resource for diagnosing CCA, and bile samples can be easily obtained from patients undergoing ERCP. However, at present, some scholars have explored tumor markers in bile, but there is still no tumor marker with high sensitivity and high specificity for clinical practical application. (3) "Liquid biopsy" for gene sequencing of bile, but this method is costly and time-consuming for sequencing analysis, and its diagnostic detection efficiency has not been recognized yet.
[0004] Therefore, there is a need to provide an efficient, rapid, and economical method for detecting cholangiocarcinoma. Summary of the Invention
[0005] In view of the above analysis, embodiments of the present invention aim to provide a method for detecting cholangiocarcinoma to solve the problems of low accuracy, high cost, and low efficiency of existing cholangiocarcinoma detection methods.
[0006] Embodiments of the present invention provide a method for detecting cholangiocarcinoma, including:
[0007] Obtaining a plurality of Raman spectra of the bile to be classified, and performing data preprocessing on each Raman spectrum;
[0008] Obtaining a first classification result of the bile to be classified based on the preprocessed plurality of Raman spectra and a pre-trained first cholangiocarcinoma classification model;
[0009] Obtaining a second classification result of the bile to be classified based on the preprocessed plurality of Raman spectra and a pre-trained second cholangiocarcinoma classification model, where the second cholangiocarcinoma classification model includes a feature extraction module, a feature fusion module, and a classification module;
[0010] If the first classification result is consistent with the second classification result, output the classification result of the bile to be classified, where the classification result includes: benign biliary disease or cholangiocarcinoma.
[0011] Based on a further improvement of the above method, the obtaining of multiple Raman spectra of the bile to be classified includes:
[0012] Fix the device parameters of the stimulated Raman scattering device T-SRS, and sample multiple regions of the bile to be classified to obtain multiple Raman spectra.
[0013] Based on a further improvement of the above method, the first cholangiocarcinoma classification model is the KPCA-LDA-SVM model; the KPCA-LDA-SVM model includes a kernel principal component analysis KPCA module, an LDA classification model, and a support vector machine SVM model;
[0014] The obtaining of the first classification result of the bile to be classified based on the multiple preprocessed Raman spectra and the pre-trained first cholangiocarcinoma classification model includes:
[0015] For each Raman spectrum, intercept the spectral data between 2700 cm-1 and 3100 cm-1 as the target Raman spectrum;
[0016] For each of the target Raman spectra, use the kernel principal component analysis KPCA module for data dimensionality reduction, input the dimensionality-reduced data into the LDA classification model to extract discriminant features, and input all the obtained discriminant features into the support vector machine SVM model to obtain the first classification result.
[0017] Based on a further improvement of the above method, the data preprocessing includes:
[0018] Use a Savitzky-Golay filter for noise filtering, use a polynomial fitting method to eliminate the fluorescence background, and align the spectral displacements of the samples through dynamic time warping DTW.
[0019] Based on a further improvement of the above method, the obtaining of the second classification result of the bile to be classified based on the multiple preprocessed Raman spectra and the pre-trained second cholangiocarcinoma classification model includes:
[0020] Input each preprocessed spectral data into the feature extraction module to obtain the feature vector of the Raman spectrum, input multiple feature vectors into the feature fusion module to obtain the feature fusion vector, and the feature fusion module uses feature fusion based on the self-attention mechanism;
[0021] Input the feature fusion vector into the classification module to obtain the second classification result of the bile to be classified.
[0022] Based on the further improvement of the above method, the feature extraction module adopts a CNN model, including: an input layer, a first convolutional layer, a second convolutional layer, a third convolutional layer, and a pooling layer. The first convolutional layer is a 1D convolution with a kernel width of 3, the second convolutional layer is a 1D convolution with a kernel width of 5, the third convolutional layer is a 1D convolution with a kernel width of 7, and the activation function is LeakyReLU; the classification module adopts a random forest model.
[0023] Based on the further improvement of the above method, the number of neurons in the first convolutional layer, the second convolutional layer, and the third convolutional layer in the second cholangiocarcinoma classification model, as well as the number of trees and the maximum depth in the random forest model, are obtained based on the artificial fish swarm algorithm.
[0024] Based on the further improvement of the above method, the artificial fish swarm algorithm is executed based on the following method:
[0025] S1: Initialize the parameters of the artificial fish swarm algorithm, including: initialize the fish swarm size, the maximum number of iterations, the visual range, the crowding factor, the moving step size, and the maximum number of trial times. Each fish represents the number of neurons in the first convolutional layer, the second convolutional layer, and the third convolutional layer, and the number of trees and the maximum depth in the random forest model;
[0026] S2: Simulate the behavior of the artificial fish swarm, including: in the foraging behavior, use an adaptive step size to explore, randomly generate a new position within the visual range. If the fitness value at the new position is greater than the fitness value at the current position, move one moving step size towards the new position;
[0027] Among them, the adaptive step size is:
[0028]
[0029] Step new is the updated adaptive step size, Step initial is the moving step size at initialization, ε is the decay rate, t is the current number of iterations, and T is the maximum number of iterations;
[0030] In the aggregation behavior, for each fish, calculate the lowest fitness value among all neighbor fish within the visual range of the current i-th fish, and calculate the normalized weight of the current i-th fish based on the fitness value of the current i-th fish and the lowest fitness value.
[0031]
[0032] Calculate the weighted center based on the normalized weight of each fish and the corresponding position, that is:
[0033]
[0034] where, w i is the normalized weight of the current i-th fish, Position i is the position vector of the current i-th fish, F i is the fitness value of the current i-th fish, F min is the lowest fitness value, n is the number of all neighbor fish within the field of view, j = 1, 2, 3,..., n, N is the initialized fish swarm size, i = 1, 2, 3,..., N; when j < η(t) * N, position movement is realized according to the weighted center, and η(t) is the crowding factor after dynamic update;
[0035] In the chasing behavior, for each fish, calculate the highest fitness value among all neighbor fish within the field of view of the current i-th fish. If F max / n > η(t) * N, then calculate the normalized fitness difference based on the fitness difference between the current i-th fish and the target fish, calculate the chasing weight based on the normalized fitness difference, and realize position update based on the chasing weight;
[0036] S3: If the foraging behavior, schooling behavior, and chasing behavior are not executed, then execute the random behavior;
[0037] S4: Record the current global optimal solution and the optimal fitness, and the global optimal solution is the fish at the position corresponding to the optimal fitness;
[0038] S5: Iterative optimization, repeat steps S2 - S4 until the maximum number of iterations is satisfied;
[0039] S6: Output the optimization result, and use the optimization result as the number of neurons in the first convolutional layer, the second convolutional layer, and the third convolutional layer, as well as the number of trees and the value of the maximum depth in the random forest model.
[0040] Based on the further improvement of the above method, the realizing position movement according to the weighted center includes:
[0041]
[0042] x new is the updated position, x old is the current position, η inital is the crowding factor at initialization;
[0043] The realizing position update based on the chasing weight includes:
[0044]
[0045] x new is the updated position, x old is the current position, φ is the chasing weight, xbest is the position of the neighbor fish corresponding to the highest fitness value;
[0046] The training samples of the second cholangiocarcinoma classification model are multiple Raman spectra of each bile stored and the corresponding labels, and the labels are used to represent benign or cholangiocarcinoma.
[0047] Based on the further improvement of the above method, if the number of training samples is insufficient, the following methods can be used for expansion:
[0048] Randomly shift at least one Raman spectrum of the existing bile, input the shifted Raman spectrum into the pre-trained generative adversarial network DCGAN to generate a new Raman spectrum, and replace the original Raman spectrum with the new Raman spectrum, so as to form a new set of Raman spectra corresponding to the bile;
[0049] Or,
[0050] Add noise to at least one Raman spectrum of the existing bile, input the Raman spectrum after adding noise into the pre-trained generative adversarial network DCGAN to generate a new Raman spectrum, and replace the original Raman spectrum with the new Raman spectrum, so as to form a new set of Raman spectra corresponding to the bile;
[0051] Or,
[0052] Perform linear combination on at least two Raman spectra of the existing bile, input the Raman spectrum after linear combination into the pre-trained generative adversarial network DCGAN to generate a new Raman spectrum, and replace the original Raman spectrum with the new Raman spectrum, so as to form a new set of Raman spectra corresponding to the bile.
[0053] Compared with the prior art, the present invention can at least achieve one of the following beneficial effects:
[0054] 1. The present invention provides a cholangiocarcinoma detection method, which adopts a machine learning method, based on the Raman spectrum data of the bile to be classified, and determines the classification result based on two classification models to obtain the detection information of cholangiocarcinoma. Compared with the prior art methods such as ERCP and gene sequencing, the efficiency and accuracy of cholangiocarcinoma detection are further improved, and the detection cost is reduced.
[0055] 2. The present invention provides a cholangiocarcinoma detection method, which is implemented by using the artificial fish swarm algorithm in the training process of the cholangiocarcinoma classification model, improving the training speed of the classification model, and making adaptive improvements to the foraging, clustering, and chasing behaviors of the artificial fish swarm algorithm, so that the fish swarm algorithm can find the global optimal solution faster during the optimization process, while improving the accuracy of the solution. The optimized fish swarm algorithm can also better balance global search and local search and avoid falling into local optimal solutions.
[0056] In the present invention, the above-mentioned technical solutions can also be combined with each other to achieve more preferred combination solutions. Other features and advantages of the present invention will be described in the following specification. Moreover, some advantages can be made obvious from the specification or understood by implementing the present invention. The objectives and other advantages of the present invention can be achieved and obtained from the content specifically pointed out in the specification and the drawings. Description of the Drawings
[0057] The drawings are only for the purpose of showing specific embodiments and are not considered as a limitation to the present invention. Throughout the drawings, the same reference signs represent the same components.
[0058] Figure 1 It is an exemplary diagram of a method for detecting cholangiocarcinoma in an embodiment of the present invention. Detailed Embodiments
[0059] The following will specifically describe the preferred embodiments of the present invention with reference to the drawings. The drawings form a part of this application and are used together with the embodiments of the present invention to explain the principle of the present invention, rather than to limit the scope of the present invention.
[0060] A specific embodiment of the present invention discloses a method for detecting cholangiocarcinoma, as Figure 1 shown, including:
[0061] S1: Obtain multiple Raman spectra of the bile to be classified and perform data preprocessing on each Raman spectrum.
[0062] Among them, the obtaining of multiple Raman spectra of the bile to be classified includes: fixing the device parameters of the stimulated Raman scattering device T-SRS and sampling multiple regions of the bile to be classified to obtain multiple Raman spectra.
[0063] It can be understood that bile is a complex liquid composed of cholesterol, bile salts, proteins, metabolites, etc., and its composition may be unevenly distributed due to different tumor locations, inflammation degrees or pathological stages. For example, tumor-related metabolites (such as abnormal lipids, etc.) may be enriched in certain regions. Therefore, different sampling points may cause local differences in the collected Raman spectra. In addition, the parameter settings of the T-SRS device directly affect the spectral resolution, signal-to-noise ratio and detection sensitivity, thereby resulting in spectral differences. Therefore, when obtaining Raman spectra in the present invention, it is necessary to standardize the device parameters of the stimulated Raman scattering device T-SRS, including fixing the laser power, integration time, scanning mode, etc., so as to ensure that the differences between multiple spectra mainly come from the samples themselves, and by analyzing the Raman spectra of multiple regions, the accuracy of the classification results can be further improved.
[0064] Exemplarily, at least 4 Raman spectra need to be obtained.
[0065] To further improve the data quality, after obtaining the Raman spectroscopy data, it is also necessary to process the data of each spectrum. Specifically, it includes:
[0066] Denoising processing. The Savitzky-Golay filter can be used for noise filtering, which can smooth the spectral curve and retain high-frequency features; the wavelet transform method can also be used, which can decompose the signal and remove high-frequency noise components.
[0067] Baseline correction. Polynomial fitting or Asymmetric Least Squares (ALS) is used to eliminate baseline drift, thereby eliminating the fluorescence background.
[0068] Outlier processing. Box plots or Z-score methods are used to identify and remove spectra with abnormal signal intensities (such as extreme values caused by laser power fluctuations). Abnormal values can be removed if necessary.
[0069] Spectral alignment. Dynamic Time Warping (DTW) or feature peak matching is used to align the spectral displacements of different samples, eliminating the influence of equipment drift or sample differences.
[0070] Standardization / Normalization. Z-score standardization or Min-Max normalization is performed on each Raman spectrum to eliminate the dimensional difference.
[0071] S2: Obtain the first classification result of the bile to be classified based on multiple preprocessed Raman spectra and a pre-trained first cholangiocarcinoma classification model.
[0072] Among them, the first cholangiocarcinoma classification model is the KPCA-LDA-SVM model, which includes a Kernel Principal Component Analysis (KPCA) module, an LDA classification model, and a Support Vector Machine (SVM) model.
[0073] The obtaining of the first classification result of the bile to be classified based on multiple preprocessed Raman spectra and a pre-trained first cholangiocarcinoma classification model includes:
[0074] For each Raman spectrum, the spectral data between 2700 cm-1 and 3100 cm-1 is intercepted as the target Raman spectrum;
[0075] For each of the target Raman spectra, the Kernel Principal Component Analysis (KPCA) module is used for data dimensionality reduction, and the dimensionality-reduced data is input into the LDA classification model to extract discriminant features. All the obtained discriminant features are input into the Support Vector Machine (SVM) model to obtain the first classification result.
[0076] By analyzing the Raman spectra of patient bile samples, it is found that lipids and proteins are key metabolites in tumor diagnosis and detection. Therefore, the range between 2700 cm-1 and 3100 cm- in the Raman spectrum is selected. 1Establish a first cholangiocarcinoma classification model based on the spectral data therebetween.
[0077] To complete the training of the KPCA-LDA-SVM model, the present invention collected 172 Raman spectra of 43 benign disease samples and 132 Raman spectra of 33 CCA samples, and processed each Raman spectral data using the preprocessing given in step S1, and used the 4 Raman spectral data corresponding to each bile sample (i.e., between 2700 cm- 1 and 3100 cm- 1 spectral data) and the corresponding annotation (i.e., benign bile duct disease / cholangiocarcinoma) as a training sample. For the specific implementation method of training the KPCA-LDA-SVM model, such as the setting of the cumulative variance contribution rate, the setting of parameters such as the penalty coefficient and kernel function in the SVM model, can be set according to actual needs, and the present invention does not limit it here.
[0078] If the model prediction accuracy obtained based on the above sample data is not good, data augmentation can also be performed based on the existing sample data. The augmentation methods include: randomly offsetting at least one Raman spectrum of the existing bile, inputting the offset Raman spectrum into the pre-trained generative adversarial network DCGAN to generate a new Raman spectrum, and replacing the original Raman spectrum with the new Raman spectrum to form a new set of Raman spectra corresponding to the bile; or, adding noise to at least one Raman spectrum of the existing bile, inputting the Raman spectrum with added noise into the pre-trained generative adversarial network DCGAN to generate a new Raman spectrum, and replacing the original Raman spectrum with the new Raman spectrum to form a new set of Raman spectra corresponding to the bile; or, performing a linear combination of at least two Raman spectra of the existing bile, inputting the linearly combined Raman spectrum into the pre-trained generative adversarial network DCGAN to generate a new Raman spectrum, and replacing the original Raman spectrum with the new Raman spectrum to form a new set of Raman spectra corresponding to the bile.
[0079] S3: Obtain a second classification result of the bile to be classified based on the preprocessed multiple Raman spectra and the pre-trained second cholangiocarcinoma classification model. The second cholangiocarcinoma classification model includes a feature extraction module, a feature fusion module, and a classification module.
[0080] The obtaining of the second classification result of the bile to be classified based on the preprocessed multiple Raman spectra and the pre-trained second cholangiocarcinoma classification model includes:
[0081] Input each preprocessed spectral data into the feature extraction module to obtain the feature vector of the Raman spectrum, and input multiple feature vectors into the feature fusion module to obtain a feature fusion vector. The feature fusion module uses feature fusion based on the self-attention mechanism;
[0082] Input the feature fusion vector into the classification module to obtain the second classification result of the bile to be classified.
[0083] Among them, the feature extraction module uses a CNN model, including: an input layer, a first convolutional layer, a second convolutional layer, a third convolutional layer, and a pooling layer. The first convolutional layer is a 1D convolution with a kernel width of 3, the second convolutional layer is a 1D convolution with a kernel width of 5, the third convolutional layer is a 1D convolution with a kernel width of 7, and the activation function is LeakyReLU; the classification module uses a random forest model.
[0084] The number of neurons in the first convolutional layer, the second convolutional layer, and the third convolutional layer in the second cholangiocarcinoma classification model, as well as the number of trees and the maximum depth in the random forest model, are obtained based on the artificial fish swarm algorithm, so as to improve the training efficiency of the second cholangiocarcinoma classification model.
[0085] The artificial fish swarm algorithm is executed based on the following method:
[0086] M1: Initialize the parameters of the artificial fish swarm algorithm, including: initialize the fish swarm size, the maximum number of iterations, the visual range, the crowding factor, the moving step size, and the maximum number of trial times. Each fish represents the number of neurons in the first convolutional layer, the second convolutional layer, and the third convolutional layer, the number of trees and the maximum depth in the random forest model;
[0087] M2: Simulate the behavior of the artificial fish swarm, including: in the foraging behavior, use an adaptive step size to explore, randomly generate a new position within the visual range. If the fitness value at the new position is greater than the fitness value at the current position, move one moving step size towards the new position;
[0088] Among them, the adaptive step size is:
[0089]
[0090] Step new is the updated adaptive step size, Step initial is the moving step size at initialization, ε is the attenuation rate, t is the current number of iterations, and T is the maximum number of iterations; it is realized that a larger step size is adopted in the initial stage, so that the artificial fish can quickly cover a larger search space, thereby accelerating the global search speed. As the number of iterations increases, the step size is gradually reduced, so that the artificial fish can perform more precise local search when approaching the optimal solution.
[0091] In the clustering behavior, for each fish, calculate the lowest fitness value among all neighbor fish within the visual range of the current i-th fish, and calculate the normalized weight of the current i-th fish based on the fitness value of the current i-th fish and the lowest fitness value.
[0092]
[0093] Calculate the weighted center based on the normalized weight and the corresponding position of each fish, that is:
[0094]
[0095] Where w i is the normalized weight of the current i-th fish, Position i is the position vector of the current i-th fish, F i is the fitness value of the current i-th fish, F min is the lowest fitness value, n is the number of all neighbor fish within the field of view, j = 1, 2, 3,..., n, N is the initialized fish swarm size, i = 1, 2, 3,..., N; when j < η(t) * N, perform position movement according to the weighted center, η(t) is the crowding factor after dynamic update, that is:
[0096]
[0097] x new is the updated position, x old is the current position, η inital is the crowding factor at initialization;
[0098] In the chasing behavior, for each fish, calculate the highest fitness value among all neighbor fish within the field of view of the current i-th fish. If F max / n > η(t) * N, then calculate the normalized fitness difference based on the fitness difference between the current i-th fish and the target fish, calculate the chasing weight based on the normalized fitness difference, and perform position update based on the chasing weight, that is:
[0099]
[0100] x new is the updated position, x old is the current position, φ is the chasing weight, x best is the position of the neighbor fish corresponding to the highest fitness value;
[0101] M3: If the foraging behavior, schooling behavior, and chasing behavior are not executed, then execute the random behavior;
[0102] M4: Record the current global optimal solution and the optimal fitness. The global optimal solution is the fish at the position corresponding to the optimal fitness;
[0103] M5: Iterative optimization, repeat steps M2 - M4 until the maximum number of iterations is satisfied;
[0104] M6: Output the optimized results, and use the optimized results as the number of neurons in the first convolutional layer, the second convolutional layer, and the third convolutional layer, as well as the number of trees and the maximum depth value in the random forest model.
[0105] Among them, the fitness function value in the artificial fish swarm algorithm is calculated based on the accuracy. It can be understood that the position of each fish represents a set of parameters (i.e., the number of neurons in the first convolutional layer, the second convolutional layer, and the third convolutional layer, the number of trees and the maximum depth in the random forest model). When calculating the fitness value, first construct a corresponding second cholangiocarcinoma classification model based on this set of parameters, and then use the sample data to train it. To complete the training of the second cholangiocarcinoma classification model, the present invention collected 172 Raman spectra of 43 benign disease samples and 132 Raman spectra of 33 CCA samples, and processed each Raman spectrum data using the preprocessing given in step S1. And take the 4 Raman spectrum data corresponding to each bile sample and the corresponding annotation (i.e., benign biliary disease / cholangiocarcinoma) as a training sample. For the specific implementation method of training the second cholangiocarcinoma classification model, such as the setting of training parameters such as the ratio of the training set, the test set, and the validation set, can be set according to actual needs, and the present invention does not limit it here, as long as a second cholangiocarcinoma classification model with classification function can be obtained. Then, after training it with the training set, use the test set to test it. If the test accuracy of the model obtained based on the above sample data is not good, the method of sample augmentation described in step S2 can also be used to retrain the constructed model. Finally, after obtaining the second cholangiocarcinoma classification model corresponding to this set of parameters, use the spectral data corresponding to each bile sample in the validation set as input, and count the accuracy of each sample in the validation set. If the predicted output result of this sample is consistent with the annotation of this sample in the validation set, it is correct, otherwise it is wrong, so as to obtain the accuracy of the second cholangiocarcinoma classification model corresponding to this set of parameters.
[0106] After finding the optimal set of parameters through the fish swarm algorithm, set the number of neurons in the first convolutional layer, the second convolutional layer, and the third convolutional layer in the CNN model according to this set of parameters, and set the number of trees and the maximum depth in the random forest model; then use the training sample set to train the established CNN model and random forest model, and obtain the second cholangiocarcinoma classification model after the training is completed.
[0107] S4: If the first classification result and the second classification result are consistent, output the classification result of the bile to be classified, and the classification result includes: benign biliary disease or cholangiocarcinoma.
[0108] If the first classification result and the second classification result are inconsistent, a reminder notice can be sent to inform that the classification result of the current bile to be classified cannot be obtained, and bile needs to be re-obtained for identification and classification.
[0109] Compared with the prior art, a cholangiocarcinoma detection method provided by this embodiment adopts a machine learning method. Based on the Raman spectroscopy data of bile to be classified, the classification result is determined based on two classification models to obtain the detection information of cholangiocarcinoma. Compared with methods such as ERCP and gene sequencing in the prior art, the efficiency and accuracy of cholangiocarcinoma detection are further improved, and the detection cost is reduced. During the training process of the cholangiocarcinoma classification model, the artificial fish swarm algorithm is used to improve the training speed of the classification model, and the foraging, clustering, and chasing behaviors of the artificial fish swarm algorithm are adaptively improved, enabling the fish swarm algorithm to find the global optimal solution faster during the optimization process, while improving the accuracy of the solution. The optimized fish swarm algorithm can also better balance global search and local search to avoid falling into the local optimal solution.
[0110] It is particularly worth noting that the cholangiocarcinoma detection method provided by the present invention is a process of using artificial intelligence technology to process medical information to obtain intermediate results. In the actual diagnosis process, relying on the detection results obtained by the cholangiocarcinoma detection method proposed by the present invention is only used as an intermediate result. Doctors can refer to this result during diagnosis and make a final diagnosis based on other clinical information of the patient.
[0111] Those skilled in the art can understand that all or part of the processes for implementing the methods of the above embodiments can be completed by instructing relevant hardware through a computer program, and the program can be stored in a computer-readable storage medium. Among them, the computer-readable storage medium is a disk, an optical disc, a read-only memory, or a random access memory, etc.
[0112] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention.
Claims
1. A method for detecting cholangiocarcinoma, characterized in that, Including: Obtain multiple Raman spectra of the bile to be classified, and perform data preprocessing on each Raman spectrum. Obtain the first classification result of the bile to be classified based on the preprocessed multiple Raman spectra and a pre-trained first cholangiocarcinoma classification model. Obtain the second classification result of the bile to be classified based on the preprocessed multiple Raman spectra and a pre-trained second cholangiocarcinoma classification model, where the second cholangiocarcinoma classification model includes a feature extraction module, a feature fusion module, and a classification module. If the first classification result and the second classification result are consistent, output the classification result of the bile to be classified, where the classification result includes: benign biliary diseases or cholangiocarcinoma.
2. The cholangiocarcinoma detection method according to claim 1, characterized in that, The obtaining of the multiple Raman spectra of the bile to be classified includes: Fix the device parameters of the stimulated Raman scattering device T-SRS, and sample multiple regions of the bile to be classified to obtain multiple Raman spectra.
3. A method for detecting cholangiocarcinoma according to claim 1, characterized in that, The first cholangiocarcinoma classification model is a KPCA-LDA-SVM model; the KPCA-LDA-SVM model includes a kernel principal component analysis KPCA module, an LDA classification model, and a support vector machine SVM model. The obtaining of the first classification result of the bile to be classified based on the preprocessed multiple Raman spectra and a pre-trained first cholangiocarcinoma classification model includes: For each Raman spectrum, intercept the spectral data between 2700 cm-1 and 3100 cm-1 as the target Raman spectrum. For each of the target Raman spectra, use the kernel principal component analysis KPCA module to perform data dimensionality reduction, input the reduced-dimensional data into the LDA classification model to extract discriminant features, and input all the obtained discriminant features into the support vector machine SVM model to obtain the first classification result.
4. A method for detecting cholangiocarcinoma according to claim 1, characterized in that, The data preprocessing includes: Perform noise filtering using a Savitzky-Golay filter, eliminate the fluorescence background using a polynomial fitting method, and align the spectral displacement of the samples through dynamic time warping DTW.
5. A method for detecting cholangiocarcinoma according to claim 1, characterized in that The obtaining of the second classification result of the bile to be classified based on the preprocessed multiple Raman spectra and a pre-trained second cholangiocarcinoma classification model includes: Input each preprocessed spectral data into the feature extraction module to obtain the feature vector of the Raman spectrum, input multiple feature vectors into the feature fusion module to obtain a feature fusion vector, and the feature fusion module uses feature fusion based on a self-attention mechanism. Input the feature fusion vector into the classification module to obtain the second classification result of the bile to be classified.
6. The cholangiocarcinoma detection method according to claim 5, characterized in that The feature extraction module uses a CNN model, including: an input layer, a first convolutional layer, a second convolutional layer, a third convolutional layer, and a pooling layer. The first convolutional layer is a 1D convolution with a kernel width of 3, the second convolutional layer is a 1D convolution with a kernel width of 5, the third convolutional layer is a 1D convolution with a kernel width of 7, and the activation function is selected as LeakyReLU; the classification module uses a random forest model.
7. A method for detecting cholangiocarcinoma according to claim 6, characterized in that The number of neurons in the first convolutional layer, the second convolutional layer, and the third convolutional layer in the second cholangiocarcinoma classification model, as well as the number of trees and the maximum depth in the random forest model, are obtained based on an artificial fish swarm algorithm.
8. A method for detecting cholangiocarcinoma according to claim 7, characterized in that, The artificial fish swarm algorithm is executed based on the following method: S1: Initialize the parameters of the artificial fish swarm algorithm, including: initialize the fish swarm size, maximum number of iterations, visual range, crowding factor, movement step size, maximum number of trials. Each fish represents the number of neurons in the first convolutional layer, the second convolutional layer, and the third convolutional layer, as well as the number of trees and the maximum depth in the random forest model; S2: Simulate the behaviors of the artificial fish swarm, including: in the foraging behavior, adopt an adaptive step size to explore, randomly generate a new position within the visual range. If the fitness value at the new position is greater than the fitness value at the current position, move one movement step size towards the new position; Among them, the adaptive step size is: Step new is the updated adaptive step size, Step initial is the moving step size at initialization, ε is the attenuation rate, t is the current iteration number, and T is the maximum iteration number; In the schooling behavior, for each fish, calculate the lowest fitness value among all neighboring fish within the visual range of the current i-th fish, and calculate the normalized weight of the current i-th fish based on the fitness value of the current i-th fish and the lowest fitness value; Calculate the weighted center based on the normalized weight and the corresponding position of each fish, that is: where, w i is the normalized weight of the current i-th fish, Position i is the position vector of the current i-th fish, F i is the fitness value of the current i-th fish, F min is the lowest fitness value, n is the number of all neighbor fish within the field of view, j = 1, 2, 3,..., n, N is the initialized fish swarm size, i = 1, 2, 3,..., N; when j < η(t) * N, position movement is implemented according to the weighted center, and η(t) is the crowding factor after dynamic update; In the chasing behavior, for each fish, calculate the highest fitness value among all neighbor fish within the field of view of the current i-th fish. If F max / n > η(t) * N, then calculate the normalized fitness difference based on the fitness difference between the current i-th fish and the target fish, calculate the chasing weight based on the normalized fitness difference, and update the position based on the chasing weight; S3: If the foraging behavior, schooling behavior, and following behavior are not executed, then execute the random behavior; S4: Record the current global optimal solution and the optimal fitness. The global optimal solution is the fish at the position corresponding to the optimal fitness; S5: Iteratively optimize, repeat steps S2 - S4 until the maximum number of iterations is satisfied; S6: Output the optimization result, and use the optimization result as the values of the number of neurons in the first convolutional layer, the second convolutional layer, and the third convolutional layer, as well as the number of trees and the maximum depth in the random forest model.
9. A method for detecting cholangiocarcinoma according to claim 8, characterized in that, The position movement based on the weighted center includes: x new is the updated position, x old is the current position, η inital is the crowding factor at initialization; The position update based on the following weight includes: x new is the updated position, x old is the current position, φ is the following weight, x best is the position of the neighbor fish corresponding to the highest fitness value; The training samples of the second cholangiocarcinoma classification model are multiple Raman spectra of each stored bile and the corresponding labels. The labels are used to represent benign or cholangiocarcinoma.
10. A method for detecting cholangiocarcinoma according to claim 9, characterized in that, If the number of training samples is insufficient, the following methods can be used for expansion: Randomly shift at least one Raman spectrum of the existing bile, input the shifted Raman spectrum into the pre-trained generative adversarial network DCGAN to generate new Raman spectra, and replace the original Raman spectra with the new Raman spectra, so as to form a new set of Raman spectra corresponding to this bile; Or, Add noise to at least one Raman spectrum of the existing bile, input the Raman spectrum with added noise into the pre-trained generative adversarial network DCGAN to generate new Raman spectra, and replace the original Raman spectra with the new Raman spectra, so as to form a new set of Raman spectra corresponding to this bile; Or, Perform a linear combination of at least two Raman spectra of the existing bile, input the linearly combined Raman spectrum into the pre-trained generative adversarial network DCGAN to generate new Raman spectra, and replace the original Raman spectra with the new Raman spectra, so as to form a new set of Raman spectra corresponding to this bile.
Citation Information
Patent Citations
Methods related to real-time cancer diagnostics at endoscopy utilizing fiber-optic raman spectroscopy
CN104541153A
Medical image automatic partitioning system, method and device based on multi-atlas and storage medium
CN109242865A
Health tracking device
CN113488139A
Raman spectrum classification method based on self-attention mechanism
CN115130566A
Biological sample prediction method and device and storage medium
CN115308189A