Prediction method of urinary system cancer category and evaluation method of biomarker of urinary system cancer category

By combining SERS technology and deep learning, the urine spectral characteristics are extracted using silver nanowire probes and RaNN models, and the biomarkers are evaluated in combination with SHAP algorithms, the accuracy and specificity of urinary cancer diagnosis in the prior art are solved, and efficient cancer classification and marker evaluation are achieved.

CN120408379APending Publication Date: 2025-08-01WUHAN UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510595015.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2025-04-12
Filing Date
2025-05-09
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

The existing SERS-based urinary cancer diagnosis methods have problems such as high noise, complex redundant information, diversity of biomarkers and individual differences, and deep learning models cannot fully utilize urinary SERS spectral information and cannot effectively evaluate the contribution of biomarkers.

Method used

Combining SERS technology and deep learning, silver nanowires are used as SERS probes to extract key eigenvalues of urine spectra through RaNN model, and model decisions are analyzed using SHAP algorithm to achieve prediction of urinary cancer categories and evaluation of biomarkers.

Benefits of technology

Improves the accuracy and specificity of urinary cancer diagnosis, can accurately capture key features in spectral data, quantify the contribution of biomarkers to diagnosis, and provide explainable diagnostic results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120408379A_ABST
    Figure CN120408379A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of medical artificial intelligence, and particularly discloses a prediction method of urinary system cancer categories and an evaluation method of biomarkers of the urinary system cancer categories. According to the method, a feature extraction algorithm and a Raman attention neural network are developed aiming at the SERS spectrum of urine, a healthy control group, a benign urinary disease group, a bladder cancer group, a kidney cancer group and a prostate cancer group are successfully distinguished, and the overall accuracy rate is 94.87%. In particular, the invention provides a method for evaluating biomarkers of cancer classes of the urinary system, successfully reveals that cancer diagnosis depends on differences of specific biomarker levels, and determines effective biomarkers of cancers. Research finds that most biomarkers are positively correlated with positive prediction, and the cancer specific biomarkers are particularly important, which provides important guidance for the clinical basis fuzzy problem of the current novel disease diagnosis technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of medical artificial intelligence, and particularly relates to a method for predicting urological cancer categories and a method for evaluating its biomarkers. Background Art

[0002] Urological cancers, including bladder cancer, kidney cancer, and prostate cancer, are one of the common malignant tumors worldwide, and their incidence is on the rise year by year. Traditional diagnostic methods mainly rely on imaging examinations (such as CT, MRI), urine chemical analysis, and tissue biopsies, etc. Imaging examinations can provide morphological information of tumors, but have low sensitivity to early lesions and cannot make a clear judgment on the nature of cancer. Urine chemical analysis can provide information about urine components, but its accuracy is greatly affected by factors such as biomarker concentration and individual differences of patients. Although tissue biopsies can provide relatively accurate pathological information, they are invasive, so they cannot meet the clinical needs for early, non-invasive, and accurate diagnosis.

[0003] In recent years, with the discovery of biomarkers and technological progress, biomarker-based cancer detection methods have gradually become a research hotspot. By detecting specific biomarkers in body fluids such as urine and blood, cancer can be screened and diagnosed early without invading the patient's body. Although some biomarkers (such as bladder cancer specific antigen, hyaluronic acid, prostate specific antigen, etc.) have been used for the detection of urological cancers, due to the expression of these markers being affected by multiple factors and often having overlaps and intersections, the diagnostic accuracy and specificity of single biomarkers are poor.

[0004] Surface Enhanced Raman Scattering (SERS) technology, as a highly sensitive spectral analysis method, has been widely used in the biomedical field in recent years. SERS technology can achieve precise measurement of molecular vibration information in samples with extremely low concentrations, so it is considered an ideal tool for early cancer screening. Through SERS technology, bio-molecules related to cancer in urine or blood samples can be analyzed to obtain rich spectral information, thereby achieving non-invasive and rapid diagnosis of cancer. Compared with traditional Raman spectroscopy, SERS technology greatly improves the intensity of spectral signals through the surface enhancement effect, enabling even trace bio-molecules to be detected, with extremely high sensitivity and specificity.

[0005] However, despite the great potential of SERS technology in cancer diagnosis, existing SERS-based cancer diagnosis methods still face some challenges. First, SERS spectral data contains a large amount of noise and redundant information. How to extract effective features from complex spectral data has become the key to achieving efficient diagnosis. Traditional feature extraction methods often rely on manual selection or simple mathematical models, and these methods often perform poorly when dealing with large-scale and high-dimensional data. Second, the diversity and individual differences of biomarkers affect the accuracy of the diagnostic model. How to effectively identify biomarkers closely related to cancer from a large amount of spectral data and accurately classify them is an urgent problem to be solved.

[0006] In recent years, the rapid development of deep learning technology has provided new ideas for solving these problems. The advantages of deep learning, especially in image processing and signal analysis, make it an ideal tool for processing SERS spectral data. Through deep neural network models such as convolutional neural networks (CNNs) and recurrent neural networks (RNNs), effective features can be automatically learned from complex spectral data, overcoming the limitations of manual feature selection in traditional methods. Deep learning methods can extract useful information from raw data through end-to-end training to achieve efficient classification of cancer.

[0007] However, existing deep learning-based SERS cancer diagnosis methods still have several key problems. First, most existing deep learning models are designed for other tasks and cannot fully utilize the rich biomarker information in urine SERS spectra. When used for SERS spectral recognition, the efficiency is low and the accuracy is poor. Second, existing deep learning interpretable algorithms can only provide the contribution of a single feature. However, the chemical bond compositions of different biomolecules have a very high similarity, and the corresponding biomarkers cannot be determined by a single feature. Therefore, the contributions of different biomarkers to the diagnostic results cannot be effectively evaluated, and a fine recognition of the roles of various biomarkers cannot be provided.

[0008] Therefore, developing a new type of urinary system cancer diagnosis method that combines SERS technology and deep learning, which can fully utilize the information in SERS spectral data, improve the diagnostic accuracy and specificity, and analyze the key biomarkers affecting the model decision-making, has important clinical application value. Summary of the Invention

[0009] To address the shortcomings of the existing technology, the present invention aims to provide a method for predicting the type of urinary cancer and evaluating its biomarkers. This prediction method combines surface-enhanced Raman spectroscopy (SERS) with deep learning to determine the type of urinary cancer based on the SERS spectra obtained from urine. Furthermore, the SHAP (Shapley Additive exPlanations) algorithm is used to analyze the deep learning model, and biomarkers for urinary cancers predicted by the deep learning model are screened. This invention provides immediate scientific evidence for medical decision-making, thereby improving the quality and efficiency of medical services.

[0010] The object of the present invention is achieved through the following technical solutions:

[0011] A deep learning-based prediction model for predicting urinary system cancer categories is obtained as follows:

[0012] Step 1: Using silver nanowires as SERS probes, prepare a silver nanowire solution.

[0013] Step 2: Obtain urine from healthy individuals, patients with benign diseases, bladder cancer patients, kidney cancer patients, and prostate cancer patients, and process the urine to obtain urine samples.

[0014] Step 3: measure an appropriate amount of the silver nanowire solution from step 1 and an appropriate amount of the urine sample from step 2, mix the silver nanowire solution and the urine sample thoroughly to obtain a test sample, collect SERS spectrum data of the test sample and perform preprocessing.

[0015] Step 4, extraction of key eigenvalues: for all the SERS spectral data obtained through preprocessing in step 3, the intensities corresponding to the peaks, troughs, and half-height widths of all characteristic peaks in the entire wavenumber range are extracted as key eigenvalues, that is, the intensities corresponding to the same wavenumber in different SERS spectral data are key eigenvalues, the number of key eigenvalues extracted from different SERS spectral data is the same, and the wavenumber positions and wavenumbers corresponding to the key eigenvalues extracted from each SERS spectral data are the same; and the set of key eigenvalues extracted from one SERS spectral data is denoted as X K ,

[0016] X K =[x1,x2,…,x n ] T

[0017] Wherein, K = 1, 2, 3, ..., n, and n is the total number of waves.

[0018] Step 5, Construction of the RaNN model: The RaNN model consists of an input layer, a feature transformation layer, an attention transformation layer, a fully connected layer, and an output layer in sequence. Among them, the number of neurons in the output layer is the same as the number of urinary system cancer categories.

[0019] Step 6, Training of the RaNN model, including the following steps:

[0020] S6.1. Select urine samples according to the prediction needs, and use the X K obtained in step 4 corresponding to the selected urine samples as a data set, and divide this data set into a training set and a test set according to a ratio of 8:2;

[0021] S6.2. Use the training set to train the RaNN model constructed in step 5. The number of training rounds is determined according to the accuracy of the model during the training process. When the accuracy no longer rises, stop the training.

[0022] Preferably, use the cross-entropy loss function to calculate the loss of the RaNN model for the training set prediction, use the Adam optimizer to optimize, and use a learning rate scheduler to dynamically adjust the learning rate during the training process. The RaNN model performs the training of the said training rounds, and the parameters of the model are updated through steps such as forward propagation, calculation of loss by the cross-entropy loss function, backpropagation, and optimization in each training round;

[0023] S6.3. Use the test set to test the RaNN model with updated parameters obtained by training. Compare the actual value with the prediction result of the RaNN model. When the prediction accuracy meets the requirements, obtain the RaNN prediction model, that is, the prediction model based on deep learning.

[0024] Preferably, in step 1, the preparation steps of the silver nanowire solution are as follows: (1) adding polyvinyl pyrrolidone and CuCl2 to ethylene glycol, stirring and dispersing them uniformly in an ultrasonic tank to obtain solution A; then dissolving AgNO3 in ethylene glycol to obtain solution B; then, adding the above solution A dropwise to solution B and stirring uniformly to obtain a mixed solution; (2) transferring the obtained mixed solution to a high-pressure reactor, sealing the high-pressure reactor and placing it in an oven, heating it at 160°C for 3h, and cooling it to room temperature after the reaction is completed; (3) taking 200mL of the solution in the high-pressure reactor that has cooled to room temperature; centrifuging the 200mL solution at 6000r / min for 10min, and then using The supernatant was completely removed with a pipette, 100 mL of anhydrous ethanol was added to the obtained silver nanowire precipitate, and the solution was evenly dispersed using an ultrasonic cleaner. The resulting solution was then centrifuged at 6000 r / min for 10 min. The supernatant was completely removed with a pipette, 100 mL of anhydrous ethanol was added to the obtained silver nanowire precipitate, and the solution was evenly dispersed using an ultrasonic cleaner. The above process from "centrifuging the evenly dispersed solution at 6000 r / min for 10 min" to "dispersing it evenly using an ultrasonic cleaner" was repeated three times. Finally, the obtained silver nanowire precipitate was dispersed in 500 mL of anhydrous ethanol and stored in a light-proof container for later use.

[0025] Preferably, in step 2, the process of obtaining the urine sample is: collecting urine from healthy people, patients with benign diseases, patients with bladder cancer, patients with kidney cancer and patients with prostate cancer, and centrifuging to extract the supernatant as a urine sample, wherein the centrifugation time is 15 minutes, the centrifugal speed is 3000 r / min, and the urine volume used is 1 mL; the patients with benign diseases specifically include: prostatitis, kidney stones, prostatic hyperplasia + kidney stones, prostatic hyperplasia, hydronephrosis, hydronephrosis + kidney stones, ureteral stones, hematuria, renal cysts, and urinary tract infections.

[0026] Preferably, in step 3, the specific process of collecting SERS spectral data and preprocessing is as follows: (1) thoroughly mixing 15 μL of the urine sample obtained by S2 with 30 μL of the silver nanowire solution obtained by S1 to prepare a test sample; (2) collecting SERS spectral data of the test sample; (3) preprocessing the collected SERS spectral data: performing denoising, baseline removal, standard normal transformation, smoothing and normalization in sequence to improve data quality.

[0027] Preferably, in step 4, the method for extracting the key feature value is:

[0028] For the SERS spectrum signal intensity f(x), a point x0 is a peak if and only if the following equation is satisfied:

[0029] f(x0) > f(x0 - ∈) and f(x0) > f(x0 + ∈)

[0030] where ∈ is a small interval value, indicating that the value at the x0 position is greater than the values in the left and right neighborhoods (a small window).

[0031] For the signal f(x), a point x0 is a wave trough if and only if the following formula is satisfied:

[0032] f(x0) < f(x0 - ∈) and f(x0) < f(x0 + ∈)

[0033] where ∈ is a small interval value, indicating that the value at the x0 position is less than the values in the left and right neighborhoods (a small window).

[0034] The steps for the full width at half maximum position are as follows:

[0035] Full height: The full height value is half of the maximum intensity of the wave peak:

[0036]

[0037] Intersection points: Calculate the positions of the full height values, that is, the two points that intersect with half of the wave peak intensity (one on the left side of the wave peak and the other on the right side), and these two points are the full width at half maximum positions of the characteristic peak.

[0038] Preferably, in step 5, the construction of the RaNN model includes the following steps:

[0039] Step 5.1. The feature transformation layer performs the following operations on the input X obtained from S4 and outputs z: K S5.1.1. Calculate the Pearson correlation coefficient matrix between the key feature values of each urine sample:

[0040] where x

[0041]

[0042] where x ki represents the i-th key feature value among the key feature values corresponding to the k-th urine sample, is the mean of the i-th key feature value among the key feature values corresponding to each urine sample in the training set, x kj represents the j-th key feature value among the key feature values corresponding to the k-th urine sample, is the mean of the j-th key feature value among the key feature values corresponding to each urine sample in the training set.

[0043] S5.1.2. Select the key feature values with R ij greater than the threshold q for connection; define a connection mask matrix:

[0044]

[0045] Among them, when M i,j = 1, the corresponding input key feature value x j is input into the neuron h i in the feature transformation layer; when M i,j = 0, the corresponding input key feature value x j is not input into the neuron h i .

[0046] Define the connection weight matrix:

[0047] W′ i,j = W i,j ·M i,j

[0048] Among them, W i,j is the original weight matrix of the neural network.

[0049] Then the output z of each neuron in the feature transformation layer i is calculated using the following formula:

[0050]

[0051] Among them, b i is the bias term, m is the number of features, and z i represents the content of a potential biomolecule in the urine sample.

[0052] X K inputs the key feature values into the feature transformation layer. After passing through the above feature transformation layer, the output is z, where z is a set of z i .

[0053] S5.2. The attention transformation layer performs the following operations on the output z of the feature transformation layer and outputs weighted features; the specific process is as follows:

[0054] S5.2.1. Take the output z of the feature transformation layer as the input feature, perform a linear transformation on it, and then calculate the attention weights of each linearly transformed input feature using the following formula:

[0055] attention_weights = σ(attention(z))

[0056] Among them, σ is the Sigmoid activation function, and attention(z) is the result of linearly transforming the input feature z.

[0057] S5.2.2. After obtaining the attention weights attention_weights of the input features through the processing in S5.2.1, the corresponding input features are weighted using the attention weights attention_weights to obtain the weighted feature z weighted , using the following formula:

[0058] z weighted = z · attention_weights

[0059] where z is the input feature and attention_weights is the attention weight corresponding to the input feature.

[0060] S5.3. The weighted feature z weighted is input into a fully connected layer for linear transformation to transform the weighted feature into a linear feature, using the following formula:

[0061] output1 = FC(z weighted )

[0062] where FC() is a linear operation that performs a linear transformation on the weighted feature.

[0063] The linear feature output by the fully connected layer introduces non-linearity through the ReLU activation function, and the calculation formula is as follows:

[0064] ReLU(output1) = max(0, output1)

[0065] S5.4. After the processing in S5.3, the non-linear feature x final is obtained. The non-linear feature x final is input into the output layer, and the following operations are performed:

[0066] output2 = fc_out(x final )

[0067] where fc_out() is a linear operation, and the number of its neurons is the total number of output categories (for example, 3 for a three-classification model and 5 for a five-classification model), which is used to calculate the prediction value of the urine sample to be tested for each category. Denote the result calculated by the above formula as y:

[0068] y = (y1, y2, …, y n )

[0069] where n is the total number of categories.

[0070] y is the prediction result, that is, the result y calculated by the above formula is the final output prediction result (i.e., the values corresponding to each category), and the Softmax activation function is used to calculate the probability corresponding to the prediction result:

[0071]

[0072] Among them, y i It is the prediction result of the final output corresponding to category i. The prediction result of the final output corresponding to each category is converted into a corresponding probability distribution through the Softmax activation function.

[0073] More preferably, the threshold q is tuned as a hyperparameter during the back-propagation process, and its value range is [0, 1].

[0074] Preferably, in step 6, when the output layer includes 2 neurons, the RaNN prediction model is a RaNN two-classification prediction model; when the output layer includes 3 neurons, the RaNN prediction model is a RaNN three-classification prediction model; when the output layer includes 4 neurons, the RaNN prediction model is a RaNN four-classification prediction model; when the output layer includes 5 neurons, the RaNN prediction model is a RaNN five-classification prediction model.

[0075] Based on the above prediction model, the present invention provides a method for predicting the type of urinary system cancer based on the above prediction model, comprising the following steps:

[0076] Step 1: Using silver nanowires as SERS probes, prepare a silver nanowire solution.

[0077] Step 2: obtaining a urine sample to be tested;

[0078] Step 3: measuring an appropriate amount of the silver nanowire solution from step 1 and an appropriate amount of the urine sample from step 2, respectively, mixing the silver nanowire solution and the urine sample thoroughly to obtain a test sample, collecting SERS spectrum data of the test sample and performing preprocessing;

[0079] Step 4, extraction of key eigenvalues: For all SERS spectral data obtained through preprocessing in step 3, the intensities corresponding to the peaks, troughs and half-maximum widths of all characteristic peaks in the entire wavenumber range are extracted as key eigenvalues; and the set of key eigenvalues extracted from a SERS spectral data is recorded as X K ,

[0080] X K =[x1,x2,…,x n ] T

[0081] Wherein, K = 1, 2, 3, ..., n, and n is the total number of waves.

[0082] Step 5: Input the key feature values of step 4 into the RaNN prediction model;

[0083] Step 6: Calculate and output the prediction results of urinary system cancer categories through the RaNN prediction model.

[0084] Based on the above prediction model, the present invention also provides a method for evaluating the effectiveness of urinary system cancer biomarkers, comprising the following steps:

[0085] Step 1: Using silver nanowires as SERS probes, prepare a silver nanowire solution.

[0086] Step 2: Obtain urine from healthy individuals, patients with benign diseases, and patients with urinary system cancer, and obtain urine samples through processing.

[0087] Step 3: measure an appropriate amount of the silver nanowire solution from step 1 and an appropriate amount of the urine sample from step 2, mix the silver nanowire solution and the urine sample thoroughly to obtain a test sample, collect SERS spectrum data of the test sample and perform preprocessing.

[0088] Step 4, extraction of key eigenvalues: for all SERS spectral data obtained through preprocessing in step 3, the intensities corresponding to the peaks, troughs, and half-maximum widths of all characteristic peaks in the entire wavenumber range are extracted as key eigenvalues, that is, the intensities corresponding to the same wavenumber in different SERS spectral data are key eigenvalues, the number of key eigenvalues extracted from different SERS spectral data is the same, and the wavenumber positions and wavenumbers corresponding to the key eigenvalues extracted from each SERS spectral data are the same; and the set of key eigenvalues extracted from one SERS spectral data is denoted as X K ,

[0089] X K =[x1,x2,…,x n ] T

[0090] Wherein, K = 1, 2, 3, ..., n, and n is the total number of waves.

[0091] Step 5: Use the X corresponding to the urine samples of healthy people, patients with benign diseases, and patients with a certain type of urinary system cancer extracted in step 4 K Each as a sub-dataset, the sub-datasets corresponding to healthy people and patients with benign diseases are merged to obtain 2 sub-datasets in total, and the 2 sub-datasets are divided into sub-training sets and sub-test sets in turn according to the ratio of 8:2 to obtain 2 sub-training sets and 2 sub-test sets, and then the 2 sub-training sets are merged into a training set, and the 2 sub-test sets are merged into a test set. The training set and test set are used to train the above-mentioned RaNN prediction model to obtain a RaNN binary classification model B.

[0092] Step 6: For the RaNN binary classification model B, apply the SHAP algorithm to evaluate the contribution of the wavenumber to the prediction result of the RaNN binary classification model B. The wavenumber that is positively correlated with the prediction result of positive for urinary system cancer is the positively correlated wavenumber, and the wavenumber that is negatively correlated with the prediction result of positive for urinary system cancer is the negatively correlated wavenumber.

[0093] Step 7: For a certain type of urinary system cancer, collect the SERS spectral data of the pure phase of the biomarker for the certain type of urinary system cancer in urine.

[0094] Step 8: For each biomarker of the certain type of urinary system cancer, calculate the total contribution value B of each biomarker respectively based on the positively correlated wavenumber and the negatively correlated wavenumber obtained in Step 6 i , and the calculation formula is as follows:

[0095] Biomarker contribution=∑(x i ×w i )+Σ[x j ×(-w j )]

[0096] Wherein, x i represents the characteristic peak intensity at the positively correlated wavenumber in the SERS spectrum of the pure phase urinary system cancer biomarker, w i represents the average value of the absolute value of the SHAP value of the key eigenvalue at the positively correlated wavenumber, x j represents the characteristic peak intensity at the negatively correlated wavenumber in the SERS spectrum of the pure phase urinary system cancer biomarker, w j represents the average value of the absolute value of the SHAP value of the key eigenvalue at the negatively correlated wavenumber; in B i , i is the name of the biomarker of the certain type of urinary system cancer.

[0097] Step 9: According to the total contribution value B of each biomarker obtained in Step 8 i , judge the effectiveness of the biomarker of urinary system cancer: the larger the calculated total contribution value B i of the biomarker, the higher its effectiveness;

[0098] Preferably, in Step 6, the specific process of applying the SHAP algorithm to evaluate the contribution of the wavenumber to the prediction result of the RaNN binary classification model B is as follows: First, take 20% from each of the X K included in the 2 sub-test sets in Step 5, and merge them into a test set B. For each X KFor all the key feature values included, calculate the SHAP value of each key feature value and take its absolute value; secondly, obtain the average value of the absolute values of the SHAP values calculated from all the key feature values corresponding to each wavenumber in the wavenumbers of step 4; finally, sort the wavenumbers corresponding to the obtained average values from largest to smallest; use the sorted wavenumbers as the vertical axis and the SHAP values as the horizontal axis, and plot the SHAP values calculated from all the key feature values corresponding to each wavenumber. The plotting rule is: the greater the Raman intensity of the key feature value corresponding to each SHAP value on the horizontal axis, the darker the color, and vice versa, the lighter the color; according to the obtained graph, for each row corresponding to a wavenumber, if it shows a light color on the left and a dark color on the right, it indicates that the wavenumber corresponding to this row is positively correlated with the prediction result of positive for the urinary system cancer, and conversely, if it shows a dark color on the left and a light color on the right, it indicates that the wavenumber corresponding to this row is negatively correlated with the prediction result of positive for the urinary system cancer. Accordingly, the wavenumber positively correlated with the prediction result of positive for the urinary system cancer is the positive correlation wavenumber, and the wavenumber negatively correlated with the prediction result of positive for the urinary system cancer is the negative correlation wavenumber.

[0099] Preferably, the calculation formula of the SHAP value is:

[0100]

[0101] where is the SHAP value of the key feature value x i (i.e., the contribution value of the key feature value x i to the prediction result), N is the set of all key feature values, S is the subset that does not include the key feature value x i , v(S) is the predicted value of the RaNN prediction model obtained by training with the subset S, represents the summation over all subsets S that do not include the key feature value x i , represents the probability of the subset S appearing, and [v(S∪{i}) - v(S)] represents the change in the model output after adding the key feature value x i to the existing feature subset S.

[0102] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0103] (1) The unique feature extraction algorithm can accurately capture the key features in the SERS spectral data and be used for the training of the model, greatly improving the utilization efficiency of the model for computing resources without causing a decrease in accuracy.

[0104] (2) The self-designed RaNN deep neural network fully utilizes the correlation between features in SERS spectral data, enabling it to more accurately learn the differences between the urine spectra of healthy individuals and those of patients with urinary system cancer, resulting in a higher accuracy rate for the model in the auxiliary diagnosis of urinary system cancer.

[0105] (3) The first marker association analysis algorithm solves the "black box" problem in the use of deep learning models for SERS spectral recognition and can quantify the contribution of biomarkers to the decision-making of cancer diagnosis models. Brief Description of the Drawings

[0106] Figure 1 It is a flowchart of a method for predicting urinary system cancer categories and an evaluation method for its biomarkers in the present invention.

[0107] Figure 2 It is the average SERS spectrum with standard deviation shadows of five cohorts in Example 1 and the Raman spectrum of silver nanowires.

[0108] Figure 3 It is a flowchart for predicting urinary system cancer categories.

[0109] Figure 4 It is an algorithm operation logic diagram for key feature value extraction, RaNN model, and biomarker evaluation.

[0110] Figure 5 It is a simplified diagram of key feature values extracted from the SERS spectral data of urine samples in Example 1, where the key feature x i is the intensity of the characteristic peak corresponding to the wavenumber in the SERS spectral data, and urine sample X K represents the Kth urine sample.

[0111] Figure 6 It is a schematic diagram of the feature extraction step in Example 1.

[0112] Figure 7 It is a schematic diagram of the diagnosis results of a single type of urinary system cancer in Example 2.

[0113] Figure 8 It is a schematic diagram of the collaborative diagnosis results of three types of urinary system cancers in Example 2.

[0114] Figure 9 It is a schematic diagram of the diagnosis results of invasive bladder cancer in Example 2.

[0115] Figure 10 It is a schematic diagram of the results of diagnosing urinary system cancer from hematuria patients in Example 2.

[0116] Figure 11It is a schematic diagram of the results of the (HC+BD) vs BCa model in Example 3.

[0117] Figure 12 It is a comparison chart of the SHAP analysis results of the (HC+BD) vs BCa model in Example 3.

[0118] Figure 13 It is a simplified diagram of the SHAP values corresponding to the key feature values extracted from the SERS spectral data of the urine sample in Example 1, where the urine sample X K represents the Kth urine sample.

[0119] Figure 14 It is the systematic evaluation process of biomarker contribution. Detailed implementation manners

[0120] The following Examples 1, 2, and 3 are used to further illustrate the present invention, but should not be construed as limiting the present invention. Unless otherwise specified, the technical means used in the examples are conventional means well known to those skilled in the art.

[0121] Example 1

[0122] The present invention discloses a prediction method for urinary system cancer categories, as Figure 1 shown, which specifically includes the following steps:

[0123] S1. Using silver nanowires as SERS probes, preparation of silver nanowire solution: Similar to CN115078331A, silver nanowires are synthesized by the polyol method. (1) Add 1.665 g of polyvinylpyrrolidone (PVP, molecular weight 360,000) and 0.0019 g of CuCl2 to 100 mL of ethylene glycol, stir and disperse evenly in an ultrasonic bath to obtain solution A; then dissolve 1.7 g of AgNO3 in 100 mL of ethylene glycol to obtain solution B; then, add the above solution A dropwise to solution B and stir evenly to obtain a mixed solution. (2) Transfer the obtained mixed solution to a 250 mL high-pressure reaction kettle, seal the high-pressure reaction kettle and place it in an oven, heat at 160 °C for 3 h, after the reaction is completed, cool to room temperature. (3) Take 200 mL of the solution in the high-pressure reaction kettle cooled to room temperature; centrifuge the 200 mL solution at 6000 r / min for 10 min, then use a pipette to completely remove the supernatant, add 100 mL of absolute ethanol to the obtained silver nanowire precipitate, disperse it evenly with an ultrasonic cleaner, then continue to centrifuge the dispersed solution at 6000 r / min for 10 min, and then use a pipette to completely remove the supernatant, add 100 mL of absolute ethanol to the obtained silver nanowire precipitate, disperse it evenly with an ultrasonic cleaner, repeat the process from "continue to centrifuge the dispersed solution at 6000 r / min for 10 min" to "disperse it evenly with an ultrasonic cleaner" three times, and finally disperse the obtained silver nanowire precipitate in 500 mL of absolute ethanol, store it in a light-proof container for standby.

[0124] S2. Obtaining urine samples: In the collection of clinical samples, urine from 214 healthy individuals (HC), 177 patients with benign diseases (BD), 165 patients with bladder cancer (BCa), 103 patients with kidney cancer (KCa), and 121 patients with prostate cancer (PCa) was collected, and the supernatant was extracted by centrifugation as a urine sample. Among them, the centrifugation time was 15 min, the centrifugation speed was 3000 r / min, and the urine volume used was 1 mL; the specific benign diseases of the patients were: 8 cases of prostatitis, 33 cases of kidney stones, 2 cases of prostate hyperplasia + kidney stones, 35 cases of prostate hyperplasia, 7 cases of hydronephrosis, 3 cases of hydronephrosis + kidney stones, 11 cases of ureteral stones, 71 cases of hematuria, 5 cases of renal cysts, and 2 cases of urinary tract infection.

[0125] S3. Collect SERS spectral data and perform preprocessing: (1) Mix 15 μL of the urine sample obtained in S2 with 30 μL of the silver nanowire solution obtained in S1 thoroughly to prepare a test sample; (2) Collect the SERS spectral data of the test sample: The instrument is Horiba LabRAM HR Evolution. When collecting the spectrum, the lens used is a 50x confocal lens, the laser wavelength is 532 nm, the laser power is about 13.5 mW, and the spectral acquisition range is 600 - 1800 cm -1 ; (3) Perform preprocessing on the collected SERS spectral data: Successively perform denoising, baseline removal, standard normal variate transformation, smoothing, and normalization to improve the data quality.

[0126] The average spectra with standard deviation shadows of the above five cohorts of HC, BD, BCa, KCa, and PCa are as Figure 2 shown. It is found that there is a high spectral similarity between the SERS spectra of different cohorts, making it difficult to identify differences manually, highlighting the necessity of automatic recognition by the deep learning model. In addition, the SERS spectrum of silver nanowires itself shows extremely weak intrinsic signals, indicating that the SERS spectral signals of silver nanowires themselves will not interfere with the analysis of urine samples.

[0127] After that, extract the key feature values from all spectra and use them to train and test the RaNN model. The overall process is as Figure 3 shown.

[0128] S4. Extraction of key feature values: According to the characteristics of SERS spectra, only the peak, valley, and full width at half maximum positions of a typical Raman peak have practical significance, from which information such as concentration, molecular orientation, functional groups, molecular structure, and crystallinity can be obtained. Therefore, for all spectral data (five cohorts) preprocessed in S3, the following extraction method is used to extract the intensities corresponding to the peak, valley, and full width at half maximum positions of all characteristic peaks in the entire wavenumber range as key feature values, as Figure 4 shown.

[0129] In Raman spectra, the peak corresponds to the local maximum in the signal, that is, the value at a certain position is greater than the values of its adjacent points on the left and right, which usually represents the characteristic vibration mode of certain substances in the sample. Specifically, for the signal f(x), a point x0 is a peak if and only if the following formula is satisfied:

[0130] f(x0) > f(x0 - ∈) and f(x0) > f(x0 + ∈)

[0131] where ∈ is a small interval value, indicating that the value at the position x0 is greater than the values in the left and right neighborhoods (a small window).

[0132] Similarly, a trough refers to a local minimum in a signal, that is, the value at a certain position is less than the values of its adjacent points on the left and right. For a signal f(x), a point x0 is a trough if and only if the following formula is satisfied:

[0133] f(x0) < f(x0 - ∈) and f(x0) < f(x0 + ∈)

[0134] where ∈ is a small interval value, indicating that the value at the position x0 is less than the values in the left and right neighborhoods (a small window).

[0135] The full width at half maximum (FWHM) position refers to the Raman shift at which the intensity in a Raman peak is equal to half of the peak intensity. The full width at half maximum refers to the width of the peak, which is the width where the peak intensity intersects with half of the maximum intensity. Its calculation involves the following steps:

[0136] Half height: The half-height value is half of the maximum intensity of the peak:

[0137]

[0138] Intersection points: Calculate the positions of the half-height values, that is, the two points that intersect with half of the peak intensity (one on the left side of the peak and the other on the right side), and these two points are the FWHM positions of the characteristic peak.

[0139] The total number of characteristic values in the preprocessed SERS spectral data corresponding to each urine sample is 4364. According to the above extraction method, key characteristic values are extracted from the preprocessed SERS spectral data. 468 key characteristic values are extracted from the preprocessed SERS spectral data corresponding to each urine sample (for the SERS spectral data of 780 urine samples, the total number of extracted key characteristic values is 780 * 468), and each key characteristic value corresponds to a wavenumber, that is, a total of 468 wavenumbers; within the entire wavenumber range, according to the corresponding wavenumbers from small to large, the key characteristic values are arranged in sequence. The key characteristic values extracted from the preprocessed SERS spectral data corresponding to different urine samples correspond to the same wavenumbers. As Figure 5 shown, within the entire wavenumber range of the preprocessed SERS spectral data, there are 468 wavenumbers from small to large in sequence, and the 468 key characteristic values extracted from the preprocessed SERS spectral data corresponding to each urine sample respectively correspond to the 468 wavenumbers. A schematic diagram of the extraction of key characteristic values is as Figure 6 shown. Obviously, the above extraction method of key characteristic values can accurately capture the peaks, troughs, and FWHM positions in the SERS spectral data, which is the key to realizing the auxiliary diagnosis of urinary system cancers. Moreover, after the extraction of key characteristic values, the number of characteristic values is greatly reduced, which is conducive to reducing the consumption of computing resources.

[0140] Let the set of key feature values extracted from the SERS spectral data corresponding to the K-th urine sample be denoted as X K ,

[0141] X K = [x1, x2, …, x n T

[0142] where n = 468, K = 1, 2, 3, ……, 780.

[0143] S5. Construction of the RaNN model: The RaNN model consists of an input layer (468 neurons), a feature transformation layer, an attention transformation layer, a fully connected layer, and an output layer (the number of neurons is the same as the number of classes) in sequence, as Figure 4 shown.

[0144] S5.1. The feature transformation layer performs the following operations on the X obtained from the input S4 K and outputs z:

[0145] S5.1.1. Calculate the Pearson correlation coefficient matrix between the key feature values of each urine sample:

[0146]

[0147] where x ki represents the i-th key feature value among the key feature values corresponding to the k-th urine sample, is the mean of the i-th key feature value among the key feature values corresponding to each urine sample included in the training set, x kj represents the j-th key feature value among the key feature values corresponding to the k-th urine sample, is the mean of the j-th key feature value among the key feature values corresponding to each urine sample included in the training set.

[0148] S5.1.2. Select the key feature values with a correlation greater than the threshold q (the threshold q is tuned as a hyperparameter during the backpropagation process, and its value range is [0, 1], with an initial value of 0.5) for connection; define a connection mask matrix:

[0149]

[0150] where, when M i,j = 1, the corresponding input key feature value x j is input to the neuron h i in the feature transformation layer; when M i,j = 0, the corresponding input key feature value x j is not input to the neuron h i .​

[0151] Define the connection weight matrix:

[0152] W′ i,j =W i,j ·M i,j

[0153] where W i,j is the original weight matrix of the neural network.

[0154] Then the output z of each neuron in the feature transformation layer i is calculated using the following formula:

[0155]

[0156] where b i is the bias term, m is the number of features, i.e., m = 468, and z i represents the content of a potential biomolecule in the urine sample.

[0157] X K After passing through the above feature transformation layer, the output is z i .

[0158] Since the SERS spectrum of the urine sample essentially reflects the combined contributions of the various components in the urine (such as water, urea, various proteins, and nucleic acids). On the one hand, each individual SERS spectral feature peak usually represents the overlapping contributions of multiple components. On the other hand, when the same molecule contributes to multiple wavenumbers simultaneously, there is usually a correlation between the Raman intensities of these wavenumbers. This complexity highlights the necessity of processing the key eigenvalues through the feature transformation layer to analyze the interdependencies between the key eigenvalues and improve the interpretability and usability of the data in downstream tasks.

[0159] S5.2. The attention transformation layer performs the following operations on the output z of the feature transformation layer to obtain the output features. The specific process is as follows:

[0160] S5.2.1. Take the output z of the feature transformation layer as the input feature, perform a linear transformation on it, and then calculate the attention weights of each linearly transformed input feature using the following formula:

[0161] attention_weights = σ(attention(z))

[0162] where σ is the Sigmoid activation function and attention(z) is the result of linearly transforming the input feature z.

[0163] S5.2.2. Obtain the attention weights attention_weights of the input features through the processing in S5.2.1, and then use the attention weights attention_weights to weight the corresponding input features to obtain the weighted feature z weighted , using the following formula:

[0164] z weighted = z · attention_weights

[0165] where z is the input feature and attention_weights is the attention weight corresponding to the input feature.

[0166] The essence of urine biopsy is to detect the content changes of specific biomolecules in urine. To highlight these differences, an attention mechanism is introduced to ensure that the model can focus on the classification of the most informative key features associated with biomolecules, enhance the model's ability to capture cohort-specific biochemical differences, and thus improve its diagnostic accuracy and interpretability.

[0167] S5.3. Input the weighted feature z weighted into the fully connected layer for linear transformation to transform the weighted feature into a linear feature, using the following formula:

[0168] output1 = FC(z weighted )

[0169] where FC() is a linear operation that performs a linear transformation on the weighted feature.

[0170] The linear feature output by the fully connected layer introduces nonlinearity through the ReLU activation function, and the calculation formula is as follows:

[0171] ReLU(output1) = max(0, output1)

[0172] S5.4. After the processing in S5.3, obtain the nonlinear feature x final , and input the nonlinear feature x final into the output layer to perform the following operations:

[0173] output2 = fc_out(x final )

[0174] where fc_out() is a linear operation, and the number of its neurons is the total number of output categories (for example, 3 for a three-classification model and 5 for a five-classification model), which is used to calculate the predicted value of the urine sample to be tested for each category. Denote the result calculated by the above formula as y:

[0175] y = (y1, y2, …, y)

[0176] Among them, n is the total number of categories.

[0177] y is the prediction result, that is, the result y calculated by the above formula is the finally output prediction result (i.e., the values corresponding to each category), and the Softmax activation function is used to calculate the probability corresponding to the prediction result:

[0178]

[0179] Among them, y i is the finally output prediction result corresponding to category i, and the finally output prediction result corresponding to each category is converted into a corresponding probability distribution through the Softmax activation function.

[0180] S6. Training of the RaNN model: (1) Select urine samples according to the prediction needs, and use the X corresponding to the selected urine samples obtained in S4 K as a data set, and divide this data set into a training set and a test set according to a ratio of 8:2; (2) Use the training set to train the RaNN model constructed in S5. The number of training rounds is determined according to the accuracy rate of the model during the training process. When the accuracy rate no longer rises, the training stops. Among them, the cross-entropy loss function is used to calculate the loss of the RaNN model for the training set prediction (i.e., the difference between the predicted probability distribution and the true probability distribution), the Adam optimizer is used for optimization, and the learning rate scheduler is used to dynamically adjust the learning rate during the training process. The RaNN model is trained in multiple said training rounds, and the parameters of the model are updated through steps such as forward propagation, calculation of loss by the cross-entropy loss function, backpropagation, and optimization in each training round; (3) Use the test set to test the RaNN model with updated parameters obtained by training, compare the actual value with the prediction result of the RaNN model, and when the prediction accuracy meets the requirements, obtain the RaNN prediction model.

[0181] Example 2

[0182] Apply the obtained RaNN prediction model to the auxiliary diagnosis of bladder cancer, renal cancer, or prostate cancer respectively, that is, distinguish the BCa, KCa, and PCa cohorts from the HC cohort and the BD cohort, namely three groups of tasks: BCa vs BD vs HC, KCa vs BD vs HC, and PCa vs BD vs HC. For the BCa vs BD vs HC group task, the set X of key feature values extracted from the SERS spectral data corresponding to the urine samples of BCa, BD, and HC KRespectively as 1 sub-dataset, 3 sub-datasets are obtained. The sub-datasets are divided into sub-training sets and sub-test sets according to the ratio of 8:2. Then the 3 sub-training sets are combined into one training set, and the 3 sub-test sets are combined into one test set. The obtained BCa-RaNN prediction model demonstrates its strong application potential in cancer diagnosis. Similarly, for the KCa vs BD vs HC and PCa vs BD vs HC group tasks, the KCa-RaNN prediction model and the PCa-RaNN prediction model are obtained respectively. As Figure 7 shown, during the training process, the cross-entropy loss rapidly decreases with the increase of training iteration times and finally tends to be stable, indicating that the model is continuously optimizing its internal parameters, thereby gradually improving its ability to accurately identify cancer categories. The subsequent increase in accuracy further supports this. After sufficient training and optimization, the BCa-RaNN prediction model, the KCa-RaNN prediction model, and the PCa-RaNN prediction model achieved high diagnostic accuracies of 96.40%, 94.95%, and 96.08% respectively in the diagnosis tasks of BCa, KCa, and PCa. These results show that the RaNN prediction model of Example 1 can effectively capture and analyze complex SERS spectral features. The BCa-RaNN prediction model, the KCa-RaNN prediction model, and the PCa-RaNN prediction model are all RaNN three-classification models.

[0183] Since the biomarkers of different types of urinary system cancers may overlap, diagnosing a single type of cancer alone may result in the test samples being judged as multiple cancers simultaneously. To address this potential diagnostic confusion, X corresponding to the urine samples of HC, BD, BCa, KCa, and PCa extracted by S4 K respectively as 1 sub-dataset, 5 sub-datasets are obtained. The sub-datasets are divided into sub-training sets and sub-test sets according to the ratio of 8:2. Then the 5 sub-training sets are combined into one training set, and the 5 sub-test sets are combined into one test set. The obtained training set and test set are used for training and testing to obtain a RaNN five-classification model. As Figure 8As shown, the training process of the RaNN five-classification model went through 300 rounds in total. During this process, the loss rate of the RaNN five-classification model on the training set gradually decreased and then stabilized, while the accuracy rate gradually increased and then stabilized. Finally, in the confusion matrix, the average accuracy rate for predicting the five types of HC, BD, BCa, KCa, and PCa was 94.87%. This shows that the RaNN prediction model in Example 1 can not only achieve a high accuracy rate, but also maintain a robust performance under different dataset partitions. Moreover, since the RaNN five-classification model can classify and predict BCa, KCa, and PCa, the RaNN five-classification model can distinguish these three types of cancer, that is, it can initially determine the onset area of urinary system tumors.

[0184] Prediction of Invasive Bladder Cancer

[0185] Invasive bladder cancer refers to cancer cells that have invaded the deep layer of the bladder wall and usually require more aggressive treatment methods such as surgical resection, radiotherapy, or chemotherapy. Diagnosing whether bladder cancer is invasive cancer has important clinical significance because it helps to evaluate the severity of the disease and provides a basis for selecting appropriate treatment strategies. Therefore, the X corresponding to the urine samples of BCa extracted in S4 K is divided into two sub-datasets according to invasive and non-invasive, and the invasive sub-dataset and the non-invasive sub-dataset are sequentially divided into a sub-training set and a sub-test set according to a ratio of 8:2. The two obtained sub-training sets are combined into a training training set, and the two obtained sub-test sets are combined into a test test set. Then, a RaNN binary classification model is constructed and trained according to S5 and S6. The results are as Figure 9 shown. The model training process went through 60 rounds in total, and the accuracy rate generally increased and then stabilized during the training process. Finally, in the confusion matrix, the overall accuracy rate for invasive and non-invasive predictions was 100.00%. This shows that the RaNN prediction model of the present invention also has a very high accuracy rate for identifying whether bladder cancer is invasive bladder cancer.

[0186] Predicting BCa, KCa, and PCa Based on SERS Spectral Data of Urine Samples with Hematuria Symptoms

[0187] Patients with urinary system cancer often have hematuria symptoms. Research shows that the presence of red blood cells and their lysates in urine will significantly interfere with the detection of biomarkers. For the X corresponding to the urine samples of the five categories of HC, BD, BCa, KCa, and PCa after extraction in S4 K , use the X corresponding to the urine samples with hematuria symptoms in the urine samples of BD (confirmed by hospital diagnosis whether there are hematuria symptoms) KAs the Hematuria dataset, use X corresponding to urine samples with hematuria symptoms in urine samples of BCa K As the new BCa dataset, use X corresponding to urine samples with hematuria symptoms in urine samples of KCa K As the new KCa dataset, use X corresponding to urine samples with hematuria symptoms in urine samples of PCa K As the new PCa dataset, divide the Hematuria dataset, the new BCa dataset, the new KCa dataset, and the new PCa dataset into training sets and test sets according to the ratio of 8:2 respectively. Then combine the obtained training sets into a total training set and combine the obtained test sets into a total test set. Use the total training set and the total test set to construct and train a RaNN four-classification model according to S5 and S6. The results are as Figure 10 shown. During the training process, a total of 100 rounds are experienced. The loss rate of the RaNN four-classification model on the training set decreases overall and then remains stable, and the accuracy rate increases overall and then remains stable. Finally, in the confusion matrix, the overall accuracy rate predicted by the RaNN four-classification model is 98.00%. This shows that the RaNN prediction model in Example 1 will not be seriously interfered by the hematuria phenomenon when diagnosing urinary system cancers.

[0188] Example 3

[0189] A method for evaluating biomarkers related to urinary system cancers

[0190] Analyzing the contribution of key features is crucial for understanding cancer category prediction models because the analysis results suggest potential biomarkers. Specifically, it includes the following steps:

[0191] Step 1: Use the X corresponding to the urine samples of HC, BCa, and BD extracted in S4 K Each as 1 sub-dataset, combine the sub-datasets corresponding to HC and BD, and a total of 2 sub-datasets are obtained. Divide these 2 sub-datasets into sub-training sets and sub-test sets in sequence according to the ratio of 8:2, obtaining 2 sub-training sets and 2 sub-test sets. Then combine the obtained 2 sub-training sets into a training set and combine the 2 sub-test sets into a test set. Use the training set and the test set to construct and train according to S5 and S6 to obtain a RaNN binary classification model B((HC + BD) vs BCa), that is, a non-bladder cancer vs bladder cancer classification model. The results are as Figure 11 shown. During the model training process, a total of 100 rounds are experienced. During the training process, the loss rate of the model on the training set decreases overall and then remains stable, and the accuracy rate increases overall and then remains stable. Finally, as shown in the confusion matrix, the overall accuracy rate for non-bladder cancer and bladder cancer prediction is 95.51%.

[0192] Step 2: For the RaNN binary classification model B, the SHAP algorithm is respectively applied to evaluate the contribution of wavenumbers to the prediction results of the RaNN binary classification model. As Figure 12 shown. To understand which wavenumbers contribute the most to the prediction results of the model, first, calculate the SHAP value of each key feature value x i in the test set - for the RaNN binary classification model B of (HC + BD) vs BCa, 20% of X K corresponding to the urine samples of HC, 20% of X K corresponding to the urine samples of BD, and 20% of X K corresponding to the urine samples of BCa are used as a test set B; then, sort them according to the average value obtained from the absolute values of the SHAP values calculated for all key feature values corresponding to each wavenumber as described in S4 of Example 1. The specific process is as follows: First, take the absolute value of the SHAP values corresponding to all key feature values in the above test set; then, according to each wavenumber described in S4 of Example 1, calculate the average value of the absolute values of all SHAP values corresponding to each wavenumber, and finally, sort the corresponding wavenumbers from largest to smallest according to the obtained average value. Figure 12 For the left and right figures of, in the vertical axis, the corresponding average values of the wavenumbers decrease from top to bottom. In the horizontal axis, with the vertical line at SHAP value = 0 as the boundary, for each row, the wavenumbers corresponding to the light - colored on the left and dark - colored on the right SHAP values are the wavenumbers positively correlated with the positive prediction (i.e., the prediction result is BCa), and the wavenumbers corresponding to the dark - colored on the left and light - colored on the right SHAP values are the wavenumbers negatively correlated with the positive prediction.

[0193] The calculation formula of the SHAP value (see Figure 4 ) is:

[0194]

[0195] where is the SHAP value of the key feature value x i (i.e., the contribution value of the key feature value x i to the prediction result), N is the set of all key feature values, S is the subset that does not include the key feature value x i , υ(S) is the predicted value of the RaNN prediction model obtained by training with the subset S, represents the summation over all subsets S that do not include the key feature value x i , represents the probability of the subset S appearing, and [v(S∪{i}) - v(S)] represents the change in the model output after adding the key feature value x i to the existing feature subset S.

[0196] ​Figure 13 It is a simplified diagram of the SHAP value calculated from the key feature values extracted from the SERS spectral data corresponding to the urine sample in Example 1 according to the above calculation formula. The SHAP value can be used not only for the local interpretation of a single prediction (e.g., why a certain sample is predicted as class A), but also for global interpretation (e.g., which wavenumbers are the most important in all samples). When used for local interpretation, the SHAP value can assign the contribution degree of each key feature x i to each individual predicted value, helping to understand why the model makes a certain specific prediction. When used for global interpretation, by calculating the SHAP value of each key feature value x i , it can be summarized which wavenumbers have the greatest impact on the decision-making of the prediction model, contributing to the identification of the key features important for the prediction model across the entire dataset.

[0197] Analyzing the contribution of the key feature values corresponding to the wavenumbers related to biomarkers not only helps to reveal the potential mechanisms driving the prediction, but also establishes a key connection between the wavenumbers in the SERS spectral data and the known biomarkers of a cancer category. This connection enhances the biological relevance and interpretability of the prediction model.

[0198] Step 3: Collect the SERS spectral data of bladder tumor antigen (BTA), hyaluronic acid, hyaluronidase, albumin, and urea in urine, that is, obtain the SERS spectral data of the pure phases of these 5 biomarkers, as Figure 14 shown. Among them, BTA, hyaluronic acid, and hyaluronidase are specific biomarkers for bladder cancer.

[0199] Step 4: The wavenumbers positively correlated with the positive prediction (bladder cancer) obtained in Step 2 are positive correlation wavenumbers, and the wavenumbers negatively correlated with the positive prediction are negative correlation wavenumbers. For a known biomarker to be determined (or, controversial) that is positively correlated with the positive prediction, in the SERS spectral data of the pure phase of this biomarker, the characteristic peak intensity at the positive correlation wavenumbers is higher, and the characteristic peak intensity at the negative correlation wavenumbers is lower. To quantify this, sum the characteristic peak intensities at the positive correlation wavenumbers in the SERS spectral data of the pure phase of the biomarker, and use the average value of the absolute values of the SHAP values corresponding to the positive correlation wavenumbers as the weight to obtain the positive contribution, that is, the sum of the characteristic peak intensities at the positive correlation wavenumbers in the SERS spectral data of the pure phase of the biomarker × weight (positive value). Similarly, sum the characteristic peak intensities at the negative correlation wavenumbers in the SERS spectrum of the pure phase of the biomarker, and use the negative value of the average value of the absolute values of the SHAP values corresponding to the negative correlation wavenumbers as the weight to obtain the negative contribution, that is, the sum of the characteristic peak intensities at the negative correlation wavenumbers in the SERS spectral data of the pure phase of the biomarker × weight (negative value). The total contribution of the wavenumbers to each biomarker is obtained by adding the positive and negative contributions, asFigure 14 as shown

[0200] The total contribution of wavenumber to each biomarker within the entire wavenumber range is calculated by the following formula (see Figure 4 ):

[0201] Biomarker contribution=∑(x i ×w i )+Σ[x j ×(-w j )]

[0202] where x i represents the characteristic peak intensity at the positively correlated wavenumber in the SERS spectrum of the biomarker in the pure phase, w i represents the average value of the absolute value of the SHAP value of the key eigenvalue at the positively correlated wavenumber, x j represents the characteristic peak intensity at the negatively correlated wavenumber in the SERS spectrum of the biomarker in the pure phase, w j represents the average value of the absolute value of the SHAP value of the key eigenvalue at the negatively correlated wavenumber.

[0203] It was found that for the prediction model of (HC + BD) vs BCa, the contributions of urea and albumin to positive prediction were the lowest, and even the content of albumin was negatively correlated with positive prediction. This may be because benign diseases also increase the content of urea and albumin in urine, so the content of urea and albumin in urine samples of BCa is not significantly higher than that in urine samples of BD. However, the contributions of hyaluronic acid, hyaluronidase and BTA to positive prediction are all relatively high. Therefore, for bladder cancer, "hyaluronic acid, hyaluronidase, BTA" are more effective bladder cancer biomarkers.

[0204] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited by the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.

Claims

1. A deep learning-based prediction model for predicting urinary system cancer categories, obtained as follows: Step 1: Using silver nanowires as SERS probes, preparing a silver nanowire solution; Step 2: obtaining urine from healthy individuals, patients with benign diseases, patients with bladder cancer, patients with kidney cancer, and patients with prostate cancer, and processing the urine to obtain urine samples; Step 3: measuring an appropriate amount of the silver nanowire solution from step 1 and an appropriate amount of the urine sample from step 2, respectively, mixing the silver nanowire solution and the urine sample thoroughly to obtain a test sample, collecting SERS spectrum data of the test sample and performing preprocessing; Step 4, extraction of key characteristic values: For all the preprocessed SERS spectral data in Step 3, extract the intensities corresponding to the peaks, valleys, and full-width at half-maximum positions of all characteristic peaks within the entire wavenumber range as the key characteristic values, that is, the intensities corresponding to the same wavenumber in different SERS spectral data are the key characteristic values. The number of key characteristic values extracted from different SERS spectral data is the same, and the wavenumber positions and the number of wavenumbers corresponding to the key characteristic values extracted from each SERS spectral data are the same; moreover, denote the set of key characteristic values extracted from one SERS spectral data as X K , X K = [x1, x2, …, x n T ​ Among them, K=1,2,3,……,n, n is the total number of waves; Step 5: Construction of the RaNN model: The RaNN model consists of an input layer, a feature transformation layer, an attention transformation layer, a fully connected layer, and an output layer. The number of neurons in the output layer is the same as the number of urological cancer categories. Step 6: Training the RaNN model includes the following steps: Step 6.1: Select a urine sample according to the prediction requirement, and use the X corresponding to the selected urine sample obtained in Step 4 K as a data set, and divide this data set into a training set and a test set according to a ratio of 8:2; Step 6.2: Use the training set to train the RaNN model constructed in step 5. The number of training rounds is determined by the accuracy of the model during training. When the accuracy stops improving, training is stopped. Step 6.3: Use the test set to test the RaNN model after the training parameters are updated, compare the actual value with the prediction result of the RaNN model, and if the prediction accuracy meets the requirements, obtain the RaNN prediction model, that is, the prediction model based on deep learning.

2. The prediction model according to claim 1, wherein A cross-entropy loss function is used to calculate the loss of the RaNN model for the training set prediction. The Adam optimizer is used for optimization. A learning rate scheduler is used to dynamically adjust the learning rate during training. The RaNN model is trained for the training rounds. Each training round updates the model parameters through forward propagation, cross-entropy loss function loss calculation, backpropagation, and optimization.

3. The prediction model according to claim 1, characterized in that, In step 1, the preparation steps of the silver nanowire solution are as follows: (1) adding polyvinyl pyrrolidone and CuCl2 to ethylene glycol, stirring and dispersing them uniformly in an ultrasonic bath to obtain solution A; then dissolving AgNO3 in ethylene glycol to obtain solution B; then, adding the solution A dropwise to the solution B and stirring uniformly to obtain a mixed solution; (2) transferring the obtained mixed solution to a high-pressure reactor, sealing the high-pressure reactor and placing it in an oven, heating it at 160°C for 3h, and cooling it to room temperature after the reaction is completed; (3) taking 200mL of the solution in the high-pressure reactor that has cooled to room temperature; and adding the 200mL The solution was centrifuged, and then the supernatant was completely removed. 100 mL of anhydrous ethanol was added to the resulting silver nanowire precipitate, and the solution was evenly dispersed using an ultrasonic cleaner. The resulting solution was then centrifuged again, and the supernatant was completely removed using a pipette. 100 mL of anhydrous ethanol was added to the resulting silver nanowire precipitate, and the solution was evenly dispersed using an ultrasonic cleaner. The aforementioned process from "continuously centrifuging the evenly dispersed solution" to "evenly dispersing the solution using an ultrasonic cleaner" was repeated three times. Finally, the resulting silver nanowire precipitate was dispersed in anhydrous ethanol and stored in a light-proof container for later use.

4. The prediction model according to claim 1, wherein In step 2, the urine sample acquisition process is as follows: urine from healthy people, patients with benign diseases, bladder cancer patients, kidney cancer patients, and prostate cancer patients is collected, and the supernatant is extracted as a urine sample by centrifugation, wherein the centrifugation time is 15 minutes, the centrifugal speed is 3000 r / min, and the urine volume used is 1 mL.

5. The prediction model according to claim 1, wherein In step 3, the specific process of collecting SERS spectral data and preprocessing is as follows: (1) 15 μL of the urine sample obtained in step 2 is thoroughly mixed with 30 μL of the silver nanowire solution obtained in step 1 to prepare a test sample; (2) SERS spectral data of the test sample is collected; (3) the collected SERS spectral data is preprocessed by performing denoising, baseline removal, standard normal transformation, smoothing, and normalization in sequence; Preferably, in step 4, the method for extracting the key feature value is: For the intensity f(x) of the SERS spectrum signal, a point x0 is a peak if and only if the following equation is satisfied: f(x0)>f(x0-∈) and f(x0)>f(x0+∈) Among them, ∈ is a small interval value, indicating that the value at position x0 is greater than the values of the left and right neighbors; For a signal f(x), a point x0 is a trough if and only if the following equation is satisfied: f(x0) <f(x0-∈)且f(x0)<f(x0+∈) Among them, ∈ is a small interval value, indicating that the value at position x0 is smaller than the values of the left and right neighbors; The steps for the FWHM position are as follows: Half Height: The half-height value is half the maximum intensity of the peak: Intersection point: The position where the half-height value is calculated, that is, the two points that intersect with half of the peak intensity (one on the left side of the peak and the other on the right side). These two points are the half-height width positions of the characteristic peak.

6. The prediction model according to claim 1, wherein In step 5, the construction of the RaNN model includes the following steps: Step 5.

1. The feature transformation layer performs the following operations on the input X obtained in Step 4 and outputs z: K ​ Step 5.1.

1. Calculate the Pearson correlation coefficient matrix between the key eigenvalues of each urine sample: where, x ki represents the i-th key feature value among the key feature values corresponding to the k-th urine sample, is the mean of the i-th key feature value among the key feature values corresponding to each urine sample in the urine samples included in the training set, x kj represents the j-th key feature value among the key feature values corresponding to the k-th urine sample, is the mean of the j-th key feature value among the key feature values corresponding to each urine sample in the urine samples included in the training set; Step 5.1.2, select R ij Connect the key feature values greater than the threshold q; define a connection mask matrix: Among them, when M i,j = 1, the corresponding input key feature value x j is input into the neuron h i in the feature transformation layer; when M i,j = 0, the corresponding input key feature value x j is not input into the neuron h i in the feature transformation layer; Define the connection weight matrix: W′ i,j = W i,j · M i,j Among them, W i,j is the original weight matrix of the neural network; Then the output z of each neuron in the feature transformation layer i is calculated using the following formula: where b i is the bias term, m is the number of features, and z i represents the content of a potential biomolecule in the urine sample; X K The key feature value is input into the feature transformation layer, and after passing through the feature transformation layer, the output is z, where z is a set of z i ;; Step 5.

2. The attention transformation layer performs the following operations on the output z of the feature transformation layer and outputs the weighted feature z weighted ; specifically, the process is as follows: Step 5.2.

1. Take the output z of the feature transformation layer as the input feature, perform a linear transformation on it, and then calculate the attention weight of each linearly transformed input feature using the following formula: attention_weights=σ(attention(z)) Among them, σ is the Sigmoid activation function, and attention(z) is the result of linear transformation of input feature z; Step 5.2.2: Obtain the attention weights attention_weights of the input features through the processing in Step 5.2.1, and then use the attention weights attention_weights to weight the corresponding input features to obtain the weighted feature z weighted , using the following formula: z weighted = z · attention_weights Among them, z is the input feature, attention_weights is the attention weight corresponding to the input feature; Step 5.3: Input the weighted feature z weighted into the fully connected layer for linear transformation to transform the weighted feature into a linear feature, using the following formula: output1 = FC(z weighted ) Among them, FC() is a linear operation, which performs a linear transformation on the weighted features; The linear features output by the fully connected layer are transformed into nonlinearity through the ReLU activation function. The calculation formula is as follows: ReLU(output1)=max(0,output1) Step 5.4: After being processed in Step 5.3, the non-linear feature x is obtained final , and the non-linear feature x final is input into the output layer and the following operations are performed: output2 = fc_out(x final ) Among them, fc_out() is a linear operation, and its number of neurons is the total number of output categories, which is used to calculate the predicted value of each category of the urine sample to be tested; the result calculated by the above formula is recorded as y: y = (y1, y2, …, y n ) Where n is the total number of categories; y is the prediction result, that is, the result y calculated by the above formula is the final output prediction result (that is, the value corresponding to each category), and the Softmax activation function is used to calculate the probability corresponding to the prediction result: where y i is the predicted result of the final output corresponding to class i, and the predicted result of the final output corresponding to each class is converted into a corresponding probability distribution through the Softmax activation function.

7. The prediction model according to claim 1, wherein In step 6, when the output layer includes 2 neurons, the RaNN prediction model is a RaNN two-class prediction model; when the output layer includes 3 neurons, the RaNN prediction model is a RaNN three-class prediction model; when the output layer includes 4 neurons, the RaNN prediction model is a RaNN four-class prediction model; when the output layer includes 5 neurons, the RaNN prediction model is a RaNN five-class prediction model.

8. A method for predicting urinary system cancer categories based on the deep learning-based prediction model according to any one of claims 1 to 7, comprising the following steps: Step 1: Using silver nanowires as SERS probes, preparing a silver nanowire solution; Step 2: obtaining a urine sample to be tested; Step 3: measuring an appropriate amount of the silver nanowire solution from step 1 and an appropriate amount of the urine sample from step 2, respectively, mixing the silver nanowire solution and the urine sample thoroughly to obtain a test sample, collecting SERS spectrum data of the test sample and performing preprocessing; Step 4. Extraction of key characteristic values: For all the SERS spectral data obtained through preprocessing in Step 3, extract the intensities corresponding to the peaks, valleys, and full width at half maximum positions of all the characteristic peaks within the entire wavenumber range as the key characteristic values; and denote the set of key characteristic values extracted from one SERS spectral data as X K , X K = [x1, x2, …, x n T ​ Among them, K=1,2,3,……,n, n is the total number of waves; Step 5: inputting the key feature values of step 4 into the RaNN prediction model according to any one of claims 1 to 7; Step 6: Calculate and output the prediction results of urinary system cancer categories through the RaNN prediction model.

9. A method for evaluating the effectiveness of a urinary system cancer biomarker, comprising the following steps: Step 1: Using silver nanowires as SERS probes, preparing a silver nanowire solution; Step 2: obtaining urine from healthy individuals, patients with benign diseases, and patients with urinary system cancer, and processing the urine to obtain urine samples; Step 3: Measure appropriate amounts of the silver nanowire solution in Step 1 and the urine sample in Step 2 respectively. After fully mixing the silver nanowire solution and the urine sample, a test sample is obtained. Collect the SERS spectral data of the test sample and perform preprocessing; Step 4, extraction of key characteristic values: For all the SERS spectral data obtained through preprocessing in Step 3, extract the intensities corresponding to the peaks, valleys, and full-width at half-maximum positions of all characteristic peaks within the entire wavenumber range as the key characteristic values, that is, the intensities corresponding to the same wavenumber in different SERS spectral data are the key characteristic values. The number of key characteristic values extracted from different SERS spectral data is the same, and the wavenumber positions and the number of wavenumbers corresponding to the key characteristic values extracted from each SERS spectral data are the same. Moreover, denote the set of key characteristic values extracted from one SERS spectral data as X K , X K = [x1, x2, …, x n T ​ Wherein, K = 1, 2, 3, ……, n, where n is the total number of wave numbers; Step 5: Use the X corresponding to the urine samples of healthy people, patients with benign diseases, and patients with a certain type of urinary system cancer extracted in Step 4 K Each is used as a sub-dataset. The sub-datasets corresponding to healthy people and patients with benign diseases are combined to obtain a total of 2 sub-datasets. According to the ratio of 8:2, the 2 sub-datasets are sequentially divided into a sub-training set and a sub-test set to obtain 2 sub-training sets and 2 sub-test sets. Then, the obtained 2 sub-training sets are combined into one training set, and the 2 sub-test sets are combined into one test set. Use the training set and the test set to train the RaNN prediction model described in any one of claims 1-7 to obtain a RaNN binary classification model B; Step 6: For the RaNN binary classification model B, apply the SHAP algorithm to evaluate the contribution of the wave numbers to the prediction result of the RaNN binary classification model B. The wave numbers that are positively correlated with the prediction result of positive for urinary system cancer are positive correlation wave numbers, and the wave numbers that are negatively correlated with the prediction result of positive for urinary system cancer are negative correlation wave numbers; Step 7: For a certain type of urinary system cancer, collect the SERS spectral data of the pure phase of the biomarker for the certain type of urinary system cancer in urine; Step 8. For each biomarker of a certain type of urinary system cancer, calculate the total contribution value B of each biomarker respectively based on the positively correlated wavenumber and negatively correlated wavenumber obtained in Step 6 i , and the calculation formula is as follows: Biomarker contribution=∑(x i ×w i )+∑[x j ×(-w j )] Among them, x i represents the characteristic peak intensity at the positive correlation wavenumber in the SERS spectrum of the pure-phase urinary system cancer biomarker, w i represents the average value of the absolute value of the SHAP value of the key eigenvalue at the positive correlation wavenumber, x j represents the characteristic peak intensity at the negative correlation wavenumber in the SERS spectrum of the pure-phase urinary system cancer biomarker, w j represents the average value of the absolute value of the SHAP value of the key eigenvalue at the negative correlation wavenumber; B i where i is the name of the biomarker of a certain type of urinary system cancer; Step 9. According to the total contribution value B of each biomarker obtained in Step 8 i , determine the effectiveness of the biomarkers for urinary system cancer: The larger the total contribution value B of the biomarker calculated i , the higher its effectiveness.

10. The evaluation method according to claim 9, wherein In step 6, the specific process of using the SHAP algorithm to evaluate the contribution of the wavenumber to the prediction result of the RaNN binary classification model B is as follows: First, take 20% from the X included in each of the 2 sub-test sets in step 5, and merge them into a test set B. For all the key feature values included in each X in the test set B, calculate the SHAP value of each key feature value and take its absolute value; K in the test set B, calculate the SHAP value of each key feature value and take its absolute value; K Calculate the SHAP value of each key feature value and take its absolute value; Secondly, obtain the average value of the absolute values of the SHAP values calculated from all the key feature values corresponding to each wave number in the wave numbers in Step 4. Finally, sort the wave numbers corresponding to the obtained average values from largest to smallest according to the average values; Using the sorted wave numbers as the vertical axis and the SHAP values as the horizontal axis, plot the SHAP values calculated from all the key feature values corresponding to each wave number. The plotting rule is: the greater the Raman intensity of the key feature value corresponding to each SHAP value on the horizontal axis, the darker the color, and vice versa, the lighter the color. According to the obtained graph, for each horizontal row corresponding to a wave number, if it shows light color on the left and dark color on the right, it indicates that the wave number corresponding to this horizontal row is a wave number that is positively correlated with the prediction result of positive for urinary system cancer. On the contrary, if it shows dark color on the left and light color on the right, it indicates that the wave number corresponding to this horizontal row is a wave number that is negatively correlated with the prediction result of positive for urinary system cancer. Accordingly, the wave numbers that are positively correlated with the prediction result of positive for urinary system cancer are positive correlation wave numbers, and the wave numbers that are negatively correlated with the prediction result of positive for urinary system cancer are negative correlation wave numbers; Preferably, the calculation formula of the SHAP value is: Among them, is the SHAP value of the key feature value x i , that is, the contribution value of the key feature value x i to the prediction result. N is the set of all key feature values, and S is a subset that does not include the key feature value x i . v(S) is the predicted value of the RaNN prediction model obtained by training with the subset S. denotes the summation over all subsets S that do not include the key feature value x i . denotes the probability of the subset S appearing, and [v(S∪{i}) - v(S)] represents the change in the model output after adding the key feature value x i to the existing feature subset S.

Citation Information

Patent Citations

  • Spectroscopy and artificial intelligence interactive serum analysis method and application thereof

    CN115078331A