A bladder cancer fine typing method and system based on nucleic acid aptamer differential analysis

Through the method of nucleic acid aptamer differential analysis, the sequencing data of bladder cancer cells is obtained, abnormal samples are removed and standardized, differentially expressed aptamers are selected, and combined with the probe package, fine typing of bladder cancer is achieved, which solves the problem of inaccurate molecular typing of bladder cancer, provides personalized treatment plans, and improves diagnostic accuracy and treatment effects.

CN119724344BActive Publication Date: 2025-10-17ZHEJIANG UNIV OF TECH +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411769306.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-04
Publication Date
2025-10-17
Estimated Expiration
2044-12-04

AI Technical Summary

Technical Problem

Existing molecular typing technologies for bladder cancer lack a unified, mature, and practically applicable solution, and a single marker nucleic acid aptamer is insufficient to achieve accurate typing of bladder cancer.

Method used

Through the method based on nucleic acid aptamer differential analysis, sequencing data is obtained, abnormal samples are removed and standardized, differentially expressed aptamers are selected, linear models are used for fitting and typing, and nucleic acid aptamer probe packages are used to perform fine typing of bladder cancer cells.

Benefits of technology

It achieves more detailed and accurate bladder cancer classification, can provide patients with personalized treatment plans, improve diagnostic accuracy, reduce misdiagnosis and missed diagnosis, and optimize treatment paths.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119724344B_ABST
    Figure CN119724344B_ABST
Patent Text Reader

Abstract

The application relates to a bladder cancer fine typing method and system based on nucleic acid aptamer differential analysis, which comprises the following steps: incubating a sample containing bladder cancer cells through a probe package, and performing high-throughput gene sequencing on the incubated sample to obtain sequencing data; removing abnormal samples from the sequencing data to obtain nucleic acid aptamer data; standardizing the nucleic acid aptamer data to obtain standard data including a health sample group and a case sample group; selecting differential expression aptamers from the standard data; and typing a to-be-tested sample through the differential expression aptamers. The application can identify the expression difference of specific nucleic acid aptamers in bladder cancer cells, and further realizes fine typing.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of bioinformatics, and particularly relates to a bladder cancer fine typing method and system based on nucleic acid aptamer differential analysis. BACKGROUND

[0002] Bladder cancer is a common malignant tumor in the urinary system. With the continuous progress of molecular biology technology, the molecular typing research of bladder cancer has gradually emerged. Molecular typing of bladder cancer helps to predict the patient's response to treatment and assess the prognosis of the disease. However, the molecular typing research of bladder cancer is still in its infancy, and compared with other cancers, there is still a lack of unified, mature and practically valuable typing scheme. Therefore, it is crucial to establish a more perfect molecular typing system to guide the treatment of bladder cancer.

[0003] Nucleic acid aptamer is a single-stranded DNA or RNA molecule with unique three-dimensional structure, which can specifically bind to various targets including small molecules, proteins and cancer cells, and is widely used in disease biomarker screening and identification. In the field of precise division of bladder cancer, nucleic acid aptamer shows great application potential. As a new type of molecular tool, it can significantly improve the recognition ability of tumor cells, provide important support for the development of individualized treatment plan, and provide more accurate diagnosis and treatment options for bladder cancer patients.

[0004] The current Cell-SELEX is a screening technology targeting whole cells, which is used to obtain DNA aptamers that specifically bind to cell surface proteins. This technology can identify the natural conformation of membrane proteins and achieve multi-target aptamer screening. In the study of bladder cancer, several nucleic acid aptamers have been reported, such as EBA aptamer that specifically binds to EN2, SPL3C aptamer that binds to CKAP4, and TB-5 aptamer that binds to NCL. However, due to the high heterogeneity of bladder cancer, these nucleic acid aptamers targeting a single marker are not sufficient to achieve precise typing of bladder cancer. SUMMARY

[0005] The present application provides a bladder cancer fine typing method and system based on nucleic acid aptamer differential analysis to solve the defects of the prior art.

[0006] The present application provides a bladder cancer fine typing method based on nucleic acid aptamer differential analysis, comprising:

[0007] S1: obtaining sequencing data, the sequencing data is obtained by incubating a sample containing bladder cancer cells with a probe package, and then performing high-throughput gene sequencing on the incubated sample;

[0008] S2: removing abnormal samples from the sequencing data to obtain nucleic acid aptamer data;

[0009] S3: standardizing the aptamer data to obtain standard data including a healthy sample group and a case sample group;

[0010] S4: selecting differentially expressed aptamers from the standard data;

[0011] S5: typing a to-be-tested sample by using the differentially expressed aptamers.

[0012] According to the bladder cancer fine typing method based on aptamer differential analysis provided by the application, step S2 further includes:

[0013] S21: receiving sequencing data based on a molecular identity tag;

[0014] S22: screening the sequencing data to obtain filtered data;

[0015] S23: merging the filtered data corresponding to each sample to obtain group data;

[0016] S24: screening the group data to obtain aptamer data.

[0017] According to the bladder cancer fine typing method based on aptamer differential analysis provided by the application, step S22 specifically includes:

[0018] S221: filtering aptamers with a length less than or equal to a first preset threshold in the sequencing data to obtain first filtered data;

[0019] S222: calculating the longest common subsequence between aptamers and primers in the sequencing data by using a dynamic programming algorithm;

[0020] S223: removing primer dimers with a primer length and the longest common subsequence greater than a second preset threshold in the first filtered data to obtain second filtered data, and outputting the second filtered data as filtered data.

[0021] According to the bladder cancer fine typing method based on aptamer differential analysis provided by the application, step S222 further includes:

[0022] S2221: constructing a two-dimensional array;

[0023] S2222: traversing the bases of the sequences of aptamers and primers in the sequencing data to fill the two-dimensional array to obtain a two-dimensional matrix;

[0024] S2223: outputting the last array of the two-dimensional matrix as the longest common subsequence.

[0025] The application provides a bladder cancer fine typing method based on nucleic acid aptamer differential analysis, and the step S24 further comprises the following steps:

[0026] S241: calculating a CPM value of each aptamer according to the nucleic acid aptamer abundance of the group data;

[0027] S242: removing samples with a CPM value less than or equal to 1 in the group data to obtain the nucleic acid aptamer data.

[0028] The application provides a bladder cancer fine typing method based on nucleic acid aptamer differential analysis, and the step S3 further comprises the following steps:

[0029] S31: dividing samples in the nucleic acid aptamer data into a healthy sample group and a case sample group;

[0030] S32: selecting a median value for the healthy sample group and the case sample group respectively to obtain a reference sample;

[0031] S33: calculating an expression difference value and an expression level value of each aptamer according to the reference sample;

[0032] S34: calculating a normalization factor according to the expression difference value;

[0033] S35: normalizing the nucleic acid aptamer data according to the normalization factor to obtain standard data.

[0034] The application provides a bladder cancer fine typing method based on nucleic acid aptamer differential analysis, and the expression difference value in the step S33 has the following expression:

[0035] M = log2(CPM sample ) - log2(CPM reference );

[0036] Wherein M is an expression difference value of a single aptamer between a reference sample and other samples, CPM sample is a CPM value of the other samples, and CPM reference is a CPM value of the reference sample;

[0037] The expression level value in the step S33 has the following expression:

[0038] A = 0.5 * (log2(CPM sample ) + log2(CPM reference ));

[0039] Wherein A is an expression level value of a single aptamer between a reference sample and other samples;

[0040] The normalization factor in the step S34 has the following expression:

[0041] scaling factor=2 mean(M) ;

[0042] Wherein, scaling factor is a standardization factor, mean(M) is the mean of the expression difference value of the screened sample;

[0043] The expression of the standard data in step S35 is:

[0044] CPM normalized =log2(CPM original ×scaling factor+1);

[0045] Wherein, CPM normalized is the sample CPM value in the standard data, CPM original is the original CPM value of a single aptamer.

[0046] According to the bladder cancer fine classification method based on nucleic acid aptamer difference analysis provided by the application, step S4 further comprises:

[0047] S41: fitting the health sample group and the case sample group respectively by a linear model to obtain a fitted health sample group and a fitted case sample group;

[0048] S42: calculating the expression result difference of the corresponding groups of the fitted health sample group and the fitted case sample group;

[0049] S43: selecting the aptamer group corresponding to the maximum expression result difference as the differential expression aptamer.

[0050] According to the bladder cancer fine classification method based on nucleic acid aptamer difference analysis provided by the application, the expression of the expression result difference in step S42 is:

[0051] C=β1-β2;

[0052] Wherein, C is the expression result difference of the current group, β1 is the model coefficient corresponding to the aptamer belonging to the fitted health sample group of the current group, and β2 is the model coefficient corresponding to the aptamer belonging to the case health sample group of the current group.

[0053] The application further provides a bladder cancer fine classification system based on nucleic acid aptamer difference analysis, comprising:

[0054] A receiving module is used for receiving sequencing data, wherein the sequencing data is obtained by incubating a sample containing bladder cancer cells with a probe package and performing high-throughput gene sequencing on the incubated sample;

[0055] The screening module is used for removing abnormal samples from the sequencing data to obtain aptamer data;

[0056] The standardization module is used for standardizing the aptamer data to obtain standard data including a healthy sample group and a case sample group;

[0057] The selection module is used for selecting differentially expressed aptamers from the standard data;

[0058] The application module is used for typing a to-be-tested sample by using the differentially expressed aptamers.

[0059] The application provides a bladder cancer fine typing method and system based on aptamer differential analysis. The application can identify the expression difference of specific aptamers in bladder cancer cells by using the method based on aptamer differential analysis, and further realizes fine typing. The typing method is more detailed and accurate, and is helpful for doctors to make more accurate treatment plans for patients. In the data processing process, the application adopts strict screening and standardization steps to ensure the accuracy and reliability of the sequencing data, and can reduce the possibility of misdiagnosis and missed diagnosis, and improve the accuracy of diagnosis. In addition, the application can provide personalized treatment plans for patients by identifying differentially expressed aptamers in bladder cancer. Since different patients may have different differentially expressed aptamers, the personalized treatment plans made according to the differences may be more effective. The typing method of the application is fast and efficient, and can quickly provide detailed typing information about bladder cancer patients, so as to help relevant personnel make decisions faster and optimize the treatment path of patients. BRIEF DESCRIPTION OF DRAWINGS

[0060] In order to more clearly illustrate the technical solutions in the application or prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description.

[0061] Figure 1 A fine typing method of bladder cancer based on aptamer differential analysis is provided for the embodiments of the application.

[0062] Figure 2 A fine typing system of bladder cancer based on aptamer differential analysis is provided for the embodiments of the application.

[0063] Figure 3 A distribution diagram of differentially expressed aptamers is provided for the embodiments of the application.

[0064] Figure 4A schematic diagram of the binding situation of the nucleic acid aptamer seq12a sequence provided by the embodiment of the present application for flow binding is shown in the figure.

[0065] Figure 5 A schematic diagram of another distribution of the differentially expressed aptamer provided by the embodiment of the present application is shown in the figure.

[0066] Reference signs:

[0067] 100, receiving module; 200, screening module; 300, standardization module; 400, selection module; 500, application module. DETAILED DESCRIPTION

[0068] In order to make the objects, technical solutions and advantages of the present application clearer, the technical solutions in the present application will be described clearly and completely below in combination with the drawings in the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments, and they should not be understood as limiting the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts fall within the scope of protection of the present application. In the description of the present application, it should be understood that the terms used are only for the purpose of description, and should not be understood as indicating or implying relative importance.

[0069] In order to better understand the present application, the terms appearing in the embodiments of the present application will be explained and described first below.

[0070] Differential analysis is a statistical technique widely used in biological and medical research, aiming to compare samples under different experimental conditions, so as to reveal significant differences at the level of genes, proteins or other molecules. Through differential analysis, biomarkers with expression changes between disease and healthy state, different treatment groups, or time points can be found. Common applications of differential analysis include differential analysis of gene expression (such as RNA-Seq data), quantitative differential analysis in proteomics, and comparison of metabolite levels in metabolomics. Based on statistical models and hypothesis testing methods, differential analysis can quantify the differences between different groups and assess their significance, which is crucial for revealing potential biological mechanisms, screening disease markers, and identifying therapeutic targets. Modern differential analysis often combines multivariate analysis and data regularization techniques to improve robustness to noise and data imbalance, thereby enhancing the credibility of the results.

[0071] Aptamers are single-stranded DNA or RNA molecules with unique three-dimensional structures that can specifically bind to various targets, including small molecules, proteins, and cancer cells. They are widely used in the screening and identification of disease biomarkers. Aptamers have shown great potential in the precision diagnosis and treatment of bladder cancer. As a new molecular tool, they can significantly improve the ability to identify tumor cells, provide important support for the development of personalized treatment plans, and offer more precise diagnosis and treatment options for bladder cancer patients.

[0072] However, due to the high heterogeneity of bladder cancer, these aptamers targeting single markers are not sufficient for accurate diagnosis of bladder cancer. Therefore, for highly heterogeneous tumors such as bladder cancer, differential analysis of aptamers based on efficient data extraction can not only reveal the specificity of aptamers for different subtypes but also provide support for the potential mechanisms of molecular recognition, thereby achieving more accurate cancer diagnosis and the development of personalized treatment plans. This differential analysis based on aptamer data provides new tools and strategies for cancer molecular typing and is expected to enhance the application value of aptamers in precision cancer medicine.

[0073] The following combination Figures 1 to 5 Embodiments of the present invention are described.

[0074] like Figure 1 As shown, the present invention provides a method for fine typing of bladder cancer based on nucleic acid aptamer differential analysis, comprising:

[0075] S1: Acquiring sequencing data, wherein the sequencing data is obtained by incubating a sample containing bladder cancer cells with a probe package and performing high-throughput gene sequencing on the incubated sample.

[0076] In step S1, a universal bladder cancer enrichment library probe package, in this example, the C+CU library probe package, is first used to incubate bladder cancer cells and normal cells. The nucleic acid aptamers bound to the cells are isolated and subjected to high-throughput sequencing to obtain molecular ID tags. This technology assigns a unique identity tag to each nucleic acid aptamer in the experiment to prevent abnormal aptamer quantity caused by specific amplification during high-throughput sequencing.

[0077] S2: removing abnormal samples from the sequencing data to obtain nucleic acid aptamer data.

[0078] Wherein, step S2 further includes:

[0079] S21: Receive sequencing data based on the molecular ID tag.

[0080] In the first specific embodiment of the present application, after obtaining accurate experimental data based on the molecular identity tag, the binding aptamer abundance of each sample is calculated according to the experimental library and the UMI connected to the back end of the primer of different samples, wherein the front and back end primers of the library are AAGGAGCAGCGTGGAGGATA (SEQ ID NO: 1) and CTCATGGACGTGCTGGTGAC (SEQ ID NO: 2) respectively, and the UMIs corresponding to the six cancer cell samples are CCTGTCTTGTCTGCCTACCT (SEQ ID NO: 3), CCTGTCTTGTCTGCCTAGAT (SEQ ID NO: 4), CCTGTCTTGTCTGCCTTACT (SEQ ID NO: 5), CCTGTCTTGTCTGCCTTGAT (SEQ ID NO: 6), CCTGTCTTGTCTGCCTTCGT (SEQ ID NO: 7), and CCTGTCTTGTCTGCCTGCAT (SEQ ID NO: 8) respectively, and the UMIs corresponding to the six normal cell samples are CCTGTCTTGTCTGCCTGAGT (SEQ ID NO: 9), CCTGTCTTGTCTGCCTCTAT (SEQ ID NO: 10), CCTGTCTTGTCTGCCTCGGT (SEQ ID NO: 11), CCTGTCTTGTCTGCCTAAGT (SEQ ID NO: 12), CCTGTCTTGTCTGCCTATCA (SEQ ID NO: 13), and CCTGTCTTGTCTGCCTATGC (SEQ ID NO: 14) respectively.

[0081] In another specific embodiment of the present application, after obtaining accurate experimental data based on molecular identity tag, the calculation of the binding aptamer abundance of each sample is realized according to the experimental library and the UMI connected to the back end of the primer of different samples. The primers of the front and back ends of the two experiments carried out in this embodiment are AAGGAGCAGCGTGGAGGATA (SEQ ID NO: 1) and CTCATGGACGTGCTGGTGAC (SEQ ID NO: 2), and the UMIs corresponding to the five healthy human urine samples in the first patient and healthy human urine sample experiment are CCTGTCTTGTCTGCCTACCT (SEQ ID NO: 3), CCTGTCTTGTCTGCCTAGAT (SEQ ID NO: 4), CCTGTCTTGTCTGCCTTACT (SEQ ID NO: 5), CCTGTCTTGTCTGCCTTGAT (SEQ ID NO: 6), CCTGTCTTGTCTGCCTTCGT (SEQ ID NO: 7), CCTGTCTTGTCTGCCTGCAT (SEQ ID NO: 8), CCTGTCTTGTCTGCCTGAGT (SEQ ID NO: 9), CCTGTCTTGTCTGCCTCTAT (SEQ ID NO: 10), CCTGTCTTGTCTGCCTCGGT (SEQ ID NO: 11), CCTGTCTTGTCTGCCTAAGT (SEQ ID NO: 12), and the UMIs corresponding to the patient urine sample are CCTGTCTTGTCTGCCTATCA (SEQ ID NO: 13), CCTGTCTTGTCTGCCTATGC (SEQ ID NO: 14), CCTGTCTTGTCTGCCTTTCA (SEQ ID NO: 15), CCTGTCTTGTCTGCCTTTGC (SEQ ID NO: 16), CCTGTCTTGTCTGCCTTTAG (SEQ ID NO: 17), CCTGTCTTGTCTGCCTGCGC (SEQ ID NO: 18), CCTGTCTTGTCTGCCTGACA (SEQ ID NO: 19), CCTGTCTTGTCTGCCTGTCT (SEQ ID NO: 20), CCTGTCTTGTGCCTCATA (SEQ ID NO: 21), and CCTGTCTTGTGCCTCAGC (SEQ ID NO: 22).The corresponding UMIs of the two repeated experiments of the three preoperative urine samples in the second patient preoperative and postoperative urine sample experiment are CCTGTCTTGTCTGCCTACCT (SEQ ID NO: 3), CCTGTCTTGTCTGCCTAGAT (SEQ ID NO: 4), CCTGTCTTGTCTGCCTTACT (SEQ ID NO: 5), CCTGTCTTGTCTGCCTTGAT (SEQ ID NO: 6), CCTGTCTTGTCTGCCTTCGT (SEQ ID NO: 7), and CCTGTCTTGTCTGCCTGCAT (SEQ ID NO: 8), and the corresponding UMIs of the postoperative urine samples are CCTGTCTTGTCTGCCTGAGT (SEQ ID NO: 9), CCTGTCTTGTCTGCCTCTAT (SEQ ID NO: 10), CCTGTCTTGTCTGCCTCGGT (SEQ ID NO: 11), CCTGTCTTGTCTGCCTAAGT (SEQ ID NO: 12), CCTGTCTTGTCTGCCTATCA (SEQ ID NO: 13), and CCTGTCTTGTCTGCCTATGC (SEQ ID NO: 14).

[0082] S22: filtering the sequencing data to obtain filtered data.

[0083] Specifically, step S22 includes:

[0084] S221: filtering aptamers with a length less than or equal to a first preset threshold in the sequencing data to obtain first filtered data.

[0085] After obtaining accurate experimental data based on molecular identity tags in step S1, in step S22, the binding nucleic acid aptamer abundance of each sample is calculated according to the UMIs of the experimental library and the different samples connected to the primer back end, and then based on the nucleic acid aptamer abundance, for each aptamer, the aptamer with a length less than or equal to 20 is preliminarily excluded, and the influence of primer dimers on the experimental results needs to be removed subsequently.

[0086] Specifically, the high-throughput sequencing data result is a sequence containing an aptamer, a primer, a sample tag (UMI) and an identity tag, and a sequencing copy number. Step S21 is specifically: first, normalize the copy number of each sequence to 1, and obtain accurate screening enrichment library data; classify the normalized data of each sample in the enrichment library through the library primer and different sample tags; and obtain statistical data based on the molecular identity tag by re-counting the same aptamer in each sample, and the final count is the accurate copy number of the aptamer of each sample in the screening enrichment library.

[0087] S222: calculating the longest common subsequence between the aptamer and the primer in the sequencing data by a dynamic programming algorithm.

[0088] A small amount of base mutation may occur in the primer dimer during amplification, and all dimers cannot be accurately found by directly using the string matching method, so the dynamic programming algorithm is used to calculate the longest common subsequence between the aptamer and the primer in steps S222 to S223.

[0089] In step S222, the method further comprises:

[0090] S2221: constructing a two-dimensional array.

[0091] S2222: filling the two-dimensional array by traversing the bases of the sequences of the aptamer and the primer in the sequencing data to obtain a two-dimensional matrix.

[0092] S2223: outputting the last array of the two-dimensional matrix as the longest common subsequence.

[0093] In steps S2221 to S2223, first, a two-dimensional array L is established, wherein L[i][j] represents the length of the longest common subsequence of the first i characters of the first sequence and the first j characters of the second sequence; the first row and the first column are initialized to 0, because the length of the longest common subsequence of any sequence and an empty sequence is 0; secondly, the bases of the two sequences are traversed, if the characters are the same, i.e. x[i-1]=y[j-1], then L[i][j]=L[i-1][j-1]+1, if the characters are different, L[i][j]=max(L[i-1][j],L[i][j-1]) is taken; and finally, the right lower corner L[m][n] of the matrix is the length of the longest common subsequence of the two sequences.

[0094] S223: removing the primer dimer whose primer length and the longest common subsequence are greater than a second preset threshold value in the first filtered data to obtain second filtered data, and outputting the second filtered data as the filtered data.

[0095] In step S223, based on the longest common subsequence calculated in step S222, the ratio of the longest common subsequence to the primer length is calculated, and then the threshold is set to 0.9. If it is greater than 0.9, it is a primer dimer and is removed.

[0096] S23: Merge the filtered data corresponding to each sample to obtain group data.

[0097] The nucleic acid aptamer data of each sample filtered in step S22 is merged. If the nucleic acid aptamer does not exist in a certain sample, it is counted as 0. Finally, raw data similar to genomic difference analysis is obtained. The difference is that genomic data is the number of genes corresponding to different samples, and nucleic acid aptamer data is the abundance of nucleic acid aptamer corresponding to different samples.

[0098] S24: Screen the group data to obtain nucleic acid aptamer data.

[0099] Step S24 further comprises:

[0100] S241: Calculate the CPM value of each aptamer according to the nucleic acid aptamer abundance of the group data;

[0101] S242: Remove samples with CPM values less than or equal to 1 in the group data to obtain the nucleic acid aptamer data.

[0102] The purpose of step S24 is to further filter the nucleic acid aptamer data and exclude the influence of low-abundance non-specific aptamer on analysis. Specifically, the CPM of each aptamer is calculated, i.e. the number of counts per million, i.e. the original reads counts divided by the total reads number multiplied by 1*10 6 , and the aptamer with CPM greater than 1 in at least K / 2 samples is screened out, wherein K represents the number of samples, which helps to remove the aptamer of abnormal samples and ensures the reliability of subsequent analysis.

[0103] S3: Standardize the nucleic acid aptamer data to obtain standard data including a healthy sample group and a case sample group.

[0104] Step S3 further comprises:

[0105] S31: Divide the samples in the nucleic acid aptamer data into a healthy sample group and a case sample group.

[0106] S32: Select the median of the healthy sample group and the case sample group respectively to obtain a reference sample.

[0107] In steps S31-S32, the samples are first divided into two groups, representing the healthy sample group and the cancer sample group, respectively, and a median sample is selected as the reference sample in each group, i.e., for each aptamer, the expression values of all samples are sorted in ascending order, and the median is selected, if the number of samples is odd, the middle number is selected, and if the number of samples is even, the average of the two middle numbers is selected.

[0108] S33: According to the reference sample, the expression difference value and the expression level value of each aptamer are calculated.

[0109] In step S33, the expression of the expression difference value is:

[0110] M = log2(CPM sample ) - log2(CPM reference );

[0111] Wherein, M is the expression difference value of a single aptamer between the reference sample and other samples, CPM sample is the CPM value of other samples, and CPM reference is the CPM value of the reference sample.

[0112] The expression of the expression level value in step S33 is:

[0113] A = 0.5 * (log2(CPM sample ) + log2(CPM reference ));

[0114] Wherein, A is the expression level value of a single aptamer between the reference sample and other samples.

[0115] The M value above represents the expression difference of each aptamer between the reference sample and other samples, and the A value represents the average expression level of each aptamer in the two samples.

[0116] S34: Calculate the normalization factor according to the expression difference value.

[0117] In step S34, the expression of the normalization factor is:

[0118] scaling factor = 2 mean(M) ;

[0119] Wherein, scaling factor is the normalization factor, and mean(M) is the mean of the expression difference values of the screened samples.

[0120] In step S34, the M values in the first 30% and the last 30% are removed, the mean of the remaining M values is calculated, and then the mean is converted into a normalization factor by the above formula, i.e., by converting the mean into a logarithmic form.

[0121] S35: normalizing the aptamer data according to the normalization factor to obtain standard data.

[0122] In step S35, the expression of the standard data is:

[0123] CPM normalized = log2(CPM original × scaling factor + 1);

[0124] wherein CPM normalized is the sample CPM value in the standard data, and CPM original is the original CPM value of a single aptamer.

[0125] In step S35, by multiplying each sample CPM by the normalization factor in step S34, the expression of different samples can be compared under the same reference.

[0126] S4: selecting differentially expressed aptamers from the standard data.

[0127] In step S4, further comprising:

[0128] S41: fitting the healthy sample group and the case sample group respectively by a linear model to obtain a fitted healthy sample group and a fitted case sample group.

[0129] The purpose of step S41 is to construct a design matrix to represent the group information of the experiment, and specifically to use a linear model to fit the aptamer data to evaluate the expression levels of different aptamers between different groups. After fitting, the model will generate a set of parameters for each gene, including the expression level of each group and the corresponding standard error.

[0130] S42: calculating the expression result difference of the corresponding groups of the fitted healthy sample group and the fitted case sample group. In step S42, the expression of the expression result difference is:

[0131] C = β1- β2;

[0132] wherein C is the expression result difference of the current group, β1 is the model coefficient of the aptamer belonging to the fitted healthy sample group of the current group, and β2 is the model coefficient of the aptamer belonging to the case healthy sample group of the current group.

[0133] The purpose of step S42 is to calculate the expression difference of each set of aptamers under the selected comparison, wherein β1 and β2 are the model coefficients of the two groups of samples, respectively, and then the standard error is adjusted by the Bayesian method to improve the stability and reliability of the results.

[0134] In addition, the significance level of the corresponding group of the fitted health sample group and the fitted case sample group is calculated in step S42, and the expression of the significance level is obtained by t test:

[0135]

[0136] wherein, is the sample mean of the fitted health sample group, is the sample mean of the fitted case sample group, is the sample variance of the fitted health sample group, is the sample variance of the fitted case sample group, n1 is the total number of samples of the fitted health sample group, and n2 is the total number of samples of the fitted case sample group.

[0137] S43: selecting the aptamer group corresponding to the maximum difference in expression result as the differentially expressed aptamer.

[0138] In step S43, the significant differentially expressed aptamer is extracted from the model result obtained in step S42, and the expression change, statistical significance and other information thereof are obtained. Visualization tools such as volcano plot and heat map can also be used to display the distribution of the differentially expressed aptamer.

[0139] S5: typing the to-be-tested sample by using the differentially expressed aptamer.

[0140] In a specific embodiment, 88 different nucleic acid aptamers are obtained based on the nucleic acid aptamer difference analysis, and the difference value is the logarithmic ratio value between the cancer cell sample and the normal cell sample. Figure 3 As shown in FIG. 8A, the color of each point in the volcano plot corresponds to the average value of the cancer repeat sample, and the aptamers with large and significant differences between the cancer cell sample and the normal cell sample on both sides can be used to distinguish the two groups of samples. Figure 3 As shown in FIG. 8B, the differentially expressed aptamer obtained after standardizing the data of the two groups of samples can significantly distinguish the cancer cell sample from the normal cell sample, and the results show that the binding mode of the C+CU probe package in YTS-1 and SV-HUC-1 is significantly different, and the differentially expressed aptamer extracted from the probe package aptamer can significantly distinguish the bladder cancer cell YTS-1 from the normal epithelial tissue cell SV-HUC-1.

[0141] Specifically, Figure 3 FIG. 8A is a sequencing data difference analysis, and the circular color represents the logarithmic average abundance of the cancer cell sample. The redder the color, the higher the aptamer abundance in the cancer sample, and the bluer the color, the lower the aptamer abundance in the cancer sample. Figure 3 FIG. 8B is the standardized abundance of 88 different nucleic acid aptamers on each bladder cancer cell YTS-1 and normal cell SV-HUC-1 sample. The redder the color, the higher the copy number, and the bluer the color, the lower the copy number.

[0142] In a specific embodiment, in order to verify whether the extracted aptamer can truly distinguish cancer cells from normal cells, the present invention further synthesized the seq12a sequence with the largest binding difference in YTS-1 and SV-HUC-1 cells for flow cytometry binding, and found that it binds to bladder cancer cells but not to normal cells. The equilibrium dissociation constant of seq12a was measured, and its Kd was 0.79±0.18nM, indicating that the differential analysis based on nucleic acid aptamers can distinguish cancer cells from normal cells and is conducive to the selection of specific nucleic acid aptamers, such as Figure 4 As shown, Figure 4 A in the middle is the flow cytometry binding of nucleic acid aptamer seq12a on YTS-1 and SV-HUC-1 cells, where Random is a random control sequence. Figure 4 Where B is the equilibrium dissociation constant of aptamer seq12a.

[0143] In another specific embodiment, the present invention cooperates with the nucleic acid aptamer differential analysis of the probe package to conduct a preliminary study on the molecular typing of clinical urine samples. The specific results are as follows Figure 5 As shown, Figure 5 A in the middle is differential nucleic acid aptamer analysis, Figure 5 Figure B in the middle shows the sequence copy number analysis of differential nucleic acid aptamers in urine samples from five bladder cancer patients and five healthy subjects. The redder the color, the higher the copy number, and the bluer the color, the lower the copy number. Figure 5 Middle C is differential aptamer analysis, Figure 5 Figure D in the middle shows the sequence copy number analysis of differential nucleic acid aptamers in preoperative and postoperative urine samples from three bladder cancer patients. The redder the color, the higher the copy number, and the bluer the color, the lower the copy number.

[0144] like Figure 2 As shown, the present invention also provides a bladder cancer fine typing system based on nucleic acid aptamer differential analysis, comprising:

[0145] Receiving module 100: is used to receive sequencing data, wherein the sequencing data is obtained by incubating a sample containing bladder cancer cells with a probe package and performing high-throughput gene sequencing on the incubated sample.

[0146] Screening module 200: used to remove abnormal samples from the sequencing data to obtain nucleic acid aptamer data.

[0147] The standardization module 300 is used to perform standardization processing on the nucleic acid aptamer data to obtain standard data including a healthy sample group and a case sample group.

[0148] Selection module 400 is used to select differentially expressed aptamers based on the standard data.

[0149] The application module 500 is used for typing the sample to be tested by the differentially expressed aptamer.

[0150] The apparatus embodiments described above are merely illustrative, wherein the units described as separate components can or can not be physically separate, and the components displayed as units can or can not be physical units, i.e., can be located in one place or distributed on multiple network units. Part or all of the modules can be selected to achieve the purpose of the embodiment scheme according to actual needs. Those skilled in the art can understand and implement without creative labor.

[0151] From the above description of the embodiments, those skilled in the art can clearly understand that the embodiments can be realized by means of software plus a necessary general hardware platform, and of course can also be realized by hardware. Based on such understanding, the above technical solutions can be embodied in the form of a software product, which can be stored in a computer readable storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, etc., and includes a plurality of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute the methods described in the various embodiments or some parts of the embodiments.

[0152] The application discloses a bladder cancer fine typing method and system based on nucleic acid aptamer difference analysis, and mainly aims to realize preliminary molecular typing and clinical monitoring of bladder cancer based on nucleic acid aptamer difference analysis and a bladder cancer nucleic acid aptamer probe package. The bladder cancer nucleic acid aptamer probe package contains a plurality of nucleic acid aptamers for recognizing different levels of bladder cancer cell lines. The different aptamers can be selected from high-throughput sequencing results by nucleic acid aptamer difference analysis, and can distinguish bladder cancer cells and normal cells. The significantly up-regulated aptamer seq12 is combined with bladder cancer cells and not combined with normal cells, has good specificity and affinity (Kd=0.79±0.18nM), and finally, the application also attempts to realize typing of clinical patients and healthy people and clinical course monitoring by using the nucleic acid aptamer difference analysis and the probe package. Through the typing technology based on the nucleic acid aptamer, the early detection rate and treatment effect of bladder cancer can be improved, so that the recurrence risk and mortality of patients can be reduced.

[0153] It should be pointed out finally that the above embodiments are only used to illustrate the technical solutions of the present application, but not to limit the same; and although the present application has been described in detail with reference to the foregoing embodiments, it should be appreciated by those skilled in the art that the technical solutions recorded in the foregoing embodiments can be modified, or some technical features thereof can be replaced equivalently; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for fine typing of bladder cancer based on nucleic acid aptamer differential analysis, characterized in that: include: S1: Acquiring sequencing data, wherein the sequencing data is obtained by incubating a sample containing bladder cancer cells with a probe package and performing high-throughput gene sequencing on the incubated sample; S2: removing abnormal samples from the sequencing data to obtain nucleic acid aptamer data; S3: normalizing the nucleic acid aptamer data to obtain standard data including a healthy sample group and a case sample group; Wherein, step S3 further comprises: S31: dividing the samples in the nucleic acid aptamer data into a healthy sample group and a case sample group; S32: selecting the median of the healthy sample group and the case sample group respectively to obtain a reference sample; S33: calculating the expression difference value and expression level value of each aptamer based on the reference sample; S34: calculating a normalization factor based on the expression difference value; S35: normalizing the nucleic acid aptamer data based on the normalization factor to obtain standard data; The expression of the expression difference value in step S33 is: M=log2(CPM sample )-log2(CPM reference ); Where M is the expression difference value of a single aptamer between the reference sample and other samples, CPM sample is the CPM value of other samples, CPM reference is the CPM value of the reference sample; The expression of the expression level value in step S33 is: A=0.5×)log2(CPM sample )+log2(CPM reference )); Where A is the expression level value of a single aptamer between the reference sample and other samples; The expression of the normalization factor in step S34 is: scaling factor=2 mean(M) ; Among them, scaling factor is the normalization factor, mean (M) is the mean of the expression difference values ​​of the screened samples; The expression of the standard data in step S35 is: CPM normalized =log2(CPM original ×scaling factor+1); Among them, CPM normalized is the sample CPM value in the standard data, CPM original is the original CPM value of a single aptamer; S4: selecting differentially expressed aptamers based on the standard data; Wherein, step S4 further comprises: S41: fitting the healthy sample group and the case sample group respectively by a linear model to obtain a fitted healthy sample group and a fitted case sample group; S42: calculating the difference in expression results between the fitted healthy sample group and the fitted case sample group; S43: selecting the aptamer group corresponding to the maximum difference in expression results as the differentially expressed aptamer; S5: Typing the sample to be tested using the differentially expressed aptamers.

2. The method for fine typing of bladder cancer based on nucleic acid aptamer differential analysis according to claim 1, characterized in that: Step S2 further comprises: S21: receiving sequencing data based on molecular ID tags; S22: screening the sequencing data to obtain filtered data; S23: Merge the filtered data corresponding to each sample to obtain group data; S24: Screening the group data to obtain nucleic acid aptamer data.

3. The method for fine typing of bladder cancer based on nucleic acid aptamer differential analysis according to claim 2, characterized in that: Step S22 further includes: S221: Filtering aptamers in the sequencing data whose aptamer length is less than or equal to a first preset threshold to obtain first filtered data; S222: Calculating the longest common subsequence between the aptamer and the primer in the sequencing data by a dynamic programming algorithm; S223: Remove primer dimers whose primer lengths and the longest common subsequence are greater than a second preset threshold in the first filtered data to obtain second filtered data, and output the second filtered data as filtered data.

4. The method for fine typing of bladder cancer based on nucleic acid aptamer differential analysis according to claim 3, characterized in that: Step S222 further includes: S2221: Construct a two-dimensional array; S2222: traversing the bases of the sequences of the aptamers and primers in the sequencing data, filling the two-dimensional array, and obtaining a two-dimensional matrix; S2223: Output the tail array of the two-dimensional matrix as the longest common subsequence.

5. The method for fine typing of bladder cancer based on nucleic acid aptamer differential analysis according to claim 2, characterized in that: Step S22 further includes: S241: Calculating the CPM value of each aptamer based on the aptamer abundance of the group data; S242: Remove samples with CPM values ​​less than or equal to 1 from the group data to obtain the nucleic acid aptamer data.

6. The method for fine typing of bladder cancer based on nucleic acid aptamer differential analysis according to claim 1, characterized in that: The expression of the difference in the expression results in step S42 is: C=β1-β2; Among them, C is the difference in expression results of the current group, β1 is the model coefficient corresponding to the aptamer belonging to the fitting healthy sample group of the current group, and β2 is the model coefficient corresponding to the aptamer belonging to the case healthy sample group of the current group.

7. A bladder cancer fine typing system based on nucleic acid aptamer differential analysis, characterized in that: include: Receiving module: used to receive sequencing data, wherein the sequencing data is obtained by incubating a sample containing bladder cancer cells with a probe package and performing high-throughput gene sequencing on the incubated sample; Screening module: used to remove abnormal samples from the sequencing data and obtain nucleic acid aptamer data; Standardization module: used to standardize the nucleic acid aptamer data to obtain standard data including healthy sample group and case sample group; The standardization module is further used to divide the samples in the nucleic acid aptamer data into a healthy sample group and a case sample group; select the median of the healthy sample group and the case sample group respectively to obtain a reference sample; calculate the expression difference value and expression level value of each aptamer based on the reference sample; calculate the normalization factor based on the expression difference value; and normalize the nucleic acid aptamer data based on the normalization factor to obtain standard data; The expression of the expression difference value is: M=log2(CPM sample )-log2(CPM reference ); Where M is the expression difference value of a single aptamer between the reference sample and other samples, CPM sample is the CPM value of other samples, CPM reference is the CPM value of the reference sample; The expression of the expression level value is: A=0.5×(log29CPM sample )+log2(CPM reference )); Where A is the expression level value of a single aptamer between the reference sample and other samples; The expression of the normalization factor is: scaling factor=2 mean(M) ; Among them, scaling factor is the normalization factor, mean (M) is the mean of the expression difference values ​​of the screened samples; The expression of the standard data is: CPM normalized =log2(CPM original ×scaling factor+1); Among them, CPM normalized is the sample CPM value in the standard data, CPM original is the original CPM value of a single aptamer; Selection module: used for selecting differentially expressed aptamers based on the standard data; The selection module is further configured to fit the healthy sample group and the case sample group respectively through a linear model to obtain a fitted healthy sample group and a fitted case sample group; calculate the difference in expression results between the fitted healthy sample group and the fitted case sample group; and select the aptamer group corresponding to the maximum difference in expression results as the differentially expressed aptamer; Application module: used to perform typing on the sample to be tested using the differentially expressed aptamer.

Citation Information

Patent Citations

  • Nucleic acid aptamers for detecting bladder cancer and application thereof

    CN109266654A

  • Nucleic acid aptamer and application thereof

    CN114317545A