Application of reagent for detecting lung cancer specific cfDNA fragment characteristic combination in peripheral blood in preparation of lung cancer prediction product

By performing low-deep sequencing of peripheral blood cfDNA to extract characteristic values and combining machine learning algorithms, a lung cancer risk prediction model was constructed, solving the radiation risk and low sensitivity problems of traditional screening methods, and achieving efficient early-stage lung cancer screening and prediction.

CN120366451APending Publication Date: 2025-07-25SHENZHEN HAPLOX MEDICAL TESTING LABORATORY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510249894.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

In the prior art, traditional lung cancer screening methods such as chest X-rays and low-dose spiral CT have problems with the risk of radiation exposure and many false positive results. The cancer screening strategy based on ctDNA mutation detection is insufficient in sensitivity, making it difficult to achieve efficient early-stage lung cancer recognition.

Method used

By conducting low-deep whole genome sequencing of peripheral blood cfDNA, the frequency of the 5’ end 6bp motif of the cfDNA fragment, the length distribution index and coverage characteristic value of the cfDNA fragment in the 10kb window of the genome were extracted, and the lung cancer risk prediction model was constructed in combination with machine learning algorithms, and early lung cancer screening and prediction were used to use the cfDNA fragment omic characteristics.

Benefits of technology

The constructed lung cancer risk prediction model has an AUC of more than 0.94, a sensitivity of more than 0.84, and a specificity of more than 0.89, which can accurately determine lung cancer samples and is suitable for risk assessment and prediction of early cancers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120366451A_ABST
    Figure CN120366451A_ABST
Patent Text Reader

Abstract

The invention discloses application of a reagent for detecting lung cancer specific cfDNA fragment characteristic combination in peripheral blood in preparation of a product for predicting lung cancer. The invention provides a plurality of lung cancer specific cfDNA fragment feature combinations, relates to three dimensions of 5 '-terminal 6bp motif frequency of a cfDNA fragment, a length distribution index of the cfDNA fragment of a genome 10kb window and cfDNA fragment coverage of the genome 10kb window, and is also combined with a machine learning algorithm to construct a lung cancer risk prediction model, the AUC is higher than 0.94, the sensitivity is higher than 0.84, and the specificity is higher than 0.89. The method is suitable for risk assessment and prediction of early cancers, has excellent classification performance, and can realize accurate determination of lung cancer samples.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of biomedical diagnostic technology, and in particular, to the use of a reagent for detecting a characteristic combination of lung cancer-specific cfDNA fragments in peripheral blood in the preparation of a product for lung cancer prediction. Background Art

[0002] Lung cancer is one of the malignant tumors with the highest morbidity and mortality in the world. Its high lethality is mainly attributed to its hidden early symptoms and difficulty in detection. Although traditional screening methods such as chest X-rays and low-dose spiral CT (LDCT) can improve the early detection rate to a certain extent, they have certain limitations, such as radiation exposure risks and a large number of false positive results. Therefore, finding more sensitive and specific biomarkers is crucial to improving the clinical management of lung cancer patients.

[0003] With the advent of the big data era, artificial intelligence algorithms, especially machine learning and deep learning technologies, have been widely used in the medical and health fields. Those skilled in the art have built a prediction model for lung cancer based on gene sequencing technology in order to more accurately identify early lung cancer patients, assess disease progression, and predict treatment effects.

[0004] In recent years, with the development of liquid biopsy technology, cfDNA has gradually attracted widespread attention as a new type of non-invasive biomarker. cfDNA refers to small fragments of DNA released into the blood circulation by normal cells or diseased tissues, which can carry changes in genetic information about primary or metastatic lesions. Studies have shown that cfDNA levels are usually elevated in cancer patients and may contain specific molecular features such as mutations and methylation changes, which are closely related to the occurrence and development of tumors. Since the level of ctDNA molecules derived from tumors in peripheral blood is extremely low, and not all ctDNA fragments carry mutations, cancer screening strategies based on ctDNA mutation detection often require extremely high sequencing depth, which poses a huge challenge to the sensitivity of screening. In comparison, the fragment omics characteristics of cfDNA are highly correlated with tissue source type and pathological state, with higher signal abundance, which can realize cancer screening and tissue tracing based on low-depth sequencing. The fragment omics characteristics of cfDNA are cancer molecular markers that have received much attention in recent years, including the length distribution of cfDNA fragments, terminal motif frequency, fragment breakpoint coordinates, nucleosome imprints, and different topological structures. There is an urgent need to establish a machine learning model based on cfDNA data to accurately predict lung cancer. Summary of the invention

[0005] In order to solve the problems existing in the prior art, the present invention provides the use of a reagent for detecting a characteristic combination of lung cancer-specific cfDNA fragments in peripheral blood in the preparation of a product for predicting lung cancer.

[0006] An object of the present invention is to provide an application of a reagent for detecting a characteristic combination of lung cancer-specific cfDNA fragments in peripheral blood in the preparation of a product for lung cancer prediction.

[0007] Another object of the present invention is to provide a method for constructing a lung cancer risk prediction model.

[0008] Still another object of the present invention is to provide a system for predicting lung cancer risk.

[0009] Still another object of the present invention is to provide a computer device.

[0010] Still another object of the present invention is to provide a computer-readable storage medium.

[0011] In order to achieve the above objects, the present invention is implemented by the following solutions:

[0012] The present invention performs low-depth whole-genome sequencing on peripheral blood cfDNA, extracts and integrates cfDNA fragment characteristic values in 3 dimensions, and provides a new marker combination that can be used for early screening and prediction of lung cancer. The cfDNA fragment characteristic values provided by the present invention are the 6bp motif frequency 6bp-f at the 5' end of the cfDNA fragment, the length distribution index F-Index_bin of the cfDNA fragments in the 10kb window of the genome, and the coverage Coverage_bin of the cfDNA fragments in the 10kb window of the genome.

[0013] The 6bp-f is the occurrence frequency of each 6bp motif in the terminal motif set; the terminal motif set includes Set 1 and Set 2, and Set 1 is composed of the 5' end 6bp motifs of the nucleotide sequences of each cfDNA fragment; Set 2 is composed of the reverse complementary sequences of the 3' end 6bp motifs of the nucleotide sequences of each cfDNA fragment; the nucleotide sequence of the cfDNA fragment is obtained by aligning paired-end reads to the human reference genome; the paired-end reads are obtained by the following method: obtaining the paired-end sequencing data of peripheral blood cfDNA of a sample to be tested, and retaining the paired-end reads aligned to the human reference genome. Based on the paired-end reads, information of the cfDNA fragment is obtained, and each pair of paired-end reads corresponds to a cfDNA fragment. The information of the cfDNA fragment includes but is not limited to the position of the cfDNA in the genome, the nucleotide sequence of the cfDNA, and the length of the cfDNA.

[0014] The 10kb windows of the genome are obtained by the following method: the genomic regions with redundant regions removed are non-overlappingly segmented into intervals of 10kb to obtain M 10kb windows of the genome; the redundant regions include the short arm regions, telomere regions and centromere regions of 5 chromosomes, and the 5 chromosomes are chromosome 13, chromosome 14, chromosome 15, chromosome 21 and chromosome 22.

[0015] The Coverage_bin is calculated from the 10kb windows of the genome and formula (2);

[0016] Formula (2): Coverage_bin k = N k / N total ;

[0017] where Coverage_bin k represents the coverage of cfDNA fragments in the k-th 10kb-bins, 1 ≤ k ≤ M; N k represents the number of cfDNA fragments aligned to the k-th 10kb-bins; N total represents the number of all cfDNA fragments aligned to the reference genome based on paired-end reads;

[0018] The F-Index_bin is calculated from the 10kb windows of the genome and formula (3);

[0019] Formula (3):

[0020] where i represents the length of the cfDNA fragment; P is a constant value, 160bp - 170bp; X' i represents the proportion of cfDNA fragments with length i among all cfDNA fragments in the k-th 10kb window of the genome.

[0021] Preferably, P is 166bp.

[0022] Use of a reagent for detecting a characteristic combination of lung cancer-specific cfDNA fragments in peripheral blood in the preparation of a product for lung cancer prediction, characterized in that the lung cancer-specific cfDNA fragment characteristic combination includes the frequency of the 5'-terminal 6bp motif of the cfDNA fragment combination (1), the length distribution index of the cfDNA fragments in the genomic 10kb window combination (1), and the coverage of the cfDNA fragments in the window Chr2:111440000 - 111449999.

[0023] The cfDNA fragment combination (1) consists of cfDNA fragments with 6-bp motifs of AAGGGT, ACCTCT, AAGTGC, AAAGTG, ACCTGT, ACCTAG, AAGAGG, ACCATG, and TTTAGG at the 5'-end;

[0024] The genomic 10-kb window combination (1) consists of the Chr7:36690000-36699999 window, the Chr6:21130000-21139999 window, the Chr10:19940000-19949999 window, the Chr3:143480000-143489999 window, the Chr9:33660000-33669999 window, and the Chr1:184650000-184659999 window.

[0025] The application of a reagent for detecting the combination of lung cancer-specific cfDNA fragment characteristics in peripheral blood in the preparation of a product for lung cancer prediction, wherein the combination of lung cancer-specific cfDNA fragment characteristics includes the frequency of the 6-bp motifs at the 5'-end of the cfDNA fragment combination (2), the length distribution index of the cfDNA fragments in the genomic 10-kb window combination (2), and the coverage of the cfDNA fragments in the genomic 10-kb window combination (3);

[0026] The cfDNA fragment combination (2) consists of cfDNA fragments with 6-bp motifs of AAGGGT, ACCTCT, TTTAGG, AAGTGC, ACCACT, AAGAGG, ATGGGA, AAGGGG, AAGCAG, TTTAAG, and CACCAT at the 5'-end;

[0027] The genomic 10-kb window combination (2) consists of the Chr11:4700000-4709999 window, the Chr13:47010000-47019999 window, the Chr1:173190000-173199999 window, the Chr10:19940000-19949999 window, and the Chr1:184650000-184659999 window;

[0028] The genomic 10-kb window combination (3) consists of the Chr7:73130000-73139999 window, the Chr10:125260000-125269999 window, the Chr6:163640000-163649999 window, and the Chr2:234550000-234559999 window.

[0029] Use of a reagent for detecting a combination of characteristics of lung cancer-specific cfDNA fragments in peripheral blood in the preparation of a product for predicting lung cancer, characterized in that the combination of characteristics of the lung cancer-specific cfDNA fragments includes the frequency of the 5'-terminal 6bp motif of the cfDNA fragment combination (3); the cfDNA fragment combination (3) consists of cfDNA fragments with 5'-terminal 6bp motifs of ACCATG, ACCTGT, ACCTCT, CAGTTG, ACCCAA, ACTAAA, ACCTAG, CAGTCA, ACCCTA, ACCCTT, ACTCGT, TTTAGG, ACCCTG, ACTCTG, ACCACA, AAGTGC, ACCTAC, ACTCAC, AAGGGT, and TTAAGG.

[0030] Use of a reagent for detecting a combination of characteristics of lung cancer-specific cfDNA fragments in peripheral blood in the preparation of a product for predicting lung cancer, characterized in that the combination of characteristics of the lung cancer-specific cfDNA fragments includes the frequency of the 5'-terminal 6bp motif of the cfDNA fragment combination (4), the length distribution index of the cfDNA fragments of the genomic 10kb window combination (4), and the coverage of the cfDNA fragments of the genomic 10kb window combination (5);

[0031] The cfDNA fragment combination (4) consists of cfDNA fragments with 5'-terminal 6bp motifs of AGTACT, ATGAGG, GTTGGA, CCCAAA, ACGATA, ACCAGT, GGCTAT, and AAGTGC;

[0032] The genomic 10kb window combination (4) consists of the Chr4:113740000-113749999 window and the Chr11:3980000-3989999 window;

[0033] The genomic 10kb window combination (5) consists of the Chr17:15620000-15629999 window, the Chr5:17950000-17959999 window, the Chr5:8670000-8679999 window, the Chr8:21550000-21559999 window, the Chr3:58170000-58179999 window, the Chr1:120080000-120089999 window, the Chr1:121070000-121079999 window, the Chr7:73130000-73139999 window, the Chr4:8180000-8189999 window, and the Chr1:61070000-61079999 window.

[0034] A method for constructing a lung cancer risk prediction model, using any of the described lung cancer-specific cfDNA fragment feature combinations as input features, lung cancer risk probability as the output result, and AUC as the performance index to train the classification model to obtain the lung cancer risk prediction model.

[0035] Preferably, the classification model includes random forest, support vector machine, gradient boosting machine or LASSO regression.

[0036] More preferably, the classification model is a gradient boosting machine.

[0037] Preferably, the cutoff value of the classification model is set to 0.5, and the determination criterion is: if the lung cancer risk probability ≥ 0.5, the sample is determined to be positive (i.e., a lung cancer sample); if the lung cancer risk probability < 0.5, the sample is determined to be negative (i.e., a non-lung cancer sample); the training is performed based on 10-fold repeated cross-validation.

[0038] A system for predicting lung cancer risk, including a data acquisition module, an analysis module and an output module; the data acquisition module is used to acquire the lung cancer-specific cfDNA fragment feature combination of the sample to be tested, and the lung cancer-specific cfDNA fragment feature combination is any of the described lung cancer-specific cfDNA fragment feature combinations;

[0039] The analysis module is the lung cancer risk prediction model obtained by any of the described construction methods, using the lung cancer-specific cfDNA fragment feature combination of the sample to be tested acquired by the data acquisition module as the input variable, and inputting it into the lung cancer risk prediction model to obtain the lung cancer risk probability;

[0040] The output module is used to output the lung cancer risk probability.

[0041] Preferably, the analysis module also determines the type of the sample to be tested based on the lung cancer risk probability, and the determination criterion is: if the lung cancer risk probability ≥ 0.5, the sample to be tested is a lung cancer sample; if the lung cancer risk probability < 0.5, the sample to be tested is a non-lung cancer sample; the output module is also used to output the type of the sample to be tested.

[0042] Preferably, the system further includes a storage module, and the storage module is used to store the lung cancer-specific cfDNA fragment feature combination of the sample to be tested obtained by the data acquisition module and the lung cancer risk probability obtained by the analysis module.

[0043] A computer device, including a memory and a processor, characterized in that a computer program executable on the processor is stored on the memory; when the computer program is executed by the processor, the operations of the system for predicting lung cancer risk are implemented.

[0044] A computer-readable storage medium stores a computer program executable by a processor. When the computer program is executed by the processor, the operations of the system for predicting lung cancer risk are implemented.

[0045] Compared with the prior art, the present invention has the following beneficial effects:

[0046] The present invention provides multiple combinations of lung cancer-specific cfDNA fragment features, and also constructs a lung cancer risk prediction model in combination with a machine learning algorithm. The AUC is higher than 0.94, the sensitivity is higher than 0.84, and the specificity is higher than 0.89. It is suitable for risk assessment and prediction of early cancers, has excellent classification performance, and can accurately determine lung cancer samples. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 It is an AUC graph of the lung cancer risk prediction model constructed by the present invention based on a machine learning algorithm in combination with lung cancer-specific cfDNA fragment feature combinations. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0048] The present invention will be further elaborated in detail below in conjunction with the accompanying drawings of the specification and specific embodiments. The embodiments are only used to explain the present invention and are not used to limit the scope of the present invention. The test methods used in the following embodiments are all conventional methods unless otherwise specified; the materials, reagents, etc. used are all reagents and materials that can be obtained from commercial channels unless otherwise specified.

[0049] Example 1 Obtaining Lung Cancer-Specific cfDNA Fragment Feature Combinations and Constructing a Lung Cancer Risk Prediction Model Based on Random Forest

[0050] I. Subject Information

[0051] A total of 398 peripheral blood samples of lung cancer patients (i.e., lung cancer samples) and 168 peripheral blood samples of healthy people (i.e., healthy samples) were included in this example. Each peripheral blood sample was from a different individual and all were from the sample library of Shenzhen Hypros Biotechnology Co., Ltd. Informed consent of the subjects was obtained before sample testing.

[0052] II. Determination of Lung Cancer-Specific cfDNA Fragment Feature Combinations in Peripheral Blood

[0053] 1. Peripheral Blood cfDNA Sequencing

[0054] cfDNA was extracted from the peripheral blood samples of the subjects using the circulating Nucleic Acid Kit, and the cfDNA library was constructed using the library construction kit QIAseq cfDNA All-in-One Kit (24) (Cat.No. / ID: 180023). During the library construction process, cfDNA was not fragmented. The Illumina Novaseq sequencing platform was used to perform paired-end sequencing on the cfDNA library, and the sequencing read length was ~150 bp. After the sequencing was completed, the original fastq sequencing data was obtained. The Fastp software was used to perform quality control processing on the original fastq data based on the default parameters, removing adapters, data with base quality values lower than Q30, sequences with a high N content (i.e., removing read sequences with an N-Content greater than 2%), and PCR-introduced duplicate sequences, resulting in paired-end reads. Each cfDNA fragment matched a pair of paired-end reads (read1 and read2).

[0055] Using the bwa software, the quality-controlled paired-end reads were aligned to the human reference genome hg38 (Genome Reference Consortium human genome build 38, GRCh38) to obtain the Bam file of the alignment results. The picard tool was used to preprocess the Bam file, removing the paired-end reads with multiple alignments, removing the paired-end reads that did not align to the reference genome, and retaining the paired-end reads with a quality value greater than 30, resulting in a clean bam file.

[0056] 2. Obtaining cfDNA fragment characteristics

[0057] The present invention relates to a total of 4 types of cfDNA fragment characteristics: the global length distribution index of cfDNA fragments, the 6-bp motif at the 5' end of cfDNA fragments, the cfDNA fragment coverage of the 10-kb window of the genome, and the length distribution index of cfDNA fragments in the 10-kb window of the genome.

[0058] (1) Calculation of the characteristic value all-F-Index of the global length distribution index of cfDNA fragments

[0059] Extract the alignment coordinates, orientation, length, and base sequence information of each pair of paired-end reads from the clean bam file obtained in the previous step. In this analysis step, only retain the paired-end reads aligned to the autosomal regions 1 to 22, and remove the paired-end reads aligned to the X and Y chromosomes, mitochondrial genome, or other supplementary contig regions to obtain a bed file of the cfDNA fragment distribution on the genome (i.e., the genomic bed file). This file contains cfDNA fragment information in the genome, including: cfDNA fragment sequence, length, and position information on the chromosome.

[0060] Calculate the all-F-Index based on the genomic bed file and formula (1).

[0061] Formula (1):

[0062] Among them, i represents the length of the cfDNA fragment; P is a constant value of 166bp, which is the general length of cfDNA fragments in the plasma of healthy individuals; X i represents the proportion of cfDNA fragments with length i among all cfDNA fragments.

[0063] (2) Calculation of the eigenvalue of the 5'-terminal 6bp motif of cfDNA fragments

[0064] The eigenvalue of the 5'-terminal 6bp motif of cfDNA fragments is the 5'-terminal 6bp motif frequency (6bp-f) of cfDNA fragments. This value reflects the influence of nuclease activity and tissue type on the terminal motif pattern of cfDNA fragments. The statistical method is as follows:

[0065] For each cfDNA fragment, each pair of paired-end reads consists of read1 on the forward strand and read2 on the reverse strand. By aligning to the reference genome, the forward complete sequence of the cfDNA fragment (i.e., the nucleotide sequence from the 5'-end to the 3'-end) is obtained. The 6bp sequence at the 5'-end is defined as End_motif_F, and the 6bp sequence at the 3'-end is defined as End_motif_R. After that, the sequence of End_motif_F is not transformed, and all End_motif_F of cfDNA fragments form set 1; the sequence of End_motif_R is transformed by reverse complementary base conversion to obtain the reverse complementary sequence of End_motif_R, and all reverse complementary sequences of End_motif_R of cfDNA fragments form set 2. After combining set 1 and set 2, the set of 6bp motifs at the 5'-end of all cfDNA fragments is obtained, denoted as the terminal motif set. Theoretically, there are at most 4096 different 6bp motifs, such as {AAAAAA,AAAAAT,...,ACGACG,CTTGTT,...,GGGGGG}. By counting the occurrence frequency of each 6bp motif in the terminal motif set, 6bp-f is obtained.

[0066] (3) Calculation of the coverage characteristic value of cfDNA fragments in 10kb windows of the genome

[0067] The calculation method of the coverage characteristic value Coverage_bin of cfDNA fragments in 10kb windows of the genome is as follows:

[0068] For human chromosomes 13, 14, 15, 21, and 22, since the short arm regions of these 5 chromosomes contain nucleolus organizer regions and are rich in a large number of repetitive ribosomal RNA genes, it is difficult for genome sequencing technology to accurately analyze the complete sequences therein. Therefore, based on the bed file of the cfDNA fragment distribution on the genome obtained in this embodiment, the short arm regions of these 5 chromosomes as well as the telomere and centromere regions are removed, and the remaining genomic regions are non-overlappingly segmented according to 10kb intervals (such as chr1:20000~chr1:29999) to obtain M 10kb windows of the genome (i.e., 10kb-bins). Theoretically, at most 287465 10kb-bins can be obtained, that is, the maximum value of M is 287465.

[0069] Coverage_bin is calculated according to the genome bed file, 10kb-bins and formula (2).

[0070] Formula (2): Coverage_bin k =N k / N total ;

[0071] Among them, Coverage_bin k represents the coverage of cfDNA fragments in the k-th 10kb-bins, where 1 ≤ k ≤ M; N k represents the number of cfDNA fragments aligned to the k-th 10kb-bins. Among them, reads spanning adjacent two 10kb-bins are regarded as being aligned to both 10kb-bins simultaneously. When calculating the coverage, both 10kb-bins on both sides increase by one count. When calculating the F-Index_bin k subsequently, this read is also included in the calculation; N total represents the number of all cfDNA fragments aligned to the reference genome based on paired-end reads.

[0072] (4) Calculation of the length distribution index F-Index_bin of cfDNA fragments in the 10kb window of the genome

[0073] The F-Index_bin of cfDNA fragments in the k-th 10kb-bins is calculated according to the genome bed file, 10kb-bins, and formula (3).

[0074] Formula (3):

[0075] Among them, i represents the length of the cfDNA fragment; P is a constant value of 166bp, and this value is the general length of cfDNA fragments in the plasma of healthy people; X’ i represents the proportion of cfDNA fragments with length i among all cfDNA fragments in the k-th 10kb-bins.

[0076] Through the calculation and statistics of this embodiment, for each sample, 1 all-F-Index value, and at most 4096 6bp-f values, 287465 Coverage_bin values, and 287465 F-Index_bin values are obtained.

[0077] 3. Determine the lung cancer-specific cfDNA fragment feature combination based on RandomForest

[0078] (1) Feature matrix generation

[0079] According to the above step, all cfDNA fragment feature values of each sample are obtained and a feature matrix is generated. Each row represents a sample, and the columns represent each cfDNA fragment feature and its feature value. All feature values are normalized and normalized transformation.

[0080] (2) Data division

[0081] Randomly divide the feature matrix obtained in the previous step into a training set and a validation set according to the samples. The training set contains the feature matrices of 239 lung cancer samples and the feature matrices of 101 healthy samples. The validation set samples contain the feature matrices of the remaining 159 lung cancer samples and the feature matrices of 67 healthy samples.

[0082] (3) Feature selection

[0083] Use the feature matrix of the training set as the input features, and perform feature selection based on a machine learning model. Specifically: Apply the Recursive Feature Elimination (RFE) algorithm combined with a model constructed by RandomForest for feature screening. When constructing the model, the model parameters are selected for 10-fold repeated cross-validation, and the number of repeated sampling iterations is 10. Use the accuracy of predicting whether a patient has lung cancer as the result evaluation parameter, and return the optimal number of features and feature combinations.

[0084] The top 20 feature combinations ranked by relative weight are shown in Table 1, which are the lung cancer-specific cfDNA fragment feature combinations. There are 6bp-f of the 5'-end 6bp motifs (AAGGGT, ACCTCT, AAGTGC, AAAGTG, ACCTGT, ACCTAG, AAGAGG, ACCATG, and TTTAGG) of 13 cfDNA fragments, the length distribution indices of cfDNA fragments in 6 genomic 10kb windows (Chr7:36690000-36699999 window, Chr6:21130000-21139999 window, Chr10:19940000-19949999 window, Chr3:143480000-143489999 window, Chr9:33660000-33669999 window, and Chr1:184650000-184659999 window) (Findex_bin126860, Findex_bin108225, Findex_bin169468, Findex_bin63460, Findex_bin157002, and Findex_bin18465 in sequence), and the cfDNA fragment coverage (Coverage_bin36038) of 1 genomic 10kb window (Chr2:111440000-111449999 window).

[0085] Table 1 Lung cancer-specific cfDNA fragment feature combinations obtained based on RandomForest

[0086]

[0087]

[0088] 4. Construction of Lung Cancer Risk Prediction Model Based on RandomForest and Performance Evaluation

[0089] (1) Model Construction

[0090] For the training set of this embodiment, the lung cancer-specific cfDNA fragment feature combinations of each sample in the training set are obtained as shown in Table 1, and used as the input features of the model (RandomForest), with the lung cancer risk probability as the output result, and the output result range is [0, 1]. The greater the lung cancer risk probability, the higher the possibility that the sample is a lung cancer sample. Combining the actual status (lung cancer and non-lung cancer) of each sample in the training set, 10-fold repeated cross-validation and 10 repeated iterations are performed, and the optimal training model is output based on the performance of the AUC index.

[0091] During the 10-fold cross-validation process, the cutoff value is 0.5, and the judgment criterion is: if the lung cancer risk probability ≥ 0.5, the sample is judged as positive (i.e., a lung cancer sample); if the lung cancer risk probability < 0.5, the sample is judged as negative (i.e., a non-lung cancer sample).

[0092] (2) Performance Evaluation

[0093] Formula (4): Sensitivity = TP / (TP + FN);

[0094] Formula (5): Specificity = TN / (TN + FP);

[0095] Formula (6): PPV = TP / (TP + FP);

[0096] Formula (7): NPV = TN / (TN + FN);

[0097] Formula (8): Accuracy = (TP + TN) / (TP + FN + TN + FP).

[0098] Use the validation set samples of this embodiment to evaluate the performance of the optimal training model obtained in step (1) of this embodiment, and calculate 5 indicators according to formulas (4) to (8): sensitivity (Sensitivity), specificity (Specificity), positive predictive value (Positive Predictive Value, PPV), negative predictive value (Negative Predictive Value, NPV), and comprehensive accuracy (Accuracy).

[0099] As Figure 1As shown, the best AUC of the lung cancer risk prediction model constructed based on the RandomForest algorithm in 10-fold repeated cross-validation is 0.953. In addition, the sensitivity of this model is calculated to be 0.874, the specificity is 0.896, the positive predictive value is 0.952, the negative predictive value is 0.75, and the comprehensive accuracy is 0.881. It shows that the lung cancer risk prediction model constructed based on the RandomForest algorithm has excellent classification performance and can accurately determine whether a sample is a lung cancer sample.

[0100] Example 2 Obtaining a lung cancer-specific cfDNA fragment feature combination and constructing a lung cancer risk prediction model based on Gradient Boosting Machine (GBM)

[0101] I. Subject Information

[0102] Same as Example 1.

[0103] II. Determination of the lung cancer-specific cfDNA fragment feature combination in peripheral blood

[0104] 1. Peripheral blood cfDNA sequencing

[0105] Same as Example 1.

[0106] 2. Obtaining cfDNA fragment features

[0107] Same as Example 1.

[0108] 3. Determining the lung cancer-specific cfDNA fragment feature combination based on GBM

[0109] Basically the same as Example 1, except that: in the feature selection in the third step, a recursive feature elimination algorithm is applied to combine the model constructed by GBM for feature screening. The top 20 feature combinations ranked by relative weight are shown in Table 2, which are the lung cancer-specific cfDNA fragment feature combinations. There are 6bp-f of 6bp motifs (AAGGGT, ACCTCT, TTTAGG, AAGTGC, ACCACT, AAGAGG, ATGGGA, AAGGGG, AAGCAG, TTTAAG, and CACCAT) at the 5' ends of 11 cfDNA fragments, the length distribution indices of cfDNA fragments in 5 genomic 10kb windows (Chr11:4700000-4709999 window, Chr13:47010000-47019999 window, Chr1:173190000-173199999 window, Chr10:19940000-19949999 window, and Chr1:184650000-184659999 window) (Findex_bin181322, Findex_bin212386, Findex_bin17319, Findex_bin169468, and Findex_bin18465 in sequence), and the coverage of cfDNA fragments in 4 genomic 10kb windows (Chr7:73130000-73139999 window, Chr10:125260000-125269999 window, Chr6:163640000-163649999 window, and Chr2:234550000-234559999 window) (Coverage_bin130504, Coverage_bin180000, Coverage_bin122476, and Coverage_bin48349 in sequence).

[0110] Table 2 Lung cancer-specific cfDNA fragment feature combinations obtained based on GBM

[0111]

[0112]

[0113] 4. Construction of a lung cancer risk prediction model based on GBM and performance evaluation

[0114] (1) Model construction

[0115] Basically the same as Example 1, except that: GBM is used for model construction.

[0116] (2) Performance evaluation

[0117] According to the method of Example 1, the performance of the optimal training model obtained in step (1) of this example is evaluated using the validation set samples of this example.

[0118] As Figure 1 shown, the best AUC of the lung cancer risk prediction model constructed based on the GBM algorithm in 10-fold repeated cross-validation is 0.961. In addition, the sensitivity of this model is calculated to be 0.874, the specificity is 0.94, the positive predictive value is 0.972, the negative predictive value is 0.759, and the comprehensive accuracy is 0.894. It shows that the lung cancer risk prediction model constructed based on the GBM algorithm has excellent classification performance and can accurately determine whether the sample is a lung cancer sample.

[0119] Example 3 Obtaining a lung cancer-specific cfDNA fragment feature combination and constructing a lung cancer risk prediction model based on support vector machine (svmLinear)

[0120] I. Subject information

[0121] The same as in Example 1.

[0122] II. Determination of the lung cancer-specific cfDNA fragment feature combination in peripheral blood

[0123] 1. Peripheral blood cfDNA sequencing

[0124] The same as in Example 1.

[0125] 2. Obtaining cfDNA fragment features

[0126] The same as in Example 1.

[0127] 3. Determining the lung cancer-specific cfDNA fragment feature combination based on svmLinear

[0128] Basically the same as in Example 1, the difference is that in the feature selection of the third step, the recursive feature elimination algorithm is applied to screen features in combination with the model constructed by svmLinear. The top 20 feature combinations ranked by relative weight are shown in Table 3, which are the lung cancer-specific cfDNA fragment feature combinations, and there are 6bp-f of the 5'-end 6bp motifs (ACCATG, ACCTGT, ACCTCT, CAGTTG, ACCCA A, ACTAAA, ACCTAG, CAGTCA, ACCCTA, ACCCTT, ACTCGT, TTTAGG, ACCCTG, ACTCTG, ACCACA, AAGTGC, ACCTAC, ACTCAC, AAGGGT and TTAAGG) of 20 cfDNA fragments.

[0129] Table 3 Lung cancer-specific cfDNA fragment feature combination obtained based on svmLinear

[0130] Serial Number Type Feature Name Description Relative Weight 1 6bp Motif at the 5' End of cfDNA Fragment ACCATG 6bp-f of ACCATG 0.953 2 6bp Motif at the 5' End of cfDNA Fragment ACCTGT 6bp-f of ACCTGT 0.949 3 6bp Motif at the 5' End of cfDNA Fragment ACCTCT 6bp-f of ACCTCT 0.947 4 6bp Motif at the 5' End of cfDNA Fragment CAGTTG 6bp-f of CAGTTG 0.946 5 6bp Motif at the 5' End of cfDNA Fragment ACCCAA 6bp-f of ACCCAA 0.945 6 6bp Motif at the 5' End of cfDNA Fragment ACTAAA 6bp-f of ACTAAA 0.944 7 6bp Motif at the 5' End of cfDNA Fragment ACCTAG 6bp-f of ACCTAG 0.943 8 6bp Motif at the 5' End of cfDNA Fragment CAGTCA 6bp-f of CAGTCA 0.943 9 6bp Motif at the 5' End of cfDNA Fragment ACCCTA 6bp-f of ACCCTA 0.942 10 6bp Motif at the 5' End of cfDNA Fragment ACCCTT 6bp-f of ACCCTT 0.942 11 6bp Motif at the 5' End of cfDNA Fragment ACTCGT 6bp-f of ACTCGT 0.942 12 6bp Motif at the 5' End of cfDNA Fragment TTTAGG 6bp-f of TTTAGG 0.941 13 6bp Motif at the 5' End of cfDNA Fragment ACCCTG 6bp-f of ACCCTG 0.941 14 6bp Motif at the 5' End of cfDNA Fragment ACTCTG 6bp-f of ACTCTG 0.941 15 6bp Motif at the 5' End of cfDNA Fragment ACCACA 6bp-f of ACCACA 0.941 16 6bp Motif at the 5' End of cfDNA Fragment AAGTGC 6bp-f of AAGTGC 0.940 17 6bp Motif at the 5' End of cfDNA Fragment ACCTAC 6bp-f of ACCTAC 0.940 18 6bp Motif at the 5' End of cfDNA Fragment ACTCAC 6bp-f of ACTCAC 0.938 19 6bp Motif at the 5' End of cfDNA Fragment AAGGGT 6bp-f of AAGGG 0.937 20 6bp Motif at the 5' End of cfDNA Fragment TTAAGG 6bp-f of TTAAGG 0.937

[0131] 4. Construction of Lung Cancer Risk Prediction Model Based on svmLinear and Performance Evaluation

[0132] (1) Model Construction

[0133] It is basically the same as Example 1, except that: svmLinear is used for model construction.

[0134] (2) Performance Evaluation

[0135] According to the method of Example 1, the performance of the optimal training model obtained in step (1) of this example is evaluated using the validation set samples of this example.

[0136] As Figure 1 shown, the best AUC of the lung cancer risk prediction model constructed based on the svmLinear algorithm in 10-fold repeated cross-validation is 0.948. In addition, the sensitivity of this model is calculated to be 0.849, the specificity is 0.91, the positive predictive value is 0.957, the negative predictive value is 0.718, and the comprehensive accuracy is 0.867. It shows that the lung cancer risk prediction model constructed based on the svmLinear algorithm has excellent classification performance and can accurately determine whether the sample is a lung cancer sample.

[0137] Example 4 Obtaining Lung Cancer-Specific cfDNA Fragment Features and Constructing a Lung Cancer Risk Prediction Model Based on the Regression Algorithm (LASSO)

[0138] I. Subject Information

[0139] It is the same as Example 1.

[0140] II. Determination of Lung Cancer-Specific cfDNA Fragment Feature Combinations in Peripheral Blood

[0141] 1. Peripheral Blood cfDNA Sequencing

[0142] It is the same as Example 1.

[0143] 2. Obtaining cfDNA Fragment Features

[0144] It is the same as Example 1.

[0145] 3. Determining Lung Cancer-Specific cfDNA Fragment Feature Combinations Based on LASSO

[0146] It is basically the same as Example 1, except that in the feature selection in the third step, a recursive feature elimination algorithm is applied to combine with the model constructed by LASSO for feature screening. The top 20 feature combinations ranked by relative weight are shown in Table 4, which are the lung cancer-specific cfDNA fragment feature combinations. There are 6bp-f of the 5'-end 6bp motifs (AGTACT, ATGAGG, GTTGGA, CCCAAA, ACGATA, ACCAGT, GGCTAT, and AAGTGC) of 8 cfDNA fragments, the length distribution indices of cfDNA fragments in 2 genomic 10kb windows (Chr4:113740000-113749999 window and Chr11:3980000-3989999 window) (Findex_bin80314 and Findex_bin181250 in sequence), and the coverage of cfDNA fragments in 10 genomic 10kb windows (Chr17:15620000-15629999 window, Chr5:17950000-17959999 window, Chr5:8670000-8679999 window, Chr8:21550000-21559999 window, Chr3:58170000-58179999 window, Chr1:120080000-120089999 window, Chr1:121070000-121079999 window, Chr7:73130000-73139999 window, Chr4:8180000-8189999 window, and Chr1:61070000-61079999 window) (Coverage_bin250615, Coverage_bin89755, Cove rage_bin88827, Coverage_bin141279, Coverage_bin54929, Coverage_bin12008, Coverage_bin12107, Coverage_bin130504, Coverage_bin69758, and Coverage_bin6107 in sequence).

[0147] Table 4 Lung cancer-specific cfDNA fragment feature combinations obtained based on LASSO

[0148]

[0149]

[0150] 4. Construction of a lung cancer risk prediction model based on LASSO and performance evaluation

[0151] (1) Model construction

[0152] It is basically the same as Example 1, except that: for the LASSO regression algorithm, alpha is set to seq(0, 1, by = 0.05), lambda is set to 10^seq(-2, 2, length = 100), and the AUC is used as the performance index.

[0153] (2) Performance evaluation

[0154] According to the method of Example 1, the performance of the optimal training model obtained in step (1) of this example is evaluated using the validation set samples of this example.

[0155] As Figure 1 shown, the best AUC of the lung cancer risk prediction model constructed based on the LASSO algorithm in 10-fold repeated cross-validation is 0.959. In addition, the sensitivity of this model is calculated to be 0.893, the specificity is 0.91, the positive predictive value is 0.959, the negative predictive value is 0.782, and the overall accuracy is 0.782. It shows that the lung cancer risk prediction model constructed based on the LASSO algorithm has excellent classification performance and can accurately determine whether a sample is a lung cancer sample.

[0156] Example 5 A lung cancer prediction system based on a combination of lung cancer-specific cfDNA fragment features

[0157] This example provides a lung cancer prediction system based on a combination of lung cancer-specific cfDNA fragment features, including a data acquisition module, an analysis module, a storage module, and an output module.

[0158] The data acquisition module is used to obtain the combination of lung cancer-specific cfDNA fragment features of the test sample, and the combination of lung cancer-specific cfDNA fragment features is the combination of lung cancer-specific cfDNA fragment features obtained in Example 1.

[0159] The analysis module is the lung cancer risk prediction model constructed in Example 1. Using the combination of lung cancer-specific cfDNA fragment features of the test sample obtained by the data acquisition module as the input variable, it is input into the lung cancer risk prediction model to obtain the lung cancer risk probability, and the type of the test sample is determined based on the lung cancer risk probability. The determination criterion is: if the lung cancer risk probability ≥ 0.5, the test sample is a lung cancer sample; if the lung cancer risk probability < 0.5, the test sample is a non-lung cancer sample.

[0160] The storage module is used to store the combination of lung cancer-specific cfDNA fragment features of the test sample obtained by the data acquisition module, the lung cancer risk probability obtained by the analysis module, and the type of the test sample.

[0161] The output module is used to output the lung cancer risk probability and the type of the test sample.

[0162] Example 6: A lung cancer prediction system based on a combination of lung cancer-specific cfDNA fragment features

[0163] This example provides a lung cancer prediction system based on a combination of lung cancer-specific cfDNA fragment features, including a data acquisition module, an analysis module, a storage module, and an output module.

[0164] The data acquisition module is used to obtain the combination of lung cancer-specific cfDNA fragment features of the sample to be tested, and the combination of lung cancer-specific cfDNA fragment features is the combination of lung cancer-specific cfDNA fragment features obtained in Example 2.

[0165] The analysis module is the lung cancer risk prediction model constructed in Example 2. Using the combination of lung cancer-specific cfDNA fragment features of the sample to be tested obtained by the data acquisition module as the input variable, it is input into the lung cancer risk prediction model to obtain the lung cancer risk probability, and based on the lung cancer risk probability, the type of the sample to be tested is determined. The determination criterion is: if the lung cancer risk probability ≥ 0.5, the sample to be tested is a lung cancer sample; if the lung cancer risk probability < 0.5, the sample to be tested is a non-lung cancer sample.

[0166] The storage module is used to store the combination of lung cancer-specific cfDNA fragment features of the sample to be tested obtained by the data acquisition module, the lung cancer risk probability obtained by the analysis module, and the type of the sample to be tested.

[0167] The output module is used to output the lung cancer risk probability and the type of the sample to be tested.

[0168] Example 7: A lung cancer prediction system based on a combination of lung cancer-specific cfDNA fragment features

[0169] This example provides a lung cancer prediction system based on a combination of lung cancer-specific cfDNA fragment features, including a data acquisition module, an analysis module, a storage module, and an output module.

[0170] The data acquisition module is used to obtain the combination of lung cancer-specific cfDNA fragment features of the sample to be tested, and the combination of lung cancer-specific cfDNA fragment features is the combination of lung cancer-specific cfDNA fragment features obtained in Example 3.

[0171] The analysis module is the lung cancer risk prediction model constructed in Example 3. Using the combination of lung cancer-specific cfDNA fragment features of the sample to be tested obtained by the data acquisition module as the input variable, it is input into the lung cancer risk prediction model to obtain the lung cancer risk probability, and based on the lung cancer risk probability, the type of the sample to be tested is determined. The determination criterion is: if the lung cancer risk probability ≥ 0.5, the sample to be tested is a lung cancer sample; if the lung cancer risk probability < 0.5, the sample to be tested is a non-lung cancer sample.

[0172] The storage module is used to store the lung cancer-specific cfDNA fragment feature combination of the test sample obtained by the data acquisition module, the lung cancer risk probability obtained by the analysis module, and the type of the test sample.

[0173] The output module is used to output the lung cancer risk probability and the type of the test sample.

[0174] Example 8 A lung cancer prediction system based on a lung cancer-specific cfDNA fragment feature combination

[0175] This example provides a lung cancer prediction system based on a lung cancer-specific cfDNA fragment feature combination, including a data acquisition module, an analysis module, a storage module, and an output module.

[0176] The data acquisition module is used to obtain the lung cancer-specific cfDNA fragment feature combination of the test sample, and the lung cancer-specific cfDNA fragment feature combination is the lung cancer-specific cfDNA fragment feature combination obtained in Example 4.

[0177] The analysis module is the lung cancer risk prediction model constructed in Example 4. Using the lung cancer-specific cfDNA fragment feature combination of the test sample obtained by the data acquisition module as the input variable, it is input into the lung cancer risk prediction model to obtain the lung cancer risk probability, and the type of the test sample is determined based on the lung cancer risk probability. The determination criterion is: if the lung cancer risk probability ≥ 0.5, the test sample is a lung cancer sample; if the lung cancer risk probability < 0.5, the test sample is a non-lung cancer sample.

[0178] The storage module is used to store the lung cancer-specific cfDNA fragment feature combination of the test sample obtained by the data acquisition module, the lung cancer risk probability obtained by the analysis module, and the type of the test sample.

[0179] The output module is used to output the lung cancer risk probability and the type of the test sample.

[0180] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit the protection scope of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description and ideas. It is not necessary and impossible to enumerate all the implementation manners here. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the claims of the present invention.

Claims

1. Use of a reagent for detecting a characteristic combination of lung cancer-specific cfDNA fragments in peripheral blood in the preparation of a product for lung cancer prediction, characterized in that, The lung cancer-specific cfDNA fragment feature combination includes the frequency of the 6-bp motif at the 5'-end of cfDNA fragment combination (1), the length distribution index of cfDNA fragments in the genomic 10-kb window combination (1), and the coverage of cfDNA fragments in the window of Chr2: 111440000-111449999. The cfDNA fragment combination (1) consists of cfDNA fragments with 6-bp motifs AAGGGT, ACCTCT, AAGTGC, AAAGTG, ACCTGT, ACCTAG, AAGAGG, ACCATG, and TTTAGG at the 5'-end. The genomic 10-kb window combination (1) consists of the windows of Chr7: 36690000-36699999, Chr6: 21130000-21139999, Chr10: 19940000-19949999, Chr3: 143480000-143489999, Chr9: 33660000-33669999, and Chr1: 184650000-184659999.

2. Use of a reagent for detecting a characteristic combination of lung cancer-specific cfDNA fragments in peripheral blood in the preparation of a product for lung cancer prediction, characterized in that, The lung cancer-specific cfDNA fragment feature combination includes the frequency of the 6-bp motif at the 5'-end of cfDNA fragment combination (2), the length distribution index of cfDNA fragments in the genomic 10-kb window combination (2), and the coverage of cfDNA fragments in the genomic 10-kb window combination (3). The cfDNA fragment combination (2) consists of cfDNA fragments with 6-bp motifs AAGGGT, ACCTCT, TTTAGG, AAGTGC, ACCACT, AAGAGG, ATGGGA, AAGGGG, AAGCAG, TTTAAG, and CACCAT at the 5'-end. The genomic 10-kb window combination (2) consists of the windows of Chr11: 4700000-4709999, Chr13: 47010000-47019999, Chr1: 173190000-173199999, Chr10: 19940000-19949999, and Chr1: 184650000-184659999. The genomic 10-kb window combination (3) consists of the windows of Chr7: 73130000-73139999, Chr10: 125260000-125269999, Chr6: 163640000-163649999, and Chr2: 234550000-234559999.

3. Use of a reagent for detecting a characteristic combination of lung cancer-specific cfDNA fragments in peripheral blood in the preparation of a product for lung cancer prediction, characterized in that, The lung cancer-specific cfDNA fragment feature combination includes the frequency of the 6-bp motif at the 5'-end of the cfDNA fragment combination (3); the cfDNA fragment combination (3) consists of cfDNA fragments with 6-bp motifs at the 5'-end being ACCATG, ACCTGT, ACCTCT, CAGTTG, ACCCAA, ACTAAA, ACCTAG, CAGTCA, ACCCTA, ACCCTT, ACTCGT, TTTAGG, ACCCTG, ACTCTG, ACCACA, AAGTGC, ACCTAC, ACTCAC, AAGGGTTTAAGG.

4. Use of a reagent for detecting a characteristic combination of lung cancer-specific cfDNA fragments in peripheral blood in the preparation of a product for lung cancer prediction, characterized in that, The lung cancer-specific cfDNA fragment feature combination includes the frequency of the 6-bp motif at the 5'-end of the cfDNA fragment combination (4), the length distribution index of the cfDNA fragments in the genomic 10-kb window combination (4), and the coverage of the cfDNA fragments in the genomic 10-kb window combination (5). The cfDNA fragment combination (4) consists of cfDNA fragments with 6-bp motifs at the 5'-end being AGTACT, ATGAGG, GTTGGA, CCCAAA, ACGATA, ACCAGT, GGCTAT, and AAGTGC. The genomic 10-kb window combination (4) consists of the Chr4:113740000-113749999 window and the Chr11:3980000-3989999 window. The genomic 10-kb window combination (5) consists of the Chr17:15620000-15629999 window, the Chr5:17950000-17959999 window, the Chr5:8670000-8679999 window, the Chr8:21550000-21559999 window, the Chr3:58170000-58179999 window, the Chr1:120080000-120089999 window, the Chr1:121070000-121079999 window, the Chr7:73130000-73139999 window, the Chr4:8180000-8189999 window, and the Chr1:61070000-61079999 window.

5. A method for constructing a lung cancer risk prediction model, characterized in that, Using the lung cancer-specific cfDNA fragment feature combination described in any one of claims 1 to 4 as the input feature, the lung cancer risk probability as the output result, and the AUC as the performance indicator, the classification model is trained to obtain the lung cancer risk prediction model.

6. The construction method according to claim 5, characterized in that, The classification model includes random forest, support vector machine, gradient boosting machine, or LASSO regression.

7. The construction method according to claim 5, characterized in that The cutoff value of the classification model is set to 0.5; the training is performed based on 10-fold repeated cross-validation.

8. A system for predicting the risk of lung cancer, characterized in that, It includes a data acquisition module, an analysis module, and an output module. The data acquisition module is used to acquire the lung cancer-specific cfDNA fragment feature combination of the sample to be tested, and the lung cancer-specific cfDNA fragment feature combination is the lung cancer-specific cfDNA fragment feature combination described in any one of claims 1 to 4; The analysis module is a lung cancer risk prediction model obtained by the construction method described in claim 5 or 6. Using the lung cancer-specific cfDNA fragment feature combination of the sample to be tested acquired by the data acquisition module as an input variable, it is input into the lung cancer risk prediction model to obtain the lung cancer risk probability; The output module is used to output the lung cancer risk probability.

9. A computer device, comprising a memory and a processor, characterized in that, A computer program executable on the processor is stored on the memory; when the computer program is executed by the processor, the operations of the system described in claim 8 are implemented.

10. A computer-readable storage medium stores a computer program executable by a processor, characterized in that, When the computer program is executed by the processor, the operations of the system described in claim 8 are implemented.