Application of reagent for detecting liver cancer specific cfDNA fragment characteristic combination in peripheral blood in preparation of liver cancer prediction product

By performing low-deep whole genome sequencing of peripheral blood cfDNA and integrating cfDNA fragment characteristic values in multiple dimensions, combining machine learning algorithms to build a liver cancer risk prediction model, the problem of insufficient sensitivity and accuracy of early diagnosis of liver cancer in the existing technology is solved, and efficient liver cancer risk assessment and prediction is achieved.

CN120366452APending Publication Date: 2025-07-25SHENZHEN HAPLOX MEDICAL TESTING LABORATORY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510249896.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The prior art has problems of insufficient sensitivity and low accuracy in the early diagnosis of liver cancer, especially for the diagnosis of atypical liver cancer, and existing cfDNA-based models have limitations in the accuracy and reliability of liver cancer risk prediction.

Method used

By performing low-deep whole genome sequencing of peripheral blood cfDNA, the characteristic values of cfDNA fragments in multiple dimensions were extracted and integrated, including the 5’ end 6bp motif frequency of cfDNA fragments, the 5’ end 6bp motif frequency of cfDNA fragments with a length less than 150bp, the length distribution index of cfDNA fragments in the genome 10kb window, and the cfDNA fragment coverage of cfDNA fragments in the genome 10kb window, combined with machine learning algorithms to construct a liver cancer risk prediction model.

Benefits of technology

The constructed liver cancer risk prediction model has AUC above 0.97, sensitivity above 0.93, specificity above 0.88, positive prediction value above 0.82, negative prediction value above 0.96, and comprehensive accuracy above 0.91. It is suitable for risk assessment and prediction of early cancer, achieving accurate judgment of liver cancer samples.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120366452A_ABST
    Figure CN120366452A_ABST
Patent Text Reader

Abstract

The invention discloses application of a reagent for detecting liver cancer specific cfDNA fragment characteristic combination in peripheral blood in preparation of a liver cancer prediction product. The invention provides a feature combination of a plurality of liver cancer specific cfDNA fragments. The method relates to four dimensions, namely 5 '-terminal 6bp motif frequency of a cfDNA fragment, 5'-terminal 6bp motif frequency of a cfDNA fragment with the length less than 150bp, a length distribution index of the cfDNA fragment of a genome 10kb window and cfDNA fragment coverage of the genome 10kb window, and is also combined with a machine learning algorithm to construct a liver cancer risk prediction model, AUC is higher than 0.97, sensitivity is higher than 0.93, specificity is higher than 0.88, and a positive prediction value is higher than 0.82. The negative predicted value is higher than 0.96, and the comprehensive accuracy is higher than 0.91. The method is suitable for risk assessment and prediction of early cancers, has excellent classification performance, and can realize accurate determination of liver cancer samples.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of biomedical diagnosis. Specifically, it relates to the application of a reagent for detecting a combination of characteristic fragments of hepatocellular carcinoma-specific cfDNA in peripheral blood in the preparation of a product for predicting hepatocellular carcinoma. Background Art

[0002] Hepatocellular carcinoma is the sixth most common cancer globally. Hepatocellular carcinoma (HCC) accounts for 75% - 85% of primary liver cancers and is one of the most common malignant tumors in China. Although certain progress has been made in the diagnosis and treatment of hepatocellular carcinoma, there are still some deficiencies. The main reason for the low long-term survival rate of hepatocellular carcinoma is that the screening of high-risk populations for hepatocellular carcinoma has not been popularized, and the early diagnosis rate is low, resulting in 50% - 70% of patients being in the middle and advanced stages at the time of diagnosis. If early detection and early diagnosis can be achieved, radical measures such as hepatectomy can be performed, which can significantly improve the prognosis of hepatocellular carcinoma patients. Traditional methods for diagnosing hepatocellular carcinoma, such as contrast-enhanced ultrasound, AFP biomarker, enhanced CT, MRI, etc., have problems such as strong subjectivity, insufficient sensitivity, limited ability to detect early hepatocellular carcinoma, and low diagnostic accuracy for atypical hepatocellular carcinoma.

[0003] Circulating free DNA (cfDNA) is double-stranded DNA released from dead cells into peripheral blood and fragmented by DNA enzymes. cfDNA detection technology has shown certain potential in the early diagnosis of various cancers. cfDNA fragmentomics features are cancer molecular markers that have received much attention in recent years, including the length distribution of cfDNA fragments, the frequency of terminal motifs, fragment break coordinates, nucleosome imprints, and different topological structures, etc. cfDNA fragmentomics features are highly correlated with tissue source types and pathological states, have a high signal abundance in peripheral blood, and can achieve cancer screening and tissue tracing based on low-depth sequencing. With the development of bioinformatics and machine learning technologies, algorithm models can be constructed using sequencing information to predict the occurrence, development, prognosis, etc. of cancers and provide a basis for clinical diagnosis. However, existing models based on cfDNA still have certain limitations in terms of the accuracy and reliability of predicting the risk of hepatocellular carcinoma and need further optimization and improvement. Summary of the Invention

[0004] In order to solve the problems existing in the prior art, the present invention provides the application of a reagent for detecting a combination of characteristic fragments of hepatocellular carcinoma-specific cfDNA in peripheral blood in the preparation of a product for predicting hepatocellular carcinoma.

[0005] One object of the present invention is to provide the application of a reagent for detecting a combination of characteristic fragments of hepatocellular carcinoma-specific cfDNA in peripheral blood in the preparation of a product for predicting hepatocellular carcinoma.

[0006] Another object of the present invention is to provide a method for constructing a hepatocellular carcinoma risk prediction model.

[0007] Another object of the present invention is to provide a system for predicting the risk of liver cancer.

[0008] Another object of the present invention is to provide a computer device.

[0009] Another object of the present invention is to provide a computer-readable storage medium.

[0010] To achieve the above object, the present invention is implemented by the following solutions:

[0011] The present invention performs low-depth whole-genome sequencing on peripheral blood cfDNA, extracts and integrates cfDNA fragment feature values in multiple dimensions, and provides a new marker combination for early screening and prediction of liver cancer. The cfDNA fragment feature values provided by the present invention include the 6bp motif frequency 6bp-f at the 5' end of the cfDNA fragment, the 6bp motif frequency short-6bp-f at the 5' end of the cfDNA fragment with a length less than 150bp, the length distribution index F-Index_bin of the cfDNA fragment in the 10kb window of the genome, and the coverage Coverage_bin of the cfDNA fragment in the 10kb window of the genome.

[0012] The 6bp-f is the occurrence frequency of each 6bp motif in the terminal motif set; the terminal motif set includes set 1 and set 2, and set 1 is composed of the 6bp motifs at the 5' end of the nucleotide sequence of each cfDNA fragment; set 2 is composed of the reverse complementary sequences of the 6bp motifs at the 3' end of the nucleotide sequence of each cfDNA fragment; the nucleotide sequence of the cfDNA fragment is obtained by aligning paired-end reads to the human reference genome; the paired-end reads are obtained by the following method: obtaining the paired-end sequencing data of peripheral blood cfDNA of the sample to be tested, and retaining the paired-end reads aligned to the human reference genome. Based on the paired-end reads, information on the cfDNA fragment is obtained, and each pair of paired-end reads corresponds to a cfDNA fragment. The information on the cfDNA fragment includes, but is not limited to, the position of the cfDNA in the genome, the nucleotide sequence of the cfDNA, and the length of the cfDNA.

[0013] The 10kb window of the genome is obtained by the following method: dividing the genomic region after removing redundant regions into non-overlapping segments according to 10kb intervals to obtain M 10kb windows of the genome; the redundant regions include the short arm regions, telomere regions, and centromere regions of 5 chromosomes, and the 5 chromosomes are chromosome 13, chromosome 14, chromosome 15, chromosome 21, and chromosome 22.

[0014] The Coverage_bin is calculated from the 10kb window of the genome and formula (2);

[0015] Formula (2): Coverage_bin k = N k / N total ;

[0016] Wherein, Coverage_bin k represents the coverage of cfDNA fragments in the k-th 10kb-bins of the genome, 1 ≤ k ≤ M; N k represents the number of cfDNA fragments aligned to the k-th 10kb-bins; N total represents the number of all cfDNA fragments aligned to the reference genome based on paired-end reads;

[0017] The F-Index_bin is calculated from the 10kb window of the genome and formula (3);

[0018] Formula (3):

[0019] Wherein, i represents the length of the cfDNA fragment; P is a constant value, 160bp to 170bp; X' i represents the proportion of cfDNA fragments with length i among all cfDNA fragments in the k-th 10kb window of the genome.

[0020] Preferably, P is 166bp.

[0021] The short-6bp-f is the occurrence frequency of each 6bp motif in the short-terminal motif set; the short-terminal motif set is obtained by retaining the subset corresponding to all cfDNA fragments with length less than 150bp in the terminal motif set. In the present invention, for each cfDNA fragment with length less than 150bp, its 5'-terminal 6bp sequence is defined as Short_End_motif_F, and its 3'-terminal 6bp sequence is defined as Short_End_motif_R. The short-terminal motif set is obtained by merging the set 1' composed of Short_End_motif_F and the set 2' composed of the reverse complementary sequences of Short_End_motif_R.

[0022] Use of a reagent for detecting the characteristic combination of liver cancer-specific cfDNA fragments in peripheral blood in the preparation of a product for predicting liver cancer, characterized in that the liver cancer-specific cfDNA fragment characteristic combination includes the frequency of the 5'-terminal 6bp motif of the cfDNA fragment combination and the length distribution index of cfDNA fragments in the genomic 10kb window combination;

[0023] The cfDNA fragment combination consists of cfDNA fragments with a length less than 150 bp and a 5'-terminal 6-bp motif of TGCTGT, TGGCTA, CGGCCT, TGCACA, TGACCA, GGTCCA, TGCCAC, TCCACT, GCGTAA, and TGCATG;

[0024] The genomic 10-kb window combination consists of the Chr5:119070000-119079999 window, the Chr1:175850000-175859999 window, the Chr7:28430000-28439999 window, the Chr5:133820000-133829999 window, and the Chr16:12790000-12799999 window.

[0025] The application of a reagent for detecting the combination of liver cancer-specific cfDNA fragment characteristics in peripheral blood in the preparation of a product for liver cancer prediction, wherein the liver cancer-specific cfDNA fragment characteristics combination includes the frequency of the 5'-terminal 6-bp motif of the first cfDNA fragment combination, the frequency of the 5'-terminal 6-bp motif of the second cfDNA fragment combination, the length distribution index of the cfDNA fragments in the genomic 10-kb window combination, and the cfDNA fragment coverage of the Chr2:87940000-87949999 window;

[0026] The first cfDNA fragment combination consists of cfDNA fragments with a 5'-terminal 6-bp motif of ATGGGC, TGCTTC, TGATCC, ACGAGA, TGACTC, ACGAGT, and ACGATG;

[0027] The second cfDNA fragment combination consists of cfDNA fragments with a length less than 150 bp and a 5'-terminal 6-bp motif of TGTGAC, CGGCTA, GCGAAA, TGTAGC, and TGGCTC;

[0028] The genomic 10-kb window combination consists of the Chr7:28430000-28439999 window and the Chr16:12790000-12799999 window.

[0029] The application of a reagent for detecting the combination of liver cancer-specific cfDNA fragment characteristics in peripheral blood in the preparation of a product for liver cancer prediction, wherein the liver cancer-specific cfDNA fragment characteristics combination includes the frequency of the 5'-terminal 6-bp motif of the first cfDNA fragment combination and the frequency of the 5'-terminal 6-bp motif of the second cfDNA fragment combination;

[0030] The first cfDNA fragment combination consists of cfDNA fragments with a 6-bp motif at the 5'-end being TGATCC, TGTTCC, TGCTTC, ACGAAG, ACGAGA, TGACTC, CAGTTC, CAGTCC, AAGAGC, TGATCG, and CAGTCA;

[0031] The second cfDNA fragment combination consists of cfDNA fragments with a length less than 150 bp and a 6-bp motif at the 5'-end being CGGTGG, CGGCCT, CGGCTA, and TGCCTG.

[0032] The application of a reagent for detecting the characteristic combination of liver cancer-specific cfDNA fragments in peripheral blood in the preparation of a product for liver cancer prediction, characterized in that the liver cancer-specific cfDNA fragment characteristic combination includes the frequency of the 6-bp motif at the 5'-end of the first cfDNA fragment combination, the frequency of the 6-bp motif at the 5'-end of the second cfDNA fragment combination, and the cfDNA fragment coverage of the genomic 10-kb window combination;

[0033] The first cfDNA fragment combination consists of cfDNA fragments with a 6-bp motif at the 5'-end being TGCTTC, CAGTCA, TGTTCC, ACGAAG, and ATCGGG;

[0034] The second cfDNA fragment combination consists of cfDNA fragments with a length less than 150 bp and a 6-bp motif at the 5'-end being TTATGG, TGTTCG, and TGTTCA;

[0035] The genomic 10-kb window combination consists of the Chr2:87940000-87949999 window, Chr6:24860000-24869999 window, Chr5:550000-559999 window, Chr18:26740000-26749999 window, Chr10:11170000-11179999 window, Chr15:62300000-62309999 window, and Chr1:231610000-231619999 window.

[0036] A method for constructing a liver cancer risk prediction model, using any of the above-mentioned liver cancer-specific cfDNA fragment characteristic combinations as input features, using the liver cancer risk probability as the output result, and using AUC as the performance index to train a classification model to obtain the liver cancer risk prediction model.

[0037] Preferably, the classification model includes random forest, support vector machine, gradient boosting machine, or LASSO regression.

[0038] More preferably, the classification model is a random forest.

[0039] Preferably, the cutoff value of the classification model is set to 0.5, and the judgment criterion is as follows: if the liver cancer risk probability ≥ 0.5, the sample is judged as positive (i.e., a liver cancer sample); if the liver cancer risk probability < 0.5, the sample is judged as negative (i.e., a non-liver cancer sample); the training is performed based on 10-fold repeated cross-validation.

[0040] A system for predicting the risk of liver cancer, comprising a data acquisition module, an analysis module, and an output module; the data acquisition module is used to acquire the combination of liver cancer-specific cfDNA fragment features of the sample to be tested, and the combination of liver cancer-specific cfDNA fragment features is any one of the combinations of liver cancer-specific cfDNA fragment features;

[0041] The analysis module is a liver cancer risk prediction model obtained by any one of the construction methods, and uses the combination of liver cancer-specific cfDNA fragment features of the sample to be tested acquired by the data acquisition module as an input variable, and inputs it into the liver cancer risk prediction model to obtain the liver cancer risk probability;

[0042] The output module is used to output the liver cancer risk probability.

[0043] Preferably, the analysis module further determines the type of the sample to be tested based on the liver cancer risk probability, and the judgment criterion is as follows: if the liver cancer risk probability ≥ 0.5, the sample to be tested is a liver cancer sample; if the liver cancer risk probability < 0.5, the sample to be tested is a non-liver cancer sample; the output module is further used to output the type of the sample to be tested.

[0044] Preferably, the system further includes a storage module, and the storage module is used to store the combination of liver cancer-specific cfDNA fragment features of the sample to be tested obtained by the data acquisition module and the liver cancer risk probability obtained by the analysis module.

[0045] A computer device, comprising a memory and a processor, characterized in that a computer program executable on the processor is stored on the memory; when the computer program is executed by the processor, the operations of the system for predicting the risk of liver cancer are implemented.

[0046] A computer-readable storage medium stores a computer program capable of being executed by a processor, and when the computer program is executed by the processor, the operations of the system for predicting the risk of liver cancer are implemented.

[0047] Compared with the prior art, the present invention has the following beneficial effects:

[0048] The present invention provides multiple combinations of liver cancer-specific cfDNA fragment features, and also constructs a liver cancer risk prediction model by combining machine learning algorithms. The AUC is higher than 0.97, the sensitivity is higher than 0.93, the specificity is higher than 0.88, the positive predictive value is higher than 0.82, the negative predictive value is higher than 0.96, and the comprehensive accuracy is higher than 0.91. It is applicable to risk assessment and prediction of early cancers, has excellent classification performance, and can accurately determine liver cancer samples. Description of the Drawings

[0049] Figure 1 It is an AUC graph of the liver cancer risk prediction model constructed by the present invention based on a machine learning algorithm combined with a combination of liver cancer-specific cfDNA fragment features. Detailed Embodiments

[0050] The present invention will be further elaborated in detail below in conjunction with the drawings in the specification and specific embodiments. The embodiments are only used to explain the present invention and are not used to limit the scope of the present invention. The test methods used in the following embodiments are all conventional methods unless otherwise specified; the materials, reagents, etc. used are commercially available reagents and materials unless otherwise specified.

[0051] Example 1 Obtaining a combination of liver cancer-specific cfDNA fragment features and constructing a liver cancer risk prediction model based on Random Forest

[0052] I. Subject Information

[0053] In this example, peripheral blood samples of 121 liver cancer patients (i.e., liver cancer samples) and peripheral blood samples of 213 healthy people (i.e., healthy samples) were included. Each peripheral blood sample was from a different individual and all were from the sample library of Shenzhen Hypereal Biotechnology Co., Ltd. Informed consent of the subjects was obtained before sample testing.

[0054] II. Determination of a combination of liver cancer-specific cfDNA fragment features in peripheral blood

[0055] 1. Sequencing of peripheral blood cfDNA

[0056] cfDNA was extracted from the peripheral blood samples of the subjects using the circulating Nucleic Acid Kit, and the cfDNA library was constructed using the library construction kit QIAseq cfDNA All-in-One Kit (24) (Cat.No. / ID: 180023). During the library construction process, cfDNA was not fragmented. The Illumina Novaseq sequencing platform was used to perform paired-end sequencing on the cfDNA library, and the sequencing read length was ~150 bp. After the sequencing was completed, the original fastq sequencing data was obtained. The Fastp software was used to perform quality control processing on the original fastq data based on the default parameters, removing adapters, data with base quality values lower than Q30, sequences with a high proportion of N (i.e., removing read sequences with an N-Content greater than 2%), and PCR-introduced duplicate sequences, resulting in paired-end reads. Each cfDNA fragment matched a pair of paired-end reads (read1 and read2).

[0057] Using the bwa software, the quality-controlled paired-end reads were aligned to the human reference genome hg38 (Genome Reference Consortium human genome build 38, GRCh38), and a Bam file of the alignment results was obtained. The picard tool was used to preprocess the Bam file, removing the paired-end reads with multiple alignments, removing the paired-end reads that did not align to the reference genome, and retaining the paired-end reads with a quality value greater than 30, resulting in a clean Bam file.

[0058] 2. Obtaining cfDNA fragment characteristics

[0059] This invention involves a total of 5 types of cfDNA fragment characteristics: the global length distribution index of cfDNA fragments, the 6-bp motif at the 5' end of cfDNA fragments, the cfDNA fragment coverage of the 10-kb genomic window, the length distribution index of cfDNA fragments in the 10-kb genomic window, and the 6-bp motif at the 5' end of cfDNA fragments with a length less than 150 bp.

[0060] (1) Calculation of the characteristic value all-F-Index of the global length distribution index of cfDNA fragments

[0061] Extract the alignment coordinates, directions, lengths, and base sequence information of each pair of paired-end reads from the clean bam file obtained in the previous step. In this analysis step, only retain the paired-end reads aligned to the autosomal regions 1 to 22, and remove the paired-end reads aligned to the X and Y chromosomes, mitochondrial genome, or other supplementary contig regions to obtain a bed file of the cfDNA fragment distribution on the genome (i.e., the genomic bed file). This file contains cfDNA fragment information in the genome, including: cfDNA fragment sequence, length, and position information on the chromosome.

[0062] Calculate the all-F-Index based on the genomic bed file and formula (1).

[0063] Formula (1):

[0064] Among them, i represents the length of the cfDNA fragment; P is a constant value of 166bp, which is the general length of cfDNA fragments in the plasma of healthy individuals; X i represents the proportion of cfDNA fragments with length i among all cfDNA fragments.

[0065] (2) Calculation of the characteristic value of the 6bp motif at the 5' end of the cfDNA fragment

[0066] The characteristic value of the 6bp motif at the 5' end of the cfDNA fragment is the 6bp motif frequency (6bp-f) at the 5' end of the cfDNA fragment. This value reflects the influence of nuclease activity and tissue type on the motif pattern at the end of the cfDNA fragment. The statistical method is as follows:

[0067] For each cfDNA fragment, each pair of paired-end reads consists of read1 on the forward strand and read2 on the reverse strand. By aligning to the reference genome, the forward complete sequence of the cfDNA fragment (i.e., the nucleotide sequence from the 5'-end to the 3'-end) is obtained. The 6bp sequence at the 5'-end is defined as End_motif_F, and the 6bp sequence at the 3'-end is defined as End_motif_R. After that, the sequence of End_motif_F is not transformed, and all End_motif_F of cfDNA fragments form set 1; the sequence of End_motif_R is subjected to reverse complementary transformation of bases to obtain the reverse complementary sequence of End_motif_R, and all reverse complementary sequences of End_motif_R of cfDNA fragments form set 2. After combining set 1 and set 2, the set of 6bp motifs at the 5'-ends of all cfDNA fragments is obtained, denoted as the terminal motif set. Theoretically, there are at most 4096 different 6bp motifs, such as {AAAAAA,AAAAAT,...,ACGACG,CTTGTT,...,GGGGGG}. By counting the occurrence frequency of each 6bp motif in the terminal motif set, 6bp-f is obtained.

[0068] (3) Calculation of the coverage characteristic value of cfDNA fragments in a 10kb window of the genome

[0069] The calculation method of the coverage characteristic value Coverage_bin of cfDNA fragments in a 10kb window of the genome is as follows:

[0070] For human chromosomes 13, 14, 15, 21, and 22, since the short arm regions of these 5 chromosomes contain nucleolus organizer regions and are rich in a large number of repetitive ribosomal RNA genes, it is difficult for genome sequencing technology to accurately analyze the complete sequences therein. Therefore, based on the bed file of the cfDNA fragment distribution on the genome obtained in this embodiment, the short arm regions of these 5 chromosomes as well as the telomere and centromere regions are removed, and the remaining genomic regions are non-overlappingly segmented according to 10kb intervals (such as chr1:20000~chr1:29999) to obtain M genomic 10kb windows (i.e., 10kb-bins). Theoretically, at most 287465 10kb-bins can be obtained, that is, the maximum value of M is 287465.

[0071] Coverage_bin is calculated according to the genomic bed file, 10kb-bins and formula (2).

[0072] Formula (2): Coverage_bin k =N k / N total ;

[0073] Among them, Coverage_bin k represents the coverage of cfDNA fragments in the k-th 10kb-bins, where 1 ≤ k ≤ M; N k represents the number of cfDNA fragments aligned to the k-th 10kb-bins. Among them, reads spanning two adjacent 10kb-bins are regarded as being aligned to both 10kb-bins simultaneously. When calculating coverage, both 10kb-bins on both sides increase by one count. When calculating F-Index_bin k this read is also included in the calculation; N total represents the number of all cfDNA fragments aligned to the reference genome based on paired-end reads.

[0074] (4) Calculation of the length distribution index F-Index_bin of cfDNA fragments in the 10kb window of the genome

[0075] The F-Index_bin of cfDNA fragments in the k-th 10kb-bins is calculated according to the genome bed file, 10kb-bins and formula (3).

[0076] Formula (3):

[0077] Among them, i represents the length of the cfDNA fragment; P is a constant value of 166bp, which is the general length of cfDNA fragments in the plasma of healthy people; X’ i represents the proportion of cfDNA fragments with length i among all cfDNA fragments in the k-th 10kb-bins.

[0078] Through the calculation and statistics of this embodiment, for each sample, 1 all-F-Index value, and at most 4096 6bp-f values, 287465 Coverage_bin values and 287465 F-Index_bin values are obtained.

[0079] (5) Calculation of the 5'-end 6bp motif (short-6bp-f) of cfDNA fragments with length less than 150bp

[0080] For each cfDNA fragment with a length less than 150 bp, the 6-bp sequence at its 5'-end is defined as Short_End_motif_F, and the 6-bp sequence at its 3'-end is defined as Short_End_motif_R. A subset corresponding to all cfDNA fragments with a length less than 150 bp is extracted from the set of end motifs, denoted as the short-end motif set, which is obtained by merging set 1' composed of Short_End_motif_F and set 2' composed of the reverse complementary sequences of Short_End_motif_R. The occurrence frequency of each 6-bp motif in the short-end motif set is counted to obtain short-6bp-f.

[0081] Through the calculation and statistics of this example, for each sample, 1 all-F-Index, up to 4096 6bp-f, 4096 short-6bp-f, 287465 Coverage_bin, and 287465 F-Index_bin are obtained.

[0082] 3. Determine the characteristic combination of liver cancer-specific cfDNA fragments based on RandomForest

[0083] (1) Feature matrix generation

[0084] According to the above step, all cfDNA fragment characteristic values of each sample are obtained and a feature matrix is generated. Each row represents a sample, and the columns represent each cfDNA fragment characteristic and its characteristic value. All characteristic values are normalized and transformed into a normal distribution.

[0085] (2) Data division

[0086] The feature matrix obtained in the above step is randomly divided into a training set and a validation set according to the samples. The training set contains the feature matrices of 73 liver cancer samples and the feature matrices of 128 healthy samples, and the validation set samples contain the feature matrices of the remaining 48 liver cancer samples and the feature matrices of 85 healthy samples.

[0087] (3) Feature selection

[0088] Using the training set as the input features, feature selection is performed based on a machine learning model. Specifically: the Recursive feature elimination (RFE) algorithm combined with a model constructed by RandomForest is used for feature screening. When constructing the model, the model parameters are selected for 10-fold repeated cross-validation, and the number of repeated sampling iterations is 10. The accuracy of predicting whether a patient has liver cancer is used as the result evaluation parameter, and the optimal number of features and feature combination are returned.

[0089] The top 15 feature combinations ranked by the obtained relative weights are shown in Table 1, which are the cfDNA fragment feature combinations specific to liver cancer. There are 5' terminal 6bp motifs (Short_TGCTGT, Short_TGGCTA, Short_CGGCCT, Short_TGCACA, Short_TGACCA, Short_GGTCCA, Short_TGCCAC, Short_TCCACT, Short_GCGTAA, and Short_TGCATG) of 10 cfDNA fragments with lengths less than 150bp, namely short-6bp-f, and length distribution indices (Findex_bin99867, Findex_bin17585, Findex_bin126034, Findex_bin101342, and Findex_bin241300) of cfDNA fragments in 5 genomic 10kb windows (Chr5:119070000-119079999 window, Chr1:175850000-175859999 window, Chr7:28430000-28439999 window, Chr5:133820000-133829999 window, and Chr16:12790000-12799999 window).

[0090] Table 1 Feature combinations of cfDNA fragments specific to liver cancer obtained based on RandomForest

[0091]

[0092] 4. Construction of a liver cancer risk prediction model and performance evaluation based on RandomForest

[0093] (1) Model construction

[0094] For the training set of this embodiment, the cfDNA fragment feature combinations specific to liver cancer of each sample in the training set as shown in Table 1 are respectively obtained as the input features of the model (RandomForest), and the liver cancer risk probability is used as the output result. The range of the output result is [0,1]. The greater the liver cancer risk probability, the higher the possibility that the sample is a liver cancer sample. Combining the actual status (liver cancer and non-liver cancer) of each sample in the training set, 10-fold repeated cross-validation and 10 repeated iterations are performed, and the optimal training model is output based on the AUC index performance.

[0095] During the 10-fold cross-validation process, the cutoff value is 0.5, and the judgment criterion is: if the liver cancer risk probability ≥ 0.5, the sample is judged as positive (i.e., a liver cancer sample); if the liver cancer risk probability < 0.5, the sample is judged as negative (i.e., a non-liver cancer sample).

[0096] (2) Performance evaluation

[0097] Formula (4): Sensitivity = TP / (TP + FN);

[0098] Formula (5): Specificity = TN / (TN + FP);

[0099] Formula (6): PPV = TP / (TP + FP);

[0100] Formula (7): NPV = TN / (TN + FN);

[0101] Formula (8): Accuracy = (TP + TN) / (TP + FN + TN + FP).

[0102] Use the validation set samples of this embodiment to evaluate the performance of the optimal training model obtained in step (1), and calculate 5 indicators according to Formulas (4) to (8): sensitivity (Sensitivity), specificity (Specificity), positive predictive value (Positive Predictive Value, PPV), negative predictive value (Negative Predictive Value, NPV), and comprehensive accuracy (Accuracy).

[0103] As Figure 1 shown, the best AUC of the liver cancer risk prediction model constructed based on the RandomForest algorithm in 10-fold repeated cross-validation is 0.984. In addition, the sensitivity of this model is calculated to be 0.958, the specificity is 0.941, the positive predictive value is 0.902, the negative predictive value is 0.976, and the comprehensive accuracy is 0.947. It shows that the liver cancer risk prediction model constructed based on the RandomForest algorithm has excellent classification performance and can accurately determine whether the sample is a liver cancer sample.

[0104] Example 2 Obtaining a combination of liver cancer-specific cfDNA fragment features and constructing a liver cancer risk prediction model based on gradient boosting machine (GBM)

[0105] I. Subject information

[0106] Same as Example 1.

[0107] II. Determination of a combination of liver cancer-specific cfDNA fragment features in peripheral blood

[0108] 1. Peripheral blood cfDNA sequencing

[0109] Same as Example 1.

[0110] 2. Obtaining cfDNA fragment features

[0111] The same as in Example 1.

[0112] 3. Determining a combination of liver cancer-specific cfDNA fragment features based on GBM

[0113] Basically the same as in Example 1, except that in the feature selection in the third step, a recursive feature elimination algorithm combined with a model constructed by GBM is used for feature screening. The top 15 feature combinations ranked by relative weights are shown in Table 2, which are the liver cancer-specific cfDNA fragment feature combinations. There are 6bp-f of the 5'-end 6bp motifs (ATGGGC, TGCTTC, TGATCC, ACGAGA, TGACTC, ACGAGT, and ACGATG) of 7 cfDNA fragments, short-6bp-f of the 5'-end 6bp motifs (Short_TGTGAC, Short_CGGCTA, Short_GCGAAA, Short_TGTAGC, and Short_TGGCTC) of 5 cfDNA fragments with lengths less than 150bp, length distribution indices (Findex_bin126034 and Findex_bin241300 in sequence) of cfDNA fragments in 2 genomic 10kb windows (Chr7:28430000-28439999 window and Chr16:12790000-12799999 window), and cfDNA fragment coverage (Coverage_bin33688) in 1 genomic 10kb window (Chr2:87940000-87949999 window).

[0114] Table 2 Liver cancer-specific cfDNA fragment feature combinations obtained based on GBM

[0115]

[0116]

[0117] 4. Constructing a liver cancer risk prediction model based on GBM and performance evaluation

[0118] (1) Model construction

[0119] Basically the same as in Example 1, except that GBM is used for model construction.

[0120] (2) Performance evaluation

[0121] According to the method in Example 1, the performance of the optimal training model obtained in step (1) of this example is evaluated using the validation set samples of this example.

[0122] As Figure 1 shown, the best AUC of the liver cancer risk prediction model constructed based on the GBM algorithm in 10-fold repeated cross-validation was 0.972. In addition, the sensitivity of this model was calculated to be 0.979, the specificity was 0.918, the positive predictive value was 0.87, the negative predictive value was 0.987, and the comprehensive accuracy was 0.94. This indicates that the liver cancer risk prediction model constructed based on the GBM algorithm has excellent classification performance and can accurately determine whether a sample is a liver cancer sample.

[0123] Example 3 Obtaining a liver cancer-specific cfDNA fragment feature combination and constructing a liver cancer risk prediction model based on support vector machine (svmLinear)

[0124] I. Subject Information

[0125] Same as Example 1.

[0126] II. Determination of the liver cancer-specific cfDNA fragment feature combination in peripheral blood

[0127] 1. Sequencing of peripheral blood cfDNA

[0128] Same as Example 1.

[0129] 2. Obtaining cfDNA fragment features

[0130] Same as Example 1.

[0131] 3. Determining the liver cancer-specific cfDNA fragment feature combination based on svmLinear

[0132] Basically the same as Example 1, with the difference that in the feature selection of the third step, a recursive feature elimination algorithm combined with a model constructed by svmLinear was used for feature screening. The feature combination with the top 15 relative weights is shown in Table 3 and is the liver cancer-specific cfDNA fragment feature combination, which includes the 6bp-f of the 5'-end 6bp motifs (TGATCC, TGTTCC, TGCTTC, ACGAAG, ACGAGA, TGACTC, CAGTTC, CAGTCC, AAGAGC, TGATCG, and CAGTCA) of 11 cfDNA fragments and the short-6bp-f of the 5'-end 6bp motifs (Short_CGGTGG, Short_CGGCCT, Short_CGGCTA, and Short_TGCCTG) of 4 cfDNA fragments with lengths less than 150bp.

[0133] Table 3 Liver cancer-specific cfDNA fragment feature combination obtained based on svmLinear

[0134] Serial Number Type Feature Name Description Relative Weight 1 6bp Motif at the 5'-End of cfDNA Fragments TGATCC 6bp-f of TGATCC 0.98 2 6bp Motif at the 5'-End of cfDNA Fragments TGTTCC 6bp-f of TGTTCC 0.978 3 6bp Motif at the 5'-End of cfDNA Fragments TGCTTC 6bp-f of TGCTTC 0.978 4 6bp Motif at the 5'-End of cfDNA Fragments ACGAAG 6bp-f of ACGAAG 0.977 5 6bp Motif at the 5'-End of cfDNA Fragments ACGAGA 6bp-f of ACGAGA 0.974 6 6bp Motif at the 5'-End of cfDNA Fragments TGACTC 6bp-f of TGACTC 0.973 7 6bp Motif at the 5'-End of cfDNA Fragments with Length Less than 150bp Short_CGGTGG short-6bp-f of Short_CGGTGG 0.973 8 6bp Motif at the 5'-End of cfDNA Fragments CAGTTC 6bp-f of CAGTTC 0.971 9 6bp Motif at the 5'-End of cfDNA Fragments with Length Less than 150bp Short_CGGCCT short-6bp-f of Short_CGGCCT 0.971 10 6bp Motif at the 5'-End of cfDNA Fragments CAGTCC 6bp-f of CAGTCC 0.970 11 6bp Motif at the 5'-End of cfDNA Fragments AAGAGC 6bp-f of AAGAGC 0.970 12 6bp Motif at the 5'-End of cfDNA Fragments with Length Less than 150bp Short_CGGCTA short-6bp-f of Short_CGGCTA 0.970 13 6bp Motif at the 5'-End of cfDNA Fragments TGATCG 6bp-f of TGATCG 0.970 14 6bp Motif at the 5'-End of cfDNA Fragments CAGTCA 6bp-f of CAGTCA 0.969 15 6bp Motif at the 5'-End of cfDNA Fragments with Length Less than 150bp Short_TGCCTG short-6bp-f of Short_TGCCTG 0.968

[0135] 4. Construction of Liver Cancer Risk Prediction Model Based on svmLinear and Performance Evaluation

[0136] (1) Model Construction

[0137] It is basically the same as Example 1, except that svmLinear is used for model construction.

[0138] (2) Performance Evaluation

[0139] According to the method of Example 1, the performance of the optimal training model obtained in step (1) of this example is evaluated using the validation set samples of this example.

[0140] As Figure 1 shown, the best AUC of the liver cancer risk prediction model constructed based on the svmLinear algorithm in 10-fold repeated cross-validation is 0.976. In addition, the sensitivity of this model is calculated to be 0.938, the specificity is 0.929, the positive predictive value is 0.882, the negative predictive value is 0.963, and the comprehensive accuracy is 0.932. It shows that the liver cancer risk prediction model constructed based on the svmLinear algorithm has excellent classification performance and can accurately determine whether the sample is a liver cancer sample.

[0141] Example 4 Obtaining Liver Cancer-Specific cfDNA Fragment Features Combination Based on Regression Algorithm (LASSO) and Constructing Liver Cancer Risk Prediction Model

[0142] I. Subject Information

[0143] The same as Example 1.

[0144] II. Determination of Liver Cancer-Specific cfDNA Fragment Features Combination in Peripheral Blood

[0145] 1. Peripheral Blood cfDNA Sequencing

[0146] The same as Example 1.

[0147] 2. Obtaining cfDNA Fragment Features

[0148] The same as Example 1.

[0149] 3. Determining Liver Cancer-Specific cfDNA Fragment Features Combination Based on LASSO

[0150] Basically the same as Example 1, with the differences being: in the feature selection of the third step, the recursive feature elimination algorithm is applied to combine with the model constructed by LASSO for feature screening. The top 15 feature combinations ranked by relative weights are shown in Table 4, which are the cfDNA fragment feature combinations specific to liver cancer. There are 6bp-f of the 6bp motifs (TGCTTC, CAGTCA, TGTTCC, ACGAAG, and ATCGGG) at the 5'-ends of 5 cfDNA fragments, short-6bp-f of the 6bp motifs (Short_TTATGG, Short_TGTTCG, and Short_TGTTCA) at the 5'-ends of 3 cfDNA fragments with lengths less than 150bp, and the cfDNA fragment coverage of 7 genomic 10kb windows (Chr2: 87940000 - 87949999 window, Chr6: 24860000 - 24869999 window, Chr5: 550000 - 559999 window, Chr18: 26740000 - 26749999 window, Chr10: 11170000 - 11179999 window, Chr15: 62300000 - 62309999 window, and Chr1: 231610000 - 231619999 window) (Coverage_bin33688, Coverage_bin108598, Coverage_bin88015, Coverage_bin260051, Coverage_bin168591, Coverage_bin236053, and Coverage_bin23161 in sequence).

[0151] Table 4 cfDNA fragment feature combinations specific to liver cancer obtained based on LASSO

[0152]

[0153]

[0154] 4. Construction of liver cancer risk prediction model based on LASSO and performance evaluation

[0155] (1) Model construction

[0156] Basically the same as Example 1, with the differences being: for the LASSO regression algorithm, alpha is set to seq(0, 1, by = 0.05), lambda is set to 10^seq(-2, 2, length = 100), and AUC is used as the performance indicator.

[0157] (2) Performance evaluation

[0158] According to the method of Example 1, the performance of the optimal training model obtained in step (1) of this example is evaluated using the validation set samples of this example.

[0159] As Figure 1 shown, the best AUC of the liver cancer risk prediction model constructed based on the LASSO algorithm in 10-fold repeated cross-validation is 0.978. In addition, the sensitivity of this model is calculated to be 0.979, the specificity is 0.882, the positive predictive value is 0.825, the negative predictive value is 0.987, and the comprehensive accuracy is 0.917. It shows that the liver cancer risk prediction model constructed based on the LASSO algorithm has excellent classification performance and can accurately determine whether a sample is a liver cancer sample.

[0160] Example 5 A liver cancer prediction system based on a combination of liver cancer-specific cfDNA fragment features

[0161] This example provides a liver cancer prediction system based on a combination of liver cancer-specific cfDNA fragment features, including a data acquisition module, an analysis module, a storage module, and an output module.

[0162] The data acquisition module is used to obtain the combination of liver cancer-specific cfDNA fragment features of the sample to be tested, and the combination of liver cancer-specific cfDNA fragment features is the combination of liver cancer-specific cfDNA fragment features obtained in Example 1.

[0163] The analysis module is the liver cancer risk prediction model constructed in Example 1. Taking the combination of liver cancer-specific cfDNA fragment features of the sample to be tested obtained by the data acquisition module as the input variable, it is input into the liver cancer risk prediction model to obtain the liver cancer risk probability, and based on the liver cancer risk probability, the type of the sample to be tested is determined. The determination criterion is: when the liver cancer risk probability ≥ 0.5, the sample to be tested is a liver cancer sample; when the liver cancer risk probability < 0.5, the sample to be tested is a non-liver cancer sample.

[0164] The storage module is used to store the combination of liver cancer-specific cfDNA fragment features of the sample to be tested obtained by the data acquisition module, the liver cancer risk probability obtained by the analysis module, and the type of the sample to be tested.

[0165] The output module is used to output the liver cancer risk probability and the type of the sample to be tested.

[0166] Example 6 A liver cancer prediction system based on a combination of liver cancer-specific cfDNA fragment features

[0167] This example provides a liver cancer prediction system based on a combination of liver cancer-specific cfDNA fragment features, including a data acquisition module, an analysis module, a storage module, and an output module.

[0168] The data acquisition module is used to obtain the hepatocellular carcinoma-specific cfDNA fragment feature combination of the sample to be tested, and the hepatocellular carcinoma-specific cfDNA fragment feature combination is the hepatocellular carcinoma-specific cfDNA fragment feature combination obtained in Example 2.

[0169] The analysis module is the hepatocellular carcinoma risk prediction model constructed in Example 2. Using the hepatocellular carcinoma-specific cfDNA fragment feature combination of the sample to be tested obtained by the data acquisition module as the input variable, it is input into the hepatocellular carcinoma risk prediction model to obtain the hepatocellular carcinoma risk probability, and based on the hepatocellular carcinoma risk probability, the type of the sample to be tested is determined. The determination criterion is: if the hepatocellular carcinoma risk probability ≥ 0.5, the sample to be tested is a hepatocellular carcinoma sample; if the hepatocellular carcinoma risk probability < 0.5, the sample to be tested is a non-hepatocellular carcinoma sample.

[0170] The storage module is used to store the hepatocellular carcinoma-specific cfDNA fragment feature combination of the sample to be tested obtained by the data acquisition module, the hepatocellular carcinoma risk probability obtained by the analysis module, and the type of the sample to be tested.

[0171] The output module is used to output the hepatocellular carcinoma risk probability and the type of the sample to be tested.

[0172] Example 7 A hepatocellular carcinoma prediction system based on a hepatocellular carcinoma-specific cfDNA fragment feature combination

[0173] This example provides a hepatocellular carcinoma prediction system based on a hepatocellular carcinoma-specific cfDNA fragment feature combination, including a data acquisition module, an analysis module, a storage module, and an output module.

[0174] The data acquisition module is used to obtain the hepatocellular carcinoma-specific cfDNA fragment feature combination of the sample to be tested, and the hepatocellular carcinoma-specific cfDNA fragment feature combination is the hepatocellular carcinoma-specific cfDNA fragment feature combination obtained in Example 3.

[0175] The analysis module is the hepatocellular carcinoma risk prediction model constructed in Example 3. Using the hepatocellular carcinoma-specific cfDNA fragment feature combination of the sample to be tested obtained by the data acquisition module as the input variable, it is input into the hepatocellular carcinoma risk prediction model to obtain the hepatocellular carcinoma risk probability, and based on the hepatocellular carcinoma risk probability, the type of the sample to be tested is determined. The determination criterion is: if the hepatocellular carcinoma risk probability ≥ 0.5, the sample to be tested is a hepatocellular carcinoma sample; if the hepatocellular carcinoma risk probability < 0.5, the sample to be tested is a non-hepatocellular carcinoma sample.

[0176] The storage module is used to store the hepatocellular carcinoma-specific cfDNA fragment feature combination of the sample to be tested obtained by the data acquisition module, the hepatocellular carcinoma risk probability obtained by the analysis module, and the type of the sample to be tested.

[0177] The output module is used to output the hepatocellular carcinoma risk probability and the type of the sample to be tested.

[0178] Example 8. A liver cancer prediction system based on a combination of liver cancer-specific cfDNA fragment features

[0179] This example provides a liver cancer prediction system based on a combination of liver cancer-specific cfDNA fragment features, including a data acquisition module, an analysis module, a storage module, and an output module.

[0180] The data acquisition module is used to obtain the combination of liver cancer-specific cfDNA fragment features of the sample to be tested, and the combination of liver cancer-specific cfDNA fragment features is the combination of liver cancer-specific cfDNA fragment features obtained in Example 4.

[0181] The analysis module is the liver cancer risk prediction model constructed in Example 4. Using the combination of liver cancer-specific cfDNA fragment features of the sample to be tested obtained by the data acquisition module as the input variable, it is input into the liver cancer risk prediction model to obtain the liver cancer risk probability, and based on the liver cancer risk probability, the type of the sample to be tested is determined. The determination criterion is: if the liver cancer risk probability ≥ 0.5, the sample to be tested is a liver cancer sample; if the liver cancer risk probability < 0.5, the sample to be tested is a non-liver cancer sample.

[0182] The storage module is used to store the combination of liver cancer-specific cfDNA fragment features of the sample to be tested obtained by the data acquisition module, the liver cancer risk probability obtained by the analysis module, and the type of the sample to be tested.

[0183] The output module is used to output the liver cancer risk probability and the type of the sample to be tested.

[0184] Finally, it should be noted that the above examples are only used to illustrate the technical solutions of the present invention rather than to limit the protection scope of the present invention. For those of ordinary skill in the art, based on the above description and ideas, other different forms of changes or modifications can also be made. It is not necessary and impossible to enumerate all the implementation manners here. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the claims of the present invention.

Claims

1. Use of a reagent for detecting a characteristic combination of hepatocellular carcinoma-specific cfDNA fragments in peripheral blood in the preparation of a product for predicting hepatocellular carcinoma, characterized in that, The specific cfDNA fragment feature combination for liver cancer includes the frequency of the 6bp motif at the 5'-end of the cfDNA fragment combination and the length distribution index of cfDNA fragments in the 10kb genomic window combination; The cfDNA fragment combination consists of cfDNA fragments with a length less than 150bp and a 6bp motif at the 5'-end of TGCTGT, TGGCTA, CGGCCT, TGCACA, TGACCA, GGTCCA, TGCCAC, TCCACT, GCGTAA, and TGCATG; The 10kb genomic window combination consists of the Chr5:119070000-119079999 window, the Chr1:175850000-175859999 window, the Chr7:28430000-28439999 window, the Chr5:133820000-133829999 window, and the Chr16:12790000-12799999 window; 2. Use of a reagent for detecting a characteristic combination of hepatocellular carcinoma-specific cfDNA fragments in peripheral blood in the preparation of a product for predicting hepatocellular carcinoma, characterized in that, The specific cfDNA fragment feature combination for liver cancer includes the frequency of the 6bp motif at the 5'-end of the first cfDNA fragment combination, the frequency of the 6bp motif at the 5'-end of the second cfDNA fragment combination, the length distribution index of cfDNA fragments in the 10kb genomic window combination, and the cfDNA fragment coverage of the Chr2:87940000-87949999 window; The first cfDNA fragment combination consists of cfDNA fragments with a 6bp motif at the 5'-end of ATGGGC, TGCTTC, TGATCC, ACGAGA, TGACTC, ACGAGT, and ACGATG; The second cfDNA fragment combination consists of cfDNA fragments with a length less than 150bp and a 6bp motif at the 5'-end of TGTGAC, CGGCTA, GCGAAA, TGTAGC, and TGGCTC; The 10kb genomic window combination consists of the Chr7:28430000-28439999 window and the Chr16:12790000-12799999 window; 3. Use of a reagent for detecting a characteristic combination of hepatocellular carcinoma-specific cfDNA fragments in peripheral blood in the preparation of a product for predicting hepatocellular carcinoma, characterized in that, The specific cfDNA fragment feature combination for liver cancer includes the frequency of the 6bp motif at the 5'-end of the first cfDNA fragment combination and the frequency of the 6bp motif at the 5'-end of the second cfDNA fragment combination; The first cfDNA fragment combination consists of cfDNA fragments with a 6bp motif at the 5'-end of TGATCC, TGTTCC, TGCTTC, ACGAAG, ACGAGA, TGACTC, CAGTTC, CAGTCC, AAGAGC, TGATCG, and CAGTCA; The second cfDNA fragment combination consists of cfDNA fragments with a length less than 150bp and a 6bp motif at the 5'-end of CGGTGG, CGGCCT, CGGCTA, and TGCCTG.

4. Use of a reagent for detecting a characteristic combination of hepatocellular carcinoma-specific cfDNA fragments in peripheral blood in the preparation of a product for predicting hepatocellular carcinoma, characterized in that, The combination of liver cancer-specific cfDNA fragment features includes the frequency of the 5'-terminal 6bp motif of the first cfDNA fragment combination, the frequency of the 5'-terminal 6bp motif of the second cfDNA fragment combination, and the cfDNA fragment coverage of the genomic 10kb window combination; The first cfDNA fragment combination consists of cfDNA fragments with 5'-terminal 6bp motifs of TGCTTC, CAGTCA, TGTTCC, ACGAAG, and ATCGGG; The second cfDNA fragment combination consists of cfDNA fragments with a length less than 150bp and 5'-terminal 6bp motifs of TTATGG, TGTTCG, and TGTTCA; The genomic 10kb window combination consists of the Chr2:87940000-87949999 window, the Chr6:24860000-24869999 window, the Chr5:550000-559999 window, the Chr18:26740000-26749999 window, the Chr10:11170000-11179999 window, the Chr15:62300000-62309999 window, and the Chr1:231610000-231619999 window.

5. A method for constructing a liver cancer risk prediction model, characterized in that, Using the combination of liver cancer-specific cfDNA fragment features described in any one of claims 1 to 4 as input features, the liver cancer risk probability as the output result, and AUC as the performance index, the classification model is trained to obtain the liver cancer risk prediction model.

6. The construction method according to claim 5, characterized in that The classification model includes random forest, support vector machine, gradient boosting machine, or LASSO regression.

7. The construction method according to claim 5, characterized in that The cutoff value of the classification model is set to 0.5; the training is performed based on 10-fold repeated cross-validation.

8. A system for predicting the risk of liver cancer, characterized in that, It includes a data acquisition module, an analysis module, and an output module; The data acquisition module is used to obtain the combination of liver cancer-specific cfDNA fragment features of the test sample, and the combination of liver cancer-specific cfDNA fragment features is the combination of liver cancer-specific cfDNA fragment features described in any one of claims 1 to 4; The analysis module is the liver cancer risk prediction model obtained by the construction method described in claim 5 or 6. Using the combination of liver cancer-specific cfDNA fragment features of the test sample obtained by the data acquisition module as the input variable, it is input into the liver cancer risk prediction model to obtain the liver cancer risk probability; The output module is used to output the liver cancer risk probability.

9. A computer device, comprising a memory and a processor, characterized in that, A computer program executable on the processor is stored on the memory; when the computer program is executed by the processor, the operations of the system described in claim 8 are implemented.

10. A computer-readable storage medium stores a computer program executable by a processor, characterized in that, When the computer program is executed by the processor, the operations of the system described in claim 8 are implemented.