Liquid biopsy-based tumor marker micro-detection and analysis method

By constructing a cross-scale graph neural network model that integrates ctDNA fragmentation features and methylation patterns, the problem of high false negative rates in existing technologies has been solved, achieving high sensitivity and high specificity in early lung cancer screening.

CN121545577BActive Publication Date: 2026-04-17THE PEOPLES HOSPITAL SHAANXI PROV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
THE PEOPLES HOSPITAL SHAANXI PROV
Filing Date
2026-01-16
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

In existing technologies, tumor marker detection methods based on ctDNA ignore the biological coupling between the end sequence of the fragment and methylation modification, resulting in a high false negative rate in early lung cancer screening and an inability to effectively distinguish between benign and malignant lung nodules.

Method used

A cross-scale graph neural network model is constructed, which integrates multi-dimensional features such as ctDNA fragment length distribution, terminal sequence preference, nucleosome localization signal and CpG island methylation status for joint modeling. Through a multi-layer message passing mechanism, the coupling pattern between fragmentation features and methylation features at the local chromatin structure level is learned, and a malignancy risk score is output.

Benefits of technology

It significantly improves the sensitivity and specificity of early lung cancer screening, reduces the false negative rate by 36%, and achieves non-invasive, high-precision early lung cancer screening.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121545577B_ABST
    Figure CN121545577B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of biological medicine, and discloses a tumor marker micro-detection and analysis method based on liquid biopsy. The method comprises the following steps: extracting peripheral blood free DNA and performing double-end sequencing; jointly extracting multi-dimensional features such as ctDNA fragment length distribution, end sequence preference, nucleosome positioning signal and CpG methylation state from the sequencing data; constructing a feature matrix under unified genomic coordinates through space-time alignment; inputting a pre-trained cross-scale graph neural network model, taking a genomic region as a node and physical and functional proximity as an edge, learning a coupling mode of multi-modal features in a local chromatin structure, and outputting a malignant risk score; and combining a threshold to distinguish early lung cancer and benign and malignant nodules. The application comprises corresponding functional units and supports efficient parallel analysis. The system significantly improves detection sensitivity and specificity, reduces the false negative rate by 36%, and realizes non-invasive and accurate early lung cancer screening.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of biomedicine, specifically relating to a method for trace detection and analysis of tumor markers based on liquid biopsy. Background Technology

[0002] With the increasing application of liquid biopsy technology in early cancer screening and auxiliary diagnosis, non-invasive detection methods based on circulating tumor DNA (ctDNA) have become an important research direction in precision medicine. Liquid biopsy, by analyzing cell-free ctDNA in the blood, can reflect the genomic, epigenetic, and fragmentation characteristics of tumors in real time, offering advantages such as minimal invasiveness, high reproducibility, and potential for dynamic monitoring. Among these advantages, the integration of multi-omics information from ctDNA, including mutation profiles, methylation status, and fragmentation patterns, is considered a key pathway to improve the sensitivity and specificity of early tumor identification.

[0003] This research, based on the combined analysis of ctDNA fragmentation characteristics and methylation patterns, aims to construct a more robust biomarker system by exploring the intrinsic correlation between breakpoint preference and abnormal CpG island methylation in tumor-derived DNA. The core of this approach lies in the fact that ctDNA released by tumor cells not only exhibits specificity in length distribution (e.g., enrichment of short fragments), but its breakpoint sequences also carry "fingerprint" information and are often accompanied by high methylation of CpG islands in promoter regions. However, existing detection methods mostly model fragment length, end sequences, and methylation data as independent features separately, lacking the ability to systematically model the synergistic mechanism of these three factors.

[0004] While current mainstream machine learning or deep learning models have been applied to tumor marker identification, they generally overlook the biological coupling between nuclease cleavage bias and methylation modification inherent in the terminal sequences of ctDNA fragments. These models typically classify based on only a single dimension, either methylation level or fragment length, resulting in insufficient ability to distinguish between benign lung nodules and malignant tumors in early lung cancer screening scenarios, leading to a high false-negative rate. Especially in the low tumor burden stage, ctDNA signals are weak; without effectively fusing the combined features of terminal sequence fingerprints and local methylation status, crucial discriminative information is easily missed. Therefore, there is an urgent need for an intelligent analysis method that can deeply integrate ctDNA fragment terminal sequences, length distribution, and CpG island methylation status to significantly improve the accuracy and reliability of early lung cancer nodule differentiation. Summary of the Invention

[0005] This invention provides a method for trace detection and analysis of tumor markers based on liquid biopsy, aiming to solve the technical problem of high false negative rates and inability to accurately distinguish between benign and malignant lung nodules in existing multi-omics data analysis algorithms for tumor markers, which ignore the intrinsic correlation between circulating tumor DNA fragmentation characteristics and methylation patterns. This invention constructs a multi-dimensional feature joint modeling architecture that integrates ctDNA fragment length distribution profiles, terminal sequence preferences, nucleosome localization signals, and CpG site methylation status, achieving highly sensitive identification and pathological characterization of early lung cancer biomarkers.

[0006] According to one aspect of the present invention, a method for trace detection and analysis of tumor markers based on liquid biopsy is provided, comprising:

[0007] Free deoxyribonucleic acid was extracted from peripheral blood samples of the subjects;

[0008] The free deoxyribonucleic acid was subjected to paired-end sequencing to obtain the raw sequencing reads;

[0009] The original sequencing reads were subjected to quality control, adapter removal, and alignment to the human reference genome to generate aligned sequencing data.

[0010] Based on the alignment of the sequencing data, ctDNA fragmentation features and methylation features were extracted respectively. The ctDNA fragmentation features include fragment length distribution, fragment terminal 5-base sequence composition, and nucleosome protective region coverage depth ratio. The methylation features include methylation level of target CpG sites, methylation density gradient, and methylation phase consistency.

[0011] Spatiotemporal alignment and coordinate mapping were performed on the fragmented and methylated features of the ctDNA to construct a multidimensional feature matrix in a unified genome coordinate system.

[0012] The multidimensional feature matrix is ​​input into a pre-trained cross-scale graph neural network model. The cross-scale graph neural network model uses genomic regions as nodes and physical proximity and functional regulatory relationships as edges. It learns the coupling pattern of fragmented features and methylation features at the local chromatin structure level through a multi-layer message passing mechanism.

[0013] The malignancy risk score for each candidate genomic region is output by the cross-scale graph neural network model.

[0014] Based on the malignancy risk score, combined with a preset threshold, it is determined whether the tested sample originates from early-stage lung cancer, and the benign or malignant nature of the lung nodules is classified.

[0015] As one embodiment of the present invention, the extraction of free deoxyribonucleic acid specifically includes: purifying free deoxyribonucleic acid from plasma using a silica gel membrane column method, with an elution volume of 30 μL, and the concentration of the obtained deoxyribonucleic acid being not less than 0.5 nanograms per μL, with the main peak of the fragment located near 165 base pairs.

[0016] As one embodiment of the present invention, the paired-end sequencing adopts a paired-end sequencing mode with a read length of 150 base pairs and a sequencing depth of not less than 30,000 times coverage.

[0017] As one embodiment of the present invention, the method for extracting the fragment length distribution is as follows: count the lengths of all successfully aligned ctDNA fragments, divide the fragments into intervals of 10 base pairs, calculate the percentage of the number of fragments in each interval relative to the total number of fragments, and form a length distribution vector.

[0018] As one embodiment of the present invention, the method for extracting the 5-base sequence composition at the end of the fragment is as follows: five bases are cut off from the 5' end and five bases from the 3' end of each ctDNA fragment, and the frequency of 4-base combinations in the 10-base window of all fragments on the positive and negative strands is counted. After one-hot encoding, the terminal sequence feature tensor is generated.

[0019] As one embodiment of the present invention, the method for calculating the coverage depth ratio of the nucleosome protected region is as follows: the nucleosome core region is defined as a sliding window of 147 base pairs in length, the ratio of the average sequencing depth within this window to the average sequencing depth of the adjacent linker region is calculated across the entire genome, and regions with a ratio greater than 1.5 are taken as nucleosome-rich regions, and the average coverage depth ratio of the nucleosome-rich regions in the promoter regions of known lung cancer-related genes is statistically analyzed.

[0020] As one embodiment of the present invention, the method for obtaining the methylation level of the target CpG site is as follows: based on the sequencing data after bisulfite treatment, the proportion of methylated cytosine readings to the total readings of each CpG site is calculated, and the proportion value is between 0 and 1.

[0021] As one embodiment of the present invention, the method for calculating the methylation density gradient is as follows: within a range of 2000 base pairs upstream and downstream of the promoter region of the target gene, a sliding window of 500 base pairs is used to calculate the average methylation level of CpG sites within the window, forming a methylation density curve, and the mean of the absolute values ​​of the first derivative of the curve is used as the methylation density gradient index.

[0022] As one embodiment of the present invention, the method for calculating the methylation phase consistency is as follows: the methylation status of two adjacent CpG sites on the same DNA molecule is jointly analyzed. If both are methylated or both are unmethylated, they are recorded as consistent. The proportion of consistent events in all co-sequencing molecules is then counted.

[0023] As one embodiment of the present invention, the spatiotemporal alignment and coordinate mapping specifically includes: sorting all features according to genomic coordinates, filling missing feature positions with nearest neighbor interpolation, and ensuring that each genomic window corresponds to a complete feature vector; the size of the genomic window is 5000 base pairs, and the step size is 1000 base pairs.

[0024] In one embodiment of the present invention, the cross-scale graph neural network model includes an input layer, three graph convolutional layers, a global attention pooling layer, and a fully connected output layer; the graph convolutional layers employ a gated update mechanism, and the node feature update formula is:

[0025] ;in For the l-th layer node eigenvectors, This represents element-wise multiplication. for The set of neighboring nodes, To modify the activation function of the linear unit, and This is the learnable parameter matrix.

[0026] As one embodiment of the present invention, the edge weight The calculation method is as follows: if the distance between the genomes of two nodes is less than 100,000 base pairs, then ,in The distance between base pairs. The attenuation constant is 50,000. The normalized value of the interaction frequency of pairs in this region during chromatin conformation capture experiments; if the distance is greater than or equal to 100,000 base pairs, then ;

[0027] In one embodiment of the present invention, the global attention pooling layer performs a weighted summation of all node features. The weights are generated by the softmax function through the dot product of the learnable query vector and the node features, and the output is a graph-level representation vector of fixed dimensions.

[0028] In one embodiment of the present invention, the malignancy risk score is a probability value output by a graph-level representation vector through a fully connected layer and a sigmoid function, with a value range of 0-1.

[0029] As one embodiment of the present invention, the preset threshold is 0.37. When the malignancy risk score is greater than or equal to the threshold, it is determined to be a malignant nodule; otherwise, it is determined to be a benign nodule.

[0030] According to another aspect of the present invention, a system for trace detection and analysis of tumor markers based on liquid biopsy is provided, comprising:

[0031] The free deoxyribonucleic acid extraction unit is used to extract free deoxyribonucleic acid from peripheral blood samples of subjects.

[0032] A high-throughput sequencing unit is used to perform paired-end sequencing on the free deoxyribonucleic acid to obtain raw sequencing reads;

[0033] The data preprocessing unit is used to perform quality control, adapter removal, and alignment to the human reference genome on the raw sequencing reads to generate aligned sequencing data.

[0034] A multidimensional feature extraction unit is used to extract ctDNA fragmentation features and methylation features based on the aligned sequencing data, respectively.

[0035] The feature alignment and matrix construction unit is used to perform spatiotemporal alignment and coordinate mapping on the ctDNA fragmentation features and methylation features to construct a multidimensional feature matrix in a unified genome coordinate system.

[0036] A cross-scale graph neural network analysis unit is used to input the multidimensional feature matrix into a pre-trained cross-scale graph neural network model and output a malignancy risk score for each candidate genomic region.

[0037] The benign / malignant discrimination unit is used to determine whether the tested sample originates from early-stage lung cancer based on the malignancy risk score and a preset threshold, and to classify the benign or malignant nature of lung nodules.

[0038] As one embodiment of the present invention, the multidimensional feature extraction unit includes a fragment length analysis module, an end sequence encoding module, a nucleosome localization calculation module, a CpG methylation quantification module, a methylation density gradient calculation module, and a methylation phase consistency evaluation module.

[0039] As one embodiment of the present invention, the cross-scale graph neural network analysis unit is deployed on a server cluster with graphics processor acceleration capabilities, supporting parallel inference tasks that process no less than 100 samples in batches, with a single sample analysis time not exceeding 45 minutes.

[0040] In one embodiment of the present invention, the system further includes a model training unit for end-to-end training of the cross-scale graphical neural network model using labeled liquid biopsy datasets of lung cancer patients and healthy controls; during training, a focus loss function is used to alleviate the class imbalance problem, and the focus parameters... The value is 2.

[0041] As one embodiment of the present invention, the labeled dataset contains no fewer than 800 early lung cancer samples and 1200 benign lung nodules or healthy human samples, all of which have been confirmed by histopathology.

[0042] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0043] 1. This invention is the first to jointly model ctDNA fragmentation features and methylation patterns at the three-dimensional chromatin structure level, overcoming the technical limitations of traditional multi-omics analysis that treats different molecular features independently. By constructing a graph neural network with genomic regions as nodes and physical and functional proximity as edges, this invention can capture the synergistic changes in fragment breakage bias and local methylation status against the background of nucleosome arrangement, thereby significantly improving the ability to identify weak signals in early lung cancer.

[0044] 2. Experiments show that in an independent validation cohort including 300 cases of early-stage lung cancer and 400 cases of benign nodules, the sensitivity of this invention reached 89.3%, the specificity reached 92.7%, and the false negative rate was reduced by 36 percentage points compared with existing detection methods based on a single methylation marker.

[0045] 3. This invention does not rely on prior imaging information and can determine benign or malignant nodules based solely on peripheral blood samples. It provides a non-invasive, high-precision, and repeatable early lung cancer screening tool for clinical use, effectively solving the core problem that existing technologies cannot accurately distinguish between benign and malignant lung nodules. Attached Figure Description

[0046] Figure 1 This is a schematic diagram of the overall technical solution architecture of the method for trace detection and analysis of tumor markers based on liquid biopsy proposed in this invention;

[0047] Figure 2 This is a schematic diagram of the core principle framework of the coupling modeling of ctDNA fragmentation features and methylation patterns at the chromatin structure level in this invention;

[0048] Figure 3 This is a flowchart illustrating the logical process of multidimensional feature extraction and spatiotemporal alignment in this invention.

[0049] Figure 4 This is a schematic diagram illustrating the node-edge construction and message passing mechanism of the cross-scale graph neural network model in this invention.

[0050] Figure 5 This is a flowchart illustrating the end-to-end analysis process from raw sequencing data to malignancy risk score generation in this invention.

[0051] Figure 6 This is a schematic diagram of the multi-level interaction relationships and data flow between the functional units of the liquid biopsy system in this invention. Detailed Implementation

[0052] Please refer to Figures 1-6 This invention provides a method for trace detection and analysis of tumor markers based on liquid biopsy. Its core lies in constructing a multidimensional joint characterization against the three-dimensional chromatin structure background by fusing fragmented features and methylation patterns of circulating tumor DNA. A cross-scale graphical neural network model is then used to achieve highly sensitive identification of early-stage lung cancer and differentiation of benign and malignant pulmonary nodules. The following detailed description of the technical implementation will follow each of the method steps S1-S8 explicitly listed in the invention description.

[0053] The method first performs step S1: extracting free deoxyribonucleic acid from the peripheral blood sample of the subject. This step is performed using the silica gel membrane column method. The specific procedure is as follows: collect 10 ml of venous blood from the subject, anticoagulate with EDTA, centrifuge at 1600 rpm for 10 minutes to separate the plasma, take the supernatant and centrifuge again at 16000 rpm for 10 minutes to remove residual cell debris, and obtain clear plasma.

[0054] The plasma sample was then added to a centrifuge column containing lysis buffer, mixed thoroughly, and centrifuged to allow free deoxyribonucleic acid to adsorb onto the silica membrane.

[0055] Wash twice each with washing buffer 1 and washing buffer 2, discarding the filtrate after each centrifugation; finally, wash with 30 μL of nuclease-free water. The concentration of the resulting free deoxyribonucleic acid solution is not less than 0.5 ng / μL, and the main peak of its fragment length distribution is located near 165 base pairs, which is consistent with the typical characteristics of nucleosome protected fragments.

[0056] The entire extraction process was completed in a biosafety level 2 laboratory, and all consumables were autoclaved to avoid contamination by exogenous deoxyribonucleic acid.

[0057] Next, step S2 is performed: paired-end sequencing of the free deoxyribonucleic acid (DNA) to obtain raw sequencing reads. This step uses paired-end sequencing mode and is performed on a high-throughput sequencing platform. Before sequencing, library construction is required for the extracted DNA, including standard procedures such as end repair, A-tailing, adapter ligation, and fragment selection. The adapters used contain unique molecular identifier sequences for subsequent removal of polymerase chain reaction (PCR) amplification duplications.

[0058] After the library passed quality control, it was loaded into the sequencer's flow cell for sequencing with 150-base-pair reads at both ends. The sequencing depth was set to cover the entire genome at least 30,000 times to ensure that low-abundance tumor-derived fragments could be effectively captured. The raw sequencing data was output in a rapid whole-genome sequencing format, including forward and reverse read files, with each read carrying quality score information.

[0059] The S3 step then proceeds: quality control, adapter removal, and alignment to the human reference genome are performed on the raw sequencing reads to generate aligned sequencing data. In the quality control phase, a sliding window method is used to assess the base quality of each read, with a window size of 4 base pairs. If the average quality score within the window is below 20, the subsequent portion of the read is truncated; reads shorter than 30 base pairs are also filtered out. Adapter removal uses a specific sequence matching algorithm to identify and remove residual sequencing adapters at both ends of the reads. The alignment process employs a modified Burroughs-Wyler transform indexing algorithm to map the processed reads to version 38 of the human reference genome, allowing a maximum of two mismatches and prohibiting cross-splicing of splice sites. The alignment results are stored in a sorted binary alignment format, including the chromosomal location, alignment direction, insert length, and unique molecular identifier tag for each read.

[0060] Based on this, proceed with step S4:

[0061] Based on the alignment and sequencing data, ctDNA fragmentation and methylation features are extracted. This step is further divided into six sub-operations.

[0062] The first sub-operation is to extract the fragment length distribution: traverse all successfully aligned read pairs, calculate their inserted fragment length, which is the number of bases between the 5' end of the forward read and the 5' end of the reverse read; divide the intervals into 10 base pairs, resulting in 29 intervals from 50 to 350 base pairs; count the number of fragments falling into each interval, divide by the total number of fragments to obtain the percentage, and form a 29-dimensional length distribution vector.

[0063] The second sub-operation is to extract the 5-base sequence composition of the fragment ends: for each read pair, 5 bases from the 5' end of the positive strand read and 5 bases from the 5' end of the negative strand read (corresponding to the 3' end of the original double-stranded deoxyribonucleic acid) are extracted to form a 10-base window; the frequency of 4-base combinations of all fragments in this window is counted, with a total of 1020 possible combinations; the frequency matrix is ​​one-hot encoded to generate a three-dimensional tensor of shape (total number of fragments, 10, 4).

[0064] The third sub-operation involves calculating the coverage depth ratio of nucleosome protected regions: the nucleosome core region is defined as a sliding window of 147 base pairs with a step size of 25 base pairs, scanned across the entire genome; the linker region is defined as the spacer region between adjacent nucleosome cores; the ratio of the average sequencing depth within each nucleosome core window to the average depth of the linker regions on both sides is calculated; regions with a ratio greater than 1.5 are selected as nucleosome-rich regions; further focusing is placed on the promoter regions (from 2,000 base pairs upstream to 500 base pairs downstream of the transcription start site) of known lung cancer-related genes such as EGFR, KRAS, and TP53, and the average coverage depth ratio of nucleosome-rich regions in these regions is calculated as an indicator of nucleosome localization signals.

[0065] The fourth sub-operation is to obtain the methylation level of the target CpG site: based on the sequencing data after bisulfite treatment (in this embodiment, some libraries are converted to bisulfite before sequencing), the ratio of methylated cytosine readings (i.e., readings that have not undergone C to T conversion) to the total readings is calculated for each CpG site. This ratio is strictly between 0 and 1.

[0066] The fifth sub-operation is to calculate the methylation density gradient: within a range of 2000 base pairs upstream and downstream of the promoter region of the lung cancer-related gene, with a sliding window of 500 base pairs and a step size of 100 base pairs, the average methylation level of all CpG sites within each window is calculated to form a continuous methylation density curve; the curve is numerically differentiated, and the mean of the absolute values ​​of the first derivative is obtained as the methylation density gradient index.

[0067] The sixth sub-operation is to calculate methylation phase consistency: the original single-molecule deoxyribonucleic acid sequence is reconstructed using a unique molecular identifier tag; for any two adjacent CpG sites on the same molecule (with a spacing of no more than 300 base pairs), it is determined whether their methylation status is consistent (both methylated or both unmethylated); the proportion of consistent events in all co-sequencing molecules is counted, and this proportion is the methylation phase consistency.

[0068] Next, step S5 is performed: the ctDNA fragmentation features and methylation features are spatiotemporally aligned and mapped to coordinates to construct a multidimensional feature matrix in a unified genome coordinate system.

[0069] This step first divides the entire genome into fixed windows, each 5000 base pairs in size and 1000 base pairs in step size, forming an overlapping window set. For each window, all fragmentation and methylation feature data falling within its range are collected. For the fragment length distribution vector, the length distribution statistics of all fragments within the window are taken. For the terminal sequence features, only the terminal sequences of the starting or ending fragments within the window are retained.

[0070] For the nucleosome coverage depth ratio, the weighted average of the ratios of all nucleosome-rich regions within the window is taken; for methylation characteristics, the statistical values ​​of methylation level, density gradient, and phase consistency of all CpG sites within the window are taken.

[0071] When a specific feature is missing within a window (e.g., the absence of CpG sites results in empty methylation data), nearest neighbor interpolation is used to fill the gap: the nearest non-empty window is searched forward or backward, and its corresponding feature value is used for filling. Ultimately, each genomic window corresponds to a feature vector of fixed dimensions, with the dimension formed by concatenating the feature sub-vectors to create a multidimensional feature matrix with rows equal to the total number of windows and columns equal to the total feature dimension.

[0072] Then, step S6 is executed: the multidimensional feature matrix is ​​input into a pre-trained cross-scale graph neural network model. The cross-scale graph neural network model uses genomic regions as nodes and physical proximity and functional regulatory relationships as edges. It learns the coupling pattern of fragmented features and methylation features at the local chromatin structure level through a multi-layer message passing mechanism.

[0073] The model is constructed as follows: each 5000-base-pair window in step S5 is considered a node in a graph, and the initial feature of the node is the multidimensional feature vector corresponding to that window; the establishment of edges between nodes is based on two criteria:

[0074] If the distance between the genome centers of two nodes is less than 100,000 base pairs, then an edge is established;

[0075] If two nodes are located within a known chromatin topological association domain or share an enhancer-promoter regulatory relationship, then a connection is forcibly established. The formula for calculating the edge weight η(u,v) is:

[0076] If the distance between the genomes of two nodes <100,000 base pairs, then ;in The attenuation constant is 50,000, and c is the normalized value of the interaction frequency in this region (normalized to the 0-1 interval) from the publicly available chromatin conformation capture experimental database; if If the number of base pairs is greater than or equal to 100,000, then .

[0077] The model architecture consists of an input layer, three graph convolutional layers, a global attention pooling layer, and a fully connected output layer. The graph convolutional layers employ a gated update mechanism, and the feature update formula for node v in the l-th layer is:

[0078] ;in For the l-th layer node eigenvectors, for The set of neighboring nodes, This represents element-wise multiplication. To modify the activation function of the linear unit, and The learningable parameter matrix has dimensions that are dynamically adjusted based on the input features. After three layers of graph convolution, the features of each node have been fused with fragmentation and methylation coupling information from its multi-hop neighborhood.

[0079] Then, step S7 is performed: the malignancy risk score for each candidate genomic region is output by the cross-scale graph neural network model. This score is generated through a global attention pooling layer and a fully connected output layer.

[0080] The global attention pooling layer first learns a trainable query vector. Calculate its features with each node. The dot product is then processed by the softmax function to generate normalized weights. Then, the features of all nodes are weighted and summed to obtain a fixed-dimensional graph-level representation vector. This vector The input is fed into an output module containing two fully connected layers. The first layer has an output dimension of 128, and the second layer has an output dimension of 1. Finally, after activation by a sigmoid function, a malignancy risk score is output. ,in For the last layer of linear output, The value must be strictly between 0 and 1.

[0081] Finally, step S8 is executed: Based on the malignancy risk score and a preset threshold, it is determined whether the tested sample originates from early-stage lung cancer, and the benign or malignant nature of the lung nodules is classified. In this embodiment, the preset threshold is 0.37, which is determined by maximizing the Youden index on an independent validation set. If the malignancy risk score s is greater than or equal to 0.37, the tested sample is determined to originate from early-stage lung cancer, and the corresponding lung nodule is malignant; if s is less than 0.37, it is determined to be a benign nodule or a healthy condition. The determination result is output in the form of a structured report, including the risk score value, determination conclusion, and confidence interval.

[0082] The above-described method steps fully realize an end-to-end analysis workflow from peripheral blood samples to benign / malignant differentiation. Throughout the workflow, the data flow exhibits strict temporal dependencies: the extraction quality of S1 directly affects the completeness of sequencing data in S2; the alignment accuracy of S3 determines the accuracy of feature extraction in S4; the alignment quality of S5 restricts the effectiveness of model input in S6; the model performance of S6 directly determines the reliability of the scoring in S7; and the threshold setting in S8 ultimately affects the sensitivity and specificity of clinical differentiation.

[0083] Data verification mechanisms are implemented between each step: after S1, the concentration of deoxyribonucleic acid and fragment distribution are detected; after S2, the distribution of sequencing quality scores and the diversity of unique molecular identifiers are evaluated; after S3, the alignment rate and coverage uniformity are checked; after S4, the feature distribution is verified to conform to biological priors; after S5, it is confirmed that there are no rows with all zeros in the feature matrix; and during the S6 inference process, the gradient norm is monitored to prevent explosion or disappearance.

[0084] The abnormal situation handling strategy includes: if the S1 extraction concentration is lower than the threshold, resample; if the S2 sequencing depth is insufficient, perform additional sequencing; if the S3 alignment rate is lower than 90%, check the consistency of the reference genome version; if the S4 feature loss rate exceeds 20%, mark the sample as low quality and recommend retesting.

[0085] At the system implementation level, the method is supported by an integrated analysis system. This system includes a free deoxyribonucleic acid (DNA) extraction unit, whose hardware is an automated nucleic acid extraction workstation with built-in silica membrane column consumables and a programmed liquid handling module; a high-throughput sequencing unit is a commercial sequencer equipped with a paired-end 150-base-pair sequencing kit; and a data preprocessing unit is deployed on a central processing unit server, running a customized bioinformatics workflow and integrating quality control, adapter removal, and alignment software.

[0086] The multidimensional feature extraction unit consists of multiple parallel computing modules, including a fragment length analysis module, an end sequence encoding module, a nucleosome localization calculation module, a CpG methylation quantification module, a methylation density gradient calculation module, and a methylation phase consistency evaluation module. Each module shares intermediate data through a memory-mapped file. The feature alignment and matrix construction unit is responsible for window partitioning, feature aggregation, and missing value imputation, and outputs a standardized feature matrix.

[0087] The cross-scale graph neural network analysis unit is deployed on a graphics processor accelerated server cluster, supporting parallel inference tasks that process no fewer than 100 samples in batches, with a single sample analysis taking no more than 45 minutes; the benign / malignant discrimination unit performs threshold comparison and result formatting.

[0088] The system also includes a model training unit, which uses 800 labeled early-stage lung cancer samples and 1200 benign lung nodules or healthy human samples (all confirmed by histopathology) to perform end-to-end training of the cross-scale graphical neural network model; the focus loss function is used during training. ,in As a category balance factor, The focus parameter is the predicted probability of the true class. A value of 2 is set to mitigate the class imbalance problem. All system units communicate asynchronously through message queues to ensure stability under high throughput.

[0089] The technical solution described in this embodiment reveals the molecular signal coupling pattern unique to early lung cancer at the chromatin structure level by deeply integrating ctDNA fragmentation features and methylation patterns, thereby significantly improving detection sensitivity and specificity. It solves the core problems of high false negative rates and difficulty in distinguishing between benign and malignant tumors caused by neglecting the correlation between the two types of features in existing technologies.

Claims

1. A method for microdetection and analysis of tumor markers based on liquid biopsy, characterized by, include: Free deoxyribonucleic acid was extracted from peripheral blood samples of the subjects; The free deoxyribonucleic acid was subjected to paired-end sequencing to obtain the raw sequencing reads; The original sequencing reads were subjected to quality control, adapter removal, and alignment to the human reference genome to generate aligned sequencing data. Based on the alignment of the sequencing data, ctDNA fragmentation features and methylation features were extracted respectively. The ctDNA fragmentation features include fragment length distribution, fragment terminal 5-base sequence composition, and nucleosome protective region coverage depth ratio. The methylation features include methylation level of target CpG sites, methylation density gradient, and methylation phase consistency. Spatiotemporal alignment and coordinate mapping were performed on the fragmented and methylated features of the ctDNA to construct a multidimensional feature matrix in a unified genome coordinate system. The multidimensional feature matrix is ​​input into a pre-trained cross-scale graph neural network model. The cross-scale graph neural network model uses genomic regions as nodes and physical proximity and functional regulatory relationships as edges. It learns the coupling pattern of fragmented features and methylation features at the local chromatin structure level through a multi-layer message passing mechanism. The malignancy risk score for each candidate genomic region is output by the cross-scale graph neural network model. Based on the malignancy risk score, combined with a preset threshold, it is determined whether the tested sample originates from early-stage lung cancer, and the benign or malignant nature of the lung nodules is classified.

2. The method for trace detection and analysis of tumor markers based on liquid biopsy according to claim 1, characterized in that, Free deoxyribonucleic acid (DNA) was extracted from peripheral blood samples of the subjects, including: Free deoxyribonucleic acid (DNA) was purified from plasma using a silica gel membrane column method with an elution volume of 30 μL. The resulting DNA concentration was not less than 0.5 ng / μL, and the main peak of the fragment was located near 165 base pairs.

3. The method for trace detection and analysis of tumor markers based on liquid biopsy according to claim 1, characterized in that, The free deoxyribonucleic acid was subjected to paired-end sequencing to obtain raw sequencing reads, including: Paired-end sequencing mode was used, with a read length of 150 base pairs and a sequencing depth of at least 30,000 times coverage.

4. The method for trace detection and analysis of tumor markers based on liquid biopsy according to claim 1, characterized in that, Based on the alignment and sequencing data, ctDNA fragmentation features and methylation features were extracted, including: The lengths of all successfully aligned ctDNA fragments were counted, and the fragments were divided into intervals of 10 base pairs. The percentage of fragments in each interval was calculated to form a length distribution vector. Five bases were extracted from the 5' and 3' ends of each ctDNA fragment. The frequency of 4-base combinations in the 10-base window of each fragment on both the positive and negative strands was counted. After one-hot encoding, the terminal sequence feature tensor was generated. The nucleosome core region is defined as a sliding window of 147 base pairs. The ratio of the average sequencing depth within this window to the average sequencing depth of the adjacent linker region is calculated across the entire genome. Regions with a ratio greater than 1.5 are taken as nucleosome-rich regions. The mean ratio of the coverage depth of nucleosome-rich regions in the promoter regions of known lung cancer-related genes is statistically analyzed. Based on the sequencing data after bisulfite treatment, the proportion of methylated cytosine readings to the total readings of each CpG site was calculated, with the proportion value ranging from 0 to 1. Within a 2000-base-pair range upstream and downstream of the promoter region of the target gene, with a sliding window of 500 base pairs, the average methylation level of CpG sites within the window is calculated to form a methylation density curve. The mean of the absolute values ​​of the first derivative of the curve is used as the methylation density gradient index. The methylation status of two adjacent CpG sites on the same DNA molecule is analyzed together. If both sites are methylated or both are unmethylated, they are considered consistent. The proportion of consistent events in all co-sequencing molecules is then calculated.

5. The method for trace detection and analysis of tumor markers based on liquid biopsy according to claim 1, characterized in that, The ctDNA fragmentation and methylation features are spatiotemporally aligned and mapped to coordinates to construct a multidimensional feature matrix in a unified genome coordinate system, including: The entire genome was divided into windows of 5000 base pairs each, with a step size of 1000 base pairs. Features within each window are aggregated by genomic coordinates, and missing feature locations are filled using nearest-neighbor interpolation to ensure that each window corresponds to a complete feature vector. The feature vectors of each window are concatenated to form a multidimensional feature matrix with the number of rows equal to the total number of windows and the number of columns equal to the total dimension of the features.

6. The method for trace detection and analysis of tumor markers based on liquid biopsy according to claim 1, characterized in that, Inputting the multidimensional feature matrix into a pre-trained cross-scale graph neural network model includes: Each genome window is treated as a node in a graph, and the initial features of the node are the multidimensional feature vectors of the corresponding window. If the distance between the genome centers of two nodes is less than 100,000 base pairs, then an edge is established, with the edge weight... ,in The distance between base pairs. The attenuation constant is 50,000. The normalized value of the interaction frequency of pairs in this region during chromatin conformation capture experiments; if the distance is greater than or equal to 100,000 base pairs, then ; Message passing is performed using a three-layer graph convolutional layer, and the node feature update formula is as follows: , in For the l-th layer node eigenvectors, For the l-th layer node eigenvectors, This represents element-wise multiplication. for The set of neighboring nodes, To modify the activation function of the linear unit, and This is the learnable parameter matrix.

7. The method for trace detection and analysis of tumor markers based on liquid biopsy according to claim 6, characterized in that, The malignancy risk score for each candidate genomic region is output by the cross-scale graph neural network model, including: The global attention pooling layer sums the features of all nodes by weighting them. The weights are generated by the softmax function through the dot product of the learnable query vector and the node features, and the output is a graph-level representation vector with fixed dimensions. The graph-level representation vector is input into the fully connected output layer, and the malignancy risk score with a value range of 0-1 is output through the sigmoid function.

8. The method for trace detection and analysis of tumor markers based on liquid biopsy according to claim 7, characterized in that, Based on the malignancy risk score, and combined with a preset threshold, it is determined whether the tested sample originates from early-stage lung cancer, and the benign or malignant nature of the lung nodules is classified, including: A nodule is classified as malignant if the malignancy risk score is greater than or equal to 0.37; otherwise, it is classified as benign.

9. The method for trace detection and analysis of tumor markers based on liquid biopsy according to claim 4, characterized in that, In the calculation of the nucleosome protection region coverage depth ratio, the known lung cancer-related genes include EGFR, KRAS, and TP53, and their promoter regions are defined as 2000 base pairs upstream to 500 base pairs downstream of the transcription start site.

10. The method for trace detection and analysis of tumor markers based on liquid biopsy according to claim 4, characterized in that, In the calculation of methylation phase consistency, the distance between adjacent CpG sites does not exceed 300 base pairs, and joint analysis is performed based on single-molecule deoxyribonucleic acid sequences reconstructed from unique molecular identifier tags.

Citation Information

Patent Citations

  • Preparation method of tumor tissue continuous pathological section

    CN112649267A

  • Cancer detection model and construction method and kit thereof

    CN113838533A