A multi-dimensional tumor marker combined detection device and a method for using the same

The multidimensional tumor marker joint detection device uses whole-genome ultra-low depth sequencing to simultaneously extract multidimensional tumor markers and uses an integrated learning model for risk assessment. This solves the problems of high cost, complex operation and single marker detection in existing technologies, and achieves efficient and accurate joint screening of multiple cancer types.

CN121583329BActive Publication Date: 2026-04-17SUZHOU HONGYUAN BIOTECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SUZHOU HONGYUAN BIOTECH CO LTD
Filing Date
2026-01-28
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing tumor-related molecular screening technologies suffer from problems such as high cost, complex operation, low sensitivity of single molecular marker detection, insufficient coverage of multi-dimensional molecular features, poor sample compatibility, and limited detection range, making it difficult to meet the needs of large-scale, multi-cancer joint screening.

Method used

A multidimensional tumor marker joint detection device is adopted, including a sample processing module, a sequencing module, a marker extraction module, a feature fusion and detection module, and a result output module. Through whole-genome ultra-low depth sequencing, multidimensional tumor markers such as chromosome copy variation, microsatellite instability, telomere length, nucleosome imprinting, fragment distribution, and methylation level are extracted simultaneously, and risk assessment is performed using an integrated learning model.

Benefits of technology

It enables low-cost, high-accuracy, multi-dimensional tumor marker combined detection, adapts to the needs of large-scale clinical screening, supports multi-cancer combined screening, dynamically captures changes in molecular characteristics, and improves the sensitivity and coverage of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121583329B_ABST
    Figure CN121583329B_ABST
Patent Text Reader

Abstract

This invention relates to the field of biomedical detection technology, specifically to a multi-dimensional tumor marker joint detection device and its usage method. The device includes modules for sample processing, sequencing, marker extraction, feature fusion and detection, and result output. The marker extraction module contains six detection units, corresponding to the extraction of six types of markers, including chromosomal copy variation and microsatellite instability. Method: cfDNA is extracted from bodily fluid samples. A library is directly constructed without special processing. Valid data is obtained through whole-genome sequencing. The six types of markers are simultaneously extracted through various modules of the device, and the algorithm is optimized to adapt to low-depth data. The data are input into an ensemble learning model to output a risk score. Finally, the result output module outputs relevant information. This invention requires no special experimental processing, achieving low-cost, high-accuracy joint detection, and is suitable for large-scale clinical screening needs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of biomedical detection technology, specifically relating to a multidimensional tumor marker combined detection device and its usage method. Background Technology

[0002] Early screening of tumor-related molecular characteristics is a crucial step in improving cancer prevention and control. With the development of precision medicine technology, next-generation sequencing (NGS) technology, which detects circulating cell-free DNA (cfDNA), has become a core direction in the field of molecular screening due to its non-invasive advantage. This technology provides molecular-level reference for tumor risk assessment by analyzing cfDNA molecular signals in ex vivo samples and has received widespread attention in clinical research and population screening scenarios. However, current related technologies still face many technical challenges in practical applications, hindering their large-scale promotion and the enhancement of their application value. These challenges include:

[0003] (1) Existing tumor-related molecular screening technologies mostly rely on high-depth whole-genome sequencing or targeted capture sequencing schemes. Some technologies also require additional experimental steps such as bisulfite treatment and specific probe capture, which not only increases reagent consumption and operational complexity but also keeps detection costs high for a long time. At the same time, some detection methods rely on imported reagent kits and special equipment, further raising the economic threshold for technology implementation and making it difficult to popularize large-scale molecular screening for the general population. In addition, the complex experimental procedures also prolong the detection cycle and cannot meet the actual needs of efficient screening;

[0004] (2) Current mainstream technologies have obvious limitations in molecular feature coverage: On the one hand, traditional single molecular marker detection (such as protein markers such as AFP and CEA) is easily interfered with by factors such as inflammation and benign lesions, and the signal specificity and sensitivity are insufficient, making it difficult to accurately capture weak molecular changes related to early tumors; on the other hand, existing sequencing-related technologies mostly focus on single-dimensional molecular feature analysis.

[0005] Existing relevant patent literature, publication number CN117941002A, discloses a method for detecting chromosomal and subchromosomal copy number variations (CNVs). It achieves the identification of specific molecular features through coverage calculation, normalization processing, and copy number classification, but does not involve the extraction of tumor-related molecular signals from other dimensions.

[0006] Existing patent literature, publication number CN113539355A, discloses a tissue-specific origin prediction and related risk assessment system based on cfDNA whole-genome sequencing. This system achieves analysis through nucleosome occlusion signal comparison and comparison with a cellular variation database. However, it also fails to extend the application to the combined use of multi-dimensional molecular features. Because the aforementioned technology only covers some molecular features, it is difficult to address the signal differences caused by tumor heterogeneity, resulting in insufficient ability to capture early, subtle molecular changes and failing to meet the requirements for high-accuracy screening.

[0007] (3) cfDNA has extremely low abundance, short fragments, and is easily degraded in body fluid supernatants, which imposes stringent requirements on the sample compatibility of detection technologies. Existing detection methods generally require a high amount of cfDNA extracted. When the cfDNA content in the sample is below a certain threshold (usually above 20 ng), the detection is prone to failure due to insufficient template. The cfDNA sample quality of special groups such as early-stage tumor-related individuals and the elderly is even worse and the content is even lower, further exacerbating the difficulty of detection. In addition, the problem of blood cell rupture contamination that may occur during sample transportation and storage will also affect the stability of the detection results, further limiting the clinical applicability of the technology;

[0008] (4) Most existing technologies focus only on a single screening scenario and lack the ability to support the entire lifecycle assessment process, such as dynamic monitoring of tumor-related molecular characteristics and tracking of risk changes. Although some technologies can achieve one-time detection of specific molecular characteristics, they cannot dynamically capture the changing trends of molecular characteristics and cannot provide continuous risk assessment references. At the same time, the detection scope is mostly limited to molecular characteristics related to specific cancer types and lacks comprehensive coverage of molecular signals related to multiple common solid tumors, which cannot meet the actual needs of joint screening of multiple cancer types.

[0009] In view of this, the present invention is hereby proposed. Summary of the Invention

[0010] To address the aforementioned technical problems in the prior art, this invention provides a multi-dimensional tumor marker combined detection device and its usage method, solving the problems of low sensitivity of single markers, complex operation of multi-marker detection, insufficient multi-dimensional feature extraction under low-depth sequencing, and reliance on special experimental processing in the prior art.

[0011] To achieve the above objectives, the technical solution of the present invention is as follows:

[0012] Firstly, a multi-dimensional tumor marker combined detection device includes:

[0013] The sample processing module is used to process the subject's body fluid samples, extract extracellular cell-free DNA, and construct sequencing libraries; the body fluid samples include: bile, pleural effusion, peritoneal fluid, blood, or cerebrospinal fluid;

[0014] The sequencing module is used to perform whole-genome ultra-low-depth sequencing on the sequencing library to obtain sequencing data;

[0015] The biomarker extraction module is used to simultaneously extract multidimensional tumor biomarkers from the sequencing data. The multidimensional tumor biomarkers include: chromosome copy variation, microsatellite instability, telomere length, nucleosome imprinting, fragment distribution, and methylation level.

[0016] The feature fusion and detection module is used to process the extracted multi-dimensional tumor markers and output tumor risk assessment results through a preset ensemble learning model.

[0017] The results output module is used to output the tumor risk assessment results and related biomarker characteristic information.

[0018] Furthermore, the marker extraction module includes:

[0019] The chromosome copy number variation detection unit is used to receive genome alignment fragments from sequencing data, and output copy number abnormality-related features through sliding window segmentation, coverage correction and abnormal region identification.

[0020] The microsatellite instability detection unit is used to call up a preset core microsatellite locus database and output MSI-related judgment indicators through sequence alignment, missing fragment statistics and parameter calculation.

[0021] The telomere length estimation unit is used to screen telomere repeat sequence fragments in sequencing data and outputs a relative telomere length value through coverage calculation and fragment distribution correction.

[0022] The nucleosome imprinting analysis unit is used to extract the terminal motif sequence of cfDNA fragments and construct nucleosome imprinting feature vectors through nucleosome protected region statistics and motif feature analysis.

[0023] The fragment distribution feature extraction unit is used to perform length statistics on effective cfDNA fragments in sequencing data, and output a fragment distribution feature set through interval division and key parameter calculation.

[0024] The methylation prediction unit is used to screen cfDNA fragments containing CpG sites. Through feature vector construction and hidden Markov model operation, it outputs the methylation probability of CpG sites and the average methylation level of the region.

[0025] Furthermore, the workflow of the chromosome copy variation detection unit is as follows:

[0026] The hg38 reference genome was uniformly divided using a 100kb sliding window, and the sequencing coverage of each window was calculated. The specific formula is as follows:

[0027]

[0028] in, The number of valid reads within the window. For the length of the read, For window size;

[0029] The sequencing coverage was corrected for GC content using the LOESS local regression model, and then normalized by combining the corrected coverage with the preset baseline data of healthy populations.

[0030] An improved cyclic binary segmentation algorithm was adopted, with a segmentation threshold of p < 0.001 and a minimum abnormal fragment length of 100kb. Chromosomal copy number abnormal regions were identified based on normalized coverage data.

[0031] Output the number of abnormal segments in the chromosome copy number abnormal region, the average copy number difference in the chromosome copy number abnormal region, and the length of the chromosome copy number abnormal region.

[0032] Furthermore, the workflow of the microsatellite instability detection unit is as follows:

[0033] The preset core microsatellite locus database is based on the dbSNP database. The selection criteria for the core microsatellite loci are: length ≥10bp, repeat unit is single / dinucleotide, and population heterozygosity ≥0.3. The database contains 1200 core loci, covering the entire genome and 22 autosomes.

[0034] The Smith-Waterman alignment algorithm was used to perform sequence alignment between sequencing data and core loci.

[0035] Calculate the strength of long missing segments by counting the number of long missing segments. The specific formula is as follows:

[0036]

[0037] in, This represents the number of missing segments. This represents the total number of valid segments;

[0038] Calculate the coefficient of variation of missing length The specific formula is as follows:

[0039]

[0040] in, The standard deviation of the missing length. This represents the mean length of the missing data.

[0041] Output two types of indicators: long deletion intensity and coefficient of variation.

[0042] Furthermore, the workflow of the methylation prediction unit is as follows:

[0043] Screening for CpG islands and CpG island shore regions in the genome, and extracting cfDNA fragments containing ≥1 CpG site;

[0044] Constructing a three-feature input vector ,in, The standardized segment length, The number of effective coverage reads for CpG sites. This represents the absolute value of the distance from the CpG site to the center of the fragment.

[0045] The feature vector is input into a non-homogeneous hidden Markov model, which has two states, corresponding to methylated and unmethylated states respectively. Algorithm iterative training; the emission probability calculation formula for the non-homogeneous hidden Markov model is:

[0046]

[0047] in, Input the feature value at time t. For state values, The characteristic mean of the state values, The characteristic standard deviation;

[0048] Output the methylation probability of each CpG site, filter the results of sites with a confidence level ≥ 0.8, calculate and output the average methylation level of the CpG island region.

[0049] Furthermore, the workflow of the telomere length estimation unit is as follows:

[0050] pass The algorithm filters telomere repeat sequences in sequencing data, counts the number of valid reads aligned to the region, and calculates the effective coverage of the telomere region. ;

[0051] statistics The proportion of short fragments among all valid cfDNA fragments ;

[0052] The relative telomere length is calculated using the following formula:

[0053]

[0054] in, This is the ratio of the telomere length of the sample to the baseline telomere length of healthy individuals;

[0055] Output the relative telomere length value.

[0056] Furthermore, the workflow of the nucleosome imprinting analysis unit is as follows:

[0057] Extraction of cfDNA fragments End and The end 6-mermotif sequence, through The tool performs motif recognition and filters out motifs with an E value < 1e-5;

[0058] The percentage of segments protected by nucleosomes is statistically significant.

[0059] calculate End and The enrichment degree of end motifs is calculated using the motif information entropy formula, which is as follows:

[0060]

[0061] in, Let be the probability of the i-th motif appearing. This represents the total number of motif types.

[0062] Will Segment proportion, End-motif enrichment The end motif enrichment and motif information entropy are integrated into a kernel body imprint feature vector and output.

[0063] Furthermore, the workflow of the segment distribution feature extraction unit is as follows:

[0064] The short fragment interval was defined as 100bp-166bp and the long fragment interval as 169bp-240bp. The number of valid cfDNA fragments within the two intervals was counted.

[0065] The short / long segment ratio is calculated using the following formula:

[0066]

[0067] in, This represents the total number of valid cfDNA fragments within the short fragment interval. This represents the total number of valid cfDNA fragments within the long fragment interval.

[0068] Calculate the peak fragment length and fragment length standard deviation, and then calculate the fragment length distribution entropy. The specific formula is as follows:

[0069]

[0070] in, The percentage of a single cfDNA fragment. For length is The proportion of cfDNA fragments This is the maximum threshold for segment length statistics. This is the minimum threshold for segment length statistics;

[0071] Output four types of features: short / long segment ratio, peak length, standard deviation, and distribution entropy.

[0072] Furthermore, the sample processing module includes:

[0073] The body fluid supernatant separation unit is used to perform centrifugation on body fluid samples to separate and obtain body fluid supernatant;

[0074] The cfDNA extraction unit is used to extract cfDNA by performing lysis, binding, washing, and elution operations on the body fluid supernatant using an adsorption column method.

[0075] The library construction unit is used to perform end repair, adapter ligation, and PCR amplification on the extracted cfDNA to construct a sequencing library. The process of constructing the sequencing library does not include bisulfite treatment or targeted capture operations.

[0076] Furthermore, the sequencing module includes:

[0077] The sequencing control unit is used to set the operating parameters of the sequencing platform, control the sequencing depth, sequencing read length and sequencing type, and start sequencing after performing homogenization on the sequencing library.

[0078] The data quality control unit is used to perform adapter removal, deduplication, and low-quality read removal on the raw sequencing data produced by the sequencing platform, while retaining and outputting the valid reads.

[0079] Furthermore, the feature fusion and detection module includes:

[0080] The feature standardization unit is used to receive the raw feature data output by each detection unit, execute the preset standardization algorithm, and output standardized feature values.

[0081] The feature matrix construction unit is used to classify and integrate standardized feature values ​​according to preset dimensions to construct a feature matrix.

[0082] The model computation unit is used to load a pre-trained ensemble learning model, input the feature matrix into the model to perform calculations, and output a tumor risk score.

[0083] Furthermore, the result output module includes:

[0084] The results integration unit is used to receive tumor risk scores and feature output results from each detection unit, and perform structured integration according to preset logic.

[0085] The visualization unit is used to present the integrated results in chart form, including charts related to risk scores, charts related to marker characteristics, and abnormal indicator labels.

[0086] The export unit supports exporting the presented results in a preset standardized format, adapting to clinical archives and hospital information system integration.

[0087] Secondly, a method for using a multidimensional tumor marker combined detection device, applied to the aforementioned multidimensional tumor marker combined detection device, includes:

[0088] S1. Sample processing: Process the body fluid samples of the subjects, extract extracellular cell-free DNA and construct sequencing libraries; the body fluid samples include: bile, pleural effusion, peritoneal fluid, blood or cerebrospinal fluid.

[0089] S2. Whole-genome low-depth sequencing: Perform whole-genome ultra-low-depth sequencing on the sequencing library, perform quality control processing on the generated raw sequencing data, and obtain effective sequencing data;

[0090] S3. Simultaneous extraction of multiple biomarkers: Based on the effective sequencing data, six types of tumor biomarkers, namely chromosome copy variation, microsatellite instability, telomere length, nucleosome imprinting, fragment distribution and methylation level, are extracted in parallel through six independent feature extraction processes.

[0091] S4. Feature Fusion and Risk Assessment: Standardize the extracted features of the six types of tumor markers, construct a feature matrix, input the feature matrix into a preset ensemble learning model to perform calculations, and output a tumor risk score.

[0092] S5. Results Output: Based on the tumor risk score, the risk level is determined, abnormal biomarker information and related evidence are integrated, a report is generated and can be exported.

[0093] Compared with existing technologies, the present invention provides a multi-dimensional tumor marker joint detection device and its usage method. The device includes modules for sample processing, sequencing, marker extraction, feature fusion and detection, and result output. The marker extraction module contains six detection units, corresponding to the extraction of six types of markers, including chromosomal copy variation and microsatellite instability. The method involves collecting body fluid samples to extract cfDNA, directly constructing a library without special processing, obtaining effective data through whole-genome sequencing, simultaneously extracting six types of markers through various modules of the device, optimizing algorithms to adapt to low-depth data, inputting them into an ensemble learning model to output risk scores, and finally outputting relevant information through the result output module. This invention requires no special experimental processing, achieving low-cost, high-accuracy joint detection, and is suitable for large-scale clinical screening needs. Attached Figure Description

[0094] Figure 1 This is an architectural diagram of the multi-dimensional tumor marker combined detection device provided in an embodiment of the present invention;

[0095] Figure 2 A flowchart illustrating the method of using the multidimensional tumor marker combination device provided in this embodiment of the invention. Detailed Implementation

[0096] The technical solution of the present invention will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are not all embodiments of the present invention. All other embodiments obtained by those skilled in the art without creative effort are within the protection scope of the present invention.

[0097] It should be noted that, unless otherwise specifically stated, the relative arrangement and numerical expressions of the components and steps described in these embodiments should not be construed as limiting the scope of the invention.

[0098] The following description of exemplary embodiments is merely illustrative and is not intended to limit the invention or its application or use in any way. Techniques, methods, and apparatus known to those skilled in the art may not be discussed in detail herein, but where applicable, such techniques, methods, and apparatus should be considered part of this specification.

[0099] Example 1

[0100] See Figure 1 , Figure 1 This is a schematic diagram of the architecture of a multi-dimensional tumor marker combined detection device proposed in this invention, specifically including:

[0101] M1, Sample Processing Module, is used to process the subject's body fluid samples, extract extracellular cell-free DNA, and construct sequencing libraries; body fluid samples include: bile, pleural effusion, peritoneal fluid, blood, or cerebrospinal fluid; specifically including:

[0102] M11, Supernatant Separation Unit, is used to perform centrifugation on body fluid samples to separate and obtain body fluid supernatant;

[0103] M12, the cfDNA extraction unit, is used to perform lysis, binding, washing, and elution operations on the body fluid supernatant using the adsorption column method to extract cfDNA;

[0104] Specifically, the extraction unit employs an adsorption column method, extracting cfDNA through four steps: lysis, magnetic bead binding, washing, and elution. The lysis step is performed at 56°C for ten minutes, the washing step is repeated twice using the kit's washing buffer, and the elution step uses 50 μL of elution buffer and is incubated at room temperature for two minutes. The extracted cfDNA, as determined by Nanodrop, has an A260 / A280 ratio between 1.08 and 2.0, and the extraction yield is at least 5 ng, as confirmed by Qubit quantitative PCR.

[0105] M13, the library construction unit, is used to perform end repair, adapter ligation, and PCR amplification on the extracted cfDNA to construct a sequencing library. The process of constructing the sequencing library does not include bisulfite treatment or targeted capture operations.

[0106] Specifically, adopt 5 ng of extracted cfDNA underwent end repair, adapter ligation, and polymerase chain reaction (PCR) amplification sequentially. End repair was performed at 20°C for 30 minutes, adapter ligation at 20°C for 15 minutes, and PCR amplification was performed for 3 to 5 cycles, each cycle consisting of 98°C denaturation for 10 seconds, 65°C annealing for 30 seconds, and 72°C extension for 30 seconds. The library construction process did not include bisulfite treatment or targeted capture. The amplified products were verified by 1.5% agarose gel electrophoresis, with fragment sizes concentrated between 200-400 bp. The library concentration was not less than 2 nM as determined by Qubit quantitative PCR.

[0107] M2, the sequencing module, is used to perform whole-genome ultra-low depth sequencing on sequencing libraries to obtain sequencing data; specifically, it includes:

[0108] M21, the sequencing control unit, is used to set the operating parameters of the sequencing platform, control the sequencing depth, sequencing read length and sequencing type, and start sequencing after performing homogenization on the sequencing library.

[0109] Control sequencing platform is The sequencing depth was set to 0.1X-1X (preferably 0.5X, corresponding to approximately 1.5G of data per sample), and the read length was 100bp-150bp (using paired-end sequencing mode). Before sequencing, the sequencing library was homogenized (concentration adjusted to 2nM) to ensure genome coverage uniformity ≥90%.

[0110] M22, the data quality control unit, is used to perform adapter removal, deduplication, and low-quality read removal on the raw sequencing data produced by the sequencing platform, while retaining and outputting valid reads.

[0111] The raw sequencing data were processed by adapter removal, duplicate removal (using Picard software, MarkDuplicates parameter), and low-quality read removal (reads with a Phred quality value Q < 30 and a percentage > 5% were removed) to ultimately retain the effective read rate. reads comparison rate Valid sequencing data.

[0112] M3, the biomarker extraction module, is used to simultaneously extract multidimensional tumor biomarkers from sequencing data. These multidimensional tumor biomarkers include at least chromosomal copy variation, microsatellite instability, telomere length, nucleosome imprinting, fragment distribution, and methylation levels; specifically including:

[0113] M31, the chromosome copy number variation detection unit, is used to receive genome alignment fragments from sequencing data, and output copy number variation-related features through sliding window segmentation, coverage correction, and abnormal region identification. The workflow of the chromosome copy number variation detection unit is as follows:

[0114] M311. The hg38 reference genome was uniformly divided using a 100kb sliding window, and the sequencing coverage of each window was calculated. The specific formula is as follows:

[0115]

[0116] in, The number of valid reads within the window. For the length of the read, For window size;

[0117] M312. The LOESS local regression model is used to perform GC content correction on the sequencing coverage, and then the corrected coverage is normalized by combining the preset baseline data of healthy people.

[0118] M313. An improved cyclic binary segmentation algorithm is adopted, with a segmentation threshold p < 0.001 and a minimum abnormal fragment length of 100kb. Chromosome copy number abnormal regions are identified based on normalized coverage data.

[0119] M314 outputs the number of abnormal segments in the chromosome copy number abnormality region, the average copy number difference of the chromosome copy number abnormality region, and the length of the chromosome copy number abnormality region.

[0120] M32, the microsatellite instability detection unit, is used to call upon a preset core microsatellite locus database and, through sequence alignment, missing fragment statistics, and parameter calculation, output MSI-related judgment indicators. The workflow of the microsatellite instability detection unit is as follows:

[0121] M321, the preset core microsatellite locus database is derived from the dbSNP database. The selection criteria for core microsatellite loci are length ≥10bp, repeat units are single / dinucleotides, and population heterozygosity ≥0.3. The database contains 1200 core loci, covering the entire genome and 22 autosomes.

[0122] M322: The Smith-Waterman alignment algorithm was used to perform sequence alignment between sequencing data and core loci.

[0123] Calculate the strength of long missing segments by counting the number of long missing segments. The specific formula is as follows:

[0124]

[0125] in, This represents the number of missing segments. This represents the total number of valid segments;

[0126] M323. Calculate the coefficient of variation for the missing length. The specific formula is as follows:

[0127]

[0128] in, The standard deviation of the missing length. This represents the mean length of the missing data.

[0129] M324, output long deletion intensity and coefficient of variation are two types of indicators.

[0130] M33, the telomere length estimation unit, is used to screen telomere repetitive sequence fragments in sequencing data. Through coverage calculation and fragment distribution correction, it outputs a relative telomere length value. The workflow of the telomere length estimation unit is as follows:

[0131] M331, via The algorithm filters telomere repeat sequences in sequencing data, counts the number of valid reads aligned to the region, and calculates the effective coverage of the telomere region. ;

[0132] M332, Statistics The proportion of short fragments among all valid cfDNA fragments The relative telomere length is calculated using the following formula:

[0133]

[0134] in, This is the ratio of the telomere length of the sample to the baseline telomere length of healthy individuals;

[0135] M333, outputs the relative telomere length value.

[0136] M34, the nucleosome imprinting analysis unit, is used to extract the terminal motif sequences of cfDNA fragments. Through nucleosome protected region statistics and motif feature analysis, a nucleosome imprinting feature vector is constructed. The workflow of the nucleosome imprinting analysis unit is as follows:

[0137] M341, Extraction of cfDNA fragments End and The end 6-mermotif sequence, through The tool performs motif recognition and filters out motifs with an E value < 1e-5;

[0138] M342, the percentage of segments in the nucleosome protected region;

[0139] calculate End and The enrichment degree of end motifs is calculated using the motif information entropy formula, which is as follows:

[0140]

[0141] in, Let be the probability of the i-th motif appearing. This represents the total number of motif types.

[0142] M343, will Segment proportion, End-motif enrichment The end motif enrichment and motif information entropy are integrated into a kernel body imprint feature vector and output.

[0143] M35, the fragment distribution feature extraction unit, is used to perform length statistics on valid cfDNA fragments in sequencing data. Through interval division and key parameter calculation, it outputs a fragment distribution feature set. The workflow of the fragment distribution feature extraction unit is as follows:

[0144] M351. Define the short fragment interval as 100bp-166bp and the long fragment interval as 169bp-240bp, and count the number of effective cfDNA fragments within each interval; calculate the short / long fragment ratio using the following formula:

[0145]

[0146] in, This represents the total number of valid cfDNA fragments within the short fragment interval. This represents the total number of valid cfDNA fragments within the long fragment interval.

[0147] M352. Statistically calculate the peak segment length and segment length standard deviation, and then calculate the segment length distribution entropy. The specific formula is as follows:

[0148]

[0149] in, The percentage of a single cfDNA fragment. For length is The proportion of cfDNA fragments This is the maximum threshold for segment length statistics. This is the minimum threshold for segment length statistics; , .

[0150] M353 outputs four types of features: short / long segment ratio, peak length, standard deviation, and distribution entropy.

[0151] M36, the methylation prediction unit, is used to screen cfDNA fragments containing CpG sites. Through feature vector construction and hidden Markov model calculations, it outputs the methylation probability of CpG sites and the average methylation level of the region. The workflow of the methylation prediction unit is as follows:

[0152] M361. Screen for CpG islands and CpG island shore regions in the genome and extract cfDNA fragments containing ≥1 CpG site;

[0153] Constructing a three-feature input vector ,in, The standardized segment length, The number of effective coverage reads for CpG sites. This represents the absolute value of the distance from the CpG site to the center of the fragment.

[0154] M362. Input the feature vector into a non-homogeneous hidden Markov model. The non-homogeneous hidden Markov model has two states, corresponding to the methylated state and the unmethylated state, respectively. Algorithm iterative training; the emission probability calculation formula for a non-homogeneous hidden Markov model is:

[0155]

[0156] in, Input the feature value at time t. For state values, The characteristic mean of the state values, The characteristic standard deviation;

[0157] M363 outputs the methylation probability of each CpG site, filters site results with confidence ≥ 0.8, calculates and outputs the average methylation level of the CpG island region.

[0158] M4, the feature fusion and detection module, is used to process the extracted multi-dimensional tumor markers and output tumor risk assessment results through a pre-set ensemble learning model; specifically including:

[0159] M41, Feature Standardization Unit, is used to receive the raw feature data output by each detection unit, execute a preset standardization algorithm, and output standardized feature values;

[0160] M42, Feature Matrix Construction Unit, is used to classify and integrate standardized feature values ​​according to preset dimensions to construct a feature matrix;

[0161] M43, the model computation unit, is used to load a pre-trained ensemble learning model, input the feature matrix into the model to perform calculations, and output a tumor risk score.

[0162] M5, the results output module, is used to output tumor risk assessment results and related biomarker characteristics;

[0163] M51, Result Integration Unit, is used to receive tumor risk scores and feature output results from each detection unit, and perform structured integration according to preset logic;

[0164] M52, the visualization unit, is used to present the integrated results in the form of charts, including charts related to risk scores, charts related to marker characteristics, and abnormal indicator labels;

[0165] M53, the export unit, is used to support the export of the presented results in a preset standardized format, which is compatible with clinical archives and hospital information system integration.

[0166] Example 2

[0167] See Figure 2 , Figure 2 This is a flowchart illustrating the method of using a multi-dimensional tumor marker combined detection device proposed in this invention, specifically including:

[0168] S1. Sample processing: Process the subject's body fluid samples, extract extracellular cell-free DNA and construct sequencing libraries; body fluid samples include: bile, pleural effusion, peritoneal fluid, blood or cerebrospinal fluid, etc.

[0169] Specifically, the collected bodily fluid sample is blood, specifically peripheral blood from the subject. The specific steps for processing the peripheral blood sample include:

[0170] S11. Collect 5 mL of peripheral blood from the subject, place it in an EDTA anticoagulant tube, and send it to the laboratory within 2 hours. The transport temperature should be controlled at 2-8℃.

[0171] S12. Place the peripheral blood sample into the plasma separation unit and centrifuge at 4℃, 3000r / min, and 10min to separate the plasma components.

[0172] S13. Plasma was subjected to lysis, binding, washing, and elution using an adsorption column method. The A260 / A280 ratio of the extracted cfDNA was 1.8-2.0, and the extraction amount was ≥5ng.

[0173] S14, Adopt End repair (20℃, 30 min), adapter ligation (20℃, 15 min), and 3-5 cycles of PCR amplification were performed on cfDNA. The constructed sequencing library had a concentration ≥2 nM and fragment sizes concentrated in the range of 200-400 bp. No bisulfite treatment or targeted capture was performed during the construction process.

[0174] S2. Whole-genome low-depth sequencing: Perform ultra-low-depth whole-genome sequencing on the sequencing library, and perform quality control processing on the raw sequencing data to obtain effective sequencing data; specific steps include:

[0175] S21. After homogenizing the sequencing library (adjusting the concentration to 2nM), load it into... Sequencing platform;

[0176] S22, sequencing depth 0.1X-1X (preferably 0.5X, corresponding to 1.5G of data), sequencing read length 100bp-150bp, sequencing type is paired-end sequencing;

[0177] S23. Start the sequencing program, which will run for approximately 13 hours and produce raw sequencing data;

[0178] S24. The data quality control unit performs adapter removal, duplicate removal, and low-quality read removal. The quality control standards are: effective read rate ≥90%, read alignment rate ≥85%, and coverage uniformity ≥90%, to obtain effective sequencing data.

[0179] S3. Simultaneous extraction of multiple biomarkers: Based on effective sequencing data, six types of tumor biomarkers, namely chromosome copy variation, microsatellite instability, telomere length, nucleosome imprinting, fragment distribution and methylation level, are extracted in parallel through six independent feature extraction processes.

[0180] Based on effective sequencing data, six types of tumor markers were extracted in parallel through a feature extraction process, specifically including:

[0181] S31, CNV extraction process: Reference genome sliding window segmentation → Sequencing coverage calculation → GC content correction → Baseline normalization of healthy individuals → Abnormal region identification using improved CBS algorithm → Core feature output;

[0182] S32, MSI extraction process: Core microsatellite locus alignment → Smith-Waterman algorithm sequence alignment → Long missing fragment statistics → Long missing intensity and coefficient of variation calculation → Relevant index output;

[0183] S33. Telomere length extraction process: Telomere repeat sequence identification → Calculation of effective telomere region coverage → Statistics of short fragment proportion → Correction operation → Output of relative telomere length;

[0184] S34. Nucleosome imprinting extraction process: terminal motif sequence extraction → motif identification and screening → nucleosome protected region fragment statistics → motif enrichment and information entropy calculation → feature vector construction;

[0185] S35. Fragment distribution extraction process: cfDNA fragment length statistics → short / long fragment interval division → key parameter calculation → four-dimensional feature output;

[0186] S36. Methylation extraction process: screening of cfDNA fragments containing CpG sites → construction of three feature vectors → training of non-homogeneous hidden Markov model → calculation of methylation probability → output of regional average methylation level.

[0187] S4. Feature Fusion and Risk Assessment: The extracted six types of tumor marker features are standardized to construct a feature matrix. This feature matrix is ​​then input into a pre-defined ensemble learning model to perform calculations and output a tumor risk score. Specifically, this includes:

[0188] S41. The Z-score algorithm is used to process the original feature values ​​of the six types of markers to eliminate dimensional differences;

[0189] S42. Integrate standardized feature values ​​according to preset dimensions to construct a 22-dimensional feature matrix (the number of features contributed by each marker is consistent with the device part).

[0190] S43. Input the feature matrix into the ensemble learning model (random forest + gradient booster + attention mechanism), and output a tumor risk score of 0-100 through weight allocation and feature fusion operation. Set the score ≥60 as a positive risk.

[0191] The training process for the S44 non-homogeneous hidden Markov model is as follows: using cfDNA sequencing data from healthy individuals and cancer patients as the training set, setting the initial emission probability to 0.3, iterating 100 times using the Baum-Welch algorithm, and setting the convergence threshold to... During the iteration process, the model parameters are updated and optimized using state transition probabilities. The specific formula is as follows:

[0192]

[0193] For the first The state of the next iteration to state The transition probability, For the first At time T in the next iteration, T is in state T. And at any time In state The probability, This represents the total number of samples in the training set.

[0194] S5. Results Output: Based on the tumor risk score, the risk level is determined, abnormal biomarker information and related evidence are integrated, a report is generated, and export is supported; specifically including:

[0195] S51. Integrate tumor risk scores and the characteristics of each biomarker according to the logic of "risk level - abnormal biomarker - core basis".

[0196] S52. Generate a visualized clinical report, the core content of which includes the risk level assessment results, a list of abnormal biomarkers, and relevant testing evidence. The list of abnormal biomarkers includes biomarkers and related parameters that are outside the normal range.

[0197] The visualized clinical report supports exporting reports in PDF, Excel, or DICOM formats and can be pushed to the corresponding department workstations through the hospital information system interface, adapting to clinical archiving needs.

[0198] Example 3

[0199] This embodiment uses tumor screening in healthy individuals as a specific application scenario, employing a multi-dimensional tumor marker combined detection system provided by this invention to intuitively demonstrate the actual implementation process of this detection device; specifically including:

[0200] B1. Application Scenario Setting

[0201] A tertiary hospital's health checkup center conducted an early cancer screening program, enrolling 300 participants (162 males and 138 females, aged 30-70 years) with no history of malignant tumors or a clear family history of cancer. The present invention was used for multidimensional tumor marker combined detection.

[0202] B2. Sample Collection and Preprocessing: 5 mL of peripheral blood was collected from each subject and placed in an EDTA anticoagulant tube. The sample was transported to the laboratory within 1.5 hours (transport temperature 2-8℃). The peripheral blood sample was placed in an automated liquid processing workstation and centrifuged at 4℃, 3000 rpm for 10 minutes to separate the plasma. 3 mL of plasma was used for cfDNA extraction. Extraction was performed with an elution volume of 50 μL. Qubit analysis confirmed that the cfDNA extraction amount of all samples was ≥5 ng and the A260 / A280 ratio was between 1.8 and 2.0.

[0203] B3. Library Construction: 5 ng cfDNA was used for library construction. End repair and adapter ligation were performed, followed by four cycles of PCR amplification (98℃ denaturation for 10 s, 65℃ annealing for 30 s, and 72℃ extension for 30 s). The amplified products were verified by 1.5% agarose gel electrophoresis, and the fragment sizes were concentrated in the range of 200-400 bp. The library concentration was ≥2 nM as detected by Qubit.

[0204] B4. Whole-genome low-depth sequencing: Libraries from 300 samples were homogenized at a concentration of 2 nM and loaded into the S1 flow cell of the Illumina NovaSeq 6000 sequencing platform. Sequencing was initiated with a sequencing depth of 0.5X and a read length of 150 bp in paired-end mode. After sequencing, the data was processed by the quality control unit. All samples met the requirements for valid sequencing data: effective read rate ≥ 92%, read alignment rate ≥ 88%, and coverage uniformity ≥ 91%.

[0205] B5. Simultaneous Extraction of Multiple Biomarkers: Effective sequencing data is simultaneously distributed to six detection units via a data bus, and feature extraction is performed in parallel: The CNV detection unit identifies 7 samples with abnormal copy numbers exceeding 100kb, and outputs the number of abnormal fragments, average ΔCN, and maximum fragment length; the MSI detection unit compares with the core microsatellite locus database to count the number of long missing fragments and the coefficient of variation; the methylation prediction unit screens out CpG sites with confidence ≥0.8 and calculates the average methylation level of the region; the remaining detection units output corresponding feature data according to the procedure in Example 1.

[0206] B6. Feature Fusion and Risk Assessment: The raw features output by each detection unit were standardized by Z-score to construct a 22-dimensional feature matrix, which was then input into the ensemble learning model for computation. Weights were assigned through an attention mechanism (methylation 0.25, CNV 0.22, etc.), and tumor risk scores were calculated according to the formula. Ultimately, 8 subjects scored ≥60 points (positive risk), while the remaining 292 subjects scored <60 points (negative).

[0207] B7. Results Output: Integrate the risk scores and abnormal biomarker information of all samples to generate a visual report. The report includes a risk score bar chart, a biomarker characteristic heatmap, and abnormal indicator warning labels, clearly marking the abnormal biomarkers of positive samples (such as elevated methylation levels or increased CNV abnormal fragment numbers) and related detection evidence. All reports are exported in PDF format and pushed to the health checkup center workstation via the HIS system interface for clinicians' reference.

[0208] In summary, the present invention has the following advantages:

[0209] 1. Significantly reduce testing costs, greatly lowering the economic threshold for clinical application and large-scale promotion, making it accessible to more people;

[0210] 2. Significantly improves the sensitivity and specificity of early tumor detection, effectively covering a variety of common solid tumors and reducing missed diagnoses and false negatives in early tumor detection;

[0211] 3. The operation is simple and non-invasive, requiring only a body fluid sample to complete the test, which improves the efficiency of large-scale population screening while reducing the trauma and difficulty of cooperation for patients.

[0212] 4. Strong sample compatibility, adaptable to scenarios with trace amounts of cfDNA samples, reducing test failures due to insufficient sample quantity or poor quality, and improving the success rate of clinical testing.

[0213] 5. It has comprehensive detection functions, which can support the full-cycle diagnosis and treatment management needs of early tumor screening, efficacy evaluation and recurrence warning, and is suitable for various solid tumor diagnosis and treatment scenarios;

[0214] 6. The test results are highly interpretable, clearly identifying the core influencing factors and their logic of action, helping clinicians to quickly understand the basis of the results and increasing their trust in and acceptance of the test results.

[0215] 7. The test report supports standardized format export, is compatible with commonly used clinical information systems, facilitates clinical archiving, data sharing and connection with the diagnosis and treatment process, and provides convenient support for diagnosis and treatment decisions.

[0216] The above specific embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to examples, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A multidimensional tumor marker combined detection device, characterized in that, include: The sample processing module is used to process the subject's body fluid samples, extract extracellular free DNA, and construct sequencing libraries; The body fluid samples include: bile, pleural effusion, peritoneal fluid, blood, or cerebrospinal fluid; The sequencing module is used to perform whole-genome ultra-low-depth sequencing on the sequencing library to obtain sequencing data; A biomarker extraction module is used to simultaneously extract multidimensional tumor biomarkers from the sequencing data. These multidimensional tumor biomarkers include: chromosome copy variation, microsatellite instability, telomere length, nucleosome imprinting, fragment distribution, and methylation level. The biomarker extraction module includes: The chromosome copy number variation detection unit is used to receive genome alignment fragments from sequencing data, and output copy number abnormality-related features through sliding window segmentation, coverage correction and abnormal region identification. The microsatellite instability detection unit is used to call up a preset core microsatellite locus database and output MSI-related judgment indicators through sequence alignment, missing fragment statistics and parameter calculation. The telomere length estimation unit is used to screen telomere repeat sequence fragments in sequencing data and outputs a relative telomere length value through coverage calculation and fragment distribution correction. The nucleosome imprinting analysis unit is used to extract the terminal motif sequence of cfDNA fragments and construct nucleosome imprinting feature vectors through nucleosome protected region statistics and motif feature analysis. The fragment distribution feature extraction unit is used to perform length statistics on effective cfDNA fragments in sequencing data, and output a fragment distribution feature set through interval division and key parameter calculation. The methylation prediction unit is used to screen cfDNA fragments containing CpG sites. Through feature vector construction and hidden Markov model operation, it outputs the methylation probability of CpG sites and the average methylation level of the region. The feature fusion and detection module is used to process the extracted multi-dimensional tumor markers and output tumor risk assessment results through a preset ensemble learning model. The results output module is used to output the tumor risk assessment results and related biomarker characteristics.

2. The multidimensional tumor marker combined detection device according to claim 1, characterized in that, The workflow of the chromosome copy mutation detection unit is as follows: The hg38 reference genome was uniformly divided using a 100kb sliding window, and the sequencing coverage of each window was calculated. The specific formula is as follows: in, The number of valid reads within the window. For the length of the read, For window size; The sequencing coverage was corrected for GC content using the LOESS local regression model, and then normalized by combining the corrected coverage with the preset baseline data of healthy populations. An improved cyclic binary segmentation algorithm was adopted, with a segmentation threshold of p < 0.001 and a minimum abnormal fragment length of 100kb. Chromosomal copy number abnormal regions were identified based on normalized coverage data. Output the number of abnormal segments in the chromosome copy number abnormal region, the average copy number difference in the chromosome copy number abnormal region, and the length of the chromosome copy number abnormal region.

3. The multidimensional tumor marker combined detection device according to claim 1, characterized in that, The workflow of the microsatellite instability detection unit is as follows: The preset core microsatellite locus database is based on the dbSNP database. The selection criteria for the core microsatellite loci are: length ≥10bp, repeat unit is single / dinucleotide, and population heterozygosity ≥0.

3. The database contains 1200 core loci, covering the entire genome and 22 autosomes. The Smith-Waterman alignment algorithm was used to perform sequence alignment between sequencing data and core loci. Calculate the strength of long missing segments by counting the number of long missing segments. The specific formula is as follows: in, This represents the number of missing segments. This represents the total number of valid segments; Calculate the coefficient of variation of missing length The specific formula is as follows: in, The standard deviation of the missing length. This represents the mean length of the missing data. Output two types of indicators: long deletion intensity and coefficient of variation.

4. The multidimensional tumor marker combined detection device according to claim 1, characterized in that, The workflow of the methylation prediction unit is as follows: Screening for CpG islands and CpG island shore regions in the genome, and extracting cfDNA fragments containing ≥1 CpG site; Constructing a three-feature input vector ,in, The standardized segment length, The number of effective coverage reads for CpG sites. This represents the absolute value of the distance from the CpG site to the center of the fragment. The feature vector is input into a non-homogeneous hidden Markov model, which has two states, corresponding to methylated and unmethylated states respectively. Algorithm iterative training; the emission probability calculation formula for the non-homogeneous hidden Markov model is: in, Input the feature value at time t. For state values, The characteristic mean of the state values, The characteristic standard deviation; Output the methylation probability of each CpG site, filter the results of sites with a confidence level ≥ 0.8, calculate and output the average methylation level of the CpG island region.

5. The multidimensional tumor marker combined detection device according to claim 1, characterized in that, The workflow of the telomere length estimation unit is as follows: pass The algorithm filters telomere repeat sequences in sequencing data, counts the number of valid reads aligned to the region, and calculates the effective coverage of the telomere region. ; statistics The proportion of short fragments among all valid cfDNA fragments ; The relative telomere length is calculated using the following formula: in, This is the ratio of the telomere length of the sample to the baseline telomere length of healthy individuals; Output the relative telomere length value.

6. The multidimensional tumor marker combined detection device according to claim 1, characterized in that, The workflow of the nucleosome imprinting analysis unit is as follows: Extraction of cfDNA fragments End and The end 6-mermotif sequence, through The tool performs motif recognition and filters out motifs with an E value < 1e-5; The percentage of segments within the nucleosome protected region is statistically significant. calculate End and The enrichment degree of end motifs is calculated using the motif information entropy formula, which is as follows: in, Let be the probability of the i-th motif appearing. This represents the total number of motif types. Will Segment proportion, End-motif enrichment The end motif enrichment and motif information entropy are integrated into a kernel body imprint feature vector and output.

7. The multidimensional tumor marker combined detection device according to claim 1, characterized in that, The workflow of the segment distribution feature extraction unit is as follows: The short fragment interval was defined as 100bp-166bp and the long fragment interval as 169bp-240bp. The number of valid cfDNA fragments within the two intervals was counted. The short / long segment ratio is calculated using the following formula: in, This represents the total number of valid cfDNA fragments within the short fragment interval. This represents the total number of valid cfDNA fragments within the long fragment interval. Calculate the peak segment length and segment length standard deviation, and then calculate the segment length distribution entropy. The specific formula is as follows: in, The percentage of a single cfDNA fragment. For length is The proportion of cfDNA fragments, The maximum threshold for segment length statistics. This is the minimum threshold for segment length statistics; Output four types of features: short / long segment ratio, peak length, standard deviation, and distribution entropy.

8. The multidimensional tumor marker combined detection device according to claim 1, characterized in that, The sample processing module includes: The body fluid supernatant separation unit is used to perform centrifugation on body fluid samples to separate and obtain body fluid supernatant; The cfDNA extraction unit is used to extract cfDNA by performing lysis, binding, washing, and elution operations on the body fluid supernatant using an adsorption column method. The library construction unit is used to perform end repair, adapter ligation, and PCR amplification on the extracted cfDNA to construct a sequencing library. The process of constructing the sequencing library does not include bisulfite treatment or targeted capture operations.

9. The multidimensional tumor marker combined detection device according to claim 1, characterized in that, The sequencing module includes: The sequencing control unit is used to set the operating parameters of the sequencing platform, control the sequencing depth, sequencing read length and sequencing type, and start sequencing after performing homogenization on the sequencing library. The data quality control unit is used to perform adapter removal, deduplication, and low-quality read removal on the raw sequencing data produced by the sequencing platform, while retaining valid reads and outputting them.

10. The multidimensional tumor marker combined detection device according to claim 1, characterized in that, The feature fusion and detection module includes: The feature standardization unit is used to receive the raw feature data output by each detection unit, execute the preset standardization algorithm, and output standardized feature values. The feature matrix construction unit is used to classify and integrate standardized feature values ​​according to preset dimensions to construct a feature matrix. The model computation unit is used to load a pre-trained ensemble learning model, input the feature matrix into the model to perform calculations, and output a tumor risk score.

11. The multidimensional tumor marker combined detection device according to claim 1, characterized in that, The result output module includes: The results integration unit is used to receive the tumor risk score and the feature output results of each detection unit, and perform structured integration according to preset logic. The visualization unit is used to present the integrated results in chart form, including charts related to risk scores, charts related to marker characteristics, and abnormal indicator labels. The export unit supports exporting the presented results in a preset standardized format, adapting to clinical archives and hospital information system integration.

12. A method of using a multi-dimensional tumor marker combined detection device, characterized in that, The multidimensional tumor marker combined detection device according to any one of claims 1-11 includes: S1. Sample processing: Process the subject's body fluid samples, extract extracellular cell-free DNA and construct sequencing libraries; the body fluid samples include: bile, pleural effusion, peritoneal fluid, blood or cerebrospinal fluid; S2. Whole-genome low-depth sequencing: Perform whole-genome ultra-low-depth sequencing on the sequencing library, and perform quality control processing on the generated raw sequencing data to obtain effective sequencing data; S3. Simultaneous extraction of multiple biomarkers: Based on the effective sequencing data, six types of tumor biomarkers, namely chromosome copy variation, microsatellite instability, telomere length, nucleosome imprinting, fragment distribution and methylation level, are extracted in parallel through six independent feature extraction processes. S4. Feature Fusion and Risk Assessment: Standardize the extracted features of the six types of tumor markers, construct a feature matrix, input the feature matrix into a preset ensemble learning model to perform calculations, and output a tumor risk score. S5. Results Output: Based on the tumor risk score, the risk level is determined, abnormal biomarker information and related evidence are integrated, a report is generated and can be exported.

Citation Information

Patent Citations

  • Evaluation system for predicting tissue specific source and related disease probability of cfDNA and application

    CN113539355A

  • Chromosomal and sub-chromosomal copy number variation detection

    CN117941002A

  • Construction of early liver cancer detection model containing multi-omics marker composition and kit

    CN115851951A

  • Application of unknown primary focus tumor tissue traceability detection marker and detection system

    CN118308490A