An endometrial cancer early screening system based on multi-omics liquid biopsy

By using multi-omics liquid biopsy technology, cell-free DNA and exosomes in plasma are separated and analyzed. Sequencing data are screened using dynamic weight vectors and physical features, which solves the problem of difficult tumor signal identification in early screening of endometrial cancer and improves the sensitivity and specificity of screening.

CN122428035APending Publication Date: 2026-07-21THE FIRST HOSPITAL OF LANZHOU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
THE FIRST HOSPITAL OF LANZHOU UNIV
Filing Date
2026-04-02
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing endometrial cancer screening technologies have low tumor cell-free DNA abundance and lack tissue specificity in early-stage patients, making it difficult to effectively identify weak tumor signals amidst complex hematopoietic cell background noise, resulting in insufficient screening sensitivity and difficulty in gray zone determination.

Method used

An early endometrial cancer screening system based on multi-omics liquid biopsy was adopted. The system separates cell-free DNA and exosome lysis products from plasma through the sample processing unit, and generates data by combining gene sequencing and protein detection units. The sequencing data is screened using dynamic weight vectors and physical characteristic parameters, and structural abnormality scores are calculated and risk levels are assessed.

Benefits of technology

It improves the sensitivity of abnormal signal identification and the specificity of screening results in early samples with low tumor burden, reduces the interference of non-specific background fluctuations, and enhances the fault tolerance and reliability of the screening system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122428035A_ABST
    Figure CN122428035A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of biomedical detection, and discloses an endometrial carcinoma early screening system based on multi-omics liquid biopsy, which comprises a sample processing unit, a protein detection unit, a gene sequencing unit and a data analysis unit. The system separates peripheral blood into a first component containing free DNA and a second component containing exosome products, and respectively obtains original sequencing data and transcription factor protein quantitative data. The data analysis unit calculates a dynamic weight vector according to the protein data, screens a pure read set by using fragment length and end sequence characteristics, fuses the dynamic weight vector and the nucleosome characteristics of the pure read set, calculates a structure abnormality score and outputs a risk level. The application uses transcription factor specificity to guide genome structure analysis, and effectively improves the sensitivity and specificity of early endometrial carcinoma screening by physical feature denoising.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of biomedical detection technology, specifically to an early screening system for endometrial cancer based on multi-omics liquid biopsy. Background Technology

[0002] Endometrial cancer is a common malignant tumor of the female reproductive system, and its early diagnosis is of great significance for improving patient survival and prognosis. Current clinical screening methods mainly rely on transvaginal ultrasound and serum tumor marker detection, while definitive diagnosis requires invasive fractional curettage or hysteroscopic biopsy.

[0003] However, imaging examinations have limited resolution when the tumor is small or the lesion is confined to the intima, making it difficult to accurately distinguish between benign and malignant lesions; conventional serum markers such as CA125 have low sensitivity in early-stage patients and are easily affected by non-tumor factors such as inflammation, resulting in insufficient specificity.

[0004] With the development of high-throughput sequencing technology, liquid biopsy technology based on cell-free plasma DNA (cfDNA) is gradually being applied to cancer screening. This technology identifies tumor-specific gene mutations or copy number variations by detecting trace amounts of circulating tumor DNA (ctDNA) in the blood. However, in the early stages of endometrial cancer, the tumor burden is extremely low, and the proportion of ctDNA released into the bloodstream is very small; the vast majority of ctDNA originates from the apoptosis of background hematopoietic cells. This extreme signal-to-background ratio makes it difficult for conventional mutation detection to extract effective tumor signals from complex biological noise, easily leading to false negative results.

[0005] To improve detection performance, researchers have recently focused on fragment omics features of cfDNA, such as nucleosome occupancy maps and terminal motifs. While these physical features are widely distributed across the entire genome and can provide richer signals than a single mutation point, current methods generally lack effective tissue-based origin tracing capabilities. Due to the lack of localization guidance for specific organs or tissues, existing analysis algorithms often blindly search for abnormal signals across the entire genome, unable to distinguish between chromatin structural changes caused by tumors and background changes due to physiological fluctuations or nonspecific inflammation.

[0006] Furthermore, although other omics markers such as exosomal proteins can reflect the functional state of cells, existing multi-omics detection schemes mostly adopt a parallel mode of independent detection and result superposition, which fails to deeply explore the intrinsic biological connections between different omics data.

[0007] Specifically, existing technical solutions have not yet established a coupling mechanism that uses protein expression levels to dynamically guide genomic data mining. This results in the inability to effectively utilize the tissue-specific advantages of the proteome to suppress background noise in genomic sequencing when processing early trace samples, thus limiting the practical application effectiveness of liquid biopsy technology in early screening of endometrial cancer. Summary of the Invention

[0008] To address the shortcomings of existing technologies, this invention provides an early endometrial cancer screening system based on multi-omics liquid biopsy. This system solves the problems of existing endometrial cancer screening technologies, which mainly rely on a single biomarker. Due to the low abundance and lack of tissue specificity of cell-free tumor DNA in the plasma of early-stage patients, it is difficult to effectively identify weak tumor signals in the complex background noise of hematopoietic cells, resulting in insufficient screening sensitivity and difficulty in gray zone determination.

[0009] To achieve the above objectives, the present invention is implemented through the following technical solution: an early screening system for endometrial cancer based on multi-omics liquid biopsy, the system comprising a sample processing unit, a gene sequencing unit, a protein detection unit, and a data analysis unit.

[0010] The sample processing unit is used to receive peripheral blood samples from subjects and perform physical separation, separating the peripheral blood samples into a first sample component containing cell-free plasma DNA and a second sample component containing exosome lysis products.

[0011] The gene sequencing unit is connected to the sample processing unit and is used to receive the first sample component and sequence the free DNA therein to generate raw sequencing data containing read length and sequence information.

[0012] The protein detection unit is connected to the sample processing unit and is used to receive the second sample component, determine the concentration of the target transcription factor and internal reference protein, and generate protein quantification data.

[0013] The data analysis unit is connected to both the protein detection unit and the gene sequencing unit; the data analysis unit is configured to perform the following operations: The dynamic weight vector of the target transcription factor is calculated based on the protein quantification data, and the raw sequencing data is screened according to the physical characteristic parameters to generate a pure read set; The structural anomaly score is calculated using a dynamic weight vector and a clean read set, and the risk level is output based on the structural anomaly score.

[0014] In one possible implementation, the data analysis unit includes a weight calculation module. This module calculates the ratio of the target transcription factor concentration to the internal reference protein concentration to obtain the relative abundance. Calculate the deviation of the current sample's internal reference protein concentration from the population baseline mean, and generate a dynamic gain adjustment coefficient accordingly; The relative abundance is corrected using a dynamic gain adjustment coefficient, and the corrected value is mapped to a normalized interval to generate a dynamic weight vector for the target transcription factor.

[0015] In one possible implementation, the data analysis unit includes a data filtering module. This module sets physical gating criteria including read length ranges and end sequence characteristics; Retain reads of length within a specific range (e.g., 130bp to 170bp) from the raw sequencing data; The terminal nucleotide sequences of the retained reads are extracted and compared with the hematopoietic cell background feature library. Reads that appear more frequently than a preset threshold in the background feature library are removed, and the remaining reads constitute a pure read set.

[0016] In one possible implementation, the data analysis unit includes a feature mapping module. This module acquires the coordinates of the genomic binding sites corresponding to the target transcription factor, maps the clean read set to the binding site coordinates, and calculates a window protection score centered on the binding site coordinates. The valley depth features of the window protection score are extracted, and the valley depth features are weighted using the dynamic weight vector of the target transcription factor to obtain the structural abnormality score.

[0017] Furthermore, when calculating the window protection score, the feature mapping module defines the analysis window and its neighborhood range, counts the number of read segments that completely cover the analysis window as the protection signal, counts the number of read segments whose ends fall into the neighborhood range as the breakage signal, and calculates the difference between the protection signal and the breakage signal.

[0018] In addition, when calculating the structural anomaly score, a sequencing depth normalization operation is performed: the average window protection score of the region outside the binding site coordinates is calculated as the background value, the relative drop ratio of the window protection score at the binding site coordinates relative to the background value is calculated, and the dynamic weight vector is multiplied by the relative drop ratio.

[0019] In one possible implementation, the data analysis unit includes a risk assessment module. This module divides the genome of the first sample group into multiple observation windows, calculates the entropy value of the fragment length distribution within each observation window, and calculates the average entropy value of the entire genome as the genome-wide fragment dispersion entropy value. During risk determination, a preset risk determination threshold and a system confidence interval are compared. When the structural anomaly score exceeds the high-risk threshold, the risk level is determined to be positive; When the structural abnormality score does not exceed the high-risk threshold, but the whole genome fragment dispersion entropy value is higher than the high-risk dispersion threshold, the risk level is determined to be gray zone pending investigation; otherwise, the risk level is determined to be negative.

[0020] A second aspect of the present invention provides a method for early screening of endometrial cancer based on multi-omics liquid biopsy, the method comprising the following steps: Peripheral blood samples from subjects were received and physical separation was performed to separate the peripheral blood samples into a first sample component containing cell-free plasma DNA and a second sample component containing exosome lysis products. The first sample component is received and its free DNA is sequenced to generate raw sequencing data containing read length and sequence information. The second sample component is received, the concentrations of the target transcription factor and internal reference protein are measured, and protein quantification data are generated. The dynamic weight vector of the target transcription factor is calculated based on the protein quantification data, and the raw sequencing data is screened according to the physical characteristic parameters to generate a pure read set; The structural anomaly score is calculated based on the dynamic weight vector and the pure read set, and the risk level is output based on the structural anomaly score.

[0021] This invention provides an early screening system for endometrial cancer based on multi-omics liquid biopsy. It has the following beneficial effects: 1. This invention combines the analysis of exosomal proteins and cell-free DNA in plasma, constructing a dynamic weighted vector based on the expression abundance of target transcription factors to weight and correct the nucleosome protection map of the genome. This approach leverages the tissue specificity of transcription factors to compensate for the spatial ambiguity of cell-free DNA, enabling the analysis to focus on highly active regulatory regions associated with endometrial cancer, thereby effectively improving the sensitivity of abnormal signal identification in early samples with low tumor burden.

[0022] 2. This invention employs a data screening strategy based on physical characteristics, combining fragment length distribution and terminal nucleotide motif features to selectively remove background reads originating from hematopoietic cells. This step leverages the differences in nucleosome cleavage patterns between tumor cells and hematopoietic cells to isolate high-purity single nucleosome fragments from the raw sequencing data, significantly reducing the interference of nonspecific background fluctuations on structural scoring calculations and improving the specificity of screening results.

[0023] 3. This invention establishes a two-dimensional risk assessment model. Based on conventional structural anomaly scoring, it introduces whole-genome fragment dispersion entropy as a gray-zone determination indicator. When a sample's structural anomaly score falls within the critical interval, the system uses the entropy value, reflecting the overall disorder of the genome, for auxiliary judgment. This effectively avoids the risk of missing early weak signals that might be caused by a single rigid threshold judgment, improving the fault tolerance and reliability of the screening system. Attached Figure Description

[0024] Figure 1 This is a system architecture diagram of the present invention; Figure 2This is a flowchart of the method of the present invention.

[0025] The module includes: 10. Sample processing unit; 20. Protein detection unit; 30. Gene sequencing unit; 40. Data analysis unit; 41. Weight calculation module; 42. Data filtering module; 43. Feature mapping module; and 44. Risk assessment module. Detailed Implementation

[0026] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0027] Example: Please see the appendix Figure 1 This invention provides an early screening system for endometrial cancer based on multi-omics liquid biopsy. The system includes a sample processing unit 10, a protein detection unit 20, a gene sequencing unit 30, and a data analysis unit 40.

[0028] The sample processing unit 10 is used to receive peripheral blood samples from subjects and perform physical separation, separating the samples into a first sample component containing cell-free DNA in plasma and a second sample component containing exosome lysis products.

[0029] The protein detection unit 20 is connected to the sample processing unit 10, receives the second sample component, measures the concentration of the target transcription factor and the concentration of the internal reference protein, and generates protein quantification data.

[0030] The gene sequencing unit 30 is connected to the sample processing unit 10, receives the first sample component, sequences the free DNA, and generates raw sequencing data.

[0031] The data analysis unit 40 is connected to the protein detection unit 20 and the gene sequencing unit 30, and includes a weight calculation module 41, a data filtering module 42, a feature mapping module 43 and a risk assessment module 44.

[0032] The weight calculation module 41 is used to calculate the dynamic weight vector of the target transcription factor based on the protein quantification data.

[0033] The data filtering module 42 is used to filter the raw sequencing data according to physical characteristic parameters and generate a clean read set.

[0034] The feature mapping module 43 is used to calculate the structural anomaly score based on the dynamic weight vector and the clean read set.

[0035] The risk assessment module 44 is used to output the risk level based on the structural anomaly score.

[0036] Please see the appendix Figure 2 This invention provides a method for early screening of endometrial cancer based on multi-omics liquid biopsy, comprising the following steps: S100. Construct a dual-path parallel detection channel and acquire raw biological information data; The system physically splits the collected peripheral blood samples, establishing separate protein analysis and genomic analysis channels. In the protein analysis channel, plasma exosomes are extracted and lysed, and the concentrations of target transcription factors and internal reference proteins are determined using high-sensitivity immunoassay. In the genomic analysis channel, cell-free DNA is extracted from plasma and a sequencing library is constructed. High-throughput sequencing technology is used to obtain raw sequencing data containing read lengths and sequence information.

[0037] S200, Calculate adaptive gain weights based on intrinsic parameter normalization strategy; For the protein analysis channel data, the relative abundance ratio of the target transcription factor relative to the internal reference protein is calculated. The dynamic gain adjustment coefficient is calculated using the deviation between the total amount of the internal reference protein in the sample and the mean of the population baseline. The relative abundance ratio is input into an activation function with the dynamic gain adjustment coefficient as a parameter to generate a focus weight vector for each target transcription factor. The focus weight vector is used to quantify the guiding priority of each transcription factor for subsequent genome analysis.

[0038] S300, Perform isochronous data cleaning based on physical characteristic constraints; For genome analysis channel data, physical gating criteria including read length range and terminal sequence characteristics are set. Reads from mononuclear bodies with lengths in the range of 130bp to 170bp are retained, while reads with abnormal lengths and background noise from the terminal sequence matching library are removed. A pure read dataset with biological consistency between temporal phase and exosomal proteins is constructed to eliminate temporal misalignment interference caused by differences in metabolic half-life.

[0039] S400, Perform transcription factor-guided nucleosome topological feature mapping; By calling the genome annotation database, the coordinates of chromatin binding sites corresponding to high-concern-weighted transcription factors are obtained. The clean read dataset is mapped to the binding site coordinate region, and the window protection score reflecting the nucleosome occupancy status is calculated. By combining the concern weight vector and the window protection score, the weighted structural abnormality score is calculated to quantify the matching degree of chromatin openness and protein expression level in the regulated regions of specific transcription factors.

[0040] S500 integrates multi-dimensional feature indicators and outputs risk judgment results; The system summarizes weighted structural anomaly scores, cell-free DNA mutation frequency characteristics, and genome-wide fragment dispersion entropy values. These are then compared to preset risk thresholds and system confidence intervals. A positive result is determined when the weighted structural anomaly score exceeds the high-risk threshold. When the score fails to meet the standard but the entropy value of the whole genome fragment dispersion is abnormally high, it is judged as a gray area to be investigated. Otherwise, it will be considered negative. Output the final screening result report.

[0041] As an implementation detail of the sample processing unit 10, the standardized preparation of samples is the material basis for ensuring the accuracy of subsequent multimodal analysis results.

[0042] This embodiment establishes a first sample component and a second sample component through physical diversion and differentiated extraction paths, thereby avoiding the biological background noise bias caused by heterogeneous sampling.

[0043] S110, Peripheral blood sample collection and two-stage centrifugation Whole blood was collected from subjects using vacuum blood collection tubes containing ethylenediaminetetraacetic acid (EDTA) as an anticoagulant. A time constraint mechanism was set to require processing to be completed within 4 hours after collection to prevent leukocyte apoptosis from releasing long fragments of genomic DNA that contaminate the plasma.

[0044] The pretreatment process follows a strict two-step differential centrifugation logic: In the first stage, the whole blood sample is placed at 1600×g and centrifuged at 4°C for 10 minutes to separate the supernatant plasma into a new RNase / DNase-free centrifuge tube. At this stage, a hemolysis detection logic is introduced. If the plasma turns red and the free hemoglobin concentration exceeds a preset threshold (such as 100 mg / dL), the sample is deemed invalid to prevent background interference caused by red blood cell lysis.

[0045] In the second stage, the separated plasma was centrifuged at 16000×g at 4°C for 10 minutes. The centrifugation force parameter was chosen to physically settle residual cell debris and large apoptotic bodies in the plasma while retaining exosome particles.

[0046] S120, Precision physical separation of plasma samples Cell-free plasma processed in step S110 was vortexed and then aliquoted into equal volumes using a precision pipette or automated dispensing station. The plasma was physically separated into a first sample component and a second sample component, and the dispensing error rate was recorded, requiring volume deviation to be controlled within ±1%. This provides a material conservation basis for the mathematical normalization of subsequent protein-gene multimodal data. The first sample component is directed to the genomics analysis channel, and the second sample component is directed to the proteomics analysis channel.

[0047] S130, Non-destructive extraction of the first sample component For the first sample component, a cell-free DNA extraction kit based on magnetic beads was used for extraction. Proteinase K and lysis buffer were added to the plasma, and the DNA-protein binding was disrupted by incubation at 60°C, releasing the cell-free DNA.

[0048] The free DNA was then adsorbed using silica gel membrane magnetic beads, and after washing to remove impurities, it was eluted with low TE buffer to obtain a purified free DNA solution.

[0049] The key control point in this step is to prohibit the introduction of physicochemical shearing operations such as ultrasonic fragmentation or enzyme digestion, in order to fully preserve the natural length distribution characteristics and nucleosome imprints of free DNA in plasma, ensuring that the physical fidelity of the data meets the requirements of subsequent topological analysis. The extracted product needs to be concentrated using a fluorometer. If the concentration is below 0.1 ng / μL, a low abundance enrichment retest procedure is triggered.

[0050] S140, Exosome membrane lysis and contents release of the second sample component For the second sample component, the aim is to obtain exosome contents containing nuclear transcription factors. Exosomes can be enriched by ultracentrifugation or polymer-based precipitation. For example, polyethylene glycol precipitation reagent is added, and after incubation at 4°C, the mixture is centrifuged at 10000×g. After removing the supernatant, a lysis buffer containing a nonionic surfactant is added to the exosome precipitate, along with a mixture of protease inhibitors.

[0051] Lysis was performed on ice for 20 minutes with vigorous vortexing. The surfactant disrupted the phospholipid bilayer structure of exosomes, allowing transcription factors and housekeeping proteins originally encapsulated within the membrane to be completely released into the liquid phase. The lysate was centrifuged at 14000×g to remove insoluble precipitates, and the supernatant was used as the protein sample solution for subsequent Simoa assays, ensuring that the detection signal reflects the total amount of protein carried by the exosomes rather than the protein on the membrane surface.

[0052] As a specific implementation detail of the protein detection unit 20, this embodiment constructs a high-sensitivity quantitative detection platform based on single-molecule array technology.

[0053] This unit aims to address the sensitivity bottleneck of traditional enzyme-linked immunosorbent assays (ELISA) in detecting plasma exosomal nuclear proteins (typically...). (Level), the detection limit is reached by single-molecule counting. even This provides a level, thus offering a quantitative input with statistically significant differences for subsequent adaptive gain calculations.

[0054] S210, Constructing a digital immunoassay environment conforming to a Poisson distribution. Given that the concentrations of transcription factors associated with endometrial cancer in plasma exosomes of early-stage patients are extremely low, the system employs single-molecule array detection technology.

[0055] In practice, a random binding model between the microspheres and the target protein molecules is established by adjusting the ratio of the capture antibody density on the surface of the paramagnetic microspheres to the volume of the sample to be tested, ensuring that it strictly follows a Poisson distribution. This process ensures that each microwell in the micropore array contains either 0 or 1 target molecule. Utilizing the volume confinement effect of the micropores (approximately 50 femtoliters), the fluorescence signal generated by the enzymatic reaction is confined to a very small space. By calculating the ratio of signal-bearing micropores to the total number of micropores, single-molecule digital counting is performed directly, rather than relying on the integration of analog light intensity signals. A temperature control module is required to maintain the reaction temperature fluctuation within 25 ± 0.1℃ to ensure the constancy of the enzymatic reaction rate.

[0056] S220, targeted capture and low-value imputation of endometrial cancer-specific transcription factors In selecting detection indicators, the system focuses on nuclear proteins specific to endometrial tissue and associated with malignant transformation, specifically PAX2 protein and estrogen receptor protein. For samples with potentially extremely low abundance, the system incorporates data truncation processing logic. If the detected target protein signal intensity is lower than the system detection limit, Instead of directly setting the concentration to 0, it avoids situations where the denominator is zero or the data is distorted when calculating the relative abundance ratio later.

[0057] Instead, numerical interpolation is performed using the following formula: ; in, This represents the corrected concentration estimate. This imputation strategy, based on the statistical distribution characteristics of left-censored data, minimizes bias while preserving sample information. During implementation, high-affinity monoclonal antibodies targeting the PAX2 terminal domain and ESR1 ligand-binding domain were used, and cross-reactivity validation was performed to ensure specificity greater than 99%.

[0058] S230. Quantitative internal control setting and quality gating based on exosome housekeeping protein: In order to eliminate physiological differences in the total amount of exosome secretion among individuals (such as non-specific increases in inflammatory states) and fluctuations in recovery rate during sample pretreatment, this embodiment introduces exosome housekeeping protein as a quantitative normalization internal control.

[0059] Specifically, CD63, a member of the tetraspanosome superfamily, or the tumor susceptibility gene 101 protein were selected, as these proteins exhibit relatively constant expression levels on or inside the exosome membrane. A dual-channel quality validation gating mechanism was established here: Setting the biologically effective range of internal control protein concentration If the measured value If so, it is determined that the exosome enrichment in step S140 was ineffective or the lysis was insufficient; like This suggests that the sample may have hemolysis or lipemia interference, only when At that time, the protein data of the sample is marked as valid and passed downstream.

[0060] The S240 nonlinear regression fitting and standardized data output immunoassay unit reads the fluorescence signal of the micro-well array using laser optical elements, fits the standard curve using a four-parameter logistic regression model, and converts the average enzyme molecule number / micron (AEB) into a physical mass concentration value. The regression model is as follows: ; in, This is the optical signal response value. The concentration of the protein to be tested. This is the asymptotic minimum (baseline signal). This is the asymptotic maximum value (saturation signal). The inflection point concentration, Let be the Hill slope.

[0061] The final output includes the concentration of the target transcription factor. (unit: ) and internal reference protein concentration (unit: A structured dataset, which must undergo metadata alignment verification to ensure that the data belongs to the same subject. and Not only do they originate from the same sample number, but their detection timestamp deviation does not exceed 2 hours to prevent data drift caused by reagent batch effects.

[0062] As a specific implementation detail of the gene sequencing unit 30, this embodiment aims to construct a genome data acquisition process that can faithfully reproduce the terminal features of cell-free DNA fragments in plasma and the nucleosome occupancy information.

[0063] Given the biological characteristics of cfDNA, its fragmentation pattern is jointly determined by the nucleosome protection mechanism and the specificity of nuclease cleavage during apoptosis. Therefore, the fidelity and edge integrity of sequencing data at physical resolution are key prerequisites for the effective convergence of subsequent feature extraction algorithms.

[0064] S310, Constructing a physically faithful library without amplification bias To completely avoid the accumulation of base mismatches and abundance bias caused by GC content preference during the traditional polymerase chain reaction (PCR) library construction process, this embodiment adopts a PCR-free library construction strategy to process the free DNA obtained by the sample processing unit 10 as a preferred method.

[0065] In practice, the purified cfDNA undergoes end repair and A-tail addition to enable directional ligation with pre-made sequencing adapters. During this process, fragment size selection is strictly prohibited to prevent the loss of short fragments (<100bp, typically originating from transcription factor binding sites) or long fragments (>170bp, originating from dianucleosome binding regions), thus preserving the full spectrum of length distribution information.

[0066] The adapter sequence design includes a unique molecular identifier that can distinguish the source of the sample, but in PCR-free mode, it mainly relies on the index sequence of the adapter itself to distinguish multiple samples, thereby ensuring that the ends of the sequencing reads truly reflect the original break points of DNA molecules in the plasma.

[0067] S320, high-depth whole-genome sequencing and data output The constructed library was loaded into a high-throughput sequencing platform for whole-genome sequencing. Considering the extremely high signal-to-noise ratio requirements for single-base resolution in subsequent nucleosome localization analysis (WPS calculation), i.e., the need to statistically distinguish between true nucleosome protective peaks and random background fluctuations, a minimum effective threshold for sequencing depth was set for the system. .

[0068] When the actual average sequencing depth is detected If the data sparsity of the sample is deemed too high to support high-resolution nucleosome map reconstruction, a retesting process must be triggered. The sequencing mode is configured as paired-end sequencing. The read length parameter is selected based on the average physical length of cfDNA. The paired-end 150bp sequencing strategy can not only cover the full length of most fragments, but also accurately anchor the genomic coordinates at both ends of the fragment through paired-end alignment logic, eliminating the localization ambiguity that may be caused by single-end sequencing.

[0069] The raw data output by the S330 sequencer, based on a probabilistic model for read cleaning and quality control, must undergo rigorous mathematical quality control (QC) to remove optical noise and adapter contamination.

[0070] This embodiment is based on the Phred quality scoring system and uses a probabilistic model to quantify the confidence level of bases. The quality score is defined as follows: With base sequencing error probability The logarithmic relationship is as follows: ; in, Assign a quality value to each base for the sequencer (the value is usually an integer between 0 and 40). This represents the probability that the base is misidentified. Based on this mathematical model, the following hierarchical cleaning logic is executed: Joint trimming and null value handling: A comparison algorithm is used to identify joint sequences at the ends of read segments. If any joint remnants are detected, they are removed. A length threshold discrimination logic is introduced here: a minimum retention length is set. If the length of the segment is read after trimming. This means that the biological information contained in the sequence is close to zero or cannot be uniquely located on the genome. In order to avoid introducing non-specific alignment noise, the read is directly discarded from the dataset (handling the boundary case where the denominator is close to the invalid length).

[0071] Sliding window filtering: A 4bp sliding window is used to scan the read segment, and the average quality value within the window is calculated. .like If the error rate is greater than 1%, the low-quality region is truncated to ensure that the overall confidence of the remaining sequence meets the requirements of downstream mutation analysis.

[0072] Final validity determination: High-quality read segments with less than 5% N bases (bases that cannot be determined) and more than 80% Q30 (bases with a quality value greater than 30, i.e., error rate <0.1%) are retained.

[0073] S340, Genome Alignment and Coordinate Standardization The cleaned reads were aligned to the human reference genome (hg19 or hg38 version) using an efficient alignment algorithm. The output was stored as a binary alignment file (BAM format), which detailed the start and end coordinates of each read on the chromosome and the mapping quality (MAPQ) value.

[0074] At this stage, multiple comparison filtering logic is executed: only read segments that are uniquely compared are retained, and a comparison quality threshold is set. .like This indicates that there are multiple highly similar sequences in the genome, which is a fuzzy alignment. To prevent homologous sequences from interfering with the signal recognition of transcription factor binding sites, they are removed.

[0075] Subsequently, the BAM file was indexed using a coordinate sorting tool. The physical center point of the DNA fragment was deduced based on the start and end coordinates of the reads, providing standardized geometric input for the subsequent generation of coverage depth spectra. Furthermore, despite using PCR-free library construction, tools such as Picard were still used to label and remove optical repeats to ensure that each counted DNA fragment represents an independent biological molecule, preventing inflated abundance calculations due to overlapping recognition of optical clusters by the sequencer.

[0076] The weight calculation module 41 serves as a mathematical bridge connecting proteomics data and genomics analysis. Its core function is to execute an adaptive gain control algorithm. Based on the automatic gain principle in signal processing, this algorithm aims to establish a non-linear data standardization mechanism. It adaptively adjusts the contribution weight of protein biomarker expression levels to subsequent multimodal fusion analysis, fundamentally solving the problem of signal response imbalance caused by differences in the basal exosome secretion rate among individual subjects.

[0077] S410. The benchmarked relative abundance is calculated based on the standardized dataset output by the protein detection unit 20, and the system performs internal normalization processing. Considering that the absolute concentration of the target transcription factor depends not only on the tumor burden, but also linearly on the total plasma volume and exosome recovery rate, directly using the absolute concentration value would lead to the curse of dimensionality in the multimodal data input space.

[0078] Therefore, this embodiment uses the ratio method to eliminate systematic errors and defines relative abundance. The calculation is as follows: ; in, The concentration of the target transcription factor was measured (unit: ), The concentration of the internal reference protein (unit: A smoothing factor is introduced here. (Recommended value: 10) -6 The physical significance of this (order of magnitude) is to prevent the mathematical singularity of division by zero caused by the internal reference protein measurement value approaching zero (such as extremely diluted samples or noisy data below the detection limit), and to ensure the numerical stability of the calculation process.

[0079] The output of this step The specific enrichment of tumor-associated transcription factors within a unit exosome load was characterized, i.e., specific activity.

[0080] S420, Adaptive Derivation of Dynamic Gain Coefficient and Clamping Control To address the signal masking phenomenon in low-secreting subjects that may be caused by differences in baseline exosome levels among different subjects, this embodiment designs an adaptive gain mechanism with boundary constraints.

[0081] The system first verifies the concentration of the internal reference in the current sample. The timeliness is ensured to maintain consistency with the calculation window of the population mean statistic. Subsequently, a pre-stored historical population database is retrieved to obtain the population mean of the internal reference protein concentration. Based on the internal reference concentration of the current sample and The degree of deviation is used to calculate the dynamic gain coefficient. : ; ; in, The gain sensitivity index, with a value range of [0.5, 1.0], is selected based on analysis of variance of the experimental data and is used to control the compensation level for low-secretion samples. when When it is in full compensation mode, This is a partial compensation model. This is the upper limit pin value for the gain coefficient (e.g., set to 5.0) to prevent damage caused by severe sample degradation. Extremely low values ​​lead to unrealistically high gain coefficients, thus introducing false positive signals. This is a statistic that is dynamically updated using a moving average algorithm.

[0082] The physical meaning of this formula is that for samples whose intrinsic parameter levels are significantly lower than the population mean, the system automatically assigns a gain coefficient within a safe range to restore their true signal strength. Conversely, for high-secretion samples, their signals are appropriately normalized to achieve global equalization of signal intensity.

[0083] S430, Generation of Modified Abundance The original relative abundance was corrected using the calculated dynamic gain coefficient to obtain the corrected abundance value. : ; This step transforms a simple biochemical concentration ratio into a standardized mathematical indicator that reflects the degree of pathological risk, eliminating background noise caused by individual physiological differences and making the data of different subjects comparable in the feature space.

[0084] S440, Nonlinear Weight Mapping Based on the Sigmoid Function In order to adjust the abundance values This is transformed into normalized weights that can directly participate in subsequent attention mechanisms or weighted summation operations. In this embodiment, the Sigmoid activation function is used to convert these weights into normalized weights. Mapped to Continuous interval. Weight vector. The formula for generating it is as follows: ; in, The slope parameter of the S-curve determines the sensitivity to changes in weights, and is typically set to [value missing]. To create a sensitive response characteristic similar to a soft switch; This is the center point of the threshold for clinical determination.

[0085] parameter The determination is based on the optimal Youden index obtained from ROC curve analysis of the confirmed sample set. The value represents the optimal critical point for distinguishing between benign and malignant lesions.

[0086] The output of the weight generation function This will be used as a broad-spectrum scalar and multiplied element-wise with the core body feature vector extracted by the gene sequencing unit to achieve the following technical effects: when (Note: This is a non-tumor sample.) Subsequent algorithms will automatically suppress the focus on genomic features of the sample in specific oncogenic pathways, saving computational resources and reducing the false positive rate; when (High risk warning) This indicates that the subsequent multimodal fusion network should focus on the chromatin openness characteristics of the sample with all weights. when At the same time, the system provides linear transition weights to objectively reflect the uncertainty of diagnosis and avoid the loss of edge information caused by traditional hard threshold judgment.

[0087] The core task of the data filtering module 42 is to execute the "isochronous fragment gating" strategy. This strategy, based on the principle of metabolic kinetic differences in molecular biology, aims to address the asynchrony in the half-lives of exosomal proteins and their accompanying circulating tumor DNA in vivo. Given that long-chain genomic DNA typically originates from cell necrosis and is cleared from the bloodstream at a slower rate, often representing old pathological accumulations, while exosomal and mononuclear DNA mainly originate from apoptosis or active secretion, reflecting the immediate active state of the tumor, this module uses physical feature screening to force the alignment of the observation window of genomic data with the metabolic window of proteomics data. This ensures that multimodal analysis is performed within the same biological timeframe, avoiding the failure of correlation analysis due to temporal misalignment.

[0088] S510, Physical Length Bandpass Filtering Based on Nucleosome Footprint For the standard BAM file output by gene sequencing unit 30 and compared, the system applies a length bandpass filter based on the spatial structure of nucleosomes. According to the principles of chromatin biology, the core sequence of the DNA strand wrapped around a single histone octamer is about 147 bp in length. Including the linker region, the typical length of the intact single nucleosome protective fragment is distributed near the Gaussian peak of 167 bp.

[0089] In contrast, fragments shorter than 100 bp are mostly due to transient binding protection of transcription factors, while fragments longer than 170 bp may contain binucleosomes or random breaks originating from necrotic cells. To specifically enrich mononucleosome signals that reflect nuclear localization information, an effective length range was defined. (Unit: bp). For any sequencing read Its physical length is denoted as Construct the length mask function : ; The physical significance of this formula lies in constructing a digital bandpass filter: by retaining only fragments between 130bp and 170bp, the system filters out high-molecular-weight long-fragment background noise (usually necrosis sources) and excessively degraded short-fragment random noise. The release kinetics of these specific-length DNA fragments are highly time-dependent on the secretion of tumor exosomes, thus physically eliminating spurious peaks caused by large tumor volumes but inactive metabolism (primarily necrosis), ensuring the timeliness and purity of the input data.

[0090] S520, Source Background Removal Based on Terminal Motifs Based on physical length screening, as a preferred method, the system introduces terminal motif analysis to further remove background interference from normal hematopoietic cells at the biochemical level.

[0091] The underlying technology lies in the fact that cells of different tissue types are cleaved by specific nucleases during apoptosis, leaving tissue-specific 4bp nucleotide sequence imprints at the broken ends of cfDNA fragments. The system extracts each retained read. The first four bases at the 5' end are denoted as .

[0092] Meanwhile, the system has a pre-installed background motif feature library for hematopoietic cell origin. This feature library is constructed based on statistical features from large-scale sequencing data of healthy individuals, and contains high-frequency terminal sequences significantly enriched in cfDNA derived from leukocytes and erythroid progenitor cells. A source filtering function is constructed based on a Bayesian probability model. : ; in, This represents the posterior probability that the motif belongs to the hematopoietic system. The background probability threshold is used for determination.

[0093] In this embodiment, The value was set to 0.95, determined to ensure background removal with a 95% confidence level. This means that when the probability of the terminal sequence belonging to hematopoietic background is extremely high, it is judged as background noise and hard-removed. The technical effect of this step is to significantly reduce the dilution effect of DNA released from healthy tissue (especially lysed blood cells) on the weak tumor signal, which is equivalent to performing background subtraction at the molecular level and improving the signal-to-noise ratio.

[0094] S530, the structured recombination of the isochronous pure dataset, combined with the above-mentioned physical and biochemical dimension screening logic, the system performs strict logical AND operation, that is, only when... and At that time, the read segment is marked as valid and retained. The system will then reassemble all filtered read segments into an "isochronous clean dataset," denoted as... This dataset serves as the sole input source for the subsequent collaborative analysis module, and its data structure is defined as a set of quintuple vectors: s.t. ; in, Chromosome numbering, As the starting coordinates of the genome, For positive and negative chain directions, The length of the segment. The construction of this structured dataset, which provides nucleotide sequence information, ensures that the nucleosome features subsequently input into the graph neural network are based entirely on high signal-to-noise ratio data from tumor sources that are in the active secretory phase. This effectively avoids the false negative problem caused by spatiotemporal misalignment in traditional liquid biopsies at the data source.

[0095] The feature mapping module 43, as the core unit of the multimodal fusion computing architecture, performs "transcription factor-guided nucleosome topology analysis". Based on the principle of cross-modal parameter mapping, this algorithm uses the protein-terminal weight vector generated by the preceding module as a spatial constraint to reduce the dimensionality of the high-dimensional feature space across the entire genome, accurately locates transcription factor binding sites (TFBS) with high biological relevance, and quantitatively assesses the chromatin open state and the actual occupancy of protein factors at these sites.

[0096] S610, Adaptive Region of Interest (ROI) Locking Based on Weight Vector The system receives the normalized weight vector output by the weight calculation module 41. ( (This refers to the number of transcription factor types being monitored). To avoid wasting computational resources and accumulating false positives due to blind scanning of the entire genome, this embodiment employs a dynamic threshold gating strategy. The system first calls a pre-configured transcription factor motif database and, in conjunction with genome annotation information, extracts the coordinates of the potential binding centers for each target transcription factor.

[0097] During this process, the system executes weight-based spatial filtering logic: For any i Transcription factors, only if their corresponding weights Exceeding the preset activation threshold Only then will the system activate topological analysis of its relevant genomic coordinates.

[0098] In this embodiment, The threshold is set to 0.1, selected based on the statistical lower limit of background noise levels. This aims to filter out transcription factors that are not significantly expressed or effectively detected in exosomes, thus focusing computational power on high-risk biological pathways. For activated transcription factors, the system generates a set of target coordinates. ,in In addition to motif matching points, the analysis should also be filtered through genome annotation to limit it to promoter and enhancer regions, in order to ensure that the analyzed sites have a clear transcriptional regulatory function.

[0099] S620, Discretization and Quantization of Window Protection Scoring For locked sets Each coordinate point in The system uses the isochronous clean dataset output by the data filtering module 42 to calculate the Windowed Protection Score (WPS) near the site. WPS, as a signal processing indicator, is based on the physical principle of nuclease differential cleavage of chromatin. DNA regions that are tightly bound by nucleosomes or transcription factors are not only difficult to cut, but sequencing reads can also completely span these regions. Conversely, exposed connection regions or open regions caused by factor substitution are prone to breakage, leading to accumulation of read ends in these areas.

[0100] Defined by coordinates Set the window radius parameter to the center of the analysis window. In this embodiment, The preferred setting is 120bp, corresponding to a full window width of 240bp. This size is chosen to cover the spatial scale of a standard kernel body (approximately 147bp) and its adjacent connected regions. For any discrete position within the window... WPS value The calculation formula is as follows: ; in, Characterizing the protection signal, defined as full coverage (span) Centered on, with a length of The number of segments read from the window; The characterization of a break signal is defined as the 5th or 3rd end of the read segment falling into a position that is... Centered on, with radius The cumulative number of reads within a micro-neighborhood (e.g., 5 bp).

[0101] This formula avoids the risk of numerical divergence caused by the denominator approaching zero in low sequencing depth regions by using difference operations without division. In the output results, positive peaks represent nucleosome occupancy, and negative valleys represent nucleosome deletion or chromatin openness caused by transcription factor binding.

[0102] S630, Weighted Aggregation of Multimodal Structural Anomaly Scores To integrate protein expression levels and chromatin openness into a single quantitative metric, this embodiment constructs a "structural anomaly score." The system first performs Savitzky-Golay smoothing filtering (polynomial order) on the calculated raw WPS waveform signal. Frame length This process removes random Poisson noise introduced during sequencing. Subsequently, at the target transcription factor binding center... Extract the depth of the feature valley.

[0103] Define local valley depth The signal drop at this site relative to the background of the adjacent wing nucleosome: ; in, Center point to the left and right to The average WPS value of the outer region physically represents the local background nucleosome density around the site. The function acts as a one-sided rectification operator, ensuring that only concave features representing open / missing features are extracted, ignoring non-specific positive oscillations.

[0104] Finally, the system utilizes the protein-terminal weight vector. By multiplying and coupling these genomic features, a final single-point structural anomaly score is generated. : ; This aggregation formula integrates physical information from three dimensions: This serves as a priori gain coefficient from proteomics, acting as a soft gating mechanism. Physically and causally, this means that even if a weak chromatin opening signal is observed in the sequencing data, if the corresponding transcription factor protein does not show abnormal enrichment in exosomes (i.e., ...), the signal will not be considered valid. This signal will be used as background noise to suppress it; conversely, if the protein is extremely abnormal, the weight of this genomic feature will be significantly amplified. As a physical characteristic at the genome level, it directly reflects the degree of chromatin accessibility; For local sequencing depth The confidence level correction term. This is the original coverage read number for this site; adding 1 is to handle the boundary case where the depth is 0.

[0105] The reason for introducing the logarithmic term is to follow the diminishing information gain property of the law of large numbers: to give appropriate confidence rewards to high-depth regions, while suppressing linear overfitting caused by PCR amplification bias.

[0106] Through the above calculations, the feature mapping module outputs a set of weighted anomaly score vectors. These vectors highly condense the dual pathological evidence of abnormal protein expression and abnormal genome opening, serving as strong feature inputs for subsequent classification models to determine tissue origin and benign / malignant risk.

[0107] As the final node in the multimodal data processing workflow, the risk assessment module 44 not only performs data aggregation calculations, but its core function lies in configuring the "dual fail-safe decision" logic. This addresses the heterogeneity inherent in tumor biology, particularly the fact that some tumor subtypes may not significantly secrete exosomes, leading to the need for protein-guided structural scoring in preceding steps. In the event of false negative low values, this module introduces genome-wide dispersion, which is completely independent of protein expression pathways, as an orthogonal validation metric to construct a three-state logic gate circuit, aiming to completely avoid the risk of missed detection caused by signal silencing.

[0108] S710, Independent Calculation of Whole Genome Fragment Discreteness To assess the overall stability of the genome from a macroscopic physical perspective, the system calculates the dispersion of whole-genome fragments. This step is based on the biophysical principle that cfDNA released by healthy cells mainly originates from programmed apoptosis and is precisely regulated by nucleosomes and nucleases, with its fragment length distribution exhibiting highly regular peaks (concentrated around 167 bp). Due to genomic instability and an increased proportion of necrosis, tumor cells release DNA fragments with significantly random and disordered length distribution.

[0109] The system divides the whole genome into The first non-overlapping observation window, in this embodiment, is set to a size of 1Mb. This scale is chosen to balance the robustness of local statistics with computational efficiency. For the first... Each observation window counts the segment length falling within the interval. Normalized frequency distribution within ,in This is a variable for fragment length. The effective interval is preferably set to [100, 250] (unit: bp) to cover the main distribution range of mononuclear bodies and their variants.

[0110] Calculate the local Shannon entropy of this window. : ; in, It is a very small positive number (e.g., 10). -9 ), used to prevent when a certain length frequency Logarithmic operations sometimes result in mathematical singularities. The larger the value, the more uniform and disordered the distribution of segment lengths in that region, i.e., the higher the degree of disorder.

[0111] Subsequently, the system calculates the average entropy value of the entire genome as an indicator of total dispersion. : ; This indicator It can directly quantify the thermodynamic disorder of plasma DNA libraries without relying on any specific mutation sites or protein markers, serving as a non-specific, broad-spectrum indicator of malignancy risk.

[0112] S720, the multi-feature linear fusion and total score calculation system receives the set of structural anomaly scores from specific mapping block 43, and obtains the total structural score through weighted summation and aggregation. Simultaneously, as a conventional technical method, the system retrieves the single nucleotide variant load output from the variant detection unit in parallel. With copy number variation load .for and The specific calculation method can be implemented by those skilled in the art based on existing standard processes such as GATK or CNVkit. The specific algorithm details are well-known in the field and will not be elaborated here.

[0113] To synthesize multi-dimensional evidence, this embodiment constructs a linear weighted fusion model to generate a comprehensive risk score. : ; In this formula: The Sigmoid normalization function aims to map raw scores of different dimensions to... The standardized interval is used to prevent a single feature value from dominating the overall decision-making process.

[0114] These are the feature weight coefficients, and they satisfy the normalization condition. In this embodiment, as a preferred parameter configuration, the following settings are configured: The rationale for this asymmetric weighting is based on epigenetic characteristics guided by exosome proteins ( It integrates both protein and gene evidence, and has higher sensitivity and specificity in early tumor detection than traditional single-dimensional SNV or CNV features, thus giving it a dominant position in this respect.

[0115] 、 、 The input values ​​represent structural anomalies, mutational loads, and copy number anomalies, respectively. These inputs need to be preprocessed with zero-mean standardization before calculation to eliminate baseline differences.

[0116] S730, decision output of tri-state logic gate circuit To address the limitations of traditional binary classification models (positive / negative) in complex clinical scenarios, this module designs a comprehensive risk score-based approach. With genome-wide dispersion The dual-input, three-state decision logic.

[0117] The system presets two key thresholds: Risk scoring threshold The maximum point of the Youden index, based on the receiver operating characteristic (ROC) curve of the training set, is used to distinguish significantly high-risk samples.

[0118] High-risk threshold for dispersion Determined based on the 99th percentile of baseline data from a large number of healthy individuals, it is used to define whether there are statistically significant abnormalities or disturbances in the genome.

[0119] The system performs rigorous logical operations to generate a final diagnostic conclusion. : High-risk confirmation: ; When the composite score exceeds the threshold, regardless of the peak value, it is directly judged as positive. This logic corresponds to typical tumor samples with active exosome secretion or extremely high variant burden, where the multimodal evidence chain is closed and the confidence level is highest.

[0120] Safe and low risk: IF( ) ( THEN "Low risk"; The system excludes risk only when the specificity score is normal and the genome-wide disorder is within the normal range. This double-negation logic ensures that the sample has neither specific epigenetic / genetic abnormalities nor increased non-specific background noise, thereby greatly reducing the false negative rate.

[0121] Gray zone warning IF( ) ( THEN Uncertainty Warning When the specificity score is low (indicating no specific target was detected), but the whole genome shows abnormally high disorder, the system does not directly determine it as negative, but instead forcibly outputs a gray zone warning. This logical branch forms the core safety mechanism of this invention, which covers special tumor subtypes that are negative for exosome secretion but positive for cfDNA necrosis and release, or situations where the sample is severely degraded due to improper handling. By blocking the direct negative judgment, it triggers subsequent manual review or resampling procedures.

[0122] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. An early screening system for endometrial cancer based on multi-omics liquid biopsy, characterized in that, include: The sample processing unit is used to receive peripheral blood samples from subjects and perform physical separation, separating the peripheral blood samples into a first sample component containing cell-free plasma DNA and a second sample component containing exosome lysis products. A gene sequencing unit, connected to the sample processing unit, is used to receive the first sample component and sequence the free DNA therein to generate raw sequencing data containing read length and sequence information. A protein detection unit, connected to the sample processing unit, is used to receive the second sample component, determine the concentration of the target transcription factor and the internal reference protein, and generate protein quantification data. The data analysis unit is connected to both the protein detection unit and the gene sequencing unit. The data analysis unit is used to calculate the dynamic weight vector of the target transcription factor based on the protein quantification data, and to filter the raw sequencing data according to physical characteristic parameters to generate a pure read set; The data analysis unit is also used to calculate a structural anomaly score based on the dynamic weight vector and the clean read set, and output a risk level based on the structural anomaly score.

2. The early endometrial cancer screening system based on multi-omics liquid biopsy according to claim 1, characterized in that, The data analysis unit includes a weight calculation module, which is configured to perform the following steps: The relative abundance is obtained by calculating the ratio of the concentration of the target transcription factor to the concentration of the internal reference protein. Obtain the population baseline mean of the internal reference protein concentration, calculate the deviation of the current sample's internal reference protein concentration from the population baseline mean, and generate a dynamic gain adjustment coefficient based on the deviation. The relative abundance is corrected using the dynamic gain adjustment coefficient, and the corrected value is mapped to a normalized interval to generate a dynamic weight vector for the target transcription factor.

3. The early endometrial cancer screening system based on multi-omics liquid biopsy according to claim 1, characterized in that, The data analysis unit includes a data filtering module, which is configured to perform the following steps: A physical gating criterion is set, which includes the read segment length range and the end sequence characteristics; Reads with lengths between 130bp and 170bp were retained from the raw sequencing data. The terminal nucleotide sequences of the retained reads are extracted, compared with a preset hematopoietic cell background feature library, and reads whose terminal nucleotide sequences appear more frequently than a preset threshold in the background feature library are removed. The remaining reads are then recombined into the pure read set.

4. The early endometrial cancer screening system based on multi-omics liquid biopsy according to claim 1, characterized in that, The data analysis unit includes a feature mapping module, which is configured to perform the following steps: Obtain the coordinates of the genomic binding site corresponding to the target transcription factor; Map the clean read set to the binding site coordinates, and calculate the window protection score centered on the binding site coordinates; The valley depth features of the window protection score are extracted, and the valley depth features are weighted using the dynamic weight vector of the target transcription factor to obtain the structural abnormality score.

5. The early endometrial cancer screening system based on multi-omics liquid biopsy according to claim 4, characterized in that, The feature mapping module is configured to: Define the analysis window and its neighborhood range; The number of read segments that completely cover the analysis window is counted as a protection signal; The number of read segments whose ends fall within the neighborhood range is used as the break signal; The difference between the protection signal and the fracture signal is calculated and used as the window protection score.

6. The early endometrial cancer screening system based on multi-omics liquid biopsy according to claim 4, characterized in that, The feature mapping module is also configured to perform sequencing depth normalization when calculating the structural anomaly score: The average window protection score of the region outside the coordinates of the binding site is calculated as the background value; Calculate the relative drop ratio of the window protection score at the coordinates of the binding site relative to the background value; The structural anomaly score is obtained by multiplying the dynamic weight vector by the relative drop ratio.

7. The early endometrial cancer screening system based on multi-omics liquid biopsy according to claim 1, characterized in that, The data analysis unit includes a risk assessment module, which is configured to perform the following steps: The genome of the first sample component is divided into multiple observation windows. The entropy value of the fragment length distribution within each observation window is calculated, and the average entropy value of the whole genome is calculated as the whole genome fragment dispersion entropy value. Compare with preset risk assessment thresholds and system confidence intervals; When the structural anomaly score exceeds the high-risk threshold, the risk level is determined to be positive; When the structural abnormality score does not exceed the high-risk threshold, but the whole genome fragment dispersion entropy value is higher than the high-risk dispersion threshold, the risk level is determined to be gray zone pending investigation; Otherwise, the risk level is determined to be negative.

8. The early endometrial cancer screening system based on multi-omics liquid biopsy according to claim 1, characterized in that, The risk assessment module also incorporates cell-free DNA mutation frequency characteristics and copy number variation load characteristics when outputting the risk level. The risk assessment module uses a linear weighted fusion model to aggregate the structural anomaly score, the free DNA mutation frequency feature, and the copy number variation load feature into a comprehensive risk score, wherein the weight coefficient of the structural anomaly score is greater than the sum of the weight coefficients of the other features.

9. The early endometrial cancer screening system based on multi-omics liquid biopsy according to claim 1, characterized in that, The sample processing unit is configured to perform a two-stage centrifugation separation on the peripheral blood sample. The first stage centrifugation separates the plasma, and the second stage centrifugation removes cell debris from the plasma and physically divides the plasma into the first sample component and the second sample component. The protein detection unit employs a digital immunoassay platform based on single-molecule array technology, and the internal reference protein is either CD63 protein or tumor susceptibility gene 101 protein.

10. The early endometrial cancer screening system based on multi-omics liquid biopsy according to claim 1, characterized in that, The gene sequencing unit uses a PCR-free library preparation strategy to construct sequencing libraries and performs paired-end whole-genome sequencing, with the sequencing depth threshold set to be no less than 30-fold.