Full exome burden and serum marker-based method and system for evaluating risk of miscarriage
Patent Information
- Application Number
- CN202610615406.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-07
- Publication Date
- 2026-08-21
- Estimated Expiration
- 2046-05-07
AI Technical Summary
现有临床筛查手段存在显著局限:染色体核型分析仅能检出大片段染色体异常,无法捕捉罕见单核苷酸变异及小片段插缺对胚胎发育通路的累积扰动;针对抗磷脂综合征、甲状腺功能异常等的单项血清检测仅覆盖有限的病理机制,且各项检测结果彼此独立、缺乏整合建模框架,难以实现多维度生物学信号的协同判读;传统候选基因检测方法依赖先验致病基因列表,在面对遗传异质性极高的复发性流产时,漏检率高且可解释性不足
1.本发明提供了全外显子负荷与血清标志物的流产风险评估方法及系统,构建了四维变异综合致病评分体系,将群体频率、临床数据库收录状态、功能危害性预测及表型匹配度四个维度的信息整合为单一连续评分,并在此基础上以基因为单位进行致病负荷聚合,将WES数据中分散、异质的罕见变异信号转化为连续数值型特征,从根本上解决了单点变异分析在小样本稀疏数据场景下统计功效不足的问题,使原本难以检出的弱效罕见变异信号得以在基因及通路层级累积呈现。
Smart Images

Figure CN122177471B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of bioinformatics and precision reproductive medicine, specifically to a method and system for assessing miscarriage risk using whole exon load and serum biomarkers. Background Technology
[0002] Spontaneous abortion, especially recurrent abortion, is a major challenge in clinical reproductive medicine, accounting for approximately 10% to 15% of clinically confirmed pregnancies. Of these, recurrent abortions with normal karyotypes account for over 50% of all cases. The etiology involves the interaction of multiple systems, including genetic susceptibility, abnormal immune regulation, and endocrine disorders, with a single cause explaining less than 30% of cases. Current clinical screening methods have significant limitations: chromosomal karyotype analysis can only detect large chromosomal abnormalities and cannot capture the cumulative disturbances to embryonic development pathways caused by rare single nucleotide variants and small insertions / deletions; single serum tests for antiphospholipid syndrome, thyroid dysfunction, etc., only cover limited pathological mechanisms, and the results of each test are independent and lack an integrated modeling framework, making it difficult to achieve collaborative interpretation of multidimensional biological signals; traditional candidate gene detection methods rely on prior lists of pathogenic genes, resulting in high false negative rates and insufficient interpretability when facing highly genetically heterogeneous recurrent abortions.
[0003] With the widespread adoption of whole-exome sequencing (WES) technology, unbiased capture sequencing of whole-genome protein-coding regions has become an important tool for the clinical diagnosis of genetic diseases. However, existing WES analysis workflows face three core technical bottlenecks when dealing with the complex and multifactorial clinical phenotype of recurrent miscarriage. First, pathogenic variants associated with miscarriage in WES data are mostly rare, with weak effect sizes and high heterogeneity among samples. Traditional single-point variant association analysis severely lacks statistical power in small clinical cohorts, making it difficult to reliably detect signals. Second, existing gene load analysis methods typically treat genes as independent statistical units, ignoring the synergistic pathogenic mechanisms at the biological pathway level. This fails to integrate scattered single-gene signals into biologically meaningful pathway-level functional perturbation features, resulting in poor model interpretability and limited generalization ability. Third, there is a lack of a unified probabilistic modeling framework for genetic data and serum immune metabolic detection data. Existing methods either simply splice the two together and model them in segments, or use only a single mic data set for prediction, failing to achieve mutual constraints and synergistic inference between two types of heterogeneous signals within the same decision space.
[0004] To address the aforementioned technological gaps, while existing methodologies include double-ridge penalty frameworks such as PHARAOH for pathway-level genetic burden modeling, their penalty mechanisms are based on linear model designs. When directly transferred to the binary classification of miscarriage, the collinearity generated by shared genes between pathways is nonlinearly amplified in the nonlinear space of logit links, leading to increased variance in model coefficient estimates and a significant decrease in cross-validation stability, failing to meet clinical-level stability requirements. Therefore, developing a method that can transform sparse and rare WES variant signals into pathway-level functional perturbation features and co-model them with conventional serum detection indicators within a unified probabilistic framework is an urgent technological need for achieving high-precision individualized risk assessment of recurrent miscarriage with normal karyotypes. Summary of the Invention
[0005] To overcome the shortcomings of existing technologies, this invention proposes a method for assessing miscarriage risk based on whole-exome load and serum biomarkers. This method outputs probability values and probability stratification labels related to miscarriage risk from biological sample data, and includes the following steps: S1: Each variant in the whole exome sequencing data is independently scored from at least two dimensions to obtain a comprehensive pathogenicity score; S2: Calculate the comprehensive pathogenicity score of variations within the same gene region according to the preset aggregation rules to construct a gene pathogenicity load matrix; S3: Map the gene pathogenicity load to biological pathways to obtain a comprehensive pathway score, and construct a generalized linear model with logit as the link function; In the loss function of the generalized linear model, a two-ridge penalty term is set for gene weights and pathway coefficients. The two-ridge penalty term is weighted by the diagonal elements of the Fisher information matrix corresponding to each parameter in the current iteration. The diagonal elements of the Fisher information matrix are dynamically recalculated as the model output probability is updated during the iteration process, so that the penalty intensity is adaptively adjusted with the probability curvature under the logit link. The model is solved by jointly optimizing the regularization strength parameters of the two-layer ridge penalty terms, and significant pathway features are output. S4: Standardize and preprocess serum test data; S5: The significant pathway features output from step S3 and the standardized serum data output from step S4 are jointly modeled under the same logistic regression probability framework to output the probability value of miscarriage-related risks, and the samples are probabilistically stratified based on the probability reference cutoff value determined on the training set.
[0006] Furthermore, the diagonal elements of the Fisher information matrix described in step S3 are the variance factors of the logit model. As the core structure, in which The variance factor represents the output probability of sample j in the current iteration of the model. The variance factor reduces the contribution of samples with output probabilities close to 0 or 1 to the penalty term, and strengthens the contribution of samples with output probabilities close to 0.5 to the penalty term.
[0007] Furthermore, the loss function described in step S3 has the following form: ; or ; in For genes The weight parameters, Let be the regression coefficient of pathway k. and They are respectively and The corresponding diagonal elements of the Fisher information matrix, and For two sets of regularization intensity parameters, the and Determined through K-fold cross-validation and joint grid search.
[0008] Furthermore, the multi-dimensional scoring in step S1 specifically involves: independently scoring from four dimensions—population frequency, clinical database inclusion status, functional hazard prediction, and phenotypic matching degree—and then adding the four scores together to obtain the comprehensive pathogenicity score.
[0009] Furthermore, in step S2, only variations that meet any of the following conditions are retained to participate in intragene aggregation: Located within the vicinity of exon regions or splice regions as defined by standard gene annotation; Located within an intron region but marked as pathogenic or potentially pathogenic in clinical databases; The gene in question has a clear Mendelian genetic disease association record in the OMIM database.
[0010] Furthermore, the standardization preprocessing described in step S4 includes sequentially performing missing value imputation, extreme value truncation, and Z-score standardization on the serum test data; The Z-score standardized mean and standard deviation are calculated based solely on the training set samples and stored as fixed parameters.
[0011] Furthermore, the probability reference threshold value mentioned in step S5 is determined by maximizing the Youden index on the training set; samples with a risk probability higher than the threshold value are classified into the high probability group, and samples with a risk probability not higher than the threshold value are classified into the low probability group.
[0012] Furthermore, the method also includes outputting a dual-track coordination coefficient, which is defined as the cosine of the angle between the genetic pathway weighted contribution vector and the serum weighted contribution vector, used to quantify the directional consistency between genetic signals and serum signals.
[0013] Furthermore, the method also includes a new sample inference step: persistently storing the parameter set determined after training in the form of fixed parameters; For new input samples, the risk probability is calculated using the fixed set of parameters without re-performing step S3 parameter optimization, and the probability stratification results are output.
[0014] A whole-exome load and serum biomarker miscarriage risk assessment system, used to implement the above-described method, comprising: The variant multidimensional pathogenicity scoring module is used to assign multidimensional scores to each variant and generate a comprehensive pathogenicity score; The gene pathogenicity load aggregation module is used to aggregate the comprehensive pathogenicity scores of variants by gene interval and generate a gene pathogenicity load matrix. The pathway-level adaptive double ridge penalty modeling module is used to construct and solve the Fisher information-weighted double ridge penalty logit generalized linear model described in step S3, and output significant pathway features. The serum biomarker standardization and integration module is used for standardized preprocessing of serum test data; The genetic-serum dual-track joint probability calculation and sample stratification module is used to jointly model significant pathway features and standardized serum data under the same logistic regression probability framework, output risk probability values, and perform probability stratification of samples.
[0015] Compared with the prior art, the beneficial effects of the present invention are: 1. This invention provides a method and system for assessing miscarriage risk using whole-exome burden and serum biomarkers. It constructs a four-dimensional variant comprehensive pathogenicity scoring system, integrating information from four dimensions—population frequency, clinical database inclusion status, functional hazard prediction, and phenotypic matching degree—into a single continuous score. Based on this, pathogenicity burden is aggregated at the gene level, transforming scattered and heterogeneous rare variant signals in WES data into continuous numerical features. This fundamentally solves the problem of insufficient statistical power of single-point variant analysis in small-sample sparse data scenarios, enabling weak and rare variant signals that were originally difficult to detect to accumulate and appear at the gene and pathway levels.
[0016] 2. This invention provides a method and system for assessing miscarriage risk based on whole-exome load and serum biomarkers. Through a pathway-level adaptive double-ridge penalty modeling mechanism, two layers of regularization are applied simultaneously to the loss function of the logit-linked generalized linear model: gene weight ridge penalty and pathway coefficient ridge penalty. The penalty terms are dynamically weighted and scaled using the diagonal elements of the Fisher information matrix for each parameter at the current iteration step, allowing the penalty intensity to adaptively adjust with probability curvature. This mechanism overcomes the coefficient estimation bias caused by curvature changes in the logit nonlinear space when the penalty intensity is fixed, effectively alleviating the high collinearity problem caused by shared genes between pathways. Compared to the existing PHARAOH linear double-ridge scheme directly transferred to binary classification scenarios, this invention improves model stability by approximately 62% (cross-validation standard deviation decreased from 0.13 to 0.05) and AUC increased from 0.79 to 0.91 in 5-fold cross-validation of a 36-case validation cohort, demonstrating that this mechanism has significant and unexpected technical advantages within the logit nonlinear framework.
[0017] 3. This invention provides a method and system for assessing miscarriage risk using whole-exome load and serum biomarkers. It jointly models seven serum biomarkers and genetic pathway characteristics within the same logistic regression probability framework, enabling genetic susceptibility signals and immune metabolic status signals to mutually constrain and synergistically participate in coefficient estimation, rather than simply splicing and segmenting the data. Validation experiments show that the combined model's AUC (0.91) is significantly better than the single model using only WES genetic data (AUC=0.72) and the single model using only the seven serum biomarkers (AUC=0.76), demonstrating that the synergistic modeling of these two heterogeneous signals within a unified framework produces a synergistic gain effect that surpasses their individual predictive capabilities.
[0018] 4. This invention provides a method and system for assessing miscarriage risk using whole-exome burden and serum biomarkers. The output dual-track synergy coefficient C quantifies the consistency between genetic pathway signals and serum detection signals in the direction of probability judgment. It can distinguish between "genetically dominant" and "serum-dominant" risk samples at the individual level, providing operable data for the precise selection of clinical intervention pathways and significantly improving the model's biological interpretability and clinical applicability. Furthermore, the model parameter set supports fixed storage and independent deployment; new sample inferences do not require retraining, ensuring the comparability of probability values for different batches of samples and possessing good accessibility for large-scale clinical applications. Attached Figure Description
[0019] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0020] Figure 1 This is a schematic diagram of the overall process of risk assessment methodology; Figure 2 This is a schematic diagram of a multidimensional pathogenicity scoring system. Figure 3 This is a schematic diagram of adaptive double ridge penalty modeling at the pathway level; Figure 4 This is a schematic diagram of genetic-serum dual-track probability calculation and sample stratification. Detailed Implementation
[0021] The technical solution of the present invention will be more clearly and completely explained below with reference to the accompanying drawings and through the description of preferred embodiments of the present invention.
[0022] Terminology Explanation: WES: Whole exome sequencing. VCF: Variant call format. NGS: Next-generation sequencing.
[0023] SNV: Single nucleotide variant. Indel: Insertion / deletion variant. PE150: Paired-end 150 bp sequencing read.
[0024] BWA-MEM: Genome sequence alignment algorithm. GATK: Genome analysis toolkit.
[0025] HPO: Human Phenotypic Ontology. gnomAD: Genome Aggregation Database. HGMD: Human Genetic Mutation Database. DM is its pathogenicity marker, indicating a confirmed pathogenic mutation; DM? indicates a suspected pathogenic mutation.
[0026] ClinVar: Clinical Variability Database. OMIM: Online Database of Human Mendelian Genetics.
[0027] ACMG: American Association for Medical Genetics and Genomics. ANNOVAR: A gene variation function annotation tool.
[0028] VEP: A tool for predicting variant effects. RefSeq: A database of reference sequences maintained by the National Center for Biotechnology Information (NCBI).
[0029] Ensembl: A genome database maintained by the European Institute for Bioinformatics.
[0030] ExAC: Exome Aggregation Consortium, used for population frequency assessment of variants, is now integrated into gnomAD.
[0031] ESP6500: The National Heart, Lung, and Blood Institute's exome sequencing project, which includes exome data from approximately 6,500 samples.
[0032] InterVar: An automated pathogenicity interpretation tool based on ACMG / AMP guidelines.
[0033] BED: Standard file format for genomic region coordinates. STR: Short tandem repeat sequence.
[0034] SIFT: A tool for predicting the impact of amino acid substitutions on protein function based on sequence conservation.
[0035] PolyPhen-2: A tool for predicting the functional impact of missense mutations using structural and evolutionary information.
[0036] MutationTaster: A tool that integrates multiple sequence features to predict the pathogenicity of variants.
[0037] M-CAP: A tool for predicting the pathogenicity of rare missense variants designed specifically for Mendelian genetic diseases.
[0038] REVEL: A tool for predicting the pathogenicity of rare missense variants based on ensemble learning methods.
[0039] AlphaMissense: A tool developed by DeepMind for predicting the pathogenicity of missense variants based on deep learning of protein structure.
[0040] dPsi: A scoring tool for predicting the impact of intron mutations on splicing efficiency.
[0041] dbscSNV RF Databases and tools for predicting SNV-affected splice sites using the random forest algorithm.
[0042] KEGG: Kyoto Encyclopedia of Genes and Genomes, providing a systematic database of biological pathways, metabolic networks, and gene function annotations.
[0043] BioCarta: A biological pathway database that maps human signal transduction and metabolic pathways.
[0044] LASSO: Minimal absolute contraction and selection operator. PHARAOH: A double-ridge penalized linear modeling framework for genetic load association analysis based on pathway hierarchy.
[0045] AUC: Area under the receiver operating characteristic curve, a statistical indicator used to evaluate the overall discriminative ability of a binary classification model. The closer the value is to 1, the better the model performance.
[0046] ROC: Receiver Operating Characteristic curve, plotted as sensitivity versus 1-specificity, used to evaluate the performance of classification models.
[0047] CV: Cross-validation. ANA: Antinuclear antibody. ANA-M: Antinuclear antibody IgM subtype.
[0048] ANA-NT: Antinuclear antibody karyotype. HbA1C: Glycated hemoglobin. FRT4: Ferritin. TOTT4: Total thyroxine.
[0049] CMA: Chromosome microarray analysis. ELISA: Enzyme-linked immunosorbent assay. OR: Odds ratio. CI: Confidence interval.
[0050] like Figure 1 As shown, the overall process of the miscarriage risk assessment method of this invention consists of two parallel input tracks: genetic data and serum data, and six sequentially executed processing steps. The genetic data input track includes whole-exome sequencing data (e.g., VCF files) and subject phenotypic information (e.g., HPO phenotypic coding). After step S1 (variate pathogenicity scoring), step S2 (gene load aggregation), and step S3 (pathway modeling), significant pathway features are output. The serum data input track includes M serum test values. After standardization and integration in step S4, a standardized serum feature vector is output. In step S5, the features from both tracks are integrated into the same logistic regression probability framework for joint modeling, outputting the miscarriage-related risk probability and probability stratification results for each sample. Step S6 is a new sample inference module, which uses the parameter set fixed after training to directly calculate the risk probability for any new input sample without re-performing parameter optimization. The system's final output includes four types of results: sample probability stratification (high probability group and low probability group), regression coefficients and odds ratios of each significant pathway, regression coefficients and contribution rankings of each serum biomarker, and dual-track synergistic coefficients and genetic / serum dominant typing that quantify the consistency of the two types of signal directions.
[0051] In practical applications, the etiology of spontaneous abortion (especially recurrent abortion) is complex, involving the interaction of multiple systems such as genetics, immunity, and endocrine function. Traditional screening methods are insufficient to comprehensively capture the potential susceptibility background of the mother. To address this, the inventors of this application, through a series of inventive efforts, have proposed a method and system for calculating the probability of miscarriage risk by combining whole-exome sequencing (WES) gene pathogenicity burden with serum biomarkers. This method achieves high-precision and highly interpretable individualized risk stratification through an end-to-end process: multi-dimensional variation scoring → gene burden aggregation → pathway-level Fisher weighted double regularization modeling → standardized integration of serum biomarkers → combined genetic-serum probability calculation.
[0052] This method uses whole-exome sequencing data, subject phenotypic information, and serum test values as input data, and performs the following five steps sequentially. The following describes the specific implementation process of each step in detail with an exemplary implementation method (using VCF files as the form of whole-exome sequencing data and HPO phenotypic codes as the form of subject phenotypic information).
[0053] Step S1 is used to perform a refined pathogenicity assessment of each variant detected in the whole-exome sequencing data. The whole-exome sequencing data can be any variant list data that conforms to industry standards. In a specific embodiment, a VCF file (containing SNV / Indel) processed by the GATK Best Practices workflow is used as an example. The subject phenotypic information can be any phenotypic description data that conforms to the standard phenotypic coding system. In a specific embodiment, a sample clinical phenotype in HPO coding format is used as an example.
[0054] Functional annotation was performed using ANNOVAR or VEP tools, integrating the following database information: gene structure was derived from the RefSeq / Ensembl gene model; population frequencies were derived from gnomAD (EAS, NFE, AFR, etc. subgroups), 1000 Genomes, ExAC, and esp6500; clinical significance was derived from ClinVar (Pathogenic / Likely Pathogenic, etc.), HGMD (DM / DM tag), and InterVar (ACMG classification); functional prediction was integrated from SIFT, PolyPhen-2, MutationTaster, M-CAP, REVEL, AlphaMissense, dPsi (spidex), and dbscSNV. RF There are a total of 8 tools.
[0055] The four-dimensional scoring logic is as follows: Population frequency scoring is based on variations in gnomAD and 1000. The maximum allele frequency in each subgroup of the Genomes database is assigned a reverse score: 0 = 100 points, greater than 0 but not exceeding 0.001 = 80 points, greater than 0.001 but not exceeding 0.005 = 60 points, greater than 0.005 but not exceeding 0.01 = 40 points, greater than 0.01 but not exceeding 0.05 = 20 points, and greater than 0.05 = 0 points. Clinical database inclusion scores are calculated as follows: HGMD and ClinVar double positive (DM and Pathogenic) → 100 points, only one of these criteria or presence in the internal positive database → 80 points, DM? or LikelyPathogenic → 40 points, and all others → 0 points. Functional hazard scoring adds 10 points for each of the eight prediction tools met, up to a maximum of 80 points. Phenotypic matching scoring uses the Phenolyzer algorithm to calculate the semantic similarity between the subject's HPO-coded clinical phenotype and genes, multiplying the original matching score by 100 and rounding down to the nearest integer, ranging from 0 to 100 points. The output of step S1 is a comprehensive pathogenicity score Score(v) for each variant, which is equal to the sum of the scores of the four items mentioned above, ranging from 0 to 360.
[0056] like Figure 2 As shown, each variant record in the WES VCF file is scored for pathogenicity across four independent dimensions. The variant record serves as the input node, leading to four scoring modules: ① Population frequency scoring module: Scoring is based on allele frequencies in the gnomAD and 1000 Genomes databases, with lower frequencies receiving higher scores. Six grading intervals are set, with a maximum score of 100. ② Clinical database scoring module: Scoring is based on the inclusion status of HGMD and ClinVar. Pathogenicity markers from both databases receive 100 points, for a maximum score of 100. ③ Functional hazard scoring module: Integrating eight prediction tools, including SIFT, PolyPhen-2, and REVEL, each reaching a pathogenicity threshold adds 10 points, for a maximum score of 80. ④ Phenotypic matching scoring module: Semantic similarity calculation between the subject's HPO code and genes is performed using the Phenolyzer algorithm, with a maximum score of 100. The above four scoring results are incorporated into the bottom comprehensive scoring node. The four scores are added together to obtain the comprehensive pathogenicity score, which ranges from 0 to 360 points. This score serves as the input data for gene load aggregation in step S2.
[0057] Step S2 involves gene pathogenicity load aggregation. The comprehensive pathogenicity score (Score(v)) of all variants within the same gene region that passed the validity filter is arithmetically summed to form a continuous numerical feature at the gene level, constructing a sample × gene sparse load matrix. Gene regions are defined using the RefSeq standard gene BED file (including ±20 bp splice regions). Validity filtering retains only variants meeting any of the following conditions for aggregation: located within 20 bp upstream or downstream of an exon or splice region defined by RefSeq or Ensembl standard gene annotations; located within an intron region but already marked as DM by HGMD or Pathogenic by ClinVar; or the gene containing the variant has a clear Mendelian genetic disease association record in the OMIM database. The sample at the th The pathogenicity burden on each gene is calculated using the following formula: ; The summation range is all variants v in sample j that have passed the above validity filtering and are located within gene i. Score(v) is the comprehensive pathogenicity score of variant v. Matrix construction generates an N×G sparse matrix (N is the number of samples, G≈20,000 is the number of protein-coding genes), and missing values in the loading matrix are uniformly filled with 0; then, intra-column Z-score standardization is applied to each column (gene) to eliminate the difference in basic variation rate between genes.
[0058] Step S3 involves adaptive double-ridge penalty modeling at the pathway hierarchy. Gene pathogenicity load is mapped to a comprehensive pathway score based on the biological pathway database. A generalized linear model is constructed using a two-level regularization weighted by the Fisher information matrix to extract biologically significant risk pathway features. The pathway database uses KEGG (v2023) and BioCarta (v2022) pathway annotation files to construct a gene-pathway membership matrix. For the k-th pathway, its comprehensive pathway score on sample j is calculated. Defined as the sum of the products of the pathogenicity load of all genes within the pathway and their corresponding weight parameters: ; Using the scores of all pathways as input variables, a generalized linear model of the logit link function is established: ; The model loss function introduces two layers of regularization terms based on the negative log-likelihood. The first layer is a gene weight ridge penalty to suppress functional redundancy within pathways, and the second layer is a pathway coefficient ridge penalty to alleviate collinearity between pathways. The two penalty terms are weighted and scaled using the diagonal elements of the Fisher information matrix corresponding to each parameter in the current iteration step, allowing the penalty strength to dynamically and adaptively adjust with the probability curvature. This overcomes the coefficient estimation bias caused by curvature changes in the logit nonlinear space when the penalty strength is fixed. The loss function is defined as: ; in: Let j be the binary data label for the pregnancy outcome of the j-th sample; This represents the probability value output by the current iteration of the logit model. For genes Weight parameters within its respective pathway; The regression coefficient for pathway k; Gene weights The corresponding diagonal elements of the Fisher information matrix are explicitly defined as follows: ;
[0059] Reflecting the local uncertainty in gene weight estimation under the current probability curvature; the partial derivative term refers to about The partial derivatives, according to Can be explicitly expanded as (When gene i belongs to pathway k) or 0 (when gene i does not belong to pathway k); Pathway coefficient The corresponding diagonal elements of the Fisher information matrix are explicitly defined as follows: ; This reflects the local uncertainty in the estimation of path coefficients under the current probability curvature.
[0060] and During the iteration process The system dynamically recalculates each parameter update, ensuring that the penalty intensity is adaptively adjusted based on the current probability curvature with each parameter update.
[0061] Two sets of regularization strength parameters and 5-fold cross-validation in 10 -2 Up to 10 6 The optimal parameter combination is determined by joint grid search optimization within the range, with the goal of minimizing the negative log-likelihood loss of the training set. Retain regression coefficients Non-zero pathways are considered significant risk pathways. Typically, 10 to 50 significant pathways are selected, and the statistical significance of each pathway can be numerically verified using a permutation test. It should be noted that the pathway regression coefficients in step S3... This is only used for the screening of significant pathway features within this step; its function is to identify the set S of significant pathway indices with non-zero coefficients through optimization under the Fisher information-weighted double ridge penalty constraint. After the significant pathway screening is completed in step S3... That is, they will no longer participate in subsequent modeling. Step S5 will introduce serum biomarkers onto the significant pathway subset S and re-estimate the regression coefficients of the pathway characteristics within the same logistic regression probability framework. ;therefore and These are two different sets of parameters belonging to the two independent estimation processes of step S3 and step S5, respectively. They are generally different in value and should not be confused with each other.
[0062] like Figure 3 As shown, the input consists of a column-standardized sample × gene sparse load matrix. The pathway mapping layer aggregates the gene loads into a comprehensive score for each pathway based on the KEGG and BioCarta databases, covering a total of 186 pathways. Example pathways include complement and coagulation cascades, thyroid hormone synthesis, and ferroptosis regulation. The pathway scores then enter the double-ridge penalty loss function module: the first layer is a gene weight ridge penalty, which applies to the weight parameters of each gene within the pathway to suppress the interference of functionally redundant genes within the pathway on the model coefficient estimation; the second layer is a pathway coefficient ridge penalty, which applies to the regression coefficients of each pathway to alleviate collinearity caused by shared genes between pathways. Both penalty terms are weighted and scaled using the diagonal elements of the Fisher information matrix, as shown by the dashed box in the figure. This mechanism dynamically recalculates the weighting factor based on the current probability output during each parameter iteration, allowing the penalty intensity to adaptively adjust with the probability curvature, overcoming the coefficient estimation bias in the logit nonlinear space with a fixed penalty intensity. Finally, the two sets of regularization parameters are jointly optimized through 5-fold cross-validation, and the pathways with non-zero regression coefficients are retained as significant pathway features and output to step S5.
[0063] Step S4 involves the standardized integration of serum biomarkers, preprocessing the serum test data to ensure consistency with genetic characteristic scales. The serum test data can be any one or more routine clinical serological indicators covering apoptosis, autoimmunity, glucose metabolism, iron metabolism, thyroid function, and other biological dimensions related to miscarriage. The number and specific combination of these indicators can be flexibly selected according to the clinical scenario without affecting the overall implementation of the method and system of this invention. As a specific embodiment, seven serum test data items—nucleosomes (NUCLEOSOME), antinuclear antibodies (ANA), antinuclear antibody IgM (ANA-M), antinuclear antibody karyotype (ANA-NT), glycated hemoglobin (HbA1C), ferritin (FRT4), and total thyroxine (TOTT4)—are used as exemplary inputs, covering five detection dimensions: apoptosis-related indicators, autoimmunity-related indicators, glucose metabolism-related indicators, iron metabolism-related indicators, and thyroid function-related indicators. In other embodiments, the number M of the serum test data items can be any integer not less than one. Specific items can also be flexibly selected from the above five biological dimensions or expanded with other relevant serological indicators based on clinical data accessibility. The joint probability formula in step S5 adaptively adjusts the upper limit of the summation according to the number of input items M to ensure that the system can output effective probability values under different serum input configurations. The processing steps are as follows: missing values are filled with the mean of the training set; 99th percentile Winsorization is performed on each indicator to suppress extreme outliers; Z-score standardization is applied. ; in and The calculations are performed only on the training set samples and stored with fixed parameters. When standardizing new input samples, these fixed training set parameters are used instead of being recalculated to ensure the comparability of output probability values for different batches of input samples. The output is an M-dimensional standardized vector z. jm (j is the sample number, m=1,...,M), and is used in conjunction with pathway features for modeling in step S5.
[0064] Step S5 involves combined genetic and serum probabilities calculation and sample stratification, fusing genetic and serum characteristics to output the miscarriage-related risk probability value for each sample. The significant pathway features selected in step S3 (k∈S) and the M serum biomarkers z after standardization in step S4 jm Joint modeling within the same logistic regression probabilistic framework allows for mutual constraint and collaborative computation of genetic load signals and serum detection signals. The joint probability calculation formula is as follows: ; in: is the intercept term; S is the set of significant pathway indices whose regression coefficients are not zero after screening in step S3; The overall score of sample j on the k-th pathway is equal to the sum of the products of the pathogenicity load of each gene in that pathway and its weight parameter; z jm Let m be the standardized value of the m-th serum test data of sample j; and The regression coefficients for pathway characteristics and serum data, respectively, are used together in coefficient estimation within the same probabilistic framework, ensuring that the two types of input signals are mutually constrained rather than simply spliced together and modeled in segments. A probability reference threshold is determined on the training set based on the Youden index. The Youden index is defined as: ; Take The largest τ value is used as ; Higher than The samples were assigned to the high-probability group. Not higher than The samples are assigned to the low-probability group. Step S5 simultaneously outputs the following dual-track reference data: the genetic pathway data track outputs the regression coefficients of each significant pathway. And the estimated advantage ratio, sorted in descending order of contribution, for Higher than The samples were further labeled with the top three contributing pathways and their core genes; the serum data track output the regression coefficients of seven serum biomarkers. And standardized relative contribution ranking; the dual-track synergy coefficient C quantifies the synergistic effect of genetic pathway signals and serum detection signals in the direction of probability calculation, and is defined as the pathway-weighted contribution vector. With serum-weighted contribution vector The cosine value of the included angle: ; Where the j-th component of u is , The j-th component is When C is greater than a preset threshold and the two types of signals are in the same direction, a positive co-signal is output, indicating that genetic susceptibility signals and immune-metabolic signals mutually reinforce each other in the probability judgment of the same batch of samples. The above dual-track reference data and probability stratification results together constitute the complete output of the system, which is used to support the differentiation of data features between high-probability samples and low-probability samples in the dimensions of genetic burden and serum detection, and to provide actionable intervention clues for clinical practice.
[0065] like Figure 4As shown, the genetic pathway feature track from step S3 and the serum biomarker feature track from step S4 are combined and jointly modeled within the same logistic regression framework. This allows the two types of signals to mutually constrain each other and collaboratively participate in coefficient estimation, rather than simply splicing them together for segmented modeling. After joint modeling, the optimal probability cutoff value τ* is determined on the training set by maximizing the Youden index (the sum of sensitivity and specificity minus 1). Based on this, samples are divided into a high-probability group (risk probability higher than τ*) and a low-probability group (risk probability not higher than τ*). For samples classified into the high-probability group, the system further outputs the top three pathways that contribute the most to the probability of that sample and their core gene annotations, providing actionable genetic clues for clinical intervention. In addition, step S5 calculates and outputs the dual-track synergy coefficient C, as shown in the dotted box below the figure: when C is higher than the preset threshold, a positive synergy label is output, indicating that the genetic susceptibility signal and the serum immune metabolism signal are highly consistent in their probability judgment direction for the same batch of samples; when C is lower than the preset threshold, the relative intensity of the two tracks is further compared, and the typing labels of "genetic dominance" or "serum dominance" are output respectively, providing a reference for the selection of individualized precision intervention paths.
[0066] Through the collaborative work of the above five steps, this invention constructs an end-to-end, multi-omics fusion, pathway-driven miscarriage risk probability calculation system. This system deeply integrates a multi-dimensional WES variant pathogenicity scoring mechanism, a gene-based pathogenicity burden aggregation strategy, a Fisher-weighted double regularization modeling framework based on biological pathway hierarchies, a standardized integration process for clinical serum immune-metabolic biomarkers, and a genetic-serum dual-track joint probabilistic decision-making mechanism for binary clinical outcomes. It can transform sparse, heterogeneous, and rare variant signals in the whole exome into biologically meaningful pathway-level functional perturbation features without relying on a priori lists of pathogenic genes or single-point mutation assumptions, and organically integrates these features with conventional serum detection indicators.
[0067] This invention abandons the limitation of treating genes as independent units in traditional gene load analysis, instead introducing a structured modeling approach based on pathway hierarchy. By applying Fisher information matrix weighted double-ridge regularization with both gene weights and pathway coefficients, it effectively alleviates the high collinearity problem caused by shared genes between pathways, thus stably extracting synergistic pathogenic signals associated with miscarriage phenotypes from high-dimensional sparse genetic data. Simultaneously, the system uses serum biomarkers as independent but complementary predictive dimensions, achieving joint inference of genetic susceptibility and immune-metabolic status within a unified probabilistic framework. The directional consistency of the two types of signals is quantified using a dual-track synergistic coefficient C, significantly improving the model's biological rationality and clinical interpretability. This method is not only applicable to risk assessment of recurrent miscarriage in patients with normal karyotypes, but its modular design can also be flexibly extended to risk modeling scenarios of complex pregnancy complications or multi-system diseases caused by the synergistic effects of multiple factors, providing a novel analytical paradigm for precision reproductive medicine that combines theoretical rigor with clinical applicability.
[0068] The features and performance of the present invention will be further described in detail below with reference to embodiments.
[0069] As a specific embodiment, in order to verify the effectiveness and clinical applicability of the method for calculating the probability of miscarriage risk by combining whole exon gene load and serum markers as described in this invention, the inventors constructed a retrospective case-control cohort and strictly followed the five steps defined in this invention to perform the full-process analysis.
[0070] This study included 36 women of childbearing age with a clear record of pregnancy outcomes, divided into two groups: the case group (n=18) consisted of subjects with a history of ≥2 ultrasound-confirmed early spontaneous abortions (gestation ≤10 weeks), and whose most recent miscarriage product was confirmed to have a normal karyotype (euploid) by chromosomal microarray analysis (CMA); the control group (n=18) consisted of healthy women who had successfully conceived to full term (≥37 weeks) and had no history of miscarriage. All subjects signed informed consent forms, and the study protocol was approved by the local ethics committee, complying with the ethical guidelines of the Declaration of Helsinki. Baseline characteristics of the two groups are shown in Table 1. Age (P=0.62) and BMI (P=0.41) were comparable, and parity and gravidity differed significantly between the two groups (P<0.001), meeting the study design expectations.
[0071] Table 1: Baseline Characteristics Statistics of the Sample Cohort in This Study
[0072] Regarding DNA source, the case group received chorionic villus tissue after miscarriage (STR testing ruled out maternal contamination), while the control group received peripheral blood. Sequencing was performed using the Illumina NovaSeq 6000 platform for whole-exome capture sequencing (AgilentSureSelect All Exon v7+UTR), with an average sequencing depth ≥100× and >95% target region coverage ≥20×. In terms of bioinformatics workflow, raw data underwent alignment (BWA-MEM), deduplication (Picard), and variant identification (HaplotypeCaller) using the GATK Best Practices workflow, outputting a standard VCF file as input data for step S1.
[0073] Fasting venous blood was collected from all subjects before pregnancy or during early pregnancy (≤8 weeks) to detect seven indicators: NUCLEOSOME, ANA, ANA-M, ANA-NT, HbA1C, FRT4, and TOTT4. The detection methods were chemiluminescence immunoassay or ELISA, and the results were recorded according to standard clinical procedures. This embodiment uses a complete seven-item serum input protocol (M=7), consistent with the standard configuration described in claim 1. In scenarios with limited data acquisition, this invention also supports five or six items (M≥5), with the upper limit of the summation in the joint probability formula in step S5 adjusted accordingly, while the system still outputs valid probability values. Missing values were filled with the mean of the training set, and extreme values (>99th percentile) underwent Winsorization, followed by Z-score standardization according to the protocol in step S4. ,in and The results were fixed and stored in the training set and then used for standardization in all batches. Serum biomarker detection results for some subjects (before standardization) are shown in Table 2.
[0074] Table 2: Examples of serum biomarker test results for some subjects (before standardization) ;
[0075] In the gene pathogenicity burden construction phase, for each sample, the WES variants were calculated using the four-dimensional scoring system defined in step S1 to obtain a comprehensive pathogenicity score (Score(v)): the population frequency score was based on the inverse assignment of population frequencies from gnomAD EAS; the clinical database inclusion score was calculated with HGMD+ClinVar double positivity receiving 100 points; the functional hazard score was accumulated based on the achievement of 8 functional prediction tools (maximum 80 points); the phenotype matching score was obtained by matching the "recurrent miscarriage" HPO term (HP:0005268) with genes using Phenolyzer, and linearly scaled to 0–100 points. Only variants located in RefSeq gene exons or ±20 bp splice regions, or those located in introns but marked as pathogenic by ClinVar, and whose genes had clear reproductive / developmental phenotypic records in OMIM, were retained for aggregation according to the formula in step S2. The final gene load matrix was 36×19,842, with missing values filled with 0 and columns standardized. Table 3 shows a comparison of pathogenic load scores on key genes between some case groups and the control group, where the class field T represents the case group (RPL) and N represents the control group.
[0076] Table 3: Comparison of pathogenicity burden scores on key genes between the case group and the control group (partial examples)
[0077] DYNC2H1, ALOX15, C3, and F5 belong to the "complement and coagulation cascade" pathway (KEGG: hsa04610), which was significantly enriched in the pathway analysis of this invention; THRB is a thyroid hormone receptor associated with TOTT4 abnormalities; HBA1 and SERPINA1 served as negative control genes, showing no significant differences between the two groups. The Mann-Whitney U test was performed on the above genes, and the statistical results are shown in Table 4.
[0078] Table 4: Statistical Test Results (Mann-Whitney U Test)
[0079] DYNC2H1, ALOX15, and C3 showed significantly high loads in the case group (P<0.05) and were retained for subsequent pathway modeling feature screening. F5, although belonging to the same pathway, did not reach statistical significance due to high-scoring variations in some controls (possibly benign polymorphism). No difference was observed in genes in the negative control group, validating the system's specificity. This invention successfully transformed scattered rare variations into biologically significant continuous load features through multi-dimensional scoring and gene aggregation. DYNC2H1 and ALOX15, as genes known to be associated with embryonic lethality, showed significantly high loads in the case group, while the control group showed almost zero loads. The core complement pathway gene C3 also showed a consistent increasing load trend, indicating that this invention can not only reproduce known pathogenic genes but also discover cooperative perturbation signals at the pathway level. In contrast, irrelevant genes (such as HBA1) showed no difference between the two groups, demonstrating the system's good signal-to-noise ratio and specificity.
[0080] In the pathway-level Fisher weighted double regularization modeling stage, 186 pathways from KEGG (v2023) were used as the pathway database; the model was set as a logit link generalized linear model, and the loss function was defined as in step S3. The optimal parameter combination is determined by minimizing the negative log-likelihood loss of the training set; feature selection preserves the estimated coefficients. The pathways were analyzed, and statistical significance was verified using the permutation test. A total of 19 significant pathways were identified (P<0.05, permutation test), including the complement and coagulation cascades, thyroid hormone synthesis, ferroptosis, and primary cilium assembly. Among these, the complement and coagulation cascades had a mean burden value 2.4 times higher in the case group than in the control group (p=0.003).
[0081] In the genetic-serum dual-track probability calculation stage, scores were calculated for 19 significant pathways. (k∈S, |S|=19) and 7 standardized serum biomarkers (M=7) The model is jointly constructed according to the formula in step S5, forming a total of 26-dimensional feature vectors.
[0082] Maximize the Youden exponent on the training set to determine the probability reference cutoff value. =0.41. In a 5-fold cross-validation of a cohort of 36 cases, the joint model of this invention had an AUC of 0.91 (95% CI: 0.84–0.97), sensitivity of 87.5%, specificity of 85.7%, and Youden index of 0.73.
[0083] Key contributing factors include: in terms of pathways, the regression coefficient of "complement and coagulation cascade" The corresponding OR was 4.2 (p = 0.008); in serum, the corresponding coefficient for ANA-M was... The corresponding OR was 3.8 (p=0.012), and the OR corresponding to an increase in NUCLEOSOME was 3.1 (p=0.021).
[0084] In the dual-track synergy coefficient calculation stage, according to the definition in step S5, the pathway-weighted contribution vector u and the serum-weighted contribution vector v are calculated for all 36 samples, and the synergy coefficient is calculated according to the formula. In this embodiment, the overall cohort synergy coefficient C=0.87, which is higher than the preset threshold of 0.6, and a positive synergy marker is output, indicating that the genetic susceptibility signal and the immune-metabolic state signal are highly consistent in their probability judgment directions for this cohort of samples. The dual-track synergy data of some representative samples are shown in Table 5.
[0085] Table 5: Calculation results of dual-track synergy coefficient for some samples
[0086] As shown in Table 5, the cooperability coefficients of high-probability case samples (such as NT231199 and NT241558) were all higher than 0.9, indicating that the genetic pathway load signal and serum immune-metabolic signal were aligned and mutually reinforcing. The cooperability coefficient of sample NT242215 was only 0.31, indicating weak cooperability, suggesting that this sample was dominated by genetic pathway load (high complement pathway load) while serum marker signals were not significant. Clinically, intervention targeting the genetic mechanism should be prioritized. Low-probability control samples (such as WES24120706) showed low signals in both tracks, with a positive cooperability marker in the low-probability direction, consistent with the final probability stratification results. The above analysis shows that the dual-track cooperability coefficient C is not only used to quantify the overall directional consistency of the two types of signals, but also to identify "genetically dominated" and "serum dominated" risk samples at the individual level, providing more granular data support for precise intervention.
[0087] To verify the technical advantages of the present invention over existing methods, the inventors conducted a multi-method comparison experiment within a 5-fold cross-validation framework using the same 36-case cohort. The results are shown in Table 6.
[0088] Table 6: Comparative experimental data of different methodological combinations (based on 5-fold cross-validation of a 36-case cohort)
[0089] The results show that only the WES model's performance decline (AUC=0.72) is consistent with the rare variant sparsity theory (low signal-to-noise ratio of single gene load); the serum model's AUC=0.76 is consistent with the reported predictive ability of ANA / HbA1C single indicators (AUC0.65–0.78); the PHARAOH direct transfer model has poor stability (CV standard deviation 0.13), which stems from the conflict between its linear assumption and the nonlinear nature of binary classification. The PHARAOH double ridge penalty was originally designed as a linear model, but direct transfer to a binary classification scenario leads to collinearity generated by shared genes between pathways being nonlinearly amplified under logit linkage, increasing the variance of model coefficient estimation, and significantly reducing cross-validation stability (standard deviation > 0.12).
[0090] This invention introduces a logit-linked Fisher information matrix diagonal element weighting mechanism into the loss function. and The explicit dynamic computation allows the double-ridge penalty to adapt to the probability curvature of the current iteration step, overcoming the coefficient estimation bias problem in the logit nonlinear space with a fixed penalty strength. Simultaneously, M serum biomarkers are jointly modeled with genetic pathway features under a unified probability framework with independent prediction dimensions, forming a dual-track constraint mechanism of "genetic susceptibility-immune metabolic state." The directional consistency of the two types of signals is quantified through a synergistic coefficient C, rather than simply splicing and segmenting the model. Experiments show that this improvement enhances model stability by approximately 62% (CV standard deviation 0.13 → 0.05), and the joint model's AUC (0.91) is significantly better than the single-mathematical model (WES 0.72, serum 0.76) and the PHARAOH direct migration scheme (0.79), demonstrating the unexpected synergistic technical effect of this invention.
[0091] Through the collaborative work of the above five steps, this invention constructs an end-to-end, multi-omics fusion, and highly interpretable system for calculating the probability of miscarriage risk. This system deeply integrates multi-dimensional WES variant pathogenicity scoring, gene pathogenicity burden aggregation, pathway-level Fisher weighted double-ridge regularization modeling, standardized integration of serum immune-metabolic biomarkers, and a genetic-serum dual-track probabilistic decision-making mechanism optimized based on the Youden index. It can efficiently utilize sparse and heterogeneous rare variant signals from the whole exome, combined with several clinically readily available serum indicators, without relying on a priori lists of pathogenic genes or single-point mutation assumptions, to achieve high-precision and high-specificity prediction of recurrent miscarriage with normal karyotypes. This system possesses strong biological interpretability (outputting significant risk pathways, contribution coefficients of each biomarker, and dual-track synergy coefficients), excellent statistical robustness (Fisher weighted double ridge effectively alleviates collinearity), high clinical applicability (requiring only routine WES and a few serum tests), and good generalization ability (modular design supports flexible replacement of scoring strategies, pathway databases, or combinations of serum biomarkers), providing a new technical approach for precision reproductive medicine that combines scientific rigor with practical value.
[0092] The above-described specific embodiments are merely preferred embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Various modifications, substitutions, and improvements made by those skilled in the art to the technical solutions of the present invention based on the provided textual description and drawings, without departing from the design concept and spirit of the present invention, should all fall within the scope of protection of the present invention. The scope of protection of the present invention is determined by the claims.
Claims
1. A method for assessing miscarriage risk based on whole-exon load and serum biomarkers, wherein the method is used to output probability values and probability stratification labels related to miscarriage risk from biological sample data, characterized in that, Includes the following steps: S1: Each variant in the whole exome sequencing data is independently scored from at least two dimensions to obtain a comprehensive pathogenicity score; S2: The comprehensive pathogenicity score of variations within the same gene region is calculated according to a preset aggregation rule to construct a gene pathogenicity load matrix; the preset aggregation rule is arithmetic summation; S3: Map the gene pathogenicity load to biological pathways to obtain a comprehensive pathway score, and construct a generalized linear model with logit as the link function; In the loss function of the generalized linear model, a two-ridge penalty term is set for gene weights and pathway coefficients. The two-ridge penalty term is weighted by the diagonal elements of the Fisher information matrix corresponding to each parameter in the current iteration. The diagonal elements of the Fisher information matrix are dynamically recalculated as the model output probability is updated during the iteration process, so that the penalty intensity is adaptively adjusted with the probability curvature under the logit link. The diagonal elements of the Fisher information matrix are based on the variance factor of the logit model. As the core structure, in which The variance factor represents the output probability of sample j in the current iteration of the model. The variance factor reduces the contribution of samples with output probabilities close to 0 or 1 to the penalty term, and strengthens the contribution of samples with output probabilities close to 0.5 to the penalty term. The model is solved by jointly optimizing the regularization strength parameters of the two-layer ridge penalty terms, and significant pathway features are output. The loss function has the following form: ; in, Let j be the binary data label for the pregnancy outcome of the j-th sample. For genes The weight parameters, Let be the regression coefficient of pathway k. and They are respectively and The corresponding diagonal elements of the Fisher information matrix: ; This represents the overall pathway score of sample j on the k-th pathway. ; The pathogenic burden value of the j-th sample on the i-th gene is calculated using the following formula: ; Wherein, the summation range is all variants v in sample j that are located within gene i and filtered through validity, and Score(v) is the comprehensive pathogenicity score of variant v; ; and For two sets of regularization intensity parameters, the and Determined through K-fold cross-validation and joint grid search; S4: Standardize and preprocess the serum test data, which includes seven serum test data: serum nucleosomes, antinuclear antibodies, antinuclear antibody IgM, antinuclear antibody karyotype, glycated hemoglobin, ferritin and total thyroxine. S5: The significant pathway features output from step S3 and the standardized serum data output from step S4 are jointly modeled under the same logistic regression probability framework to output the probability value of miscarriage-related risks, and the samples are probabilistically stratified based on the probability reference cutoff value determined on the training set.
2. The method according to claim 1, characterized in that, The independent scoring in step S1 specifically involves assigning scores independently from four dimensions: population frequency, clinical database inclusion status, functional hazard prediction, and phenotypic matching degree, and then adding the four scores together to obtain the comprehensive pathogenicity score.
3. The method according to claim 1, characterized in that, In step S2, only variants that meet any of the following conditions are retained to participate in intragene aggregation: Located within the vicinity of exon regions or splice regions as defined by standard gene annotation; Located within an intron region but marked as pathogenic or potentially pathogenic in clinical databases; The gene in question has a clear Mendelian genetic disease association record in the OMIM database.
4. The method according to claim 1, characterized in that, The standardization preprocessing described in step S4 includes sequentially performing missing value imputation, extreme value truncation, and Z-score standardization on the serum test data; The Z-score standardized mean and standard deviation are calculated based solely on the training set samples and stored as fixed parameters.
5. The method according to claim 1, characterized in that, The probability reference threshold mentioned in step S5 is determined by maximizing the Youden index on the training set; samples with a risk probability higher than the threshold are classified into the high probability group, and samples with a risk probability not higher than the threshold are classified into the low probability group.
6. The method according to claim 1, characterized in that, The method also includes outputting a dual-track coordination coefficient, which is defined as the cosine of the angle between the genetic pathway weighted contribution vector and the serum weighted contribution vector, and is used to quantify the directional consistency between genetic signals and serum signals.
7. The method according to any one of claims 1-6, characterized in that, The method also includes a new sample inference step: persistently storing the set of parameters determined after training in the form of fixed parameters; For new input samples, the risk probability is calculated using the stored parameter set without re-performing step S3 parameter optimization, and the probability stratification results are output.
8. A miscarriage risk assessment system based on whole-exon load and serum biomarkers, used to implement the method according to any one of claims 1-7, characterized in that, include: The variant multidimensional pathogenicity scoring module is used to assign multidimensional scores to each variant and generate a comprehensive pathogenicity score; The gene pathogenicity load aggregation module is used to aggregate the comprehensive pathogenicity scores of variants by gene interval and generate a gene pathogenicity load matrix. The pathway-level adaptive double ridge penalty modeling module is used to construct and solve the Fisher information-weighted double ridge penalty logit generalized linear model described in step S3, and output significant pathway features. The serum biomarker standardization and integration module is used for standardized preprocessing of serum test data; The genetic-serum dual-track joint probability calculation and sample stratification module is used to jointly model significant pathway features and standardized serum data under the same logistic regression probability framework, output risk probability values, and perform probability stratification of samples.
Citation Information
Patent Citations
Multi-gene risk assessment disease prediction model and establishment method and application thereof
CN118538424A
Prediction model for curative effect prognosis of hepatocellular carcinoma TACE combined with TKIs
CN118888139A