Newborn rare disease intelligent screening and diagnosis system based on multi-omics data fusion
By constructing an intelligent screening and diagnostic system for rare neonatal diseases that integrates multi-omics data, and using Logistic regression and random forest algorithms to build a correlation model, the system solves the problems of long diagnostic time and low accuracy in traditional diagnosis, and achieves rapid and accurate diagnosis of rare neonatal diseases.
Patent Information
- Application Number
- CN202511087921.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-05
- Publication Date
- 2025-11-21
AI Technical Summary
Traditional methods for diagnosing rare neonatal diseases are time-consuming and have low accuracy, making it difficult to meet the needs for rapid and accurate diagnosis. The lack of multi-omics data fusion models also limits the accuracy of diagnosis.
A newborn rare disease intelligent screening and diagnosis system based on multi-omics data fusion was constructed, including data acquisition, data processing and diagnosis modules. Logistic regression and random forest algorithms were used to construct a clinical omics-gene metabolism correlation model to achieve cross-dimensional correlation analysis, and diagnosis was carried out by dynamically adjusting the model effectiveness strategy.
It significantly improves the accuracy of diagnosis of rare diseases in newborns, shortens the diagnosis cycle from several days to hours, and enhances the ability to adapt to the heterogeneity of rare diseases.
Smart Images

Figure CN120998461A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of medical data processing technology, and more specifically, to an intelligent screening and diagnostic system for rare neonatal diseases based on multi-omics data fusion. Background Technology
[0002] Neonatal rare diseases are characterized by strong heterogeneity in clinical manifestations and subtle early symptoms. Traditional single-dimensional diagnostic methods (such as relying solely on clinical indicators or gene testing) often face the dilemma of high missed diagnosis rates and long diagnostic cycles. With the development of omics technologies, the fusion of multi-omics data has provided new ideas for precision diagnosis. However, how to achieve efficient integration and intelligent analysis of multi-source data in clinical practice remains a key challenge in the field of pediatric rare disease diagnosis.
[0003] Current diagnostic techniques for rare neonatal diseases have significant limitations: Single clinical omics analysis is insufficient to capture the genetic nature of diseases; for example, hereditary metabolic diseases cannot be diagnosed solely based on indicators such as blood ammonia and coagulation factors. Although gene testing can locate mutation sites, it is slow to respond to phenotypic abnormalities such as metabolic disorders, and cannot meet the rapid diagnostic needs of newborns with acute illnesses. The lack of a systematic multi-omics data fusion model has resulted in the insufficient exploration of the complementary value of clinical omics, genomics, and metabolomics data, thus limiting diagnostic accuracy.
[0004] Current technologies lack a systematic solution for the rapid and accurate diagnosis of rare neonatal diseases by deeply integrating clinical omics, genomics, and metabolomics data. This makes it difficult for clinicians to achieve accurate subtyping through multi-dimensional data correlation analysis in the early stages of the disease, thus delaying intervention. Therefore, we propose an intelligent screening and diagnostic system for rare neonatal diseases based on multi-omics data fusion. Summary of the Invention
[0005] The purpose of this invention is to provide an intelligent screening and diagnostic system for rare neonatal diseases based on multi-omics data fusion, so as to solve the technical problems of traditional diagnostic methods, which are time-consuming, have low accuracy, and cannot meet the needs of rapid and accurate diagnosis of rare neonatal diseases.
[0006] To address the aforementioned technical problems, this invention provides the following technical solution: a newborn rare disease intelligent screening and diagnostic system based on multi-omics data fusion, comprising: The data acquisition module is used to obtain clinical omics data from the medical order database through the newborn's medical record, and to collect blood samples to obtain and label gene and metabolomics data when the medical record records that the newborn has a rare disease; The data processing module is used to preprocess clinical omics data based on neonatal medical records to obtain training and test sets, and to preprocess gene and metabolomics data based on blood samples. It also constructs a clinical omics-gene metabolism correlation model using Logistic regression and random forest algorithms. The diagnostic module, connected to the data processing module, is used to diagnose the test set based on the clinical omics-gene metabolism correlation model.
[0007] This invention constructs a three-level modular architecture: the data acquisition module integrates multidimensional data from clinical medical records, gene sequencing, and metabolite detection; the data processing module innovatively adopts a coupled model of Logistic regression and random forest to achieve cross-dimensional correlation analysis; and the diagnosis module, based on a dynamic adjustment strategy for model effectiveness, completes standardized preprocessing of multi-omics data, effectively solving the problem of missed diagnoses caused by isolated analysis of traditional multi-source data, and significantly improving the accuracy of diagnosis of rare neonatal diseases.
[0008] Preferably, the data acquisition module includes: The automatic neonatal medical record acquisition system is used to collect clinical omics data from neonatal medical records from the medical order database; The automated clinical omics detection system is used to obtain gene data and metabolomics data based on the blood sample test results collected by the neonatal medical records, gene detection system, and metabolomics detection system. A gene testing system used to detect rare hereditary disease genes in newborn blood samples via PCR-Sanger sequencing; A metabolomics detection system is used to detect organic acid metabolites and polar amino acids in neonatal blood samples through blood tandem mass spectrometry analysis.
[0009] Preferably, the data processing module includes: The clinical training set generation system is used to clean and normalize clinical omics data, and then generate a clinical training set after screening by a Logistic regression model. A clinical test set generation system is used to construct multi-channel data and obtain the clinical test set through the test set of the random forest algorithm; A clinical-gene metabolism correlation model generation system is used to construct clinical omics-gene metabolism correlation models using Logistic regression and random forest algorithms.
[0010] Preferably, the clinical training set generation system includes: The data processing unit is used to clean and normalize the clinical omics data in neonatal medical records, and then filter them through a Logistic regression model to obtain clinical training set data. The training set generation unit is used to construct multi-channel data, classify the clinical training set data according to the training set of the Logistic regression model, and generate the training set. The label generation unit is used to determine the sample group based on the clinical training set data and generate labels.
[0011] Preferably, the clinical omics data in neonatal medical records are cleaned and normalized, and then filtered using a logistic regression model to obtain the clinical training set data, as follows: The minimum-maximum normalization method was used to clean the clinical omics data in neonatal medical records to obtain the clinical omics dataset. The formula is as follows: In the formula, The original data, and These are the minimum and maximum values of the feature, respectively. The data is after normalization; The formula for the Logistic regression model is: In the formula, For the input feature vector, For model parameter vectors, This indicates that in the input feature vector and model parameter vector The probability that a sample belongs to the positive class.
[0012] Preferably, the clinical test set generation system includes: The test set generation unit is used to construct multi-channel data, annotate the clinical training set data, maintain the consistency between the clinical grouping and the clinical grouping of the training set, classify the test set of the random forest algorithm, and generate the test set. The clinical test set judgment unit is used to validate the clinical-gene metabolism correlation model based on the clinical test set and obtain the validation results.
[0013] Preferably, the clinical-gene metabolism correlation model generation system includes: The model building unit is used to screen gene data and metabolomics data through logistic regression to obtain clinical-gene-metabolic correlation data, and to build a clinical-gene-metabolic correlation model by combining the clinical-gene-metabolic correlation data with clinical training set data. The clinical-gene metabolism correlation model validation unit is used to judge samples in the clinical test set according to the clinical-gene metabolism correlation model, obtain the judgment result, compare the judgment result with the label, and obtain the validation result. The model application unit is used to determine whether gene data and metabolomics data are helpful for the diagnosis of rare neonatal diseases based on the clinical-gene metabolism correlation model combined with the clinical grouping of the clinical test set.
[0014] Preferably, the Logistic regression screening process is achieved by maximizing the likelihood function. for: ; in, For the sample size, For the sample The true label, The input feature vector, For model parameter vectors, Indicates that under a given feature and parameters Below, sample The probability of belonging to the positive class.
[0015] Preferably, the following method is used to determine whether gene data and metabolomics data are helpful in the diagnosis of rare neonatal diseases: Let the clinical group label be The gene data feature vector is The feature vector of metabolomics data is Construct a logistic regression model: In the formula, and These are the weight vectors for gene data and metabolomics data, respectively. For bias terms; Set threshold ,like If the data is helpful for diagnosis, the model is marked as "effective" if the result is yes; otherwise, the model is deleted.
[0016] Preferably, the diagnostic module includes: The clinical diagnostic execution unit is used to perform rapid diagnosis based on the clinical omics-gene metabolism correlation model only when the model is marked as "effective". The disease diagnosis unit is used to build an independent diagnostic model based on gene data and metabolomics data to make a diagnosis when the model is marked as "invalid".
[0017] Compared with the prior art, the beneficial effects of the present invention are: 1. This invention constructs a three-level modular architecture: the data acquisition module integrates multidimensional data from clinical medical records, gene sequencing, and metabolite detection; the data processing module innovatively adopts a coupled model of Logistic regression and random forest to achieve cross-dimensional correlation analysis; and the diagnosis module, based on a dynamic adjustment strategy for model effectiveness, completes standardized preprocessing of multi-omics data, effectively solving the problem of missed diagnoses caused by isolated analysis of traditional multi-source data, and significantly improving the accuracy of diagnosis of rare neonatal diseases.
[0018] 2. This invention also leverages the rapid predictive capabilities of a clinical-gene metabolism correlation model. When the model determines the result to be "effective," the traditional step-by-step testing process can be skipped, and diagnosis can be completed based on a multi-omics fusion model. This mechanism shortens the gene testing cycle from several days to hours, saving valuable time for early intervention in neonatal acute rare diseases.
[0019] 3. This invention also utilizes a dynamic evaluation mechanism for model application units, employing likelihood functions to screen key features and setting probability thresholds to determine data validity. The system can automatically identify the differences in multi-omics feature weights for different rare diseases. When the core model's diagnostic efficacy for a specific disease is insufficient, it automatically switches to a gene-metabolism independent diagnostic model, effectively addressing the problem of weak model generalization ability and significantly improving the system's adaptability to the heterogeneity of rare diseases. Attached Figure Description
[0020] Figure 1 This is a system architecture diagram of the present invention. Detailed Implementation
[0021] like Figure 1 As shown, the present invention relates to an intelligent screening and diagnostic system for rare neonatal diseases based on multi-omics data fusion, comprising a data acquisition module, a data processing module, and a diagnostic module; The data acquisition module is used to obtain clinical omics data from the medical order database through neonatal medical records; The data acquisition module is also used to analyze neonatal medical records. If the neonatal medical record records that the newborn has a rare disease, a blood sample of the newborn is collected, and gene data and metabolomics data are obtained through the newborn blood sample and the data is labeled. The gene data refers to gene sequencing data from Sanger sequencing and clinical gene data, and the metabolomics data refers to organic acid metabolites and polar amino acid data. In embodiments of the present invention, the clinical omics data includes the following parameters: routine coagulation, routine urine, electrolytes, coagulation factors, cerebrospinal fluid cells, liver ultrasound, liver imaging diagnosis, blood biochemistry, kidney ultrasound, blood ammonia, coagulation factor activity and coagulation function, T and B lymphocyte subsets, T lymphocyte proliferation, T cell receptor diversity, vitamin D, vitamin E, vitamin K, gene detection and metabolomics detection; In embodiments of the present invention, the data acquisition module includes an automatic neonatal medical record acquisition system, an automatic clinical omics detection system, a gene detection system, and a metabolomics detection system; The automatic neonatal medical record acquisition system is used to collect clinical omics data from neonatal medical records from the medical order database; The clinical omics data refers to clinical omics detection indicators and clinical omics detection values. The automated clinical omics detection system is used to obtain gene data and metabolomics data based on the blood sample test results collected by the neonatal medical records, gene detection system, and metabolomics detection system. The gene detection system is used to detect rare hereditary disease genes in newborn blood samples via PCR-Sanger sequencing; Among them, the genetic rare disease genes include hemophilia, cystic fibrosis, myasthenia gravis, thalassemia, mucopolysaccharidosis, cerebrospinal diseases, Gaucher disease, Fabry disease, van Annegold syndrome, and mucopolysaccharide storage disease genes; The metabolomics detection system is used to detect organic acid metabolites and polar amino acids in neonatal blood samples by blood tandem mass spectrometry analysis.
[0022] The data processing module is used to preprocess clinical omics data based on neonatal medical records to obtain training and test sets, and to preprocess gene and metabolomics data based on blood samples. It also constructs a clinical omics-gene metabolism correlation model using Logistic regression and random forest algorithms. In embodiments of the present invention, the data processing module includes a clinical training set generation system, a clinical test set generation system, and a clinical-gene metabolism correlation model generation system; The clinical training set generation system is used to clean and normalize clinical omics data, and generate a clinical training set after screening by a Logistic regression model. The clinical training set generation system includes a data processing unit, a training set generation unit, and a label generation unit. The data processing unit is used to clean and normalize the clinical omics data in neonatal medical records, and then filter them through a Logistic regression model to obtain clinical training set data. Specifically, the clinical omics data in neonatal medical records underwent data cleaning and normalization, and then were filtered using a logistic regression model to obtain the clinical training set data, as detailed below: First, the clinical omics data in neonatal medical records were cleaned to remove missing and outlier values, resulting in a clinical omics dataset. Data normalization was performed using the min-max normalization method, with the following formula: ; In the formula, The original data, and These are the minimum and maximum values of the feature, respectively. The data is normalized. Data normalization is used to standardize the clinical omics dataset, and then multi-channel clinical omics data is constructed from the normalized clinical omics dataset. The formula for the Logistic regression model is: In the formula, The input feature vector (i.e., clinical omics data) is used. For model parameter vectors, This indicates that in the input feature vector and model parameter vector Below, the probability that the sample belongs to the positive class (having a rare disease); The training set generation unit is used to construct multi-channel data, classify the clinical training set data according to the training set of the Logistic regression model, and generate a training set. The label generation unit is used to determine the sample group based on the clinical training set data and generate labels.
[0023] The clinical test set generation system is used to construct multi-channel data and obtain the clinical test set through the test set of the random forest algorithm; The clinical test set generation system includes a test set generation unit and a clinical test set judgment unit. The test set generation unit is used to construct multi-channel data, annotate the clinical training set data, maintain the consistency between the clinical grouping and the clinical grouping of the training set, classify the test set of the random forest algorithm, and generate the test set. In the random forest algorithm, a single decision tree The prediction process can be represented as: ; In the formula, For the input sample, As a category, For the first Decision trees for samples Prediction categories, This is the indicator function. The final prediction result of a random forest is obtained by combining the prediction results of multiple decision trees, using either a voting method (for classification tasks) or an averaging method (for regression tasks). The clinical test set judgment unit is used to validate the clinical-gene metabolism correlation model based on the clinical test set and obtain the validation results.
[0024] The clinical-gene metabolism correlation model generation system is used to construct clinical omics-gene metabolism correlation models using Logistic regression and random forest algorithms. The clinical-gene metabolism correlation model generation system includes a model construction unit, a clinical-gene metabolism correlation model validation unit, and a model application unit. The model building unit is used to screen gene data and metabolomics data through Logistic regression to obtain clinical-gene-metabolic correlation data, and to combine the clinical-gene-metabolic correlation data with clinical training set data to build a clinical-gene-metabolic correlation model. The logistic regression screening process can be achieved by maximizing the likelihood function. for: ; in, For the sample size, For the sample The true label, The input feature vector (here, a feature vector composed of gene data and metabolomics data). For model parameter vectors, Indicates that under a given feature and parameters Below, sample The probability of belonging to the positive class (having a rare disease). This is determined by solving... Obtain the optimal parameters This allows for the screening of disease-related gene and metabolomics data features. The clinical-gene metabolism correlation model validation unit is used to judge the samples in the clinical test set according to the clinical-gene metabolism correlation model and obtain the judgment result; the judgment result is compared with the label to obtain the validation result; The clinical-gene metabolism correlation model validation unit is also used to display the validation results on the display interface.
[0025] The model application unit is used to determine whether gene data and metabolomics data are helpful for the diagnosis of rare neonatal diseases based on the clinical-gene metabolism correlation model combined with the clinical grouping of the clinical test set. The following algorithm is used to determine whether gene data and metabolomics data are helpful in diagnosing rare neonatal diseases: Let the feature vector of gene data be... The feature vector of metabolomics data is Clinical group label is (1 indicates having a rare disease, 0 indicates not having it), construct a logistic regression model: ; In the formula, and These are the weight vectors for gene data and metabolomics data, respectively. This is the bias term. The predicted probability is calculated. Set threshold (like ),like If the data is helpful for diagnosis, then the formula is as follows: The probability threshold is set manually. Based on the above judgment results, if the judgment result is yes, the model is marked as "valid"; if it is no, the model is deleted. Among them, the rare neonatal diseases mentioned are hemophilia, cystic fibrosis, myasthenia gravis, mucopolysaccharidosis, cerebrospinal diseases, Gaucher disease, Fabry disease, van Annegold syndrome, and mucopolysaccharide storage diseases; The diagnostic module is connected to the data processing module and is used to diagnose the test set based on the clinical omics-gene metabolism correlation model. In an embodiment of the present invention, the diagnostic module includes a clinical diagnostic execution unit and a disease diagnostic unit; The clinical diagnostic execution unit is used to obtain clinical groupings of the clinical test set based on the clinical-gene metabolism correlation model, and performs rapid diagnosis of rare neonatal diseases only when the model is marked as "effective" through the clinical-gene metabolism correlation model. The disease diagnosis unit is used to establish a diagnostic model for rare neonatal diseases based on gene data and metabolomics data. When the model is marked as "invalid", the disease diagnosis unit is automatically triggered to diagnose rare neonatal diseases based on gene data and metabolomics data through the rare neonatal disease diagnosis model.
[0026] The embodiments disclosed in this invention are preferred embodiments, but are not limited thereto. Those skilled in the art can easily understand the spirit of this invention based on the above embodiments and make different extensions and variations, but as long as they do not depart from the spirit of this invention, they are all within the protection scope of this invention.
Claims
1. A newborn rare disease intelligent screening and diagnostic system based on multi-omics data fusion, characterized in that, include: The data acquisition module is used to obtain clinical omics data from the medical order database through the newborn's medical record, and to collect blood samples to obtain and label gene and metabolomics data when the medical record records that the newborn has a rare disease; The data processing module is used to preprocess clinical omics data based on neonatal medical records to obtain training and test sets, and to preprocess gene and metabolomics data based on blood samples. It also constructs a clinical omics-gene metabolism correlation model using Logistic regression and random forest algorithms. The diagnostic module, connected to the data processing module, is used to diagnose the test set based on the clinical omics-gene metabolism correlation model.
2. The intelligent screening and diagnostic system for rare neonatal diseases based on multi-omics data fusion according to claim 1, characterized in that, The data acquisition module includes: The automatic neonatal medical record acquisition system is used to collect clinical omics data from neonatal medical records from the medical order database; The automated clinical omics detection system is used to obtain gene data and metabolomics data based on the blood sample test results collected by the neonatal medical records, gene detection system, and metabolomics detection system. A gene testing system used to detect rare hereditary disease genes in newborn blood samples via PCR-Sanger sequencing; A metabolomics detection system is used to detect organic acid metabolites and polar amino acids in neonatal blood samples through blood tandem mass spectrometry analysis.
3. The intelligent screening and diagnostic system for rare neonatal diseases based on multi-omics data fusion according to claim 1, characterized in that, The data processing module includes: The clinical training set generation system is used to clean and normalize clinical omics data, and then generate a clinical training set after screening by a Logistic regression model. A clinical test set generation system is used to construct multi-channel data and obtain the clinical test set through the test set of the random forest algorithm; A clinical-gene metabolism correlation model generation system is used to construct clinical omics-gene metabolism correlation models using Logistic regression and random forest algorithms.
4. The intelligent screening and diagnostic system for rare neonatal diseases based on multi-omics data fusion according to claim 3, characterized in that, The clinical training set generation system includes: The data processing unit is used to clean and normalize the clinical omics data in neonatal medical records, and then filter them through a Logistic regression model to obtain clinical training set data. The training set generation unit is used to construct multi-channel data, classify the clinical training set data according to the training set of the Logistic regression model, and generate the training set. The label generation unit is used to determine the sample group based on the clinical training set data and generate labels.
5. The intelligent screening and diagnostic system for rare neonatal diseases based on multi-omics data fusion according to claim 4, characterized in that, The clinical omics data in neonatal medical records were cleaned and normalized, and then filtered using a logistic regression model to obtain the clinical training set data, as follows: The minimum-maximum normalization method was used to clean the clinical omics data in neonatal medical records to obtain the clinical omics dataset. The formula is as follows: In the formula, The original data, and These are the minimum and maximum values of the feature, respectively. The data is after normalization; The formula for the Logistic regression model is: In the formula, For the input feature vector, For model parameter vectors, This indicates that in the input feature vector and model parameter vector The probability that a sample belongs to the positive class.
6. The intelligent screening and diagnostic system for rare neonatal diseases based on multi-omics data fusion according to claim 4, characterized in that, The clinical test set generation system includes: The test set generation unit is used to construct multi-channel data, annotate the clinical training set data, maintain the consistency between the clinical grouping and the clinical grouping of the training set, classify the test set of the random forest algorithm, and generate the test set. The clinical test set judgment unit is used to validate the clinical-gene metabolism correlation model based on the clinical test set and obtain validation results.
7. The intelligent screening and diagnostic system for rare neonatal diseases based on multi-omics data fusion according to claim 6, characterized in that, The clinical-gene metabolism correlation model generation system includes: The model building unit is used to screen gene data and metabolomics data through logistic regression to obtain clinical-gene-metabolic correlation data, and to build a clinical-gene-metabolic correlation model by combining the clinical-gene-metabolic correlation data with clinical training set data. The clinical-gene metabolism correlation model validation unit is used to judge samples in the clinical test set according to the clinical-gene metabolism correlation model, obtain the judgment result, compare the judgment result with the label, and obtain the validation result. The model application unit is used to determine whether gene data and metabolomics data are helpful for the diagnosis of rare neonatal diseases based on the clinical-gene metabolism correlation model combined with the clinical grouping of the clinical test set.
8. The intelligent screening and diagnostic system for rare neonatal diseases based on multi-omics data fusion according to claim 7, characterized in that, The logistic regression screening process is achieved by maximizing the likelihood function. for: ; in, For the sample size, For the sample The true label, The input feature vector, For model parameter vectors, Indicates that under a given feature and parameters Below, sample The probability of belonging to the positive class.
9. The intelligent screening and diagnostic system for rare neonatal diseases based on multi-omics data fusion according to claim 7, characterized in that, The following methods are used to determine whether genetic and metabolomics data are helpful in diagnosing rare neonatal diseases: Let the clinical group label be The gene data feature vector is The feature vector of metabolomics data is Construct a logistic regression model: In the formula, and These are the weight vectors for gene data and metabolomics data, respectively. For bias terms; Set threshold ,like If the data is helpful for diagnosis, the model is marked as "effective" if the result is yes; otherwise, the model is deleted.
10. The intelligent screening and diagnostic system for rare neonatal diseases based on multi-omics data fusion according to claim 1, characterized in that, The diagnostic module includes: The clinical diagnostic execution unit is used to perform rapid diagnosis based on the clinical omics-gene metabolism correlation model only when the model is marked as "effective". The disease diagnosis unit is used to build an independent diagnostic model based on genetic and metabolomics data to make a diagnosis when the model is marked as "invalid".
Citation Information
Cited By
Plateau newborn rare disease multi-mode combined screening system and storage medium
CN121687550A