Causal risk early warning method and system based on insomnia metabolic characteristics, and medium

By processing data and conducting multi-omics analysis on the metabolic characteristics of insomnia, we can identify metabolites and genetic features related to insomnia. This solves the problem that existing technologies cannot integrate information on multiple metabolites, enabling causal risk warning and personalized treatment for insomnia-related diseases.

CN121191748APending Publication Date: 2025-12-23SUN YAT SEN MEMORIAL HOSPITAL SUN YAT SEN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511151518.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-18
Publication Date
2025-12-23

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately identify the complex network relationships between insomnia and metabolic disorders, lack systematic methods to integrate information from multiple metabolites, and are unable to perform downstream analysis, thus failing to provide effective means for the intervention and treatment of insomnia and related diseases.

Method used

By constructing a sample dataset, performing logarithmic transformation and standardization of metabolite concentrations, identifying metabolites using a resilient network model, and combining genome-wide association analysis and latent causal variable models, the causal risk of insomnia metabolic characteristics to the target disease was determined.

Benefits of technology

It enables accurate and efficient causal risk warning of diseases affected by insomnia metabolic characteristics, provides support for early warning and personalized treatment, and reveals the complex network relationship between insomnia and metabolic disorders.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121191748A_ABST
    Figure CN121191748A_ABST
Patent Text Reader

Abstract

The invention discloses a causal risk early warning method and system based on insomnia metabolism characteristics and a medium, and the method comprises the steps: carrying out the logarithmic transformation and standardization processing of a metabolite concentration in a sample data set, and obtaining a standardized metabolite concentration; inputting the standardized metabolite concentration into a preset elastic network model to obtain an insomnia-related metabolite, and performing weighted calculation on the metabolite to obtain an insomnia-related metabolic characteristic score; taking the metabolic characteristic score as phenotype data, and carrying out whole genome association analysis on the phenotype data and obtained genotype data of all samples to obtain first character GWAS data of insomnia metabolic characteristics; and inputting the first character GWAS data of the insomnia metabolism characteristics and the second character GWAS data of the target disease characteristic set into a potential causal variable model to calculate a genetic causal ratio, and determining a causal risk early warning result based on the genetic causal ratio, so that diseases affected by the insomnia metabolism characteristics can be accurately and efficiently determined.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of risk warning, and in particular to a causal risk warning method, system and medium based on the metabolic characteristics of insomnia. Background Technology

[0002] Insomnia is a common sleep disorder that not only leads to mental fatigue, mood swings, and cognitive decline, but can also trigger various metabolic disorders such as obesity, diabetes, and cardiovascular disease. These metabolic disorders are among the important complications of insomnia, but their specific mechanisms are not yet fully understood. Accurately and efficiently identifying diseases affected by the metabolic characteristics of insomnia is crucial for a deeper understanding of the pathophysiological mechanisms of insomnia. This will not only help reveal the complex network relationship between insomnia and metabolic disorders, but also provide strong support for early warning, personalized treatment, and precision medicine for insomnia and related diseases. By identifying disease risk factors and biomarkers related to insomnia, new targets can be provided for clinical intervention, advancing research on the prevention and treatment of related diseases.

[0003] Current technologies primarily identify insomnia-related metabolic features through single metabolites and multivariate statistical methods. However, these technologies have significant limitations. Single metabolite studies struggle to reveal the complex network relationship between insomnia and metabolic disorders, while non-targeted metabolomics results often contain numerous noisy variables, making it difficult to distinguish truly insomnia-related metabolic features. Furthermore, existing research lacks a systematic approach to integrate information from multiple metabolites, failing to construct a comprehensive metabolic feature score. More importantly, current research typically only focuses on metabolite identification, lacking downstream analysis of metabolic features, such as identifying related genetic characteristics, phenotypes influenced by insomnia-related metabolic features, and gene expression regulating these features. Therefore, it cannot provide effective means for intervention and treatment. Summary of the Invention

[0004] This invention provides a causal risk early warning method, system, and medium based on the metabolic characteristics of insomnia, which can accurately and efficiently identify diseases affected by the metabolic characteristics of insomnia.

[0005] This invention provides a causal risk early warning method based on the metabolic characteristics of insomnia, comprising:

[0006] A sample dataset was constructed based on the acquired sleep questionnaire data and metabolomics data, and the concentrations of metabolites in the sample dataset were logarithmically transformed and standardized to obtain standardized metabolite concentrations.

[0007] The standardized metabolite concentrations are input into a preset elastic network model to obtain insomnia-related metabolites. The metabolites are then weighted to obtain a metabolic characteristic score related to insomnia. The elastic network model is obtained by optimizing the model parameters using the standardized metabolite concentrations as independent variables and the insomnia score as the dependent variable, combined with a cross-validation function. The insomnia score is obtained by identifying and assigning values ​​to the insomnia classification labels of the sleep questionnaire data.

[0008] The metabolic characteristic scores were used as phenotypic data, and genome-wide association analysis was performed with the genotype data of all samples to obtain the first trait GWAS data of insomnia metabolic characteristics.

[0009] The GWAS data of the first trait corresponding to the insomnia metabolic feature and the GWAS data of the second trait of the target disease feature set are input into the latent causal variable model to calculate the genetic causal ratio, and the causal risk warning result affected by the insomnia metabolic feature is determined based on the genetic causal ratio.

[0010] This invention eliminates dimensional differences between different metabolite concentrations by performing logarithmic transformation and standardization on metabolite concentrations, improving data comparability and analytical accuracy. Through elastic network regression analysis, it effectively identifies metabolites significantly associated with insomnia. By weighting the selected metabolites, it integrates information from multiple metabolites, providing a more comprehensive metabolic profile of insomnia and offering a scientific basis for subsequent multi-omics analysis. GWAS analysis reveals the genetic basis of insomnia metabolic characteristics and provides important genetic evidence for subsequent causal inference. By combining metabolomics data with genetic data, a multi-omics integrated analysis framework is constructed, laying the foundation for a comprehensive understanding of the complex network relationship between insomnia and metabolic disorders. By assessing the causal relationship between insomnia metabolic characteristics and target diseases, it allows for a deeper exploration of the intrinsic link between insomnia and related diseases. Based on the positive and negative values ​​of the genetic causal ratio, it determines the direction and intensity of the causal influence of insomnia metabolic characteristics on target diseases, thereby achieving causal risk warning for related diseases and providing strong support for early disease warning, personalized treatment, and precision medicine. Compared with existing technologies, this application can accurately and efficiently identify diseases affected by the metabolic characteristics of insomnia.

[0011] Furthermore, the construction of the sample dataset based on the acquired sleep questionnaire data and metabolomics data specifically involves:

[0012] Acquire sleep questionnaire data and metabolomics data, and remove individuals with invalid responses from the sleep questionnaire data to obtain the first processing result;

[0013] Individuals with non-positive metabolite concentrations in the metabolomics data were removed to obtain the second processing result;

[0014] The first processing result and the second processing result are merged according to individual numbers to obtain the sample dataset.

[0015] By merging the processed sleep questionnaire data and metabolomics data by individual ID, we can ensure data consistency and integrity, providing a high-quality data foundation for subsequent genome-wide association analysis and causal inference, thereby more efficiently identifying diseases affected by insomnia metabolic characteristics.

[0016] Further, 3. The step of performing logarithmic transformation and standardization on the metabolite concentrations in the sample dataset to obtain standardized metabolite concentrations specifically involves:

[0017] Logarithmically transform the concentration of each metabolite in the sample dataset, and calculate the mean and standard deviation of the transformed concentrations of each metabolite.

[0018] The concentrations of each metabolite are converted into standardized values ​​based on the mean and the standard deviation to obtain standardized metabolite concentrations.

[0019] By performing logarithmic transformation and standardization on the concentrations of metabolites, the dimensional differences between different metabolite concentrations are eliminated, improving the comparability of data and the accuracy of analysis.

[0020] Furthermore, the metabolic characteristic score is used as phenotypic data, and a genome-wide association analysis is performed with the genotype data of all samples to obtain GWAS data of the first trait of insomnia metabolic characteristics, specifically as follows:

[0021] The metabolic characteristic scores are used as phenotypic data, and the phenotypic data is processed to obtain the target phenotypic data.

[0022] The initial genotype data of all samples is obtained, and the initial genotype data is subjected to quality control to obtain the genotype data.

[0023] The target phenotypic data and the genotype data are matched and included as covariates. A Firth-corrected generalized linear model is used to perform genome-wide association analysis to calculate the regression coefficient, standard error, and p-value for each SNP. The covariates are demographic variables such as sex, age, and the number of principal components of genes within a preset range.

[0024] Based on the regression coefficients, the standard error, and the p-value, the first trait GWAS data of insomnia metabolic characteristics are determined.

[0025] This GWAS analysis revealed the genetic basis of the metabolic characteristics of insomnia and provided important genetic evidence for subsequent causal inference.

[0026] Furthermore, after obtaining the first trait GWAS data of insomnia metabolic characteristics, the method further includes:

[0027] The SNPs in the first trait GWAS data were functionally annotated using FUMA combined with gene annotation tools to determine the genomic location and functional impact of the SNPs.

[0028] Based on the genomic location and functional impact, the pathogenic potential of the SNP is quantified by CADD scoring;

[0029] Based on the order of pathogenic potential, and combined with eQTL mapping, position mapping, and chromatin interaction mapping strategies, the first candidate gene associated with the SNP is determined;

[0030] The combined effects of the SNPs were integrated to the gene level using the MAGMA method to screen for second candidate genes associated with the phenotype of the insomnia metabolic characteristics.

[0031] The tissue-specific expression patterns of the first and second candidate genes were evaluated based on the GTEx database to identify genetic features and potential targets related to the metabolic characteristics of insomnia.

[0032] By systematically exploring the genetic basis and potential therapeutic targets of insomnia-related metabolic features through functional annotation, pathogenic potential assessment, gene mapping, and expression pattern analysis, we can provide a scientific basis for the study and intervention of the mechanisms of insomnia and related diseases.

[0033] Furthermore, after obtaining the first trait GWAS data of insomnia metabolic characteristics, the method further includes:

[0034] SNPs located within 1 Mb of the transcription start site of the target gene and significantly associated with gene expression were extracted from eQTL, mQTL and sQTL data.

[0035] Combining the SNPs with the insomnia metabolic characteristics, the p-SMR value of each QTL molecular phenotype and the insomnia metabolic characteristics is calculated using the SMR framework, and significant correlation characteristics with p values ​​lower than a first preset threshold are output.

[0036] Based on the aforementioned significant correlation features, the HEIDI test was used to screen out causal features that have a true causal relationship with the aforementioned insomnia metabolic features;

[0037] Based on the aforementioned causal characteristics, gene expression related to the aforementioned insomnia metabolic characteristics was identified, and gene expression results were obtained.

[0038] By screening out gene expression features that have a real causal relationship with the metabolic characteristics of insomnia, we can delve into the genetic regulatory mechanisms of insomnia-related metabolic features and provide a scientific basis for disease risk warning and the discovery of therapeutic targets.

[0039] Furthermore, the determination of the causal risk warning result affected by the insomnia metabolic characteristics based on the genetic causal ratio specifically includes:

[0040] If the absolute value of the genetic causal ratio is greater than the second preset threshold, and the P-value after correction by the Benjamini-Hochberg method is lower than the third preset threshold, then it is determined that there is a strong causal relationship between the first trait data and the second trait data.

[0041] If the genetic causal ratio is positive, then the causal risk warning of insomnia metabolic characteristics for the disease will be output.

[0042] This approach, through systematic integration of insomnia-related metabolic characteristics, genetic variations, and causal inference analysis, comprehensively identifies diseases significantly associated with insomnia metabolic characteristics and accurately assesses their causal relationships. It not only effectively screens key metabolites and genetic variations but also ensures the accuracy and reliability of the results through rigorous statistical correction and causal ratio analysis, thus providing strong support for accurately and efficiently identifying diseases affected by insomnia metabolic characteristics.

[0043] Another embodiment of the present invention provides a causal risk early warning system based on the metabolic characteristics of insomnia, including: a construction module, a calculation module, an analysis module and an early warning module;

[0044] The construction module is used to construct a sample dataset based on the acquired sleep questionnaire data and metabolomics data, and to perform logarithmic transformation and standardization on the concentration of metabolites in the sample dataset to obtain standardized metabolite concentrations.

[0045] The calculation module is used to input the standardized metabolite concentration into a preset elastic network model to obtain insomnia-related metabolites, and to perform weighted calculations on the metabolites to obtain a metabolic characteristic score related to insomnia. The elastic network model is obtained by optimizing the model parameters using the standardized metabolite concentration as the independent variable and the insomnia score as the dependent variable, combined with a cross-validation function. The insomnia score is obtained by identifying and assigning values ​​to the insomnia classification labels of the sleep questionnaire data.

[0046] The analysis module is used to perform genome-wide association analysis (GWAS) on the metabolic characteristic score as phenotypic data and the genotype data of all samples obtained to obtain the first trait GWAS data of insomnia metabolic characteristics.

[0047] The early warning module is used to input the first trait GWAS data corresponding to the insomnia metabolic feature and the second trait GWAS data of the target disease feature set into the latent causal variable model to calculate the genetic causal ratio, and determine the causal risk early warning result affected by the insomnia metabolic feature based on the genetic causal ratio.

[0048] This invention eliminates dimensional differences between different metabolite concentrations by performing logarithmic transformation and standardization on metabolite concentrations, improving data comparability and analytical accuracy. Through elastic network regression analysis, it effectively identifies metabolites significantly associated with insomnia. By weighting the selected metabolites, it integrates information from multiple metabolites, providing a more comprehensive metabolic profile of insomnia and offering a scientific basis for subsequent multi-omics analysis. GWAS analysis reveals the genetic basis of insomnia metabolic characteristics and provides important genetic evidence for subsequent causal inference. By combining metabolomics data with genetic data, a multi-omics integrated analysis framework is constructed, laying the foundation for a comprehensive understanding of the complex network relationship between insomnia and metabolic disorders. By assessing the causal relationship between insomnia metabolic characteristics and target diseases, it allows for a deeper exploration of the intrinsic link between insomnia and related diseases. Based on the positive and negative values ​​of the genetic causal ratio, it determines the direction and intensity of the causal influence of insomnia metabolic characteristics on target diseases, thereby achieving causal risk warning for related diseases and providing strong support for early disease warning, personalized treatment, and precision medicine. Compared with existing technologies, this application can accurately and efficiently identify diseases affected by the metabolic characteristics of insomnia.

[0049] Furthermore, the analysis module includes a processing unit, a control unit, an analysis unit, and a combination unit:

[0050] The processing unit is used to use the metabolic characteristic score as phenotypic data and to process the phenotypic data to obtain target phenotypic data.

[0051] The control unit is used to acquire the initial genotype data of all samples and to perform quality control on the initial genotype data to obtain genotype data.

[0052] The analysis unit is used to match the target phenotypic data and the genotype data, include them as covariates, and perform genome-wide association analysis using a Firth-corrected generalized linear model to calculate the regression coefficient, standard error, and p-value for each SNP. The covariates are demographic variables such as sex, age, and the number of principal components of genes within a preset range.

[0053] The combined unit is used to determine the first trait GWAS data of insomnia metabolic characteristics based on the regression coefficient, the standard error, and the p-value.

[0054] Another embodiment of the present invention provides a computer-readable storage medium item, including: a stored computer program, which, when the computer program is running, controls the device where the computer-readable storage medium is located to perform steps such as the causal risk warning method based on insomnia metabolic characteristics of the present invention. Attached Figure Description

[0055] To more clearly illustrate the technical solution of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0056] Figure 1 This is a flowchart illustrating an embodiment of the causal risk warning method based on the metabolic characteristics of insomnia provided in this application.

[0057] Figure 2 This is a flowchart illustrating one embodiment of steps S201 to S204 provided in this application;

[0058] Figure 3 This is a schematic diagram of the Manhattan plot of the genome-wide association analysis of the metabolic characteristics of insomnia provided in this application;

[0059] Figure 4 This is a schematic diagram of gene expression related to insomnia-related metabolic characteristics provided in this application;

[0060] Figure 5 This is a schematic diagram of the causal association screening results between insomnia-related metabolic characteristics and disease phenotypes provided in this application;

[0061] Figure 6 This is a schematic diagram of an embodiment of the causal risk early warning system based on the metabolic characteristics of insomnia provided in this application. Detailed Implementation

[0062] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings of the embodiments. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0063] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains; the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the application; the terms “comprising” and “having”, and any variations thereof, in the specification, claims, and foregoing description of the drawings are intended to cover non-exclusive inclusion.

[0064] In the description of the embodiments of this application, technical terms such as "first" and "second" are used only to distinguish different objects and should not be construed as indicating or implying relative importance or implicitly specifying the number, specific order, or primary and secondary relationship of the indicated technical features. In the description of the embodiments of this application, "multiple" means two or more, unless otherwise explicitly defined.

[0065] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0066] In the description of the embodiments in this application, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, the character " / " in this document generally indicates that the preceding and following related objects have an "or" relationship.

[0067] In the description of the embodiments of this application, the term "multiple" refers to two or more (including two), similarly, "multiple sets" refers to two or more (including two sets), and "multiple pieces" refers to two or more (including two pieces).

[0068] In the description of the embodiments of this application, unless otherwise expressly specified and limited, technical terms such as "installation," "connection," "joining," and "fixing" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral part; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; they can refer to the internal communication of two components or the interaction between two components. For those skilled in the art, the specific meaning of the above terms in the embodiments of this application can be understood according to the specific circumstances.

[0069] See Figure 1To accurately and efficiently identify diseases affected by insomnia metabolic characteristics, an embodiment of the present invention provides a causal risk warning method based on insomnia metabolic characteristics, including steps S101 to S104.

[0070] Step S101: Construct a sample dataset based on the acquired sleep questionnaire data and metabolomics data, and perform logarithmic transformation and standardization on the concentration of metabolites in the sample dataset to obtain standardized metabolite concentrations.

[0071] In some embodiments, constructing a sample dataset based on the acquired sleep questionnaire data and metabolomics data specifically involves: acquiring sleep questionnaire data and metabolomics data, removing individuals with invalid responses from the sleep questionnaire data to obtain a first processing result; removing individuals with non-positive metabolite concentrations from the metabolomics data to obtain a second processing result; and merging the first processing result and the second processing result by individual number to obtain the sample dataset. Specifically, firstly, sleep questionnaire data containing insomnia-related information is acquired from the UK Biobank database. This data may include an individual's insomnia status (such as classification labels like "usually," "sometimes," "never / rarely") and other variables that may affect sleep. Metabolomics data is also acquired from the UK Biobank database. Secondly, individuals who answered "Do not know" or "Prefer not to answer" are removed from the sleep questionnaire data. These answers are considered invalid because they do not provide valid information about the insomnia state. After removing these individuals, the first processing result is obtained. Next, individuals with metabolite concentrations less than or equal to zero (≤0) were removed from the metabolomics data. Non-positive metabolite concentrations may indicate measurement errors or missing data, which could affect the accuracy and reliability of subsequent analyses. After removing these individuals, the second processing result was obtained. Finally, the cleaned sleep questionnaire data (first processing result) and metabolomics data (second processing result) were merged according to the participant's individual ID to ensure that each individual's insomnia status and metabolite characteristics corresponded. The merged dataset became the sample dataset, containing valid information that could be used for subsequent analyses.

[0072] It should be noted that the UK Biobank database contains detailed health information, genetic data, and metabolomics data for over 500,000 UK residents. The metabolomics data, measured using nuclear magnetic resonance (NMR) technology, covers 251 metabolites, including lipoprotein subclasses, fatty acids, and amino acids.

[0073] By merging the processed sleep questionnaire data and metabolomics data by individual ID, we can ensure data consistency and integrity, providing a high-quality data foundation for subsequent genome-wide association analysis and causal inference, thereby more efficiently identifying diseases affected by insomnia metabolic characteristics.

[0074] In some embodiments, the step of performing logarithmic transformation and standardization on the metabolite concentrations in the sample dataset to obtain standardized metabolite concentrations specifically involves: performing a logarithmic transformation on each metabolite concentration in the sample dataset, and calculating the mean and standard deviation of each transformed metabolite concentration; and converting each metabolite concentration into a standardized value based on the mean and standard deviation to obtain the standardized metabolite concentration. Specifically, firstly, a logarithmic transformation is performed on the value of each metabolite concentration in the sample dataset using a common logarithm (log...). 10 The concentration values ​​are transformed to a relatively uniform scale, while reducing the influence of extreme values ​​and making the data closer to a normal distribution, which facilitates subsequent statistical analysis. Then, for each metabolite concentration x, its log-transformed mean μ and standard deviation σ are calculated. Based on the calculated mean μ and standard deviation σ, the log-transformed value of each metabolite concentration is converted into a standardized value to obtain the standardized metabolite concentration.

[0075] By performing logarithmic transformation and standardization on the concentrations of metabolites, the dimensional differences between different metabolite concentrations are eliminated, improving the comparability of data and the accuracy of analysis.

[0076] Step S102: Input the standardized metabolite concentration into a preset elastic network model to obtain metabolites related to insomnia, and perform weighted calculation on the metabolites to obtain a metabolic characteristic score related to insomnia. The elastic network model is obtained by optimizing the model parameters by using the standardized metabolite concentration as the independent variable and the insomnia score as the dependent variable, combined with a cross-validation function. The insomnia score is obtained by identifying and assigning values ​​to the insomnia classification labels of the sleep questionnaire data.

[0077] In some embodiments, firstly, since insomnia classification labels in sleep questionnaire data may include "Usually," "Sometimes," and "Never / rarely," these labels need to be assigned corresponding numerical values. For example, "Usually" is assigned a value of 2, "Sometimes" is assigned a value of 1, and "Never / rarely" is assigned a value of 0. These values ​​constitute the insomnia score, serving as the dependent variable in the elastic network model. Secondly, Elastic Net Regression is chosen as the model. Elastic Net Regression combines the advantages of Lasso regression (L1 regularization) and Ridge regression (L2 regularization), enabling variable selection and regularization in high-dimensional data while addressing multicollinearity issues among variables. Then, using the standardized metabolite concentration (X) as the independent variable and the insomnia score (Y) as the dependent variable, 5-fold cross-validation was performed using the cv.glmnet() function to determine the optimal regularization parameter λ value to ensure the simplification of the model. Finally, the regularization parameter λ under the 1SE rule was selected, and the elastic network model was retrained using the selected λ value to optimize the iterative model parameters and obtain the trained elastic network model.

[0078] In some embodiments, the standardized metabolite concentrations are input into a preset elastic network model to obtain metabolite concentrations with non-zero coefficients. These non-zero coefficient metabolites are those significantly associated with insomnia. A weighted calculation is performed on each associated metabolite. Since the weight of each metabolite is determined by the model's coefficients, it reflects the strength of the correlation between the metabolite and insomnia. The weighted values ​​of the associated metabolites for each individual are summed to obtain an insomnia-related metabolic characteristic score for each individual. This score comprehensively reflects the relationship between an individual's metabolic characteristics and their insomnia state.

[0079] It should be noted that elastic network regression combines the advantages of Lasso regression and Ridge regression, enabling the selection of important variables in high-dimensional data and addressing the problem of multicollinearity among variables.

[0080] Step S103: The insomnia-related metabolic characteristic score is used as phenotypic data and subjected to genome-wide association analysis with the genotype data of all samples to obtain the first trait GWAS data of insomnia metabolic characteristics.

[0081] Please refer to Figure 2 In some embodiments, step S103 includes, but is not limited to, steps S201 to S204;

[0082] Step S201: The metabolic characteristic score is used as phenotypic data, and the phenotypic data is processed to obtain the target phenotypic data;

[0083] In some embodiments, firstly, cases (insomnia patients) are screened using the Data-Field 1200 field. Participants with the Data-Field 1200 field are selected from 500,000 participants in the UKB database. Those with mismatched gender (genetic sex not matching self-reported sex), chromosomal abnormalities (sex chromosome aneuploidy), and insomnia conditions marked "Do not know" and "Prefer not to answer" are removed. The metabolic product characteristics are matched using individual IDs, resulting in 211,609 individuals. Next, the metabolic characteristic scores are matched with the IDs of the 211,609 individuals selected from the UKB database to ensure that each individual's metabolic characteristic score corresponds to their genotype data. Then, the metabolic characteristic scores are checked for missing or outlier values. Missing values ​​are filled using methods such as multiple imputation; outliers are removed or corrected depending on the specific circumstances. The metabolic characteristic scores are then standardized to obtain target phenotypic data, ensuring that the metabolic characteristic scores of different individuals are compared on the same scale.

[0084] Step S202: Obtain the initial genotype data of all samples, and perform quality control on the initial genotype data to obtain genotype data;

[0085] In some embodiments, data cleaning was performed using 211,609 individual data provided by UKB. Then, only SNPs located on autosomes (chromosomes 1-22) were retained, and low-quality SNPs were filtered out, i.e., SNPs with an INFO value less than 0.8 were removed. These SNPs had low genotypic data quality and might affect the reliability of the analysis results. Next, low-frequency variants were filtered out, i.e., SNPs with an allele frequency (MAF) less than 0.01 were removed. These low-frequency variants are sparsely distributed in the population and may have a smaller impact on the analysis results. SNPs deviating from Hardy-Weinberg equilibrium were also filtered out: SNPs with a p-value less than 1 × 10⁻¹⁵ in the Hardy-Weinberg equilibrium test were removed. These SNPs may be affected by natural selection, mutation, migration, etc., and do not conform to the basic assumptions of genetics. Finally, SNPs with a genotypic missing rate greater than 10% were removed to reduce the impact of missing data on the analysis results. This step yielded genotypic data of individuals that met the criteria (including information on individuals that met the criteria and their corresponding SNPs).

[0086] Step S203: Match the target phenotypic data and the genotype data, include them as covariates, and perform genome-wide association analysis using a Firth-corrected generalized linear model to calculate the regression coefficient, standard error, and p-value for each SNP. The covariates are demographic variables such as sex, age, and the number of principal components of genes within a preset range.

[0087] In some embodiments, firstly, it is ensured that the target phenotypic data and genotypic data are matched based on the same individual ID for subsequent association analysis. Then, demographic variables (sex, age) and the top 10 principal components (PCs) are included as covariates in the analysis. These covariates can correct for the influence of confounding factors such as population structure on the analysis results, improving the accuracy of the analysis. Next, using the UKB-recommended tool REGENIE, a Firth-corrected generalized linear model is used for GWAS analysis. Firth correction can address the bias caused by rare variants, improving the robustness of the model. Finally, the regression coefficient (beta), standard error (se), and p-value for each SNP are calculated.

[0088] It should be noted that the regression coefficient represents the strength and direction of the association between the SNP and the metabolic characteristic score; the standard error is used to measure the accuracy of the regression coefficient; and the p-value is used to determine whether the association between the SNP and the metabolic characteristic score is statistically significant.

[0089] Step S204: Based on the regression coefficients, the standard error, and the p-value, determine the first trait GWAS data of insomnia metabolic characteristics.

[0090] In some embodiments, the regression coefficients, standard errors, and p-values ​​are compiled into a data table, with each row representing a SNP and each column recording the regression coefficient, standard error, and p-value of that SNP. Then, based on the Bonferroni-corrected significance threshold (p < 5 × 10⁻⁶), the data is analyzed. -8 We screened out SNPs that were significantly associated with the metabolic characteristics of insomnia. These significantly associated SNPs constituted the GWAS data of the first trait of the metabolic characteristics of insomnia.

[0091] It should be noted that the Manhattan plot of the genome-wide association analysis of insomnia metabolic characteristics is shown in the figure below. Figure 3 As shown, the horizontal axis represents the chromosome number, and the vertical axis represents the negative logarithm of the SNP's p-value, -log 10 (P-value), the red dashed line represents the significance threshold after Bonferroni correction (p<5×10). -8 ).

[0092] This GWAS analysis revealed the genetic basis of the metabolic characteristics of insomnia and provided important genetic evidence for subsequent causal inference.

[0093] In some embodiments, after obtaining the first trait GWAS data of insomnia metabolic characteristics, the method further includes: functionally annotating the SNPs in the first trait GWAS data using FUMA combined with gene annotation tools to determine the genomic location and functional impact of the SNPs; quantifying the pathogenic potential of the SNPs using CADD scores based on the genomic location and functional impact; determining the first candidate genes associated with the SNPs according to the order of pathogenic potential, combined with eQTL mapping, position mapping, and chromatin interaction mapping strategies; integrating the combined effects of the SNPs to the gene level using the MAGMA method to screen for second candidate genes phenotypes associated with the insomnia metabolic characteristics; and evaluating the tissue-specific expression patterns of the first and second candidate genes based on the GTEx database to identify genetic features and potential targets associated with the insomnia metabolic characteristics. Specifically, firstly, significant SNPs from GWAS are input into the FUMA (Functional Mapping and Annotation) platform for functional annotation. Then, the ANNOVAR (ANNOtate VARiation) tool is used to perform genomic localization and functional impact analysis on the SNPs, helping to identify their roles in different genes, especially their potential impact in coding regions, regulatory regions, or other functional elements. Next, the functionally annotated SNPs are input into the CADD tool to calculate a CADD score for each SNP. Then, using eQTL (expression-quantitative trait loci) data, the impact of SNPs with high CADD scores on gene expression was prioritized. EQTL databases (such as the GTEx database) were used to identify gene expression changes associated with significant SNPs, recognize genes associated with insomnia metabolic characteristics, and determine the relationship between SNPs and nearby genes based on their location in the genome. Simultaneously, a location mapping strategy was used to confirm genes associated with significant SNPs and analyze their spatial distribution within the genome, studying interactions between different gene regions to reveal whether SNPs regulate gene expression through the three-dimensional structure of chromatin. Chromatin interaction data (such as Hi-C data) was used to analyze long-distance regulatory relationships between SNPs and genes, thereby identifying first-line candidate genes. Subsequently, to more comprehensively assess the impact at the gene level, the genetic effects of significant SNPs were integrated into the gene level, quantifying the genetic burden of each gene in insomnia metabolic characteristics. Based on MAGMA analysis results, second-line candidate genes associated with insomnia metabolic characteristic phenotypes were preferentially screened.Finally, the expression patterns of candidate genes in different tissues were evaluated using the GTEx database. Based on the tissue-specific expression patterns of candidate genes, key tissues and systems related to insomnia metabolic characteristics were identified, and genetic features and potential targets related to insomnia metabolic characteristics were obtained. Among them, 11 significantly associated independent genetic variations at 9 loci were identified through FUMA analysis, as shown in Table 1.

[0094] Table 1. Nine gene loci and 11 significantly associated independent genetic variations related to metabolic characteristics of insomnia.

[0095]

[0096]

[0097] It's important to note that FUMA is a tool specifically designed for GWAS result annotation and visualization, providing detailed SNP functional annotation information. ANNOVAR can pinpoint the specific location of SNPs in the genome (e.g., coding regions, regulatory regions, introns) and assess their potential functional impact (e.g., missense mutations, splice site variations). CADD (Combined Annotation Dependent Depletion) scoring is a tool that integrates multiple annotation information (e.g., evolutionary conservation, variation functional impact) to assess the pathogenic potential of SNPs. MAGMA (Multi-marker Analysis of GenoMic Annotation) is an analytical method that integrates the combined effects of SNPs at the gene level. By combining intragene and intergenetic genetic effects, it quantifies the genetic burden of each gene in insomnia-related metabolic features, thereby prioritizing candidate genes associated with the phenotype. The GTEx (Genotype-Tissue Expression) database provides gene expression data in various tissues.

[0098] By systematically exploring the genetic basis and potential therapeutic targets of insomnia-related metabolic features through functional annotation, pathogenic potential assessment, gene mapping, and expression pattern analysis, we can provide a scientific basis for the study and intervention of the mechanisms of insomnia and related diseases.

[0099] In some embodiments, after obtaining the first trait data of the insomnia metabolic characteristic, the method further includes: extracting SNPs located within 1Mb of the transcription start site of the target gene and significantly associated with gene expression from eQTL data, mQTL data, and sQTL data; combining the SNPs with the insomnia metabolic characteristic, calculating the p-SMR value of each QTL molecular phenotype with the insomnia metabolic characteristic using an SMR framework analysis, and outputting significantly associated characteristics with p values ​​lower than a first preset threshold; combining the significantly associated characteristics, screening out causal characteristics with a true causal relationship with the insomnia metabolic characteristic using a HEIDI test; and combining the causal characteristics, mining gene expression related to the insomnia metabolic characteristic to obtain gene expression results. Specifically, firstly, to further explore gene expression influencing insomnia-related metabolic characteristics and corresponding regulatory mechanisms such as methylation sites and alternative splicing sites, whole blood eQTL data (31,684 samples) were obtained from the eQTLgen database, whole blood mQTL data (1,980 samples) from the McRae database, and whole blood sQTL data (755 samples) from the GTEx database. Then, based on genome annotation information, the transcription start site (TSS) of each target gene was determined. SNPs located within a 1Mb range upstream and downstream of the target gene's TSS were extracted from the eQTL, mQTL, and sQTL data. Furthermore, SNPs significantly associated with gene expression (p-value less than 5 × 10⁻⁶) were screened. -8 SNPs were identified, and these SNPs are considered to play a potentially important role in gene expression regulation. Then, using the SMR (Summary-data-based Mendelian Randomization) framework, p-SMR values ​​were calculated between each QTL molecular phenotype (gene expression, DNA methylation, alternative splicing) and insomnia metabolic characteristics. A first pre-defined threshold (e.g., a p-value threshold corrected for FDR) was set, and significantly associated characteristics with p-SMR values ​​below this threshold were screened out. These characteristics were considered to have a significant causal relationship with insomnia metabolic characteristics. Finally, the HEIDI (Heterogeneity in Dependent Instruments) test was performed on the significantly associated characteristics to distinguish between true causality and spurious associations caused by linkage disequilibrium. A threshold for p-HEIDI values ​​was set (e.g., p-HEIDI > 0.01), and causal characteristics with p-HEIDI values ​​above this threshold were screened out. These characteristics were considered to have a true causal relationship with insomnia metabolic characteristics, rather than being caused by linkage disequilibrium. Finally, combining causal characteristics, we further explored gene expression related to insomnia metabolic characteristics. Among these, after the above processing, 17 gene expressions related to insomnia-related metabolic characteristics were identified, such as... Figure 4 As shown.

[0100] It should be noted that specific methods for mining gene expression related to the metabolic characteristics of insomnia may include: eQTL analysis: for SNPs that are significantly associated with gene expression, further analyze their impact on gene expression levels; functional annotation: in combination with the GTEx database, perform functional annotation on the screened genes, evaluate their expression patterns in different tissues, and organize the mined gene expression results into a table, including information such as gene name, SNP location, p-value, and effect size.

[0101] By screening out gene expression features that have a real causal relationship with the metabolic characteristics of insomnia, we can delve into the genetic regulatory mechanisms of insomnia-related metabolic features and provide a scientific basis for disease risk warning and the discovery of therapeutic targets.

[0102] Step S104: Input the first trait GWAS data corresponding to the insomnia metabolic characteristics and the second trait GWAS data of the target disease feature set into the latent causal variable model, calculate the genetic causal ratio, and determine the causal risk warning result affected by the insomnia metabolic characteristics based on the genetic causal ratio.

[0103] In some embodiments, to further explore diseases that can influence insomnia-related metabolic characteristics and diseases influenced by insomnia-related metabolic characteristics, firstly, SNPs significantly associated with insomnia metabolic characteristics and their statistical information (such as regression coefficients, standard errors, and p-values) are obtained from genome-wide association studies (GWAS), and phenotypic datasets from Neal Lab (containing GWAS data for other diseases or related trait data) are acquired. Then, a latent variable L is introduced, assuming that this latent variable L has a causal effect on the two traits (insomnia metabolic characteristics and the target disease) through genetic correlation, where the latent variable L represents a common genetic background or underlying biological mechanism. Then, the first trait data (GWAS results) of insomnia metabolic characteristics were integrated with the second trait GWAS data of the target disease characteristic set to ensure that the SNPs in the two trait data could correspond for subsequent causal inference analysis. For each pair of traits (insomnia metabolic characteristics and target disease), the genetic causal ratio (GCP) was calculated, and the calculated GCP values ​​were subjected to a significance test to determine whether they were statistically significant. At the same time, the Benjamini–Hochberg method was used to correct the p-value to control the false discovery rate (FDR). When the absolute value of GCP was greater than 0.6 and the corrected FDR was less than 0.05, the causal pair was considered to have a strong causal relationship. Finally, the causal direction was determined based on the positive or negative value of GCP: if the GCP was positive, it indicated that the insomnia metabolic characteristics (trait A) had a causal effect on the target disease (trait B); if the GCP was negative, it indicated that the target disease (trait B) had a causal effect on the insomnia metabolic characteristics (trait A). For trait pairs with a strong causal relationship, the causal risk warning result is determined based on the causal direction and the magnitude of the GCP value. If the insomnia metabolic characteristic has a causal effect on the target disease (positive GCP value), a warning message is output, indicating that the insomnia metabolic characteristic may increase the risk of developing the target disease. If the target disease has a causal effect on the insomnia metabolic characteristic (negative GCP value), a warning message is output, indicating that the target disease may affect the insomnia metabolic characteristic.

[0104] It should be noted that the screening results for the causal association between insomnia-related metabolic characteristics and disease phenotypes are as follows: Figure 5 As shown, the horizontal axis GCP represents the potential causal effect weight between insomnia-related metabolic characteristics and disease phenotypes. A GCP absolute value greater than 0.6 indicates a causal relationship between the two; a GCP greater than 0.6 indicates that insomnia-related metabolic characteristics lead to the occurrence of other disease phenotypes, and a GCP less than -0.6 indicates that other disease phenotypes lead to the occurrence of insomnia-related metabolic characteristics. The vertical axis is -log. 10 (P-value) represents statistical significance, where the horizontal red dashed line is the Bonferroni correction threshold (-log). 10 (3.97×10-5 Therefore, the most significant endpoints above the threshold with GCP > 0.6 were: respiratory diseases, other serious medical conditions or disabilities diagnosed by a physician, history of alendronate sodium use, basal cell carcinoma, leg pain during walking, and urinary tract stones. These outcomes cover multiple systems including cardiovascular, oncology, musculoskeletal, and genitourinary, indicating that insomnia metabolic characteristics have cross-system pathogenic potential and providing initial screening evidence for downstream pathogenic risks of insomnia-related metabolic characteristics.

[0105] It should be noted that the genetic causality ratio (GCP) represents the proportion of genetic variation in which one trait has a causal effect on another trait through the latent variable L. The specific calculation method is not the focus of this application and will not be elaborated upon here.

[0106] It should be noted that the causal risk warning results include trait pairs with a strong causal relationship, the direction of causation, GCP values, and their significance p-values. These results can be used for further biological research and clinical applications, providing a basis for the prevention and intervention of insomnia and related diseases.

[0107] This invention utilizes a combination of advanced analytical methods, including elastic network regression, genome-wide association analysis, and Mendelian randomization, to construct an innovative multi-omics integrated analysis framework, providing new technical means and research paradigms for the study of insomnia and metabolic disorders. Furthermore, the research findings of this invention have significant practical value, not only contributing to a deeper understanding of the pathophysiological mechanisms of insomnia but also providing new targets and biomarkers for the diagnosis, treatment, and intervention of insomnia and related diseases, potentially promoting clinical applications and translational medicine research in these areas.

[0108] This invention improves data comparability and analytical accuracy by performing logarithmic transformation and standardization on metabolite characteristics, eliminating dimensional differences between different metabolites. Elastic network regression analysis effectively identifies metabolites significantly associated with insomnia. Weighted calculations of the selected metabolites integrate information from multiple metabolites, providing a more comprehensive metabolic profile of insomnia and offering a scientific basis for subsequent multi-omics analysis. GWAS analysis reveals the genetic basis of insomnia metabolic characteristics and provides important genetic evidence for subsequent causal inference. Combining metabolomics and genetic data constructs a multi-omics integrated analysis framework, laying the foundation for a comprehensive understanding of the complex network relationship between insomnia and metabolic disorders. Assessing the causal relationship between insomnia metabolic characteristics and target diseases allows for a deeper exploration of the intrinsic links between insomnia and related diseases. Based on the positive and negative values ​​of the genetic causal ratio, the direction and intensity of the causal influence of insomnia metabolic characteristics on target diseases are determined, thereby enabling causal risk warnings for related diseases and providing strong support for early disease warning, personalized treatment, and precision medicine. Compared with existing technologies, this application can accurately and efficiently identify diseases affected by the metabolic characteristics of insomnia.

[0109] like Figure 6 As shown, based on the above method embodiments, corresponding apparatus embodiments are provided;

[0110] An embodiment of the present invention provides a device based on the metabolic characteristics of insomnia, comprising: a construction module 100, a calculation module 200, an analysis module 300, and an early warning module 400;

[0111] The construction module 100 is used to construct a sample dataset based on the acquired sleep questionnaire data and metabolomics data, and to perform logarithmic transformation and standardization on the concentration of metabolites in the sample dataset to obtain standardized metabolite concentrations.

[0112] The calculation module 200 is used to input the standardized metabolite concentration into a preset elastic network model to obtain insomnia-related metabolites, and to perform weighted calculations on the metabolites to obtain a metabolic characteristic score related to insomnia. The elastic network model is obtained by optimizing the model parameters using the standardized metabolite concentration as the independent variable and the insomnia score as the dependent variable, combined with a cross-validation function. The insomnia score is obtained by identifying and assigning values ​​to the insomnia classification labels of the sleep questionnaire data.

[0113] The analysis module 300 is used to perform genome-wide association analysis (GWAS) on the metabolic characteristic score as phenotypic data and the genotype data of all samples obtained, to obtain the first trait GWAS data of insomnia metabolic characteristics.

[0114] The early warning module 400 is used to input the first trait GWAS data corresponding to the insomnia metabolic feature and the second trait GWAS data of the target disease feature set into a latent causal variable model to calculate the genetic causal ratio, and determine the causal risk early warning result affected by the insomnia metabolic feature based on the genetic causal ratio.

[0115] In some embodiments, the analysis module 300 includes a processing unit, a control unit, an analysis unit, and a combination unit:

[0116] The processing unit is used to use the metabolic characteristic score as phenotypic data and to process the phenotypic data to obtain target phenotypic data.

[0117] The control unit is used to acquire the initial genotype data of all samples and to perform quality control on the initial genotype data to obtain genotype data.

[0118] The analysis unit is used to match the target phenotypic data and the genotype data, include them as covariates, and perform genome-wide association analysis using a Firth-corrected generalized linear model to calculate the regression coefficient, standard error, and p-value for each SNP. The covariates are demographic variables such as sex, age, and the number of principal components of genes within a preset range.

[0119] The combined unit is used to determine the first trait data of insomnia metabolic characteristics based on the regression coefficient, the standard error, and the p-value.

[0120] It is understood that the above-described device embodiments correspond to the method embodiments of the present invention, and can implement the causal risk warning method based on insomnia metabolic characteristics provided by any of the above-described method embodiments of the present invention.

[0121] It should be noted that the device embodiments described above are merely illustrative, and some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Furthermore, in the accompanying drawings of the device embodiments provided by this invention, the connection relationships between modules indicate that they have communication connections, which can specifically be implemented as one or more communication buses or signal lines. Those skilled in the art can understand and implement this without any creative effort.

[0122] Based on the above embodiments of the causal risk warning method based on the metabolic characteristics of insomnia, another embodiment of the present invention provides a terminal device, which includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements the causal risk warning method based on the metabolic characteristics of insomnia of any embodiment of the present invention.

[0123] For example, in this embodiment, the computer program can be divided into one or more modules, which are stored in the memory and executed by the processor to complete the present invention. The one or more modules may be a series of computer program instruction segments capable of performing a specific function, which describe the execution process of the computer program in the terminal device.

[0124] The terminal device can be a desktop computer, laptop, handheld computer, or cloud server, etc. The terminal device may include, but is not limited to, a processor and memory.

[0125] The processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor. The processor is the control center of the terminal device, connecting various parts of the terminal device via various interfaces and lines.

[0126] Based on the above-described method embodiments, another embodiment of the present invention provides a computer-readable storage medium including a stored computer program, wherein, when the computer program is executed, it controls the device where the computer-readable storage medium is located to execute the causal risk warning method based on insomnia metabolic characteristics as described in any of the above-described method embodiments of the present invention.

[0127] The modules / units integrated in the device / terminal equipment, if implemented as software functional units and sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the above embodiments of the present invention can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc.

[0128] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications are also considered to be within the scope of protection of the present invention.

Claims

1. A causal risk early warning method based on the metabolic characteristics of insomnia, characterized in that, include: A sample dataset was constructed based on the acquired sleep questionnaire data and metabolomics data, and the concentrations of metabolites in the sample dataset were logarithmically transformed and standardized to obtain standardized metabolite concentrations. The standardized metabolite concentrations are input into a preset elastic network model to obtain insomnia-related metabolites. The metabolites are then weighted to obtain a metabolic characteristic score related to insomnia. The elastic network model is obtained by optimizing the model parameters using the standardized metabolite concentrations as independent variables and the insomnia score as the dependent variable, combined with a cross-validation function. The insomnia score is obtained by identifying and assigning values ​​to the insomnia classification labels of the sleep questionnaire data. The metabolic characteristic scores were used as phenotypic data, and genome-wide association analysis was performed with the genotype data of all samples to obtain the first trait GWAS data of insomnia metabolic characteristics. The GWAS data of the first trait corresponding to the insomnia metabolic feature and the GWAS data of the second trait of the target disease feature set are input into the latent causal variable model to calculate the genetic causal ratio, and the causal risk warning result affected by the insomnia metabolic feature is determined based on the genetic causal ratio.

2. The causal risk early warning method based on insomnia metabolic characteristics according to claim 1, characterized in that, The sample dataset constructed based on the acquired sleep questionnaire data and metabolomics data is as follows: Acquire sleep questionnaire data and metabolomics data, and remove individuals with invalid responses from the sleep questionnaire data to obtain the first processing result; Individuals with non-positive metabolite concentrations in the metabolomics data were removed to obtain the second processing result; The first processing result and the second processing result are merged according to individual numbers to obtain the sample dataset.

3. The causal risk early warning method based on insomnia metabolic characteristics according to claim 1, characterized in that, The process of performing logarithmic transformation and standardization on the metabolite concentrations in the sample dataset to obtain standardized metabolite concentrations is as follows: Logarithmically transform the concentration of each metabolite in the sample dataset, and calculate the mean and standard deviation of the transformed concentrations of each metabolite. The concentrations of each metabolite are converted into standardized values ​​based on the mean and the standard deviation to obtain standardized metabolite concentrations.

4. The causal risk early warning method based on insomnia metabolic characteristics according to claim 1, characterized in that, The metabolic trait score is used as phenotypic data, and a genome-wide association analysis is performed with the obtained genotype data of all samples to obtain GWAS data of the first trait of insomnia metabolic characteristics, specifically: The metabolic characteristic scores are used as phenotypic data, and the phenotypic data is processed to obtain the target phenotypic data. The initial genotype data of all samples is obtained, and the initial genotype data is subjected to quality control to obtain the genotype data. The target phenotypic data and the genotype data are matched and included as covariates. A Firth-corrected generalized linear model is used to perform genome-wide association analysis to calculate the regression coefficient, standard error, and p-value for each SNP. The covariates are demographic variables such as sex, age, and the number of principal components of genes within a preset range. Based on the regression coefficients, the standard error, and the p-value, the first trait GWAS data of insomnia metabolic characteristics were determined.

5. The causal risk early warning method based on insomnia metabolic characteristics according to claim 4, characterized in that, Following the acquisition of the first trait GWAS data for the metabolic characteristics of insomnia, the following is also included: The SNPs in the first trait GWAS data were functionally annotated using FUMA combined with gene annotation tools to determine the genomic location and functional impact of the SNPs. Based on the genomic location and functional impact, the pathogenic potential of the SNP is quantified by CADD scoring; Based on the order of pathogenic potential, and combined with eQTL mapping, position mapping, and chromatin interaction mapping strategies, the first candidate gene associated with the SNP is determined; The combined effects of the SNPs were integrated to the gene level using the MAGMA method to screen for second candidate genes associated with the phenotype of the insomnia metabolic characteristics. The tissue-specific expression patterns of the first and second candidate genes were evaluated based on the GTEx database to identify genetic features and potential targets related to the metabolic characteristics of insomnia.

6. The causal risk early warning method based on insomnia metabolic characteristics according to claim 4, characterized in that, Following the acquisition of the first trait GWAS data for the metabolic characteristics of insomnia, the following is also included: SNPs located within 1 Mb of the transcription start site of the target gene and significantly associated with gene expression were extracted from eQTL, mQTL and sQTL data. Combining the SNPs with the insomnia metabolic characteristics, the p-SMR value of each QTL molecular phenotype and the insomnia metabolic characteristics is calculated using the SMR framework, and significant correlation characteristics with p values ​​lower than a first preset threshold are output. Based on the aforementioned significant correlation features, the HEIDI test was used to screen out causal features that have a true causal relationship with the aforementioned insomnia metabolic features; Based on the aforementioned causal characteristics, gene expression related to the aforementioned insomnia metabolic characteristics was identified, and gene expression results were obtained.

7. The causal risk early warning method based on insomnia metabolic characteristics according to claim 1, characterized in that, The determination of the causal risk warning result affected by the insomnia metabolic characteristics based on the genetic causal ratio is specifically as follows: If the absolute value of the genetic causal ratio is greater than the second preset threshold, and the P-value after correction by the Benjamini-Hochberg method is lower than the third preset threshold, then it is determined that there is a strong causal relationship between the first trait data and the second trait data. If the genetic causal ratio is positive, then the causal risk warning of insomnia metabolic characteristics for the disease will be output.

8. A causal risk early warning system based on the metabolic characteristics of insomnia, characterized in that, include: The module includes a construction module, a calculation module, an analysis module, and an early warning module. The construction module is used to construct a sample dataset based on the acquired sleep questionnaire data and metabolomics data, and to perform logarithmic transformation and standardization on the concentration of metabolites in the sample dataset to obtain standardized metabolite concentrations. The calculation module is used to input the standardized metabolite concentration into a preset elastic network model to obtain insomnia-related metabolites, and to perform weighted calculations on the metabolites to obtain a metabolic characteristic score related to insomnia. The elastic network model is obtained by optimizing the model parameters using the standardized metabolite concentration as the independent variable and the insomnia score as the dependent variable, combined with a cross-validation function. The insomnia score is obtained by identifying and assigning values ​​to the insomnia classification labels of the sleep questionnaire data. The analysis module is used to perform genome-wide association analysis (GWAS) on the metabolic characteristic score as phenotypic data and the genotype data of all samples obtained to obtain the first trait GWAS data of insomnia metabolic characteristics. The early warning module is used to input the first trait GWAS data corresponding to the insomnia metabolic feature and the second trait GWAS data of the target disease feature set into the latent causal variable model to calculate the genetic causal ratio, and determine the causal risk early warning result affected by the insomnia metabolic feature based on the genetic causal ratio.

9. The causal risk early warning system based on insomnia metabolic characteristics according to claim 8, characterized in that, The analysis module includes a processing unit, a control unit, an analysis unit, and a combination unit: The processing unit is used to use the metabolic characteristic score as phenotypic data and to process the phenotypic data to obtain target phenotypic data. The control unit is used to acquire the initial genotype data of all samples and to perform quality control on the initial genotype data to obtain genotype data. The analysis unit is used to match the target phenotypic data and the genotype data, include them as covariates, and perform genome-wide association analysis using a Firth-corrected generalized linear model to calculate the regression coefficient, standard error, and p-value for each SNP. The covariates are demographic variables such as sex, age, and the number of principal components of genes within a preset range. The combined unit is used to determine the first trait GWAS data of insomnia metabolic characteristics based on the regression coefficient, the standard error, and the p-value.

10. A computer-readable storage medium, characterized in that, include: A stored computer program, wherein, when the computer program is executed, it controls the device containing the computer-readable storage medium to perform the steps of the causal risk warning method based on the metabolic characteristics of insomnia as described in any one of claims 1-7.