Noise hearing loss genetic risk prediction method and system
By constructing gene Panel and performing linkage imbalance pruning and principal component analysis, combined with multigene risk scores, the problem of insufficient accuracy in genetic risk prediction in the prior art is solved, and the accurate assessment and protection of individualized noise hearing loss risks are achieved.
Patent Information
- Application Number
- CN202510521294.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-08-01
AI Technical Summary
The existing genetic risk prediction methods for noise hearing loss only focus on a small number of genes, and the information coverage is low, resulting in insufficient accuracy and applicability of prediction results, making it difficult to accurately identify high-risk individuals.
Gene Panel was constructed, sequencing results were obtained through sequencing, chain unbalanced pruning and principal component analysis were performed, multiple principal component data were obtained by reducing the dimensionality, and genetic risk prediction was performed by combining multigene risk scores.
It improves the accuracy, stability and generalization ability of predictions, realizes a comprehensive assessment of individualized noise hearing loss risk, and provides a scientific basis for individualized protection.
Smart Images

Figure CN120413031A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of genetic testing, and particularly to a method and system for predicting the genetic risk of noise-induced hearing loss. Background Art
[0002] Noise-Induced Hearing Loss (NIHL) refers to irreversible hearing damage caused by long-term or short-term exposure to high-intensity noise environments, which is affected by both genetic and environmental factors. Currently, genetic risk prediction methods based on gene data have become the key means for the early prevention and precise prevention and control of NIHL. However, existing genetic risk prediction studies on noise-induced hearing loss only focus on individual genes or a limited number of SNPs. But NIHL is a complex trait regulated by multiple genes, and genetic variations at a single or a small number of loci are difficult to comprehensively explain individual susceptibility, resulting in insufficient information coverage of the prediction model and difficulty in accurately identifying high-risk individuals. Before or at the early stage of noise exposure, it is impossible to effectively provide accurate preventive advice for individuals. Summary of the Invention
[0003] This application provides a method and system for predicting the genetic risk of noise-induced hearing loss, which solves the technical problem in the prior art that due to the genetic risk prediction research only focusing on a small number of genes and having low information coverage, the accuracy and applicability of the prediction results are insufficient, and achieves the technical effect of improving the prediction accuracy, stability and generalization ability, so as to realize the individual risk assessment of noise-induced hearing loss.
[0004] In view of the above problems, on the one hand, this application provides a method for predicting the genetic risk of noise-induced hearing loss. The method includes: constructing a gene Panel based on the genomic data of noise-induced hearing loss, sequencing a sample through the gene Panel to obtain a sequencing result; performing linkage disequilibrium pruning on the sequencing result to obtain SNPs; performing dimensionality reduction on the SNPs based on principal component analysis to obtain multiple principal component data; calculating a polygenic risk score according to the sequencing result; and synchronizing the polygenic risk score and the multiple principal component data to a noise-induced hearing loss prediction model for analysis to determine a genetic risk prediction report.
[0005] On the other hand, the present application also provides a genetic risk prediction system for noise-induced hearing loss. The system includes: a gene sequencing module for constructing a gene panel based on the genomic data of noise-induced hearing loss, sequencing a sample through the gene panel to obtain a sequencing result; a linkage disequilibrium pruning module for performing linkage disequilibrium pruning on the sequencing result to obtain SNPs; a principal component analysis module for performing dimensionality reduction on the SNPs based on principal component analysis to obtain a plurality of principal component data; a polygenic risk score calculation module for calculating and obtaining a polygenic risk score according to the sequencing result; and a genetic risk prediction module for synchronizing the polygenic risk score and the plurality of principal component data to a noise-induced hearing loss prediction model for analysis to determine a genetic risk prediction report.
[0006] One or more technical solutions provided in the present application have at least the following beneficial effects:
[0007] Constructing a gene panel based on the genomic data of noise-induced hearing loss, sequencing a sample through the gene panel to obtain a sequencing result, narrowing the research scope, ensuring the accuracy and pertinence of subsequent analysis, and providing a necessary genetic data basis for the entire prediction process. Performing linkage disequilibrium pruning on the sequencing result to remove the sites with linkage disequilibrium in the sequencing result, thereby obtaining a plurality of more representative and independent single nucleotide polymorphism sites (SNPs) and reducing redundant information. Performing dimensionality reduction on the SNPs through principal component analysis to obtain a plurality of principal component data, which can reduce the data dimension while retaining the main information of the data, highlighting the main features and variation information in the data, thereby improving the calculation efficiency of the prediction model, reducing the interference of noise and irrelevant information in the data to the model, and improving the accuracy and reliability of the prediction result. Calculating and obtaining a polygenic risk score according to the sequencing result, synchronizing the polygenic risk score and a plurality of principal component data to a noise-induced hearing loss prediction model for analysis to determine a genetic risk prediction report, enabling the prediction model to comprehensively consider the genetic characteristics of an individual and realizing the assessment of the genetic risk of noise-induced hearing loss for the individual.
[0008] In summary, the present application constructs a comprehensive gene panel, combines linkage disequilibrium pruning to optimize the selection of SNPs, uses principal component analysis for dimensionality reduction to improve the data processing ability, quantifies individual susceptibility through polygenic risk score calculation, and finally obtains a genetic risk prediction report by comprehensively considering the principal component data and the polygenic risk score, significantly improving the accuracy, stability and generalization ability of the prediction, more comprehensively and accurately assessing the genetic risk of noise-induced hearing loss for an individual, and thus providing a scientific basis for individualized noise-induced hearing protection.
[0009] The above description is only an overview of the technical solution of this application. In order to understand the technical means of this application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of this application more obvious and understandable, the specific implementation manners of this application are specifically given below. Description of the Drawings
[0010] Figure 1 It is a schematic flowchart of a method for predicting genetic risk of noise-induced hearing loss provided by an embodiment of this application.
[0011] Figure 2 It is a schematic diagram of multiple principal component data obtained in a system for predicting genetic risk of noise-induced hearing loss provided by an embodiment of this application.
[0012] Figure 3 It is a schematic structural diagram of a system for predicting genetic risk of noise-induced hearing loss provided by an embodiment of this application.
[0013] Description of the reference numerals: gene sequencing module 10, linkage disequilibrium pruning module 20, principal component analysis module 30, polygenic risk score calculation module 40, genetic risk prediction module 50. Detailed Description of the Invention
[0014] By providing a method and a system for predicting genetic risk of noise-induced hearing loss, the embodiments of this application solve the technical problem that in the prior art, due to the fact that genetic risk prediction research only focuses on a small number of genes and the information coverage is low, the accuracy and applicability of the prediction results are insufficient, and achieve the technical effect of improving the prediction accuracy, stability and generalization ability, thereby realizing the individualized risk assessment of noise-induced hearing loss. The data collection and acquisition involved in the embodiments of this application are all carried out on the premise of not violating relevant regulations and personal privacy.
[0015] Embodiment 1, as Figure 1 shown, the embodiments of this application provide a method for predicting genetic risk of noise-induced hearing loss, and the method includes:
[0016] Step S1: Construct a gene Panel based on the genomic data of noise-induced hearing loss, and sequence a sample through the gene Panel to obtain a sequencing result.
[0017] Specifically, a gene panel refers to a set of specific genes that have been screened and are considered highly relevant to hearing loss. By referring to relevant research literature and databases, 559 single nucleotide polymorphism (SNP) sites related to noise-induced hearing loss were identified, and a gene panel was constructed. On this basis, multiplex PCR technology combined with next-generation sequencing (NGS) technology was used to sequence the genes of the noise-exposed population (samples), and quality control was performed on the obtained SNPs to eliminate low-quality (such as insufficient sequencing depth, low base quality score, etc.) SNPs, resulting in 532 SNPs that are highly relevant to hearing loss.
[0018] By constructing a gene panel to sequence the samples, sequencing results were obtained, narrowing the research scope from the entire genome to a set of genes that may be related to noise-induced hearing loss, improving the pertinence of subsequent analysis, and reducing the complexity of data processing.
[0019] Step S2: Perform linkage disequilibrium pruning on the sequencing results to obtain SNPs.
[0020] Specifically, for the identified SNP sites, statistical analysis methods are used to detect the linkage disequilibrium between gene sites. Linkage disequilibrium refers to the phenomenon that genes at different sites are not randomly combined, but certain alleles often appear together. Finally, through linkage disequilibrium pruning, 357 low-redundancy SNPs were obtained from 532 SNPs to improve the efficiency and accuracy of subsequent prediction analysis.
[0021] By removing the sites with linkage disequilibrium, the problem of multicollinearity is reduced, and multiple more representative and independent SNPs are obtained, which helps to improve the accuracy of subsequent analysis, avoid data analysis bias caused by the mutual influence between linkage disequilibrium sites, and at the same time reduce the computational complexity.
[0022] Step S3: Perform dimensionality reduction on the SNPs based on principal component analysis to obtain multiple principal component data.
[0023] Specifically, the SNP data obtained in Step S2 is used as input, and data dimensionality reduction processing is performed using the principal component analysis algorithm. Principal component analysis constructs a new coordinate system based on the correlation between SNP data. In this coordinate system, each principal component is a linear combination of multiple SNPs. For 357 low-redundancy SNPs, through principal component analysis, multiple principal component data are screened out, and 171 principal component data can exceed 80% of the total variability based on multiple principal component data.
[0024] Through principal component analysis, it is possible to reduce the dimensionality of the data while retaining the main information of the original SNPs data, reduce the complexity of the data, improve the efficiency of subsequent calculations, and at the same time reduce the interference of noise and irrelevant information in the data on the model, making the prediction model more stable and effective.
[0025] Step S4: Calculate and obtain a polygenic risk score based on the sequencing results.
[0026] Specifically, the polygenic risk score is a quantitative indicator reflecting an individual's disease risk, calculated based on the genetic variant sites carried by the individual that are related to noise-induced hearing loss. When calculating the polygenic risk score, the gene data of the individual is calculated according to the SNPs related to noise-induced hearing loss and their corresponding weights determined in advance in the sequencing results. For example, for SNPs sites related to noise-induced hearing loss, a weight value is assigned according to their contribution to noise-induced hearing loss, and the genotypes of the individual at these sites are added up according to the weights to obtain the polygenic risk score.
[0027] Step S5: Synchronize the polygenic risk score and the multiple principal component data to a noise-induced hearing loss prediction model for analysis to determine a genetic risk prediction report.
[0028] Specifically, the calculated polygenic risk score and the multiple principal component data obtained in step S3 are input into a noise-induced hearing loss prediction model together. This prediction model is constructed based on machine learning algorithms, such as a logistic regression model, which can analyze the input data, judge the individual's genetic risk, and generate a genetic risk prediction report. By comprehensively considering the polygenic risk score and the principal component data, the noise-induced hearing loss prediction model can comprehensively evaluate the individual's genetic risk. The finally generated genetic risk prediction report can provide a basis for the prevention and early intervention of an individual's noise-induced hearing loss.
[0029] Furthermore, step S1 of the embodiment of the present application includes:
[0030] Step S11: Obtain candidate SNPs through alignment identification of the noise-induced hearing loss genomic data.
[0031] Step S12: Based on the candidate SNPs, perform on-machine sequencing according to the multiplex PCR technology combined with NGS to obtain sequencing results.
[0032] Step S13: Eliminate the candidate SNPs according to the sequencing results to obtain the remaining SNPs, perform quality control on the remaining SNPs, and construct the gene Panel.
[0033] Specifically, by referring to relevant research literature and databases, and comparing individual genomic data, 559 single nucleotide polymorphism sites (SNPs) related to noise-induced hearing loss were determined as candidate SNPs. Among the candidate SNPs, due to factors such as homology in the adjacent sequences of some sites and primer interaction, 27 SNPs could not be effectively detected in the experimental tests. Therefore, these sites were excluded to construct an initial gene Panel.
[0034] According to the gene Panel, multiplex PCR primers targeting these sites were designed. Then, the sample DNA was amplified using multiplex PCR technology. Next, the amplified products (DNA fragments containing candidate SNPs) were subjected to next-generation sequencing (NGS, Next Generation Sequencing). Meanwhile, to more comprehensively cover genetic information, sequences of 100 bp upstream and downstream of each included site were extracted and added to enhance the comprehensiveness and coverage of the detection. Then, quality control operations were performed on these remaining SNPs. SNPs with low quality (such as insufficient sequencing depth, low base quality score, etc.) were processed, and SNPs with a gene deletion rate > 10%, deviation from Hardy-Weinberg equilibrium (HWE), and minor allele frequency (MAF) < 0.5% were excluded; finally, 532 high-quality SNPs were retained as the final sites for subsequent analysis.
[0035] Furthermore, step S2 of the embodiment of the present application includes:
[0036] Based on the gene Panel, a sliding window is defined to determine fixed window data; the local linkage disequilibrium relationship is analyzed according to the fixed window data, and the window data is pruned to obtain the SNPs.
[0037] Specifically, the sliding window method is used for linkage disequilibrium pruning. The sliding window method analyzes the characteristics of a local region by defining a window of a fixed size on the genome and then sliding it forward by a certain step size. The SNPs, base pairs, or other features within the window are processed one by one to help capture local patterns or regularities. First, the size of the sliding window and the sliding step size are determined. According to the scale of the gene Panel and the expected analysis accuracy, the data range for each analysis is defined as 50 kb, and the distance for each slide is 5 bp, ensuring that the step size is less than the window to guarantee an overlapping range and prevent information omission. Then, starting from the starting position of the gene Panel, it slides according to the set window size and step size. At each slide, the gene data (SNPs, base pairs, or other features) within the window is determined as fixed window data. Using the sliding window method to determine window data can transform the overall analysis of the gene Panel into local and step-by-step analysis, thus better capturing the relationships between different regions within the gene Panel.
[0038] Local linkage disequilibrium refers to the linkage disequilibrium between loci within a fixed window of data. That is, the relationship between certain alleles in this local data range is non-random. For each fixed window of data, statistical analysis methods are used to analyze the local linkage disequilibrium between loci. For example, the PLINK tool can be used to calculate the r between SNPs. 2 The degree of linkage disequilibrium was measured by the value of 1. The gene loci in the window data were pruned according to the pruning threshold, and those redundant loci with high linkage disequilibrium were removed. 357 low-correlation or independent SNPs were retained to reduce information redundancy.
[0039] Linkage disequilibrium pruning effectively removes redundant SNPs and reduces computational complexity. The pruned SNPs set can not only represent the genetic information in the original data, but also avoid the model overfitting problem caused by redundant data, thereby improving the stability and generalization ability of the genetic risk prediction model.
[0040] Furthermore, step S3 of the embodiment of the present application includes:
[0041] Step S31: converting the genotype data of the SNPs into a standardized matrix.
[0042] Step S32: performing dimensionality reduction calculation on the standardized matrix through the covariance matrix to obtain data eigenvalues and data eigenvectors, wherein the data eigenvalues and the data eigenvectors have a corresponding relationship.
[0043] Step S33: Serially extract the SNPs according to the data eigenvalues and the data eigenvectors to obtain the plurality of principal component data.
[0044] Specifically, linkage disequilibrium pruning removes redundant SNPs, resulting in a representative set of SNP data. However, even after pruning, high-dimensional data may still exist, increasing computational complexity and affecting the training of subsequent prediction models. Therefore, principal component analysis is needed to reduce the dimensionality of SNPs and extract the most representative genetic features.
[0045] Genotype data refers to the specific combination form of an individual's genes. For SNPs, it is the combination of alleles at each locus. For example, a SNPs locus may have two alleles A and T, and an individual's genotype at this locus may be AA, AT, or TT. For the genotype data of each SNPs locus, using the 0, 1, 2 coding method, the SNPs genotype data is numericalized: AA is coded as 0, AT is coded as 1, and TT is coded as 2, representing homozygous wild type, heterozygous type, and homozygous mutant type respectively. Then, use statistical analysis software to convert these numerical data into a standardized matrix. In this process, according to the Z-score standardization formula: X new =(X - μ) / σ (where X new is the value after standardization, X is the original value, μ is the mean, and σ is the standard deviation), calculate each data point to obtain a standardized matrix, so that all data has the same mean (0) and standard deviation (1). The standardized matrix eliminates the scale effect of SNPs data, making the data distributions of different SNPs comparable and improving the accuracy of principal component analysis.
[0046] The covariance matrix is a matrix used to describe the degree of linear relationship (correlation) between two different SNPs. Based on the obtained standardized matrix, calculate its covariance matrix. This can be achieved through matrix operation algorithms. Then, perform eigenvalue decomposition on the covariance matrix to calculate the data eigenvalues and data eigenvectors. This process can use linear algebra algorithms. In this process, each data eigenvalue corresponds to a specific data eigenvector. Through covariance matrix analysis, the main variation directions between SNPs can be found, and the most important information can be screened through eigenvalues and eigenvectors to reduce the data dimension.
[0047] According to the correspondence between the obtained data eigenvalues and data eigenvectors, sort the data eigenvectors according to the size of the data eigenvalues. The larger the data eigenvalue, the greater the amount of information contained in the corresponding principal component. Then, based on the sorted eigenvectors, perform principal component extraction operations from the original SNPs data to obtain the principal component analysis results (as Figure 2 shown). According to the principal component analysis results, it is found that when using 171 principal components, more than 80% of the total variability can be explained. Although the original data may have a higher dimension or complexity, the main information and patterns can be effectively summarized by 171 principal components. These principal components capture most of the structural features in the data, reduce redundancy, and at the same time retain the main genetic variation information.
[0048] By serializing and extracting SNPs according to data eigenvalue and data eigenvector, the most representative principal component data can be effectively extracted from the original SNPs data. These principal component data retain the main information of the original SNPs data, while achieving data dimensionality reduction, reducing data complexity, and improving the efficiency and accuracy of subsequent analysis.
[0049] Furthermore, step S4 of the embodiment of the present application includes:
[0050] Step S41: Perform linear regression on the SNPs according to the sequencing result to obtain the effect value of the hearing decibel value of the SNPs.
[0051] Step S42: Screen according to the effect value to obtain the decibel contribution degree, and construct a polygenic risk calculation formula.
[0052] Step S43: Calculate through the polygenic risk calculation formula to obtain the polygenic risk score.
[0053] Specifically, a linear regression method is used to estimate the effect value of SNPs on hearing loss, and a polygenic risk calculation formula is constructed to calculate the polygenic risk score of an individual, so as to quantify the genetic susceptibility of an individual to noise exposure. Among them, the effect value represents the influence magnitude of a certain SNPs on the hearing decibel value. It is usually represented by the β value (regression coefficient), and the larger the value, the more significant the influence of the SNPs on the hearing decibel value.
[0054] To evaluate the contribution of genetic loci to the hearing decibel value, linear regression analysis is performed with SNPs as the independent variable and the hearing decibel value as the dependent variable. A linear regression model in the form of Y = β0 + β i G i + ∈ is constructed, where Y is the average hearing threshold (hearing decibel value) of high-frequency range hearing, β0 is the intercept term, G i is the genotype value of the i-th locus (taking values of 0, 1, or 2, representing no, carrying one, or two risk alleles respectively), β i is the effect estimate of this locus on the hearing decibel value, reflecting its decibel contribution degree, and ∈ represents random error. The coefficient β i of the regression equation is calculated according to the least squares method by statistical software (such as the lm function in R language) as the effect value of the hearing decibel value.
[0055] The decibel contribution degree is further screened based on the effect value, and is an index used to represent the contribution degree of each SNPs to the hearing decibel value, reflecting a measure of the relative importance of each SNPs when constructing a genetic risk assessment. The polygenic risk calculation formula is a mathematical expression for calculating the polygenic risk score of an individual based on the decibel contribution degree of SNPs. According to the effect value βi Based on the magnitude and statistical significance, loci that significantly affect the hearing decibel value are screened out. Then, a polygenic risk calculation formula is constructed according to the decibel contribution degree: where RSK j represents the genetic risk value, j represents the j-th individual, i represents the i-th locus, and E i represents the decibel contribution degree of locus i, and D ij represents the genotype value (0, 1, or 2) of the i-th locus of the j-th individual, and N j represents the number of loci of the j-th individual. According to the constructed polygenic risk calculation formula, a distribution analysis is performed on the polygenic risk scores of all individuals in the collected sample dataset, and the overall trend of genetic risk in the sample is presented in the form of a density curve. First, the polygenic risk score data of all sample individuals are sorted to form a distribution dataset, and methods such as kernel density estimation are used to plot the probability density curve to visually display the distribution of genetic risk. In the specific implementation process, the normality of the polygenic risk score is verified through the Shapiro-Wilk test, and it is found that the polygenic risk score follows a normal distribution, indicating that the risk values of most subjects are concentrated around the average value, while individuals with extremely high or low risks are in the minority. This distribution characteristic conforms to the genetic laws of most polygenic complex traits, proving the rationality of the constructed polygenic risk calculation formula.
[0056] By screening the effect values to obtain the decibel contribution degree and constructing the polygenic risk calculation formula, it is possible to focus on the SNPs that make important contributions to the hearing decibel value, making the polygenic risk calculation formula more scientific and reasonable, improving the accuracy and pertinence of the polygenic risk score calculation, and being able to better reflect the individual's noise-induced hearing loss risk based on genetic factors.
[0057] After constructing the polygenic risk calculation formula, the actual SNP data of an individual are substituted into the formula for calculation to obtain the polygenic risk score of each individual, providing a quantitative genetic risk indicator for noise-induced hearing loss for the individual.
[0058] Furthermore, the construction process of the noise-induced hearing loss prediction model described in step S5 of the embodiment of the present application includes:
[0059] Step S5-1: Use the polygenic risk score and the multiple principal component data as input features, and use the noise-induced hearing loss phenotype as the label.
[0060] Step S5-2: Train the model according to the label for the input features to generate a model training result, and evaluate the model training result to obtain a model performance score.
[0061] Step S5-3: Perform logistic regression based on the model performance score to construct the noise-induced hearing loss prediction model.
[0062] Specifically, collect a sample dataset containing polygenic risk scores, multiple principal component data, and corresponding noise-induced hearing loss phenotypes. The polygenic risk scores and multiple principal component data serve as input features, which contain gene information related to noise-induced hearing loss and processed relevant data. These data will be used by the model to predict the situation of noise-induced hearing loss. The noise-induced hearing loss phenotype refers to whether an individual has noise-induced hearing loss, which is used as a label in the model to judge the accuracy of the prediction. In the specific implementation process, 754 cases of sample gene data were collected, and the polygenic risk scores were calculated using the aforementioned polygenic risk calculation formula, and the first 30 principal component data were extracted to construct the prediction model, and the model was tested in 339 cases of samples.
[0063] Adopt different machine learning methods, including random forest, logistic regression, Xgboost, multi-layer perceptron, support vector machine, and decision tree algorithms, and divide the input features and labels into multiple training data groups to train the model. During the training process, the model will continuously adjust its own parameters according to the relationship between the input features and labels (such as the support vectors and decision boundary parameters in the support vector machine, and the model weights and bias parameters in the logistic regression), so as to obtain multiple model training results.
[0064] Use AUC, sensitivity, specificity, and accuracy to evaluate the performance of multiple trained models to obtain the model performance score. In terms of the AUC index, logistic regression performs excellently, significantly higher than other models. The AUC value of the multi-layer perceptron is also relatively high. In terms of sensitivity, the sensitivity of logistic regression is significantly higher than other models, indicating that it has a stronger ability to capture the target category. The specificity performance shows that Xgboost and the multi-layer perceptron have extremely high specificity in some groups, indicating that they are very accurate in distinguishing non-target categories. In contrast, logistic regression shows a relatively balanced performance between sensitivity and specificity. In terms of accuracy, both logistic regression and the multi-layer perceptron perform well, while the accuracy of random forest and decision tree is relatively low.
[0065] Based on the evaluation results of each index, in the NIHL prediction model constructed based on polygenic risk scores and principal components, select logistic regression as the optimal model to construct an accurate noise-induced hearing loss prediction model.
[0066] By evaluating the training results of multiple models to obtain the model performance score, the performance advantages and disadvantages of the model can be intuitively understood, which helps to construct a more accurate and reliable noise-induced hearing loss prediction model.
[0067] In summary, the method for predicting genetic risk of noise-induced hearing loss provided by the embodiments of the present application has the following beneficial effects:
[0068] A gene panel is constructed based on the genomic data of noise-induced hearing loss, and the sample is sequenced through the gene panel to obtain sequencing results, which narrow the research scope, ensure the accuracy and pertinence of subsequent analysis, and provide a necessary genetic data basis for the entire prediction process. Linkage disequilibrium pruning is performed on the sequencing results to remove the sites with linkage disequilibrium in the sequencing results, thereby obtaining multiple more representative and independent single nucleotide polymorphism sites and reducing redundant information. Dimensionality reduction is performed on multiple single nucleotide polymorphism sites through principal component analysis to obtain multiple principal component data, which can reduce the data dimension while retaining the main information of the data, highlighting the main features and variation information in the data, thereby improving the calculation efficiency of the prediction model, reducing the interference of noise and irrelevant information in the data on the model, and improving the accuracy and reliability of the prediction results. The effect value of the hearing decibel value of SNPs is analyzed through linear regression analysis, and a polygenic risk calculation formula is constructed accordingly, and the polygenic risk score is calculated according to the sequencing results. This step converts genetic information into a quantified risk score, which can scientifically reflect the risk of noise-induced hearing loss faced by an individual due to genetic factors and provide key input features for the subsequent prediction model. The polygenic risk score and multiple principal component data are synchronized to the noise-induced hearing loss prediction model for analysis to determine the genetic risk prediction report, enabling the prediction model to comprehensively consider the genetic characteristics of an individual and realize the assessment of the genetic risk of noise-induced hearing loss for an individual.
[0069] Overall, the embodiments of the present application construct a comprehensive gene panel, optimize the selection of single nucleotide polymorphism sites by combining linkage disequilibrium pruning, and use principal component analysis for dimensionality reduction to improve data processing capabilities; at the same time, the individual susceptibility is quantified through polygenic risk score calculation, and finally a genetic risk prediction report is obtained, significantly improving the accuracy, stability and generalization ability of the prediction, more comprehensively and accurately assessing the genetic risk of noise-induced hearing loss for an individual, and thus providing a scientific basis for individualized noise-induced hearing protection.
[0070] Embodiment 2, as Figure 3 shown, based on the same inventive concept as in the foregoing Embodiment 1, the embodiments of the present application provide a system for predicting genetic risk of noise-induced hearing loss, and the system includes:
[0071] A gene sequencing module 10, configured to construct a gene panel based on the genomic data of noise-induced hearing loss, and sequence a sample through the gene panel to obtain sequencing results.
[0072] The linkage disequilibrium pruning module 20 is used to perform linkage disequilibrium pruning on the sequencing results to obtain SNPs.
[0073] The principal component analysis module 30 is used to reduce the dimension of the SNPs based on principal component analysis to obtain multiple principal component data.
[0074] The polygenic risk score calculation module 40 is used to calculate and obtain a polygenic risk score according to the sequencing results.
[0075] The genetic risk prediction module 50 is used to synchronize the polygenic risk score and the multiple principal component data to a noise-induced hearing loss prediction model for analysis to determine a genetic risk prediction report.
[0076] Furthermore, the gene sequencing module 10 of the embodiment of the present application is further used to perform the following steps:
[0077] Perform alignment identification through the noise-induced hearing loss genomic data to obtain candidate SNPs; perform experimental verification on the candidate SNPs based on the multiple PCR technology combined with NGS to obtain an experimental verification result; eliminate the candidate SNPs according to the experimental verification result to obtain remaining SNPs, and perform quality control on the remaining SNPs to construct the gene Panel.
[0078] Furthermore, the linkage disequilibrium pruning module 20 of the embodiment of the present application is further used to perform the following steps:
[0079] Define a sliding window based on the gene Panel to determine fixed window data; analyze the local linkage disequilibrium relationship according to the fixed window data, and prune the window data to obtain the SNPs.
[0080] Furthermore, the principal component analysis module 30 of the embodiment of the present application is further used to perform the following steps:
[0081] Convert the genotype data of the SNPs into a standardized matrix; perform dimensionality reduction calculation on the standardized matrix through a covariance matrix to obtain data eigenvalues and data eigenvectors, and there is a corresponding relationship between the data eigenvalues and the data eigenvectors; serially extract the SNPs according to the data eigenvalues and the data eigenvectors to obtain the multiple principal component data.
[0082] Furthermore, the polygenic risk score calculation module 40 of the embodiment of the present application is further used to perform the following steps:
[0083] Perform linear regression on the SNPs according to the sequencing results to obtain the effect values of the hearing decibel values of the SNPs; screen according to the effect values to obtain the decibel contribution degree, and construct a polygenic risk calculation formula; calculate through the polygenic risk calculation formula to obtain the polygenic risk score.
[0084] Further, the system in the embodiment of the present application further includes a prediction model construction module, and the prediction model construction module is used to perform the following steps:
[0085] Use the polygenic risk score and the multiple principal component data as input features, and use the noise-induced hearing loss phenotype as a label; perform model training on the input features according to the label to generate a model training result, evaluate the model training result to obtain a model performance score; perform logistic regression based on the model performance score to construct the noise-induced hearing loss prediction model.
[0086] Through the foregoing detailed description of a method for predicting the genetic risk of noise-induced hearing loss in this specification, those skilled in the art can clearly know a system for predicting the genetic risk of noise-induced hearing loss in this embodiment. For the system disclosed in Embodiment 2, since it corresponds to the method disclosed in Embodiment 1, it has corresponding functional modules and beneficial effects. For the relevant parts, refer to the description in the method part.
[0087] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to these embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for predicting genetic risk of noise-induced hearing loss, characterized in that, The method includes: Constructing a gene panel based on genomic data of noise-induced hearing loss, sequencing a sample through the gene panel, and obtaining a sequencing result; Performing linkage disequilibrium pruning on the sequencing result to obtain SNPs; Reducing the dimension of the SNPs based on principal component analysis to obtain multiple principal component data; Calculating and obtaining a polygenic risk score according to the sequencing result; Synchronizing the polygenic risk score and the multiple principal component data to a noise-induced hearing loss prediction model for analysis to determine a genetic risk prediction report.
2. The genetic risk prediction method for noise-induced hearing loss according to claim 1, wherein Constructing a gene panel based on genomic data of noise-induced hearing loss, the method includes: Performing alignment identification through the genomic data of noise-induced hearing loss to obtain candidate SNPs; Performing experimental verification on the candidate SNPs based on the multiple PCR technology combined with NGS to obtain an experimental verification result; Eliminating the candidate SNPs according to the experimental verification result to obtain remaining SNPs, performing quality control on the remaining SNPs, and constructing the gene panel.
3. The genetic risk prediction method for noise-induced hearing loss according to claim 1, wherein Performing linkage disequilibrium pruning on the sequencing result to obtain SNPs, the method includes: Defining a sliding window based on the gene panel to determine fixed window data; Analyzing the local linkage disequilibrium relationship according to the fixed window data, pruning the window data, and obtaining the SNPs.
4. The genetic risk prediction method for noise-induced hearing loss according to claim 1, wherein Reducing the dimension of the SNPs based on principal component analysis to obtain multiple principal component data, the method includes: Converting the genotype data of the SNPs into a standardized matrix; Performing dimensionality reduction calculation on the standardized matrix through a covariance matrix to obtain data eigenvalues and data eigenvectors, and there is a corresponding relationship between the data eigenvalues and the data eigenvectors; Sequentially extracting the SNPs according to the data eigenvalues and the data eigenvectors to obtain the multiple principal component data.
5. The genetic risk prediction method for noise-induced hearing loss according to claim 1, wherein The process of calculating and obtaining a polygenic risk score, the method includes: Performing linear regression on the SNPs according to the sequencing result to obtain the effect value of the hearing decibel value of the SNPs; Screening according to the effect value to obtain the decibel contribution degree and constructing a polygenic risk calculation formula; Calculating through the polygenic risk calculation formula to obtain the polygenic risk score.
6. The genetic risk prediction method for noise-induced hearing loss according to claim 1, wherein The construction process of a noise-induced hearing loss prediction model, the method includes: Using the polygenic risk score and the multiple principal component data as input features and using the noise-induced hearing loss phenotype as a label; Training a model on the input features according to the label to generate a model training result, evaluating the model training result, and obtaining a model performance score; Performing logistic regression based on the model performance score to construct the noise-induced hearing loss prediction model.
7. A genetic risk prediction system for noise-induced hearing loss, characterized in that, The system is used to execute a method for predicting genetic risk of noise-induced hearing loss according to any one of claims 1-6, and includes: A gene sequencing module, configured to construct a gene panel based on genomic data of noise-induced hearing loss, sequence a sample through the gene panel, and obtain a sequencing result; Linkage disequilibrium pruning module, which is used to perform linkage disequilibrium pruning on the sequencing results to obtain SNPs; Principal component analysis module, which is used to perform dimensionality reduction on the SNPs based on principal component analysis to obtain multiple principal component data; Polygenic risk score calculation module, which is used to calculate and obtain a polygenic risk score according to the sequencing results; Genetic risk prediction module, which is used to synchronize the polygenic risk score and the multiple principal component data to a noise-induced hearing loss prediction model for analysis to determine a genetic risk prediction report.