Noise-induced hearing loss prediction system based on gee model and genetic signature of noisy speech

By constructing a noise-induced speech frequency hearing loss prediction system based on the GEE model, and combining genetic characteristics and environmental noise data, the system dynamically assesses the risk of individual speech frequency hearing loss, solving the problem that the time-varying nature of genetic characteristics is not considered in existing technologies, and achieving more accurate risk prediction and personalized protection.

CN121687522BActive Publication Date: 2026-04-10SHANGHAI SIXTH PEOPLES HOSPITAL
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-02-11
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing technologies fail to adequately consider the time-varying and interactive effects of genetic characteristics when predicting hearing loss at language frequencies, leading to systematic biases in predictions and an inability to provide accurate, personalized protection recommendations.

Method used

A noise-induced speech frequency hearing loss prediction system based on the GEE model and genetic characteristics is adopted. Through data acquisition, processing, feature selection, time-varying interactive modules and risk prediction, a generalized estimation equation model is constructed. Combined with genetic risk scores and environmental noise data, the risk of individual speech frequency hearing loss is dynamically assessed.

Benefits of technology

It improves the accuracy and stability of speech frequency hearing loss risk prediction, can quantify the dynamic interaction between genetic factors and environmental noise exposure, provides personalized hearing protection recommendations, and realizes the transformation from passive treatment to active prevention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121687522B_ABST
    Figure CN121687522B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of biostatistics, and discloses a noise language frequency hearing loss prediction system based on a GEE model and genetic characteristics, comprising: a data acquisition module that acquires follow-up data of workers; a data processing module that generates standardized modeling data sets; a feature screening module that screens features through statistical learning methods to obtain a key feature subset; a time-varying interaction module that constructs a generalized estimating equation model and outputs regression coefficient estimates; a risk prediction module that outputs individual risk probabilities and risk stratification results; and a decision support module that generates individual hearing protection recommendations; the present application processes follow-up data of workers exposed to occupational noise by constructing a generalized estimating equation model, thereby improving the accuracy and stability of language frequency hearing loss risk prediction, providing a basis for developing personalized hearing protection programs and recommending suitable hearing protection devices, and helping to realize the transition from passive treatment to active prevention of occupational hearing loss.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of biostatistics, more particularly, it relates to a noise-induced language frequency hearing loss prediction system based on a GEE model and genetic characteristics. BACKGROUND

[0002] Noise-induced language frequency hearing loss is one of the most common occupational diseases worldwide, which seriously affects the language communication ability and quality of life of workers. Unlike high-frequency hearing loss, language frequency hearing impairment directly affects the social communication ability and occupational adaptability of individuals, and is the focus of occupational health protection.

[0003] The existing Chinese patent with the authorization announcement number CN111584065B discloses a noise-induced hearing loss prediction and susceptible population screening method, device, terminal and medium. It includes: collecting multiple hearing feature data of noise exposure population samples and preprocessing; defining high-frequency hearing threshold notch hearing feature data based on the preprocessed data of the population samples, and constructing a high-frequency hearing threshold notch area prediction model for predicting individual susceptibility; obtaining medical feature data and hearing threshold measurement data in the to-be-tested population sample; comparing the predicted notch area value and the actual notch area value to determine the individual susceptibility of the to-be-tested population sample. Through big data prediction and long-term clinical experience, the application assists in judging hearing loss and screening susceptible and resistant individuals, fills the gap in early auxiliary diagnosis of noise-induced hearing loss, solves the problem that there is no gold standard for early diagnosis of noise-induced hearing loss at present, and helps patients to prevent early, which is the key to preventing and treating noise-induced hearing loss. The existing Chinese patent with the authorization announcement number CN111593108B discloses a detection method, kit and application of polymorphism of the 7q36.3 region related to the occurrence of noise-induced hearing loss. The method is Sequenom typing, further relates to a kit for detecting the above-mentioned susceptible site and its application. In addition, the SNP site rs10081191 in the 7q36.3 region is associated with the risk of NIHL, and the SNP and its application are disclosed. Among them, the SNP is located in the human genome region 7q36.3, and it is determined that 7q36.3 is a new susceptible gene region of NIHL, which can be effectively used for screening noise-induced hearing loss disease susceptible population and determining noise-induced hearing loss disease susceptible individuals.

[0004] However, due to the relatively slow change in hearing threshold for language frequency and the adjustment by multiple physiological factors, the influence of genetic characteristics may exhibit more complex time-varying characteristics. For example, certain genetic mutations may only manifest their effects after a certain threshold of cumulative noise exposure is reached, or exhibit different risk patterns at different age stages. The lack of such time-varying interaction effects results in systematic bias in the existing technology in predicting language frequency hearing loss, and fails to provide precise individualized protection recommendations for workers carrying specific genetic variations. SUMMARY

[0005] The present application aims to provide a noise-induced language frequency hearing impairment prediction system based on GEE model and genetic characteristics to solve the above problems.

[0006] The present application provides a noise-induced language frequency hearing impairment prediction system based on GEE model and genetic characteristics, comprising:

[0007] A data acquisition module acquires follow-up data of workers exposed to occupational noise from multiple data sources;

[0008] A data processing module, based on the follow-up data, performs data cleaning, outlier processing, missing value filling and variable standardization to generate a standardized modeling data set;

[0009] A feature selection module, based on the modeling data set, selects features significantly related to language frequency hearing loss through statistical learning methods to obtain a key feature subset;

[0010] A time-varying interaction module, based on the key feature subset and the modeling data set, constructs a generalized estimating equation model including time-varying interaction terms, and the generalized estimating equation model outputs regression coefficient estimates;

[0011] A risk prediction module, based on the regression coefficient estimates, predicts the risk of language frequency hearing loss at future time points for individuals, and outputs the risk probability and risk stratification results of individuals;

[0012] A decision support module, based on the risk probability and risk stratification results, generates hearing protection recommendations for individuals.

[0013] Further, the follow-up data includes demographic information, hearing data, genetic information and environmental noise data;

[0014] The demographic information includes the gender, age, education level, work experience, occupation type, smoking history and high-frequency hearing threshold of the worker;

[0015] The hearing data comprises a response variable; the response variable is 1 when the average hearing threshold of the worker individual at a language frequency at a time point is greater than a preset hearing loss determination threshold, and the response variable is 0 when the average hearing threshold of the language frequency is less than or equal to the preset hearing loss determination threshold; the average hearing threshold of the language frequency is obtained by adding the hearing thresholds of the individual at each frequency in a preset set of language frequency points and then dividing by the number of frequency points;

[0016] The genetic information comprises a genetic risk score and a genotype, and the genetic risk score is calculated by: obtaining single nucleotide polymorphism typing of the individual at a plurality of hearing-related gene sites, determining the number of risk alleles at each site, multiplying the number of risk alleles at the site by a preset effect weight of the site to obtain a weighted score, and accumulating the weighted scores of all sites to obtain the genetic risk score;

[0017] The environmental noise data comprises a cumulative noise exposure; the cumulative noise exposure is obtained by adding the logarithm of the length of service to the equivalent continuous A sound level, wherein the logarithm of the length of service is the length of service taking the logarithm to the base of 10 multiplied by 10.

[0018] Further, obtaining the key feature subset comprises:

[0019] A candidate feature set is constructed, and the candidate feature set comprises genetic features, audiological features, and environmental exposure and demographic features;

[0020] The candidate feature set is preliminarily screened by using a least absolute shrinkage and selection operator regression, an optimal preset regularization parameter value is determined by using a cross-validation method, the least absolute shrinkage and selection operator model is refitted on all baseline data by using the optimal preset regularization parameter value, and the features with non-zero regression coefficients are retained to form a least absolute shrinkage and selection operator screening feature set;

[0021] The importance of the candidate feature set is evaluated by using a random forest classifier, the importance of each feature is evaluated by using a Gini importance index, all features in the candidate feature set are sorted in descending order according to the Gini importance score, and a preset number of features with high importance are selected to form a random forest screening feature set;

[0022] The least absolute shrinkage and selection operator screening feature set and the random forest screening feature set are integrated, the union of the two sets is taken to construct a comprehensive feature set, the genetic features in the comprehensive feature set are biologically verified, and a key feature subset is determined.

[0023] Further, the regression coefficient estimation output by the generalized estimating equation model comprises:

[0024] define time-varying interaction terms, the time-varying interaction terms include a genotype and cumulative noise exposure interaction term obtained by multiplying a genotype variable of an individual with a cumulative noise exposure amount of the individual at a time point, and a genotype and age interaction term obtained by multiplying the genotype variable of the individual with an age of the individual at the time point;

[0025] establish a marginal mean model using a Logit link function, the Logit function mapping a probability of the individual having a language frequency hearing loss at the time point to a linear predictor;

[0026] the generalized estimating equation model specifies a working correlation matrix to describe correlations between observations of the same individual at different time points, the working correlation matrix including a correlation structure, the correlation structure including an independent structure, an exchangeable structure, a first-order autoregressive structure, and a unstructured structure, and an optimal correlation structure is selected using a quasi-likelihood criterion under an independent model assumption;

[0027] using the selected working correlation matrix, model parameters are estimated by solving the generalized estimating equation, the generalized estimating equation is solved using an iterative weighted least squares method to obtain estimated values of regression coefficient vectors, including an intercept term, an age main effect coefficient, a cumulative noise exposure main effect coefficient, a genotype main effect coefficient, a first interaction effect coefficient, and a second interaction effect coefficient.

[0028] Further, a calculation method of the linear predictor is to take the intercept term, sequentially add an age multiplied by the age main effect coefficient, a cumulative noise exposure amount multiplied by the cumulative noise exposure main effect coefficient, a genotype multiplied by the genotype main effect coefficient, a genotype and cumulative noise exposure interaction term multiplied by the first interaction effect coefficient, and a genotype and age interaction term multiplied by the second interaction effect coefficient, and finally add an inner product of an other covariate vector and the regression coefficient vector; wherein the intercept term is a base value of the marginal mean model, the other covariate vector includes gender, smoking history, and high-frequency hearing threshold, and the regression coefficient vector includes respective regression coefficients of the other covariate vector.

[0029] Further, selecting the optimal correlation structure using the quasi-likelihood criterion under the independent model assumption includes:

[0030] a calculation method of the quasi-likelihood criterion under the independent model assumption is to calculate a negative 2 times of a quasi-likelihood function value, and add a penalty term;

[0031] The quasi-likelihood function is obtained by summing quasi-likelihood contributions of all individuals, each quasi-likelihood contribution being a likelihood term multiplied by negative one-half, the likelihood term being a difference between a response variable vector of the individual and an expected response vector, the difference being left multiplied by an inverse of a working covariance matrix and right multiplied by a transpose of the difference; the expected response vector comprises expected response values of the individual at all follow-up time points, each expected response value being an expected exponential function divided by an expected term, the expected exponential function being an exponential function of an exponential predictor with a natural constant as a base, the expected term being 1 plus the expected exponential function; the working covariance matrix is obtained by left multiplying a working correlation matrix by a square root of a variance matrix and right multiplying the square root of the variance matrix, wherein the variance matrix is a diagonal matrix with diagonal elements being variances of response variables at the time points;

[0032] The penalty term is calculated by multiplying a model covariance matrix estimate under the assumption of independent structure and a model covariance matrix estimate when the working correlation matrix is used, calculating a trace of a product matrix, and multiplying the trace by a preset penalty term coefficient; the model covariance matrix estimate is a regression coefficient covariance matrix estimate calculated after the working correlation matrix fits the generalized estimating equation model;

[0033] The generalized estimating equation model is fitted using the correlation structures respectively, quasi-likelihood criterion values under independent model assumptions of the correlation structures are calculated respectively, and the correlation structure with the minimum quasi-likelihood criterion value under the independent model assumption is selected as the optimal structure.

[0034] Further, the method for calculating the individual risk probability by the risk prediction module comprises:

[0035] Obtaining variable information of a current time point of the individual to be predicted, calculating variable information of a future time point based on the variable information of the current time point, and calculating the risk probability by a Logistic function based on the variable information of the future time point and the regression coefficient estimate;

[0036] According to the genetic risk score or the genotype, the individual is divided into different risk groups, and for each risk group, a relationship curve between the cumulative noise exposure and the incidence probability is drawn.

[0037] Further, the calculation of the risk probability by the Logistic function comprises:

[0038] The variable information of the current time point comprises the genotype or the genetic risk score, the current cumulative noise exposure, the current age, and other covariates;

[0039] calculating variable information at a future time point, the variable information at the future time point comprising a future age, a future cumulative noise exposure, a future length of service, and a linear predictor at the future time point; wherein the future age is equal to the current age plus a time interval; the future length of service is equal to the current length of service plus the time interval; the future cumulative noise exposure is equal to the equivalent continuous A-weighted sound level plus a logarithm value of the future length of service, wherein the logarithm value of the future length of service is a logarithm with base 10 multiplied by 10;

[0040] the linear predictor at the future time point is equal to an intercept term estimate plus a product of an age main effect coefficient estimate and the future age, plus a product of a cumulative noise exposure main effect coefficient estimate and the future cumulative noise exposure, plus a product of a genotype main effect coefficient estimate and the genotype, plus a product of a first interaction effect coefficient estimate and the genotype and the future cumulative noise exposure, plus a product of a second interaction effect coefficient estimate and the genotype and the future age, and finally plus an inner product of the other covariate vector and a corresponding vector of regression coefficient estimates;

[0041] the method for calculating the logistic function is that the risk probability is equal to an exponential function of the linear predictor at the future time point divided by a linear predictor term at the future time point, the exponential function of the linear predictor at the future time point is an exponential function with a natural constant as a base, the linear predictor at the future time point as an index, and the linear predictor term at the future time point is the exponential function of the linear predictor at the future time point plus 1.

[0042] Further, the method for drawing the curve of the relationship between the cumulative noise exposure and the incidence probability comprises:

[0043] The risk groups comprise a low genetic risk group, a medium genetic risk group, and a high genetic risk group; wherein the low genetic risk group comprises individuals not carrying a risk allele or individuals with a genetic risk score lower than a preset low risk score threshold, the medium genetic risk group comprises individuals with a genetic risk score greater than or equal to the preset low risk score threshold and less than or equal to a preset high risk score threshold, and the high genetic risk group comprises individuals with a genetic risk score greater than the preset high risk score threshold or individuals carrying a high-risk homozygous mutation.

[0044] The fixed age, genotype, and other covariate vector are typical values, the typical values comprising a median, a mode, a minimum value, and a maximum value of the age, the genotype, and the other covariate vector, the cumulative noise exposure is changed within a reasonable range, and the incidence probability corresponding to different cumulative noise exposure values is calculated; the incidence probability is an exponential function of the linear predictor divided by a linear predictor term, the exponential function of the linear predictor is an exponential function with a natural constant as a base, the linear predictor as an index, and the linear predictor term is the exponential function of the linear predictor plus 1.

[0045] Further, the method for generating the hearing protection recommendation for the individual comprises:

[0046] The individual is divided into different risk levels based on the risk probability, and a corresponding management strategy is formulated according to the risk level, the risk level includes low risk, medium risk and high risk, the low risk is that the risk probability is less than a preset low risk probability threshold, the medium risk is that the risk probability is greater than or equal to the preset low risk probability threshold and less than a preset high risk probability threshold, and the high risk is that the risk probability is greater than or equal to the preset high risk probability threshold;

[0047] Suitable hearing protection devices are recommended according to the risk group and the current cumulative noise exposure of the individual; standard earplugs are recommended for individuals with low genetic risk and cumulative noise exposure meeting the preset low exposure, high-performance earplugs or ear muffs are recommended for individuals with medium genetic risk or cumulative noise exposure meeting the preset medium exposure, and earplug plus ear muff combination is recommended for individuals with high genetic risk or cumulative noise exposure meeting the preset high exposure.

[0048] The beneficial effects of the present application are that the present application processes the longitudinal follow-up data of occupational noise exposure workers by constructing a generalized estimating equation model, effectively solving the problem that the traditional cross-sectional study cannot capture the dynamic change trend of hearing loss. By utilizing the advantages of generalized estimating equation in processing repeated measurement data, the system fully considers the correlation between observation values of the same individual at different follow-up time points, corrects the standard error of parameter estimation by specifying the work-related matrix, avoids the statistical inference bias caused by ignoring the individual correlation, and significantly improves the accuracy and stability of the prediction of the risk of language frequency hearing loss.

[0049] The present application can dynamically evaluate the modification effect of genetic susceptibility on the dose-response relationship of noise by introducing a time-varying genetic and environmental interaction effect model. By defining time-varying interaction terms between genotype and cumulative noise exposure and genotype and age, the system can not only quantify the main effects of genetic factors and environmental noise exposure, but also accurately capture how the genetic background changes the individual's risk of disease with the accumulation of cumulative noise exposure and the growth of age. This dynamic modeling method breaks through the limitations of static risk assessment and can more truly reflect the complex pathogenesis of noise-induced hearing loss.

[0050] This invention employs a multi-stage feature selection strategy combining statistical learning and biological validation, ensuring the high efficiency and clinical interpretability of key feature subsets. By integrating the linear feature selection capabilities of minimum absolute contraction and selection operator regression with the nonlinear feature importance assessment capabilities of random forests, and combining this with biomedical literature to validate the biological function of selected genetic variation sites, the system can accurately identify features with both significant statistical predictive power and clear biological significance from high-dimensional genetic and audiological data. It provides intuitive visualization tools and quantitative decision support indicators, greatly enhancing the clinical applicability of the predictive model. By constructing nomograms, the complex generalized estimation equation model is transformed into a simple graphical tool, enabling clinicians and occupational health workers to quickly calculate an individual's disease risk through simple table lookup operations. Furthermore, by calculating the cumulative noise exposure threshold and upper limit of safe exposure time for each genotype to reach a 50% disease risk, the system transforms abstract risk probabilities into concrete safety management indicators. This provides a scientific and quantitative basis for developing personalized hearing protection plans, recommending appropriate hearing protection devices, and arranging reasonable job rotation systems, contributing to the shift from passive treatment to proactive and precise prevention of occupational hearing loss. Attached Figure Description

[0051] Figure 1 This is a module example diagram of the noise-induced speech frequency hearing loss prediction system based on GEE model and genetic characteristics of the present invention.

[0052] Figure 2 This is an example diagram showing the key feature subset obtained by the noise-induced speech frequency hearing impairment prediction system based on the GEE model and genetic characteristics of the present invention.

[0053] Figure 3 This is an example diagram showing the output regression coefficient estimation of the noise-induced speech frequency hearing loss prediction system based on the GEE model and genetic characteristics of the present invention.

[0054] Figure 4 This is an example diagram showing the risk probability and risk stratification results of the output individual of the noise-induced speech frequency hearing loss prediction system based on the GEE model and genetic characteristics of the present invention. Detailed Implementation

[0055] The subject matter described herein will now be discussed with reference to exemplary embodiments. It should be understood that these embodiments are discussed only to enable those skilled in the art to better understand and implement the subject matter described herein, and changes may be made to the function and arrangement of the elements discussed without departing from the scope of this specification. Various processes or components may be omitted, substituted, or added as needed in the examples. Furthermore, some features described in the examples may be combined in other examples.

[0056] A noise-induced language frequency hearing loss prediction system based on GEE model and genetic characteristics, as shown in Figure 1 comprises:

[0057] The data acquisition module 101 acquires follow-up data of workers exposed to occupational noise from multiple data sources, including demographic information, hearing data, genetic information, and environmental noise data.

[0058] The data acquisition module 101 extracts multi-year follow-up data of workers from the occupational health surveillance system. For each study subject, the detection data at multiple time points are recorded, with the time point index increasing sequentially from 1 until the number of follow-ups of the individual. The time points usually correspond to the annual occupational health examination time. To ensure data quality, the system excludes individuals with less than a preset minimum follow-up number threshold, with a default value of 3 times. The setting is based on the fact that at least 3 time points are needed to effectively capture the trend of change in follow-up data analysis. This threshold can be adjusted within the range of 2 to 5 times according to the specific requirements of the research design.

[0059] Demographic information includes gender, age, education level, work experience, occupation type, smoking history, and high-frequency hearing threshold of workers. The reason for collecting these demographic information as model input is that these factors have been confirmed by a large number of epidemiological studies to be related to noise-induced hearing loss. Among them, the gender factor reflects the physiological difference that men are more susceptible to noise-induced hearing loss than women due to differences in cochlear structure and hormone levels. Age is an important risk factor for hearing loss. With age, presbycusis and noise-induced hearing loss may have a superimposed effect. The inclusion of smoking history is based on the adverse effects of smoking on the blood supply of the inner ear. Smoking and noise exposure have a synergistic effect that can accelerate the progression of hearing loss. The reason for including high-frequency hearing threshold as an input feature is that noise damage first affects the high-frequency region of the cochlear basal turn. The increase in high-frequency hearing threshold is an early sign of noise-induced hearing loss and can predict the risk of subsequent language frequency hearing loss. The integration of these demographic and clinical characteristics can improve the prediction accuracy of the model and provide multi-dimensional information support for risk stratification and personalized intervention.

[0060] The genetic information includes genetic risk scores and genotypes; peripheral blood samples of workers are collected, genomic DNA is extracted, and single nucleotide polymorphisms of candidate hearing-related genes are typed. The reason for selecting single nucleotide polymorphisms as genetic markers is that single nucleotide polymorphisms are the most common form of genetic variation in the human genome, widely distributed and abundant, which can comprehensively cover the genetic variation information of candidate genes, the detection technology of single nucleotide polymorphisms is mature, the cost is relatively low, and it is suitable for genetic screening in large-scale occupational population, in addition, the genetic effect of single nucleotide polymorphisms is relatively stable and is not affected by environmental factors, which can provide reliable genetic information for individual lifelong risk assessment. The candidate genes include CDH23, GJB2, MYO7A, KCNQ4 and other known genes related to hearing function. The reason for selecting these candidate genes is that the proteins encoded by them play a key role in the structure and function of cochlear hair cells, among which the GJB2 gene encodes connexin 26, which is involved in ion circulation and signal transmission between cochlear hair cells, and mutations in this gene are the most common cause of genetic hearing loss, the CDH23 gene encodes cadherin 23, which is an important component of the intercellular junction of hair cell stereocilia, and variations in this gene are closely related to noise susceptibility, the MYO7A gene encodes myosin 7A, which is involved in the structural maintenance and material transport of hair cells, and the KCNQ4 gene encodes potassium ion channels, which regulate the electrophysiological activity of hair cells, and variations in these genes have been confirmed by multiple studies to be associated with susceptibility to noise-induced hearing loss, with clear biological function basis.

[0061] Based on the results of single nucleotide polymorphism typing, the genetic risk score is constructed. The reason for using the weighted cumulative method to construct the genetic risk score is that there are significant differences in the degree of influence of different single nucleotide polymorphism sites on hearing loss. The weighted method can give corresponding weights according to the actual effect size of each site, so as to more accurately reflect the overall genetic susceptibility of individuals. Compared with the simple counting method, the weighted score can improve the accuracy of risk prediction and the clinical application value. The calculation method of the genetic risk score is to add up the weighted scores of all single nucleotide polymorphism sites included in the score. Specifically, for each single nucleotide polymorphism site, first determine the number of risk alleles carried by the individual at this site. The risk allele number takes the value of 0, 1 or 2, which corresponds to wild-type homozygous, heterozygous mutation and risk-type homozygous three genotypes respectively. Then multiply the number of risk alleles at this site by the preset effect weight of this site to get the weighted score. The preset effect weight is determined according to the effect reported in the literature. If it is not reported in the literature, it is estimated by the logistic regression model of the training data set. When the preset effect weight needs to be estimated by the training data set, the system uses the logistic regression model for training. The training process includes the following steps. First, the individuals in the training data set are divided into case group and control group according to the language frequency hearing loss state. Then, a univariate logistic regression model is constructed for each single nucleotide polymorphism site, with the number of risk alleles at this site as the independent variable and the language frequency hearing loss state as the dependent variable. The regression coefficient is solved by the maximum likelihood estimation method. The exponential value of this regression coefficient is the odds ratio of this site, which reflects the fold effect of carrying risk alleles on hearing loss risk. Finally, the natural logarithm value of the odds ratio is standardized as the preset effect weight of this site. The default value of the preset effect weight ranges from 0.1 to 2.0. Among them, the default value of the preset effect weight of the known strong effect site such as GJB2 gene c.35delG mutation is 1.5, the default value of the preset effect weight of the medium effect site is 1.0, and the default value of the preset effect weight of the weak effect site is 0.5. Finally, the weighted scores of all sites are added to obtain the genetic risk score of the individual.

[0062] The environmental noise data includes cumulative noise exposure. For each individual at each follow-up time point, the system calculates the cumulative noise exposure. The cumulative noise exposure takes into account both the noise intensity and the exposure time. The reason for logarithmic transformation of the length of service is that the damage effect of noise on hearing does not increase linearly with time, but the damage speed is faster in the early exposure, and gradually slows down as the exposure time is prolonged. The logarithmic transformation can better reflect the non-linear dose-response relationship, and conforms to the physiological law of hearing impairment. The calculation method is to add the equivalent continuous A sound level to the logarithmic value of the length of service, wherein the logarithmic value of the length of service needs to take the logarithm of the length of service with base 10 first, and then multiply the logarithmic value by 10. The equivalent continuous A sound level is in decibel A weighting units, representing the noise exposure intensity of the individual in the working environment. The length of service is in years, representing the noise exposure time of the individual at the time point. The calculation method is based on the principle of noise dosimetry, integrates the exposure intensity and time into a single cumulative exposure index, so that individuals with different exposure modes are comparable.

[0063] The data acquisition module 101 records the pure tone hearing threshold of each individual at each follow-up time point, including the hearing threshold of both ears at 0.25, 0.5, 1, 2, 3, 4, 6, 8 kHz frequencies.

[0064] The system calculates the average hearing threshold at language frequencies, which reflects the individual's hearing level at the key frequency range for language communication. The reason for choosing the average hearing threshold at language frequencies as the main evaluation indicator is that hearing loss in the language frequency range directly affects the individual's daily language communication ability and quality of life. Compared with high-frequency hearing loss, language frequency hearing loss has a more significant negative impact on the worker's work ability and social interaction, so predicting and preventing language frequency hearing loss has important clinical significance and social value. In addition, language frequency hearing loss usually occurs after high-frequency hearing loss and is a sign of the progression of noise-induced hearing loss to a certain stage. The prediction of language frequency hearing loss can provide a critical time window for early intervention. The calculation method of the average hearing threshold at language frequencies is to add the hearing thresholds of the individual at the preset set of language frequency points and divide by the number of frequency points, that is, to take the arithmetic mean of these frequency hearing thresholds. The default value of the preset set of language frequency points is 500 Hz, 1000 Hz and 2000 Hz. The selection of this set is based on the language frequency range defined in the international audiology standard. These three frequencies cover the most important acoustic information in human language communication, among which 500 Hz mainly corresponds to the fundamental frequency information of vowels, 1000 Hz and 2000 Hz correspond to the formant information of consonants, and the combination of the three can fully reflect the hearing level required for language perception. This set can also be adjusted to include other frequency combinations within the range of 250 Hz to 4000 Hz according to the characteristics of different languages. The hearing threshold is in decibels hearing level. Based on the average hearing threshold at language frequencies, the system defines a binary response variable. When the average hearing threshold at language frequencies of the worker individual at a certain time point is greater than the preset hearing loss determination threshold, the response variable takes the value of 1, indicating that the individual has language frequency hearing loss at that time point. When the average hearing threshold at language frequencies is less than or equal to the preset hearing loss determination threshold, the response variable takes the value of 0, indicating normal hearing. The default value of the preset hearing loss determination threshold is 25 decibels hearing level. This threshold is the internationally recognized standard for diagnosing mild hearing loss by the World Health Organization and the international audiology community, representing the hearing level that begins to affect daily language communication. This threshold can be adjusted within the range of 20 to 30 decibels hearing level according to the research purpose.

[0065] The data processing module 102, based on the follow-up data, performs data cleaning, outlier processing, missing value filling and variable standardization to generate a standardized modeling data set.

[0066] The data processing module 102 first identifies and processes outliers in the hearing threshold data. The reason for using the boxplot method to detect outliers is that this method is based on the quantile distribution of the data, does not make strict assumptions about the overall distribution form of the data, and has better robustness to skewed distribution and extreme values than the method based on mean and standard deviation. It can effectively identify abnormal data points caused by measurement errors or recording errors, while avoiding misjudging real extreme individuals as outliers. For the hearing threshold data of each frequency, first calculate the first quartile, i.e. the 25th percentile, and the third quartile, i.e. the 75th percentile, then calculate the interquartile range, which is equal to the third quartile minus the first quartile. The criterion for determining outliers is that if a hearing threshold is less than the first quartile minus the product of the preset interquartile range multiple coefficient and the interquartile range, or greater than the third quartile plus the product of the preset interquartile range multiple coefficient and the interquartile range, then the hearing threshold is determined to be an outlier. The default value of the preset interquartile range multiple coefficient is 1.5. This coefficient is a standard setting in the boxplot method in statistics, used to balance the sensitivity and specificity of outlier detection. This coefficient can be adjusted within the range of 1.0 to 3.0 according to the data distribution characteristics. A smaller coefficient such as 1.0 will detect more outliers, and a larger coefficient such as 3.0 will be more conservative.

[0067] In addition, the system performs logical checks based on audiological knowledge. If the hearing threshold is less than the preset lower limit of the hearing threshold or greater than the preset upper limit of the hearing threshold, it is determined to be a measurement error. The default value of the preset lower limit of the hearing threshold is -10 decibel hearing level, and the default value of the preset upper limit of the hearing threshold is 120 decibel hearing level. These two thresholds are determined based on the measurement range of the hearing test equipment and the physiological range of human hearing. -10 decibel hearing level represents the limit of super-normal hearing, and 120 decibel hearing level represents the upper limit of extreme hearing loss. For detected outliers, the system marks them as missing values, which are filled in later through a missing value processing method.

[0068] The system also excludes individuals who do not meet the inclusion criteria, excluding individuals with middle ear dysfunction, a history of ear surgery, and a history of sudden deafness, to ensure that the hearing loss of the study subjects is mainly caused by noise exposure. Middle ear dysfunction includes type B or type C tympanometry.

[0069] For missing genetic data, the system adopts genotype imputation algorithm. The reason for adopting genotype imputation algorithm is that missing genetic data will lead to sample size reduction and statistical power decline, and genotype imputation can infer missing genotypes from known genotypes by using linkage disequilibrium relationship between genetic loci in the population, thereby maximizing the use of existing data and improving the statistical power of analysis and the reliability of results. Genotype imputation is based on linkage disequilibrium principle, and uses haplotype information of typed single nucleotide polymorphism loci to infer the genotype of missing single nucleotide polymorphism loci. The imputation process outputs the probability of each possible genotype, and the system selects the genotype with the highest probability as the imputation value. The specific training and execution process of genotype imputation includes the following steps: first, construct a haplotype reference panel using the haplotype data of the reference population, which contains the linkage disequilibrium pattern and haplotype frequency distribution between single nucleotide polymorphism loci; then, for the individual to be imputed, search for the most matched haplotype combination in the reference panel according to the typed single nucleotide polymorphism loci information of the individual by using hidden Markov model or other statistical model; then, infer the genotype probability distribution of the missing loci based on the matched haplotype combination; finally, select the genotype with the maximum posterior probability as the imputation value of the missing loci.

[0070] For missing hearing threshold data, if the hearing threshold of a certain frequency is missing but the thresholds of adjacent frequencies are complete, the system adopts linear interpolation method for imputation. The reason for adopting linear interpolation method is that the hearing threshold of human ear usually presents a continuous and smooth change trend between adjacent frequencies, and linear interpolation method can reasonably estimate the hearing threshold of the missing frequency based on the known hearing thresholds of adjacent frequencies. This method is simple and effective and does not introduce excessive assumptions, and is suitable for the case of missing a single frequency point. The calculation process of linear interpolation method is as follows: first, determine the two known frequencies adjacent to the missing frequency, which are denoted as low frequency and high frequency, respectively, and the missing frequency is located between the two known frequencies. Then calculate the difference between the missing frequency and the low frequency divided by the difference between the high frequency and the low frequency to obtain the frequency ratio coefficient. Next, calculate the difference between the hearing threshold of the high frequency and the hearing threshold of the low frequency, and multiply the difference by the frequency ratio coefficient. Finally, add the hearing threshold of the low frequency to the product to obtain the estimated value of the hearing threshold of the missing frequency.

[0071] If the data of a certain follow-up is completely missing, the system retains the individual but marks the time point as missing. In subsequent generalized estimating equation model fitting, the generalized estimating equation method can handle unbalanced follow-up data.

[0072] To eliminate the influence of different variable dimensions, the system performs Z-score standardization on continuous covariates. The reason for using Z-score standardization is that the dimensions and numerical ranges of different covariates are quite different. For example, the value range of age is usually between 20 and 60 years old, while the value range of cumulative noise exposure may be between 80 and 120 decibels. If the original values are directly used for modeling, the variables with larger numerical ranges will dominate the model fitting process, leading to biased estimates of regression coefficients and decreased model performance. Z-score standardization converts all variables to the same dimension and scale, making each variable comparable in the model, and the standardized regression coefficients can directly reflect the relative importance of each variable on the response variable. The calculation method of standardization is as follows: first, calculate the mean and standard deviation of the variable in all observations, then for each observation, subtract the mean from the observation value, and divide the difference by the standard deviation to get the standardized value. The mean of the standardized variable is 0 and the standard deviation is 1. The variables that need to be standardized include age, length of service, cumulative noise exposure, etc.

[0073] For categorical variables such as gender and smoking history, the system performs dummy variable encoding. For example, the gender variable is encoded as male with a value of 1 and female with a value of 0. Smoking history is encoded as smoking with a value of 1 and non-smoking with a value of 0.

[0074] Based on the modeling dataset, the system constructs a training dataset. The method of constructing the training dataset is to divide the modeling dataset into a training dataset and a test dataset according to a pre-set proportion. The reason for dividing the dataset is to use independent sample sets to train model parameters and validate model performance, in order to avoid overfitting and evaluate the generalization ability of the model. The specific construction method is as follows: first, set the sample proportion of the training dataset, the default value of the sample proportion is 70% of the total sample size of the modeling dataset, this proportion can be adjusted within the range of 60% to 80% of the total sample size. Then, taking individuals as the sampling unit, stratify according to the language frequency hearing loss status of individuals, and randomly select individuals and their observation data at all follow-up time points into the training dataset according to the set proportion in each layer, and the remaining individuals and their data constitute the test dataset. The reason for choosing individuals as the sampling unit rather than observation points is that generalized estimating equation models need to utilize the longitudinal correlation within individuals. If the different observation points of the same individual are split into the training set and the test set, it will destroy the correlation structure of the data and cause information leakage. The final generated training dataset contains all the necessary information for model construction, including standardized covariates, genetic information, and response variables. This training dataset will be used for subsequent feature selection, parameter estimation, and model fitting processes.

[0075] The feature selection module 103, based on the modeling dataset, selects features that are significantly related to language frequency hearing loss by statistical learning methods, and obtains a key feature subset, specifically as follows:Figure 2 The module adopts a multi-stage screening strategy. First, the least absolute shrinkage and selection operator regression is used for initial screening of features. Then, the random forest method is used to verify the importance of the features. Finally, the results of the two methods are integrated, and the final key feature subset is determined in combination with biological knowledge.

[0076] The feature screening module 103 first extracts all candidate features that can be related to language frequency hearing loss from the modeling data set to construct an initial candidate feature set. The candidate feature set includes genetic features, audiological features, and environmental exposure and demographic features. The genetic features include the genotype encoding of all single nucleotide polymorphism sites detected in the data collection module 101. The genotype of each single nucleotide polymorphism site is encoded as 0, 1 or 2 according to the number of risk alleles, corresponding to wild-type homozygous, heterozygous mutation and risk homozygous, respectively. In addition, it also includes the genetic risk score calculated according to all single nucleotide polymorphism sites. The audiological features include the individual's hearing threshold at each frequency, including the hearing threshold at 250 Hz, 500 Hz, 1000 Hz, 2000 Hz, 3000 Hz, 4000 Hz, 6000 Hz and 8000 Hz. These hearing threshold data come from the pure tone audiometry test results of the individual at the baseline follow-up, i.e. the first follow-up. Environmental exposure and demographic features include covariates such as cumulative noise exposure, age, gender, smoking history, hearing protection device usage, etc.

[0077] To ensure the effectiveness of feature screening, the system uses baseline data, i.e., the first follow-up data of each individual, for cross-sectional analysis. The reason for using baseline data for feature screening is that the purpose of feature screening is to identify key predictors related to language frequency hearing loss, and baseline data represents the initial state of individuals when they enter the study, at which time the feature values are not affected by subsequent interventions or behavior changes that may occur during follow-up, and can truly reflect the association between features and outcomes. If all follow-up data is used for feature screening, it may introduce time-dependent bias and reverse causality, for example, individuals who have already experienced hearing loss may change their protective behaviors, leading to confusion between features and outcomes. In addition, the sample size of baseline data is equal to the total number of study subjects, while the sample size of all follow-up data is the number of study subjects multiplied by the number of follow-up times. Using baseline data can significantly reduce computational complexity while ensuring statistical efficiency, improving the efficiency of feature screening. The selection of baseline data is based on the following considerations. On the one hand, the data completeness at the baseline time point is the highest, and it can provide a unified basis for feature evaluation for all individuals. On the other hand, baseline features can reflect an individual's genetic susceptibility and hearing status in the early stages of noise exposure, providing a predictive basis for the occurrence of language frequency hearing loss during subsequent follow-up. The response variable is the language frequency hearing loss status of the individual at baseline, which is determined according to whether the average hearing threshold at the language frequency exceeds the pre-set hearing loss determination threshold. The calculation method of the average hearing threshold at the language frequency and the setting of the pre-set hearing loss determination threshold are consistent with the definitions in the data acquisition module 101.

[0078] The system uses least absolute shrinkage and selection operator regression to preliminarily screen the candidate feature set. The reason for using least absolute shrinkage and selection operator regression is that the candidate feature set contains a large number of genetic features, audiological features, and environmental exposure features, the number of features can reach dozens or even hundreds, and there is a certain correlation between the features, for example, there is a correlation between the hearing thresholds of different frequencies, and there may be linkage disequilibrium between different single nucleotide polymorphism sites. Traditional feature selection methods such as stepwise regression are prone to instability and overfitting when dealing with high-dimensional correlated features, while least absolute shrinkage and selection operator regression can automatically compress the coefficients of unimportant features to zero through L1 regularization penalty, achieving simultaneous optimization of feature selection and model fitting. This method has good statistical properties and computational efficiency. Least absolute shrinkage and selection operator regression is a linear regression method with L1 regularization penalty, which can automatically select features while fitting the model, and compress the regression coefficients of unimportant features to zero. This method is particularly suitable for handling a large number of features with correlations between features, and can effectively identify key features that have independent predictive effects on the response variable.

[0079] The goal of Lasso regression is to find a set of regression coefficients that minimizes an objective function. The objective function includes two parts, the first part is model fitting error, which measures the fitting degree of the model to the training data, and the second part is regularization penalty term, which controls the model complexity and realizes feature selection. The model fitting error is measured by calculating the sum of the squares of the prediction errors of all samples. Specifically, for each sample, the predicted response variable value is calculated based on the feature values of the sample and the current regression coefficients, then the difference between the predicted value and the actual response variable value is calculated, the difference is squared, and the sum of all samples is divided by twice the number of samples to get the average fitting error. The regularization penalty term is calculated by multiplying a preset regularization parameter by the sum of the absolute values of all feature regression coefficients. The role of this penalty term is to constrain the size of the regression coefficients. The larger the preset regularization parameter, the stronger the penalty, and more feature regression coefficients will be compressed to zero.

[0080] The selection of the preset regularization parameter has a key impact on the feature selection result. The system uses the cross-validation method to determine the optimal preset regularization parameter value. The reason for using the cross-validation method to select the regularization parameter is that the regularization parameter controls the model complexity and the strictness of feature selection. A small parameter value will result in a model containing too many features, leading to overfitting. A large parameter value will result in important features being incorrectly excluded, reducing the model's prediction ability. The cross-validation method evaluates the model's performance on independent validation sets, objectively compares the generalization ability of models corresponding to different parameter values, and selects the parameter value that performs best on the validation set. This data-driven parameter selection method avoids the randomness of subjective parameter setting, ensuring the reliability of the feature selection result and the prediction performance of the model. The specific process of cross-validation is as follows: First, randomly divide the baseline dataset into a preset number of cross-validation folds of approximately equal size. The default value of the preset cross-validation fold number is 10, which is the standard setting for cross-validation in machine learning, achieving a good balance between computational efficiency and model evaluation accuracy. This number can be adjusted within the range of 5 to 20 based on the sample size, and a larger number can be chosen when the sample size is small to make full use of the data. Then, for each candidate preset regularization parameter value, perform a preset number of cross-validation rounds. In each round, use the preset number of cross-validation folds minus one subset as the training set to fit the Lasso model, and evaluate the model's prediction error on the remaining one subset as the validation set. Take the average of the prediction errors obtained in the preset number of cross-validation rounds as the cross-validation error corresponding to the preset regularization parameter value. Finally, select the preset regularization parameter value that minimizes the cross-validation error as the optimal value.

[0081] The final regression coefficient estimates are obtained by refitting the Lasso model on all baseline data using the optimal pre-specified regularization parameter value. The system retains the features with non-zero regression coefficients, which form the Lasso selected feature set. Due to the L1 regularization property of Lasso regression, the features in this set usually have strong independent predictive power and low redundancy among them. According to the pathophysiological mechanism of noise-induced hearing loss, the Lasso selected feature set usually includes the high-frequency hearing thresholds, especially at 3000 Hz, 4000 Hz, and 6000 Hz, because noise damage first affects the high-frequency region of the cochlear basal turn, and the elevation of high-frequency hearing thresholds often indicates an increased risk of subsequent hearing loss at speech frequencies. In addition, the set also includes some genetic variant sites and important covariates such as cumulative noise exposure and age.

[0082] To further verify the importance of the features selected by the Lasso method and discover important features that may be missed by the Lasso method, the system uses a random forest classifier to evaluate the importance of the candidate feature set. The reason for using random forest as a supplementary screening method is that Lasso regression is essentially a linear model and mainly identifies features that have a linear relationship with the response variable. However, in the pathogenesis of noise-induced hearing loss, there may be complex nonlinear interactions between genetic and environmental factors, such as certain genotypes that may only exhibit significant risk effects under high noise exposure conditions. These nonlinear and interactive effects are difficult to capture by linear models, while random forest, as a non-parametric method, can automatically learn the nonlinear relationships and high-order interaction patterns between features. Therefore, the combination of the two methods can more comprehensively identify important features and improve the accuracy and completeness of feature selection. Random forest is an ensemble learning method based on decision trees, which improves the stability and generalization ability of the model by constructing multiple decision trees and combining their prediction results. Compared with Lasso regression, random forest can automatically capture nonlinear relationships and high-order interactions between features, so it can evaluate the importance of features from different perspectives.

[0083] The construction process of a random forest includes multiple steps. First, the system performs random sampling with replacement from the baseline dataset to generate a preset number of bootstrap sample sets, each of which has the same size as the original baseline dataset but includes some repeated samples and omits some original samples due to the random sampling with replacement. The default value of the preset number of decision trees is 500, which is based on the research and practical experience of random forest theory. When the number of decision trees reaches 500, the prediction performance of the random forest usually stabilizes, and further increasing the number of trees has limited performance improvement but increases the computational cost. The number can be adjusted within the range of 100 to 2000 according to the size of the dataset and the computing resources.

[0084] Second, for each bootstrap sample set, the system trains a decision tree. The training of the decision tree uses a recursive splitting strategy, starting from the root node, selecting a feature and a split point at each node to divide the samples of the node into two child nodes, so that the purity of the two child nodes improves after splitting. To increase the diversity between decision trees, when splitting at each node, the system does not select the optimal splitting feature from all candidate features, but first randomly selects a preset number of node splitting features from all candidate features, and then only searches for the optimal splitting feature and split point among these randomly selected features. The default value of the preset number of node splitting features is the square root of the number of all candidate features, which is a classic configuration of the random forest algorithm, balancing the diversity between each decision tree and the prediction ability of a single decision tree. This number can also be adjusted according to the strength of the correlation between features. When the correlation between features is strong, it can be set to the logarithmic value of the number of all candidate features, and when the correlation between features is weak, it can be set to one-third of the number of all candidate features. Each decision tree grows to a preset maximum tree depth or the number of samples in a node is less than a preset minimum node sample number, and does not perform pruning. The default value of the preset maximum tree depth is not limited in depth, i.e., allowing the decision tree to grow completely until each leaf node contains samples of the same class or cannot be further split. This setting allows each decision tree to fully fit its corresponding bootstrap sample set. The depth can also be set to 10 to 30 layers according to the need to prevent overfitting.

[0085] Finally, when all decision trees are trained, the random forest makes predictions through a voting mechanism. For classification problems, each decision tree gives a class prediction for the sample to be predicted, and the random forest votes on all decision tree prediction results to select the class with the most votes as the final prediction result.

[0086] The system uses the Gini importance index to evaluate the importance of each feature in the random forest. The reason for using the Gini importance index is that it can directly quantify the contribution of the feature in the classification decision. Gini impurity is a standard index for measuring the degree of mixture of sample categories in a node. The greater the reduction of Gini impurity when a feature is used for node splitting, the stronger the distinguishing ability of the feature for sample classification. By accumulating the contribution of the feature in all decision tree nodes, Gini importance can comprehensively evaluate the overall importance of the feature in the entire random forest model. The Gini importance index is calculated based on the contribution of the feature to the purity improvement in the decision tree node splitting. For a certain feature, the system traverses all decision trees in the random forest and finds all nodes that use the feature for splitting. For each such node, the reduction of Gini impurity before and after the node splitting is calculated. This reduction reflects the degree of improvement in sample classification purity when the feature is used for splitting. Gini impurity is an index for measuring the degree of mixture of sample categories in a node. Its calculation method is to first calculate the proportion of each category sample in the node, then add the square of each category proportion, and finally subtract the sum of squares from 1. In this embodiment, the response variable is binary, i.e., hearing normal and language frequency hearing loss. Therefore, the value range of Gini impurity is 0 to 0.5. When all samples in the node belong to the same category, the Gini impurity is 0, indicating complete purity. When the number of samples of the two categories is equal, the Gini impurity is 0.5, indicating maximum mixture. The Gini impurity reduction of the feature in all splitting nodes of all decision trees is accumulated to obtain the Gini importance score of the feature. The higher the Gini importance score, the greater the contribution of the feature in the random forest model, and the more important the prediction of language frequency hearing loss.

[0087] The system sorts all features in the candidate feature set in descending order according to the Gini importance score to obtain a feature importance ranking list. This list reflects the relative importance of each feature in the random forest model. The system selects the top-ranked pre-set number of features from the list to form a random forest screening feature set.

[0088] The determination of the pre-set number of features uses the elbow rule. The specific method is to add features to the random forest model in order of feature importance ranking, draw a curve of the relationship between the number of features and the model performance index such as the area under the curve, observe the trend of the curve, and select the number of features corresponding to the position of the obvious inflection point on the curve as the pre-set number of features. The inflection point indicates that the improvement of the model performance by continuing to increase the number of features is not obvious.

[0089] If the elbow rule cannot identify a clear inflection point, the preset number of features is selected as the preset number of reserved features. The default value of the preset number of reserved features is 20 features. The number is set by considering the balance between the operability of clinical application and the complexity of the model. 20 features can cover the main genetic features, audiological features and environmental exposure features, and the model is not too complex to be difficult to explain and apply. The number can be adjusted within the range of 10 to 50 according to the actual application scenario.

[0090] The system integrates the minimum absolute shrinkage and selection operator filtered feature set and the random forest filtered feature set to construct a comprehensive feature set. The reason for adopting the union integration strategy is that the two filtering methods are based on different statistical principles and model assumptions, each has its advantages and limitations. The least absolute shrinkage and selection operator regression is based on a linear model framework and is sensitive to features with strong linear main effects, but it may miss features that only play an effect under certain conditions or through interaction. Random forest is based on a non-parametric decision tree model and can capture complex nonlinear relationships and high-order interactions, but it may not be sensitive enough to weakly effective features. By taking the union set, the system can retain important features identified by each method to avoid missing key risk factors due to the limitations of a single method. This integration strategy has been proven to improve the robustness and comprehensiveness of feature selection in the field of feature selection, thereby improving the performance of the final prediction model. The principle of integration is to take the union set of the two sets, that is, as long as a feature appears in the least absolute shrinkage and selection operator filtered feature set or the random forest filtered feature set, it will be included in the comprehensive feature set. This integration strategy can fully utilize the complementary advantages of the two methods. The least absolute shrinkage and selection operator regression is good at identifying features with independent linear effects, while the random forest is good at identifying features with nonlinear effects or interaction effects. The combination of the two can more comprehensively capture important features related to language frequency hearing loss.

[0091] For genetic features in the integrated feature set, the system performs biological validation to ensure that its association with hearing loss has biological plausibility. The reason for biological validation is that statistical screening methods can identify some statistically significant but biologically unexplained false positive associations due to random fluctuations in the sample or the particularity of the data. These false positive features, if included in the prediction model, not only reduce the model's generalization ability on new data, but also mislead clinical decision-making and scientific research. By consulting authoritative biomedical literature databases, the system can confirm whether the screened genetic variant sites have independent research support, whether they have clear molecular mechanisms and functional evidence. This validation strategy that combines statistical evidence with biological knowledge can significantly improve the reliability of feature selection, ensuring that the final model not only has statistical predictive ability, but also has biological scientificity and safety in clinical application. The system consults biomedical literature databases such as PubMed, OMIM, etc. to search for relevant research reports on each genetic variant site and hearing function, and confirms whether these variant sites have clear biological function evidence. The system prefers to retain genetic variant sites with functional validation, such as the c.35delG mutation in the GJB2 gene, which causes connexin 26 to lose function, affecting ion circulation between cochlear hair cells, and is a common cause of genetic hearing loss. For example, the p.R1588W mutation in the CDH23 gene affects the structure of cadherin 23, disrupting the adhesion between hair cells and increasing the susceptibility to noise-induced hearing loss. For genetic variant sites that have no clear functional evidence in the literature but show significant association in statistical screening, the system retains them but marks them as pending validation sites, providing candidate targets for future functional research.

[0092] Considering the significant difference in gender in noise-induced hearing loss, male workers are generally more susceptible to noise-induced hearing loss than female workers due to physiological differences and differences in occupational exposure patterns. The system performs feature screening analysis for men and women separately to identify gender-specific risk factors. Specifically, the system divides the baseline dataset into male and female subsets according to gender, and independently performs the above-mentioned least absolute shrinkage and selection operator regression screening, random forest importance evaluation, and feature set integration process for each gender subset to obtain male-specific feature sets and female-specific feature sets. The system compares the similarities and differences between the two gender-specific feature sets to identify common features that are important in both genders and specific features that are only important in one gender. Common features reflect universal risk factors for noise-induced language frequency hearing loss, while specific features suggest gender-related susceptibility mechanisms, providing a basis for developing gender-specific hearing protection strategies.

[0093] After the above multi-stage screening and verification, the system finally determines a key feature subset. The subset includes features that have a significant predictive effect on language frequency hearing loss and clear biological significance. The key feature subset usually includes the following categories of features. The first category is genetic features, including genotypes of biologically verified single nucleotide polymorphism sites, which are mainly related to hearing-related genes such as GJB2, CDH23, MYO7A, KCNQ4, etc. In addition, it also includes genetic risk scores calculated based on these key single nucleotide polymorphism sites. The calculation method of the genetic risk score is consistent with the definition in the data collection module 101, that is, the number of risk alleles of each single nucleotide polymorphism site is multiplied by the corresponding preset effect weight and then summed. The second category is audiological features, mainly including high-frequency hearing thresholds, especially at 3000 Hz, 4000 Hz and 6000 Hz. These high-frequency hearing thresholds can reflect the early response of the individual's cochlea to noise damage and are important indicators for predicting subsequent language frequency hearing loss. The third category is environmental exposure and demographic characteristics, including cumulative noise exposure, age, and other covariates that show significant associations in feature screening such as gender, smoking history, hearing protection device usage, etc.

[0094] Each feature in the key feature subset has undergone double confirmation of statistical screening and biological verification, both having statistically significant predictive power and having a reasonable biological explanation, which ensures that the subsequent predictive model has good predictive performance and clear clinical interpretability.

[0095] The time-varying interaction module 104 constructs a generalized estimating equation model including time-varying interaction terms based on the key feature subset and the modeling dataset, and the generalized estimating equation model outputs regression coefficient estimates, as shown in Figure 3 .

[0096] To capture the dynamic effect of genetic susceptibility changing with noise exposure, time-varying interaction terms are defined. The reason for introducing time-varying interaction terms is that the effect of genetic factors on noise-induced hearing loss does not exist independently, but is complexly interacted with time-varying factors such as environmental exposure and age. For example, individuals carrying risk genotypes may not be significantly different from normal individuals in low noise exposure, but as the cumulative noise exposure increases, the growth rate of their hearing loss risk may be significantly faster than that of normal individuals. This gene-environment interaction effect cannot be reflected in the traditional model containing only main effects. By introducing time-varying interaction terms, the model can quantify how genetic susceptibility dynamically changes with noise exposure and age, thereby more accurately predicting the risk of individuals with different genetic backgrounds at different exposure stages and providing a scientific basis for personalized protection. The main interaction terms include the interaction term between genotype and cumulative noise exposure and the interaction term between genotype and age. The interaction term between genotype and cumulative noise exposure is obtained by multiplying the genotype variable of an individual with the cumulative noise exposure of that individual at a certain time point. This interaction term represents the modifying effect of genetic background on the noise dose-response relationship. The interaction term between genotype and age is obtained by multiplying the genotype variable of an individual with the age of that individual at a certain time point. This interaction term represents the pattern of genetic effect changing with age. The genotype variable can be a genotype code of a single single nucleotide polymorphism site, such as 0, 1, and 2, corresponding to wild-type homozygous, heterozygous mutation, and risk homozygous, respectively, or a genetic risk score. Cumulative noise exposure and age are time-varying covariates that change with the follow-up time point.

[0097] The generalized estimating equation is a semi-parametric method for analyzing follow-up data and clustered data. The reason for using the generalized estimating equation model is that the system deals with multi-year follow-up data, and there is correlation between the observations of the same individual at different time points. This correlation violates the assumption of independence of observations in traditional regression models. If this correlation is ignored, the standard error of parameter estimation will be smaller, and statistical inference will be wrong. The generalized estimating equation model is specifically designed to handle correlated data. By using a working correlation matrix to describe the correlation structure between observations within an individual, and using a robust standard error estimation method, even if the working correlation matrix is not completely correct, the parameter estimation is still consistent, and the standard error estimation is still effective. This robustness makes the generalized estimating equation an ideal choice for analyzing follow-up data. In addition, the generalized estimating equation model focuses on marginal effects, i.e., population average effects, and its parameter interpretation is intuitive, making it suitable for developing population-oriented prevention strategies. The generalized estimating equation model does not need to specify the complete joint distribution of the response variable, but only needs to specify the marginal mean model and the working correlation matrix. Parameters are estimated by quasi-likelihood method.

[0098] For binary response variable, the system adopts Logit link function to establish marginal mean model. The reason for adopting Logit link function is that the response variable is binary variable, i.e. language frequency hearing loss state, which takes 0 or 1, and the expected value, i.e. the value range of the probability of occurrence, is limited between 0 and 1, while the value range of the linear predictor is negative infinity to positive infinity, and it is necessary to map the probability space to the real number space through the link function, and the Logit link function is the standard choice for processing binary response variable, which has simple mathematical form and good statistical properties. The value after Logit transformation is called log odds ratio, which has a clear epidemiological interpretation, and the regression coefficient represents the change of log odds ratio when the covariate increases by one unit, and the exponential value is the odds ratio. This interpretation is widely used in medical research and is easy to understand. The Logit link function maps the probability of occurrence of language frequency hearing loss of an individual at a certain time point to the linear predictor. The definition of Logit function is that for a probability p, first calculate p divided by (1 minus p), and then take the natural logarithm of the resulting quotient. The linear predictor includes main effects and time-varying interaction effects. The calculation method of the linear predictor is as follows: first, take the intercept term, then add the product of age and age main effect coefficient, the product of cumulative noise exposure and cumulative noise exposure main effect coefficient, the product of genotype and genotype main effect coefficient, the product of the interaction term of genotype and cumulative noise exposure and the first interaction effect coefficient, the product of the interaction term of genotype and age and the second interaction effect coefficient, and finally add the inner product of the other covariate vector and its corresponding regression coefficient vector. Among them, the intercept term is the base value of the marginal mean model. The age main effect coefficient reflects the independent influence of age on the risk of hearing loss. The cumulative noise exposure main effect coefficient reflects the independent influence of noise exposure on the risk of hearing loss. The genotype main effect coefficient reflects the independent influence of genetic factors on the risk of hearing loss. The first interaction effect coefficient reflects the modification of genetic background on noise susceptibility, and the second interaction effect coefficient reflects the age dependence of genetic effect. The other covariate vector includes variables such as gender, smoking history, high-frequency hearing threshold, and the corresponding regression coefficient vector includes the regression coefficients of these variables respectively.

[0099] GEE model needs to specify working correlation matrix to describe the correlation between observations of the same individual at different time points, and working correlation matrix includes correlation structure, which includes independent structure, exchangeable structure, first-order autoregressive structure and no structure.

[0100] The independent structure assumes that the observations of the same individual at different time points are independent of each other, and the correlation matrix is an identity matrix, i.e., all diagonal elements are 1, and all non-diagonal elements are 0. The exchangeable structure assumes that the observations of the same individual at any two time points have the same correlation coefficient, and all diagonal elements of the correlation matrix are 1, and all non-diagonal elements are equal to the same correlation coefficient value. The first-order autoregressive structure assumes that the correlation between adjacent time points is the strongest, and the correlation decays as the time interval increases. In the correlation matrix of the first-order autoregressive structure, all diagonal elements are 1, and the farther the element is from the diagonal, the smaller the correlation coefficient. Specifically, the correlation coefficient of adjacent time points is a certain base correlation coefficient, the correlation coefficient of one time interval is the square of the base correlation coefficient, the correlation coefficient of two time intervals is the cube of the base correlation coefficient, and so on. No structure does not impose any constraints on the correlation, and estimates all correlation coefficients, i.e., each non-diagonal element in the correlation matrix can take a different value.

[0101] The system selects an optimal correlation structure using a quasi-likelihood criterion under the independent model assumption. The reason for using the quasi-likelihood criterion under the independent model assumption for correlation structure selection is that the selection of the working correlation matrix has an important influence on the efficiency of the generalized estimating equation model. Although the parameter estimation remains consistent when the correlation structure is set incorrectly, selecting a working correlation matrix that is closer to the true correlation structure can improve the estimation efficiency, reduce the standard error, and enhance the power of statistical tests. The quasi-likelihood criterion under the independent model assumption is an information criterion that evaluates the pros and cons of different correlation structures by balancing model goodness of fit and model complexity. This criterion has a good theoretical foundation and practical performance under the generalized estimating equation framework. This criterion is a model selection criterion for generalized estimating equation models, similar to the Akaike information criterion. The calculation method of this criterion is as follows: first, calculate the negative 2 times of the quasi-likelihood function value, then add a penalty term. The quasi-likelihood function is obtained by summing the quasi-likelihood contributions of all individuals. The quasi-likelihood contribution of each individual is the product of the likelihood term and negative one-half. The likelihood term is the difference between the response variable vector of the individual and the expected response vector. The difference is left multiplied by the inverse of the working covariance matrix, and then right multiplied by the transpose of the difference. The calculation method of the penalty term is as follows: multiply the model covariance matrix estimate under the assumption of independent structure with the model covariance matrix estimate when using the working correlation matrix, then calculate the trace of the product matrix, which is the sum of the diagonal elements, and finally multiply the trace by a preset penalty term coefficient. The default value of the preset penalty term coefficient is 2. This coefficient is the standard setting of the quasi-likelihood criterion under the independent model assumption, and its penalty strength is consistent with that of the Akaike information criterion. This coefficient can be adjusted within the range of 1 to 4 according to the sample size and model complexity. A larger coefficient will tend to select a simpler correlation structure. The model covariance matrix estimate is the regression coefficient covariance matrix estimate obtained by fitting the generalized estimating equation model using the working correlation matrix. The regression coefficients include the intercept term, the age main effect coefficient, the cumulative noise exposure main effect coefficient, the genotype main effect coefficient, the first interaction effect coefficient, and the second interaction effect coefficient, as well as the corresponding regression coefficient vector.

[0102] The system fits the generalized estimating equation model using the above four correlation structures respectively, calculates the quasi-likelihood criterion value under the independent model assumption for each correlation structure, and selects the correlation structure with the smallest criterion value as the optimal structure.

[0103] Using the selected working correlation matrix, the system estimates the model parameters by solving the generalized estimating equations. The generalized estimating equations are defined as, for all individuals, the contribution of each individual is the transpose of the partial derivative matrix of the expected response vector with respect to the parameter vector, left multiplied by the inverse of the working covariance matrix, right multiplied by the difference between the response variable vector and the expected response vector, the sum of the contributions of all individuals equals the zero vector. Wherein the number of individuals represents the total number of people included in the study. The response variable vector includes the response variable values of the individual at all follow-up time points, the response variable values are equal to the language frequency hearing loss status. The expected response vector includes the expected response values of the individual at all follow-up time points, the expected response value at each time point is equal to the probability of the occurrence of language frequency hearing loss at the time point, and the specific calculation method is that the expected response is equal to the expected exponential function divided by the expected term, the expected exponential function is the exponential function with the natural constant as the base, the linear predictor is the exponential, and the expected term is 1 plus the expected exponential function. The partial derivative matrix of the expected response vector with respect to the parameter vector represents the gradient of the change of the expected response vector with the parameter vector. The working covariance matrix is obtained by left multiplying the square root of the variance matrix by the working correlation matrix and right multiplying the square root of the variance matrix, wherein the variance matrix is a diagonal matrix, the diagonal elements are the variances of the response variables at each time point, and the working correlation matrix describes the correlation between the observations at different time points of the same individual, and the correlation parameters are the parameters in the working correlation matrix.

[0104] The system solves the generalized estimating equation by using the iterative weighted least squares method. The reason for using the iterative weighted least squares method is that the generalized estimating equation is a set of nonlinear equations, which cannot obtain parameter estimates by direct solving, and needs to use an iterative algorithm to gradually approach the optimal solution. The iterative weighted least squares method is a standard algorithm for solving the generalized estimating equation. This method calculates the weight matrix and working response variable according to the current parameter estimate value in each iteration, and then updates the parameter estimate by using the weighted least squares method. This algorithm has good convergence and computational stability, and is widely used in the solution of generalized linear models and generalized estimating equations. The solution process of the iterative weighted least squares method includes the following steps. First, initialize the regression coefficient parameters and related parameters. A simple initial value such as a zero vector can usually be used. Second, perform iterative updates. In each iteration, the following operations are performed. First, calculate the expected response vector and working covariance matrix of each individual according to the current regression coefficient parameters. Then, update the regression coefficient parameters. The update method is as follows. First, calculate the information matrix. The information matrix is obtained by summing all individuals. The contribution of each individual is the transpose of the partial derivative matrix of the expected response vector with respect to the parameter, multiplied by the inverse of the working covariance matrix, and then multiplied by the partial derivative matrix. Then, calculate the inverse of the information matrix. Then, calculate the score vector. The score vector is obtained by summing all individuals. The contribution of each individual is the transpose of the partial derivative matrix of the expected response vector with respect to the parameter, multiplied by the inverse of the working covariance matrix, and then multiplied by the difference between the response variable vector and the expected response vector. Finally, add the product of the inverse of the information matrix and the score vector to the current regression coefficient parameters to obtain the updated regression coefficient parameters. Next, update the correlation parameters according to the Pearson residual. Third, repeat the iterative process of the second step until convergence. The convergence criterion is that the parameter change is less than the preset convergence threshold. The default value of the preset convergence threshold is 0.0001. The setting of this threshold is based on the standard accuracy requirement in numerical optimization theory, which can ensure the stability and accuracy of the parameter estimate. This threshold can be adjusted within the range of 0.00001 to 0.001 according to the calculation accuracy requirement. A smaller threshold will improve the accuracy but increase the calculation time. After convergence, the system obtains the regression coefficient estimate, including the intercept term, the age main effect coefficient, the cumulative noise exposure main effect coefficient, the genotype main effect coefficient, the first interaction effect coefficient and the second interaction effect coefficient, and the estimated value of the corresponding regression coefficient vector.

[0105] The system also calculates the robust standard error of the parameter, also known as sandwich estimator or empirical standard error. The reason for calculating the robust standard error is that the setting of the working correlation matrix may deviate from the true correlation structure. If the model-based standard error is used, i.e. the standard error assuming that the working correlation matrix is completely correct, the standard error estimation will be biased when the correlation structure is set incorrectly, which will affect the accuracy of hypothesis testing and confidence intervals. The robust standard error combines model information and empirical information, which can still provide consistent standard error estimation even if the working correlation matrix is set incorrectly. This robustness is an important advantage of the generalized estimating equations method, which ensures the reliability of statistical inference. The covariance matrix estimation method of the robust standard error is as follows: first, calculate the model information matrix, which is obtained by summing all individuals. The contribution of each individual is the transpose of the partial derivative matrix of the expected response vector with respect to the parameter, multiplied by the inverse of the working covariance matrix, and then multiplied by the partial derivative matrix. Then calculate the empirical information matrix, which is obtained by summing all individuals. The contribution of each individual is the transpose of the partial derivative matrix of the expected response vector with respect to the parameter, multiplied by the inverse of the working covariance matrix, multiplied by the product of the difference between the response variable vector and the expected response vector and its transpose, multiplied by the inverse of the working covariance matrix, and finally multiplied by the partial derivative matrix. Then calculate the covariance matrix, which is equal to the inverse of the model information matrix multiplied by the empirical information matrix multiplied by the transpose of the inverse of the model information matrix. The robust standard error is the square root of the diagonal elements of the covariance matrix.

[0106] The risk prediction module 105, based on the regression coefficient estimates, performs language frequency hearing loss risk prediction at future time points for individuals, outputs risk probabilities and risk stratification results for individuals, as shown in detail in Figure 4

[0107] For the individual to be predicted, the system first obtains the variable information at the current time point of the individual, including genotype or genetic risk score, current cumulative noise exposure, current age, and other covariates such as gender, smoking history, etc. The system predicts the risk probability of the individual suffering from language frequency hearing loss at a future time point, such as a preset risk prediction time window. First, the system calculates the variable information at the future time point, which includes future age, future cumulative noise exposure, future length of service, and linear predictor at the future time point. The future age is equal to the current age plus the time interval. The calculation of the future cumulative noise exposure assumes that the individual continues to work in the current noise environment. The future length of service is equal to the current length of service plus the time interval, and the future cumulative noise exposure is equal to the equivalent continuous A sound level plus the logarithm of the future length of service, which is the logarithm with base 10 multiplied by 10. Then, the system calculates the linear predictor at the future time point.

[0108] ​The calculation method of the linear predictor at the future time point is to take the intercept term estimate, plus the age main effect coefficient estimate multiplied by the future age, plus the cumulative noise exposure main effect coefficient estimate multiplied by the future cumulative noise exposure, plus the genotype main effect coefficient estimate multiplied by the genotype, plus the first interaction effect coefficient estimate multiplied by the product of the genotype and the future cumulative noise exposure, plus the second interaction effect coefficient estimate multiplied by the product of the genotype and the future age, and finally plus the inner product of the other covariate vector and its corresponding regression coefficient estimate vector. Among them, all coefficient estimates are regression coefficient estimates output by the time-varying interaction module 104. Finally, the risk probability is calculated by the inverse function of the Logit link function, that is, the Logistic function. The calculation method of the Logistic function is that the risk probability is equal to the exponential function of the linear predictor at the future time point divided by the linear predictor term at the future time point, the exponential function of the linear predictor at the future time point is the exponential function with the natural constant as the base, the linear predictor at the future time point is the exponential function, and the linear predictor term at the future time point is the exponential function of the linear predictor at the future time point plus 1. This probability represents the predicted risk of an individual developing language frequency hearing loss at a future time point.

[0109] For clinical application, the system divides individuals into different risk groups according to genetic risk scores or key genotypes, and the risk groups include a low genetic risk group, a medium genetic risk group, and a high genetic risk group. Among them, the low genetic risk group includes individuals who do not carry risk alleles or individuals whose genetic risk scores are lower than a preset low risk score threshold, the medium genetic risk group includes individuals whose genetic risk scores are between the preset low risk score threshold and the preset high risk score threshold, that is, the genetic risk score is greater than or equal to the preset low risk score threshold and less than or equal to the preset high risk score threshold; and the high genetic risk group includes individuals whose genetic risk scores are higher than the preset high risk score threshold or individuals carrying high-risk homozygous mutations, and the high-risk homozygous mutations include GJB2 gene c.35delG homozygous mutations.

[0110] The preset low risk score threshold is an individual with a preset low risk quantile threshold of the genetic risk score distribution in the training data set. The default value of the preset low risk quantile threshold is the 25th percentile, and the setting of this threshold makes the low risk group include about one-fourth of the population, representing a population with lower genetic susceptibility, and this threshold can be adjusted within the range of the 10th to the 33rd percentile according to the population distribution characteristics. The preset high risk score threshold is an individual with a preset high risk quantile threshold of the genetic risk score distribution in the training data set, and the default value of the preset high risk quantile threshold is the 75th percentile. The setting of this threshold makes the high risk group include about one-fourth of the population, representing a population with higher genetic susceptibility, and this threshold can be adjusted within the range of the 67th to the 90th percentile according to the population distribution characteristics.

[0111] For each risk group, the system plots the curve of cumulative noise exposure and the probability of onset. Specifically, fix the age, genotype, other covariate vector for typical values, including median, mode, minimum, maximum, etc., let the cumulative noise exposure change within a reasonable range, calculate the probability of onset corresponding to different cumulative noise exposure values. In an embodiment of the present application, the lower limit of the reasonable range of noise exposure is the 5th percentile of the cumulative noise exposure in the training data set, avoiding the influence of extreme low values on the readability of the curve, while ensuring coverage of most low exposure individuals; the upper limit is the 95th percentile of the cumulative noise exposure in the training data set, avoiding abnormal extrapolation caused by extreme high values, while covering most high exposure individuals. The calculation method of the probability of onset is as follows: first calculate the linear predictor, which is the sum of the intercept term estimate, the onset age term, the onset noise term, the onset genotype term, the onset first interaction term, the onset second interaction term and the onset other covariate term, wherein the onset age term is the age multiplied by the age main effect coefficient estimate, the onset noise term is the cumulative noise exposure multiplied by the cumulative noise exposure main effect coefficient estimate, the onset genotype term is the genotype multiplied by the genotype main effect coefficient estimate, the onset first interaction term is the product of the first interaction effect coefficient estimate and the product of the genotype and the cumulative noise exposure, the onset second interaction term is the product of the second interaction effect coefficient estimate and the product of the genotype and the age, and the onset other covariate term is the inner product of the other covariate vector and its corresponding regression coefficient estimate vector. The probability of onset is the exponential function of the linear predictor divided by the linear predictor term, the linear predictor is the exponential function of the natural constant, the linear predictor term is the linear predictor exponential function plus 1. Plot the curve for different genotype values corresponding to low, medium and high genetic risk groups respectively, which can intuitively display the heterogeneity of dose-response relationship under different genetic backgrounds.

[0112] To quantify the difference in noise susceptibility among different genetic backgrounds, the system calculates the cumulative noise exposure threshold at which each genotype individual reaches a fifty percent risk of onset. This threshold is defined as the cumulative noise exposure amount that makes the probability of an individual developing a language-frequency hearing loss equal to 0.5. When the probability of onset is equal to 0.5, the Logit value corresponding to this probability is equal to 0 because 0.5 divided by 0.5 is equal to 1 and the natural logarithm of 1 is equal to 0. Thus, the linear predictor is equal to 0. The expression for the linear predictor is the sum of the intercept term estimate, the onset age term, the onset noise term, the onset genotype term, the onset first interaction term, the onset second interaction term, and the onset other covariates term, where the onset age term is the age main effect coefficient estimate multiplied by the age, the onset noise term is the cumulative noise exposure main effect coefficient estimate multiplied by the cumulative noise exposure threshold, the onset genotype term is the genotype main effect coefficient estimate multiplied by the genotype, the onset first interaction term is the first interaction effect coefficient estimate multiplied by the product of the genotype and the cumulative noise exposure threshold, the onset second interaction term is the second interaction effect coefficient estimate multiplied by the product of the genotype and the age, and the onset other covariates term is the inner product of the other covariates vector and its corresponding vector of regression coefficient estimates, which is equal to 0. Rearranging this equation yields the formula for the cumulative noise exposure threshold. The cumulative noise exposure threshold is equal to the negative intercept term estimate minus the onset age term minus the onset genotype term minus the onset second interaction term minus the onset other covariates term, divided by the cumulative noise term, which is the first interaction effect coefficient estimate multiplied by the genotype plus the cumulative noise exposure main effect coefficient estimate. This formula indicates that the cumulative noise exposure threshold depends on the genotype. For different genotypes, such as genotypes with values of 0, 1, and 2 corresponding to wild-type homozygous, heterozygous mutant, and risk-type homozygous, respectively, the system calculates the respective cumulative noise exposure thresholds. The smaller the cumulative noise exposure threshold, the more sensitive the individual of that genotype is to noise and the less exposure is required to reach the same risk of onset. By comparing the differences in cumulative noise exposure thresholds among different genotypes, the system quantifies the degree of genetic susceptibility. For example, if the cumulative noise exposure threshold for a risk-type homozygous individual is 20 decibels lower than that for a wild-type homozygous individual, then carrying the risk allele significantly increases the individual's susceptibility to noise.

[0113] To facilitate the use of the prediction model by clinicians and occupational health workers, the system converts the generalized estimating equation model into a nomogram graphical representation. The reason for constructing the nomogram is that although the generalized estimating equation model has good prediction performance, its mathematical form is relatively complex, containing multiple covariates and interaction terms, and it is difficult for clinicians and occupational health workers to directly use the model formula to calculate the risk in actual application. The nomogram converts the complex mathematical model into an intuitive graphical tool, and the user does not need to understand the mathematical details of the model. Only by looking up the scores of each variable on the graph and adding them up, the individual risk prediction result can be quickly obtained. This visualization tool greatly improves the clinical usability and promotional value of the model, and is widely recognized in the application of medical prediction models. The nomogram is a graphical prediction tool that can calculate individual risk by looking up the table and simple addition. The steps for constructing the nomogram are as follows. First, determine the variable range. For each prediction variable such as age, cumulative noise exposure, genotype, high-frequency hearing threshold, etc., determine its reasonable value range in clinical practice. The method for determining the variable range is that for continuous variables, take the value range between the lower limit percentile and the upper limit percentile of the pre-set variable range of the variable in the training data set. The default value of the lower limit percentile of the pre-set variable range is the 5th percentile, and the default value of the upper limit percentile of the pre-set variable range is the 95th percentile. Such setting can exclude the influence of extreme values and cover most clinical actual situations. These two percentiles can be adjusted within the range of 1st to 10th percentiles and 90th to 99th percentiles, respectively, according to the data distribution characteristics. For categorical variables such as genotype, take all possible category values.

[0114] Second, calculate variable scores. For each variable, convert its regression coefficient to a score. Choose the variable with the largest absolute value of regression coefficient as the reference variable, and set its score range from 0 to a pre-set nomogram maximum score value. The scores of other variables are scaled proportionally. The default value of the pre-set nomogram maximum score value is 100, which is the standard configuration of the nomogram tool for clinical workers to understand and use. The value can also be set to 50 or 200 according to the accuracy requirement. For example, if the cumulative noise exposure has the largest absolute value of regression coefficient, the score calculation method of the cumulative noise exposure is to subtract the minimum value of the cumulative noise exposure from the cumulative noise exposure, then divide by the maximum value minus the minimum value of the cumulative noise exposure, and finally multiply by the pre-set nomogram maximum score value. The score calculation method of other variables is to divide the regression coefficient estimate of the variable by the regression coefficient estimate of the cumulative noise exposure, and then multiply by the score of the cumulative noise exposure. Third, calculate the total score. The total score of an individual is the sum of the scores of each variable, i.e. adding up the scores of all variables. Fourth, map the risk probability. Calculate the risk probability according to the total score. Since the linear predictor is equal to the sum of the products of all variable values and their corresponding regression coefficients, and the score is proportional to the linear predictor, a mapping relationship between the total score and the risk probability can be established. The calculation method of the risk probability is to multiply the total score by the scaling parameter a, add the scaling parameter b, take the exponential function of the sum, and then divide by 1 plus the exponential function. The scaling parameters a and b are determined by fitting.

[0115] Fifth, draw the nomogram. The nomogram includes multiple horizontal scale lines, each corresponding to a prediction variable. The uppermost is the score line, with a scale range of 0 to the pre-set nomogram maximum score value. The middle is the scale line of each prediction variable. The lowermost is the total score line and the risk probability line. When using, find the corresponding value on the scale line of each variable, draw a vertical line upwards to intersect the score line, and read the score. Add up the scores of all variables to get the total score. Find the total score value on the total score line, draw a vertical line downwards to intersect the risk probability line, and read the predicted risk probability in the pre-set risk prediction time window. The default value of the pre-set risk prediction time window is 3 years or 5 years, which are the commonly used medium-term and long-term prediction time periods in occupational health surveillance. The 3-year prediction is suitable for situations that require frequent evaluation and adjustment of protective measures, and the 5-year prediction is suitable for long-term career planning. The time window can be adjusted within 1 to 10 years according to management needs. The output of the risk prediction module 105 includes the risk probability of an individual at a specific future time point, the genetic risk stratification results such as low-risk group, medium-risk group or high-risk group, the cumulative exposure threshold of different genotypes, and the nomogram graphical tool. These results are passed to the decision support module 106.

[0116] The decision support module 106 generates hearing protection recommendations for the individual based on the risk probability and risk stratification results.

[0117] The system classifies individuals into different risk levels according to the predicted risk probability, including low risk, medium risk, and high risk, and formulates corresponding management strategies. Low risk is defined as a risk probability less than a preset low risk probability threshold, with a default value of 0.3. The threshold is set based on occupational health risk management practices and represents an acceptable low risk level. The threshold can be adjusted within the range of 0.2 to 0.4 according to industry standards and management strategies. For low-risk individuals, the protective measures are routine hearing protection, wearing standard hearing protection devices, and annual hearing monitoring. According to the routine requirements of occupational health surveillance, the intervention suggestion is health education, emphasizing the importance of correct use of hearing protection devices. Medium risk is defined as a risk probability greater than or equal to the preset low risk probability threshold and less than the preset high risk probability threshold, with a default value of 0.5. The threshold is set to represent the risk level that requires intensified intervention, corresponding to the critical point where the probability of disease reaches one-half. The threshold can be adjusted within the range of 0.4 to 0.6 according to the intensity of the prevention strategy. For medium-risk individuals, the protective measures are enhanced hearing protection, choosing hearing protection devices with better noise reduction performance, ensuring full-time wearing, and conducting semi-annual hearing monitoring to detect hearing changes in a timely manner. The intervention suggestion is to consider reducing noise exposure, such as shortening daily exposure time, arranging for job rotation, and evaluating the effectiveness of work environment noise control measures. High risk is defined as a risk probability greater than or equal to the preset high risk probability threshold. For high-risk individuals, the protective measures are mandatory high-level hearing protection, using earmuffs or earplugs combined with earmuffs, and conducting quarterly hearing monitoring to closely track hearing status. The intervention suggestion is to strongly recommend transferring to a low-noise post, and if transfer is not possible, to strictly limit exposure time and strengthen engineering control measures such as soundproofing and sound-absorbing equipment.

[0118] The system recommends appropriate hearing protection devices based on the individual's genetic risk stratification and current cumulative noise exposure level. The reason for recommending hearing protection devices based on genetic risk stratification and cumulative noise exposure level is that individuals with different genetic backgrounds have significant differences in susceptibility to noise, and individuals carrying risk genotypes have a higher risk of hearing loss under the same noise exposure conditions. Therefore, they need higher-level hearing protection. At the same time, cumulative noise exposure reflects the amount of noise an individual has already been exposed to, and the higher the exposure, the more stringent protection measures are needed to avoid further damage. By integrating genetic information and exposure information, the system can provide precise protection recommendations for each individual, achieving personalized occupational health management. This risk stratification-based protection strategy is more scientific and reasonable than the traditional one-size-fits-all approach, effectively protecting high-risk individuals while avoiding over-protection of low-risk individuals, and improving resource utilization efficiency. The noise reduction level of hearing protection devices represents their ability to reduce noise exposure, measured in decibels.

[0119] The recommendation rules are as follows. For individuals with low genetic risk and cumulative noise exposure that meets the preset low exposure, i.e., the cumulative noise exposure is less than the preset medium exposure percentile threshold of the cumulative noise exposure distribution in the training data set, the standard earplug is recommended, and the noise reduction level is in the preset standard earplug noise reduction level range. The default value of the preset medium exposure percentile threshold is the 50th percentile, i.e., the median, which divides the population's cumulative exposure level into low exposure and high exposure. This threshold can be adjusted to be within the 40th to 60th percentile range according to the exposure distribution characteristics of the population. The default value of the preset standard earplug noise reduction level range is 15 to 20 decibels, which is determined according to the typical noise reduction performance of standard earplug products and is suitable for environments with noise intensity less than the preset low noise environment threshold. The default value of the preset low noise environment threshold is 90 decibels A-weighted, which is an important dividing point for noise exposure in occupational health standards. Noise environments below this threshold are generally considered relatively safe. This threshold can be adjusted to be within the 85 to 95 decibels A-weighted range according to the occupational health standards of different countries or regions. For individuals with medium genetic risk or cumulative noise exposure that meets the preset medium exposure, i.e., the cumulative noise exposure is greater than or equal to the preset medium exposure percentile threshold and less than the preset high exposure percentile threshold, high-performance earplugs or earmuffs are recommended, and the noise reduction level is in the preset high-performance hearing protection device noise reduction level range. The default value of the preset high exposure percentile threshold is the 75th percentile, which identifies the population with higher cumulative exposure. This threshold can be adjusted to be within the 67th to 90th percentile range according to the exposure distribution characteristics of the population. The default value of the preset high-performance hearing protection device noise reduction level range is 20 to 25 decibels, which is determined according to the typical noise reduction performance of high-performance earplug or earmuff products and is suitable for environments with noise intensity greater than or equal to the preset low noise environment threshold and less than the preset high noise environment threshold. The default value of the preset high noise environment threshold is 100 decibels A-weighted, which represents a high-intensity noise environment. Measures more stringent than this threshold need to be taken. This threshold can be adjusted to be within the 95 to 105 decibels A-weighted range according to the characteristics of the work environment. For individuals with high genetic risk or cumulative noise exposure that meets the preset high exposure, i.e., the cumulative noise exposure is greater than or equal to the preset high exposure percentile threshold, earplug plus earmuff combination is recommended, and the total noise reduction level can reach the preset combination hearing protection device noise reduction level range. The default value of the preset combination hearing protection device noise reduction level range is 25 to 35 decibels, which is determined according to the additive noise reduction effect of earplug plus earmuff combination and is suitable for environments with noise intensity greater than or equal to the preset high noise environment threshold or high genetic susceptibility individuals. The system calculates the effective cumulative noise exposure after wearing the hearing protection device. The calculation method of the effective cumulative noise exposure is to subtract the noise reduction level of the selected hearing protection device from the equivalent continuous A-weighted sound level, then add the logarithm value of the service length multiplied by 10, where the logarithm is the logarithm with base 10, and the service length is in years.The effective dose is used to re-evaluate the risk and ensure that the protective measures can reduce the risk to an acceptable level.

[0120] For individuals with high genetic risk, the system calculates the upper limit of safe exposure time under the current noise environment based on their cumulative exposure threshold. According to the definition of cumulative noise exposure, the cumulative exposure threshold is equal to the equivalent continuous A-weighted sound level plus the logarithm value of the upper limit of safe exposure time multiplied by 10, where the logarithm is the logarithm with base 10. Thus, the formula for calculating the upper limit of safe exposure time is obtained. The upper limit of safe exposure time is equal to the power of 10, and the exponent of the power is the cumulative exposure threshold minus the equivalent continuous A-weighted sound level divided by 10. The unit of the upper limit of safe exposure time is years. This formula shows that under a given noise intensity, the risk of an individual reaching the upper limit of safe exposure time will reach 50% after working for a certain number of years. If the individual's current length of service is close to the upper limit of safe exposure time, for example, the current length of service is greater than the upper limit of safe exposure time multiplied by a preset length-of-service warning proportion coefficient, the system suggests considering job rotation and regular exchange of work between low-noise workers, or transferring to a position with lower noise intensity, or strictly limiting daily exposure time, such as reducing the preset standard working hours by half, which is equivalent to reducing the cumulative exposure rate by half. The default value of the preset length-of-service warning proportion coefficient is 0.8. This coefficient is set to provide early warning before the individual reaches the upper limit of safe exposure time, leaving enough intervention time. This coefficient can be adjusted within the range of 0.7 to 0.9 according to the aggressiveness of the prevention strategy. The default value of the preset standard working hours is 8 hours per day. This duration is the standard working day duration specified by the International Labor Organization and various national labor laws. This duration can be adjusted according to the work system of specific industries and positions.

[0121] The system integrates the above analysis results to generate a personalized decision support report. The report includes the following content:

[0122] The first part is individual genetic risk assessment, including genotype information, listing the key single nucleotide polymorphism sites detected and their genotypes, such as the GJB2 gene c.35delG site is wild type homozygous, genetic risk score and its percentile in the population, and genetic risk stratification results are low risk group, medium risk group or high risk group. The second part is the current hearing status assessment, including the average hearing threshold of speech frequency and its clinical significance, such as normal, mild hearing loss or moderate hearing loss, comparison with the same age and gender population, that is, the deviation of individual hearing threshold from the age-adjusted normal model, and hearing threshold change trend, that is, if there is historical data, the curve of hearing threshold change over time is displayed. The third part is future risk prediction, including preset short-term risk prediction time window incidence risk and its preset confidence level confidence interval, preset long-term risk prediction time window incidence risk and its preset confidence level confidence interval, and risk level is low risk, medium risk or high risk. The default value of the preset short-term risk prediction time window is 3 years, which is suitable for medium-term risk assessment and adjustment of preventive measures, and the time window can be adjusted within the range of 1 to 5 years according to management needs. The default value of the preset long-term risk prediction time window is 5 years, which is suitable for long-term career planning and post arrangement, and the time window can be adjusted within the range of 3 to 10 years according to management needs. The default value of the preset confidence level is 95%, which is a standard setting in statistics, corresponding to a significance level of 0.05, and can provide a high statistical confidence level. The confidence level can be adjusted within the range of 90% to 99% according to the strictness of risk assessment. The fourth part is individualized protection recommendations, including recommended types of hearing protection devices such as earplugs, earmuffs or combinations and noise reduction levels, effective exposure after wearing hearing protection devices and risk reduction effect, exposure time limit if applicable, and monitoring frequency is annual, semi-annual or quarterly hearing monitoring.

[0123] The fifth part is lifestyle and health management recommendations, including smoking cessation recommendations, as smoking and noise exposure have a synergistic effect that increases the risk of speech frequency hearing loss, avoiding ototoxic drugs such as aminoglycoside antibiotics, certain chemotherapy drugs, reducing recreational noise exposure, avoiding long-term use of high-volume earphones, reducing participation in high-noise recreational activities, regular physical examination recommendations, and monitoring cardiovascular health, as cardiovascular disease may affect the blood supply to the inner ear.

[0124] The sixth part is nomogram graphical tool, with individualized nomogram, marking the value of each variable of the individual and the corresponding score, to facilitate understanding of the source of risk. The output of the decision support module 106 is a complete individualized decision support report, which can be printed or provided in electronic form to clinicians, occupational health workers and workers themselves, for guidance on hearing protection and health management.

[0125] The above describes the embodiments of the present application, but the embodiments are not limited to the above specific embodiments, and the above specific embodiments are only illustrative but not restrictive, and the ordinary skilled in the art can make more equivalent embodiments under the inspiration of the embodiments, which are all within the protection scope of the embodiments.

Claims

1. A noise-induced hearing loss prediction system based on GEE model and genetic characteristics, characterized in that, The method comprises the following steps: a data collection module collects follow-up data of workers exposed to occupational noise from multiple data sources, the follow-up data including demographic information, hearing data, genetic information, and environmental noise data; a data processing module, based on the follow-up data, performs data cleaning, outlier processing, missing value filling, and variable standardization to generate a standardized modeling data set; a feature selection module, based on the modeling data set, selects features significantly related to language frequency hearing loss by statistical learning methods to obtain a key feature subset; a time-varying interaction module, based on the key feature subset and the modeling data set, constructs a generalized estimating equation model including time-varying interaction terms, and the generalized estimating equation model outputs regression coefficient estimates; including: The time-varying interaction terms include genotype-cumulative noise exposure interaction terms and genotype-age interaction terms, the genotype-cumulative noise exposure interaction terms are obtained by multiplying the genotype variable of an individual with the cumulative noise exposure of the individual at a certain time point, and the genotype-age interaction terms are obtained by multiplying the genotype variable of an individual with the age of the individual at a certain time point; a marginal mean model is established using a Logit link function, and the Logit function maps the probability of an individual having language frequency hearing loss at a certain time point to a linear predictor; The generalized estimating equation model specifies a working correlation matrix to describe the correlation between observations of the same individual at different time points, the working correlation matrix includes a correlation structure, the correlation structure includes an independent structure, a symmetric structure, a first-order autoregressive structure, and a structureless structure, and a quasi-likelihood criterion under the assumption of an independent model is used to select the optimal correlation structure; Using the selected working correlation matrix, the model parameters are estimated by solving the generalized estimating equation, the iterative weighted least squares method is used to solve the generalized estimating equation, and the regression coefficient estimates are obtained, including the intercept term, the age main effect coefficient, the cumulative noise exposure main effect coefficient, the genotype main effect coefficient, the first interaction effect coefficient, and the second interaction effect coefficient, and the estimated value of the corresponding regression coefficient vector; a risk prediction module, based on the regression coefficient estimates, predicts the risk of language frequency hearing loss at a future time point for an individual, and outputs the risk probability and risk stratification results of the individual; a decision support module, based on the risk probability and risk stratification results, generates hearing protection recommendations for the individual.

2. The GEE model and genetic profile based noise-induced language frequency hearing impairment prediction system according to claim 1, wherein, The demographic information includes the gender, age, education level, work experience, occupation type, smoking history, and high-frequency hearing threshold of the worker; The hearing data includes a response variable; when the language frequency average hearing threshold of a worker individual at a certain time point is greater than a preset hearing loss determination threshold, the response variable takes the value 1, and when the language frequency average hearing threshold is less than or equal to the preset hearing loss determination threshold, the response variable takes the value 0; the language frequency average hearing threshold is obtained by adding the hearing thresholds of the individual at a preset set of language frequency points and dividing by the number of frequency points; The genetic information comprises a genetic risk score and a genotype, and a method for calculating the genetic risk score comprises: obtaining single nucleotide polymorphism typing of an individual at a plurality of hearing-related gene sites, determining the number of risk alleles at each site, multiplying the number of risk alleles at the site by a preset effect weight of the site to obtain a weighted score, and accumulating the weighted scores of all sites to obtain the genetic risk score; The environmental noise data comprises a cumulative noise exposure, and the cumulative noise exposure is a sum of a logarithm value of a service length and an equivalent continuous A sound level.

3. The GEE model and genetic profile based noise induced language frequency hearing impairment prediction system as claimed in claim 1, wherein, The obtaining of the key feature subset comprises: A candidate feature set is constructed, and the candidate feature set comprises genetic features, audiological features, and environmental exposure and demographic features; A least absolute shrinkage and selection operator regression is used to preliminarily screen the candidate feature set, an optimal preset regularization parameter value is determined by a cross-validation method, the least absolute shrinkage and selection operator model is refitted on all baseline data using the optimal preset regularization parameter value, and features with non-zero regression coefficients are retained to form a least absolute shrinkage and selection operator screening feature set; A random forest classifier is used to evaluate the importance of the candidate feature set, a Gini importance index is used to evaluate the importance of each feature, all features in the candidate feature set are sorted in descending order according to the Gini importance score, and a preset number of features with high importance are selected to form a random forest screening feature set; The least absolute shrinkage and selection operator screening feature set and the random forest screening feature set are integrated, a comprehensive feature set is constructed by taking the union of the two sets, the genetic features in the comprehensive feature set are biologically verified, and a key feature subset is determined.

4. The GEE model and genetic profile based noise induced language frequency hearing impairment prediction system as claimed in claim 1, wherein, A method for calculating the linear predictor comprises: taking an intercept term, sequentially adding an age multiplied by an age main effect coefficient, a cumulative noise exposure multiplied by a cumulative noise exposure main effect coefficient, a genotype multiplied by a genotype main effect coefficient, an interaction term of the genotype and the cumulative noise exposure multiplied by a first interaction effect coefficient, an interaction term of the genotype and the age multiplied by a second interaction effect coefficient, and finally adding an inner product of an other covariate vector and a corresponding regression coefficient vector; wherein the intercept term is a basic value of a marginal mean model, the other covariate vector comprises gender, smoking history, and high-frequency hearing threshold, and the corresponding regression coefficient vector comprises regression coefficients of the other covariate vector.

5. The GEE model and genetic profile based noise induced language frequency hearing impairment prediction system as claimed in claim 1, wherein, The selection of the optimal correlation structure using the quasi-likelihood criterion under the independent model assumption comprises: A method for calculating the quasi-likelihood criterion under the independent model assumption comprises: calculating a negative 2 times of a quasi-likelihood function value, and adding a penalty term. The quasi-likelihood function is obtained by summing quasi-likelihood contributions of all individuals, each quasi-likelihood contribution being a likelihood term multiplied by negative one-half, the likelihood term being a difference between a response variable vector of the individual and an expected response vector, the difference being left-multiplied by an inverse of a working covariance matrix and right-multiplied by a transpose of the difference; the expected response vector including expected response values of the individual at all follow-up time points, each expected response value being an expected exponential function divided by an expected term, the expected exponential function being an exponential function of an exponential predictor with a natural constant as a base, the expected term being 1 plus the expected exponential function; the working covariance matrix being obtained by left-multiplying a working correlation matrix by a square root of a variance matrix and right-multiplying the square root of the variance matrix, wherein the variance matrix is a diagonal matrix, diagonal elements being variances of response variables at the time points; The penalty term is calculated by multiplying a model covariance matrix estimate under the assumption of independent structure and a model covariance matrix estimate when the working correlation matrix is used, calculating a trace of a product matrix, and multiplying the trace by a preset penalty term coefficient; the model covariance matrix estimate being a regression coefficient covariance matrix estimate calculated after the working correlation matrix fitting the generalized estimating equation model; The generalized estimating equation model is fitted using the correlation structures respectively, quasi-likelihood criterion values under independent model assumptions of the correlation structures are calculated respectively, and a correlation structure with a minimum quasi-likelihood criterion value under the independent model assumption is selected as an optimal structure.

6. The GEE model and genetic profile based noise induced language frequency hearing impairment prediction system as claimed in claim 1, wherein, The method for calculating the individual risk probability by the risk prediction module includes: Obtaining variable information of a current time point of an individual to be predicted, calculating variable information of a future time point based on the variable information of the current time point, and calculating a risk probability by a Logistic function based on the variable information of the future time point and the regression coefficient estimate; According to the genetic risk score or the genotype, the individual is divided into different risk groups, and for each risk group, a relationship curve between cumulative noise exposure and incidence probability is drawn.

7. The GEE model and genetic profile based noise-induced language frequency hearing impairment prediction system of claim 6, wherein, The calculation of the risk probability by the Logistic function includes: The variable information of the current time point includes the genotype or the genetic risk score, the current cumulative noise exposure, the current age, and other covariates; The variable information of the future time point is calculated, and the variable information of the future time point includes a future age, a future cumulative noise exposure, a future length of service, and a linear predictor of the future time point; wherein the future age is equal to the current age plus a time interval; the future length of service is equal to the current length of service plus the time interval; the future cumulative noise exposure is equal to an equivalent continuous A-weighted sound level plus a logarithm value of the future length of service, wherein the logarithm value of the future length of service is a logarithm with a base of 10 multiplied by 10; The linear predictor at the future time point is equal to the intercept term estimate plus the product of the age main effect coefficient estimate and the future age, plus the product of the cumulative noise exposure main effect coefficient estimate and the future cumulative noise exposure, plus the product of the genotype main effect coefficient estimate and the genotype, plus the product of the first interaction effect coefficient estimate and the genotype and the future cumulative noise exposure, plus the product of the second interaction effect coefficient estimate and the genotype and the future age, and finally plus the inner product of the other covariate vector and its corresponding vector of regression coefficient estimates; The calculation method of the Logistic function is that the risk probability is equal to the exponential function of the linear predictor at the future time point divided by the linear predictor term at the future time point, the exponential function of the linear predictor at the future time point is the exponential function of the linear predictor at the future time point with the natural constant as the base, and the linear predictor term at the future time point is the exponential function of the linear predictor at the future time point plus 1.

8. The GEE model and genetic profile based noise-induced language frequency hearing impairment prediction system of claim 6, wherein, The relationship curve between the cumulative noise exposure and the probability of disease is drawn, including: The risk groups include a low genetic risk group, a medium genetic risk group and a high genetic risk group; the low genetic risk group includes individuals who do not carry risk alleles or individuals whose genetic risk scores are lower than a preset low risk score threshold, the medium genetic risk group includes individuals whose genetic risk scores are greater than or equal to the preset low risk score threshold and less than or equal to a preset high risk score threshold, and the high genetic risk group includes individuals whose genetic risk scores are greater than the preset high risk score threshold or individuals carrying high-risk homozygous mutations; The fixed age, genotype and other covariate vector are typical values, including the median, mode, minimum value and maximum value of the age, genotype and other covariate vector, and the cumulative noise exposure is changed within a reasonable range to calculate the probability of disease corresponding to different cumulative noise exposure values; the probability of disease is the exponential function of the linear predictor divided by the linear predictor term, the exponential function of the linear predictor is the exponential function of the linear predictor with the natural constant as the base, and the linear predictor term is the exponential function of the linear predictor plus 1.

9. The GEE model and genetic profile based noise induced language frequency hearing impairment prediction system as claimed in claim 1, wherein, The hearing protection recommendations for the individual include: The individual is divided into different risk levels based on the risk probability, and a corresponding management strategy is formulated according to the risk level; the risk levels include low risk, medium risk and high risk, the low risk is that the risk probability is less than a preset low risk probability threshold, the medium risk is that the risk probability is greater than or equal to the preset low risk probability threshold and less than a preset high risk probability threshold, and the high risk is that the risk probability is greater than or equal to the preset high risk probability threshold; An appropriate hearing protection device is recommended according to the risk group and the current cumulative noise exposure of the individual; a standard earplug is recommended for an individual with low genetic risk and cumulative noise exposure meeting a preset low exposure, a high-performance earplug or ear cover is recommended for an individual with medium genetic risk or cumulative noise exposure meeting a preset medium exposure, and an earplug plus ear cover combination is recommended for an individual with high genetic risk or cumulative noise exposure meeting a preset high exposure.

Citation Information

Patent Citations

  • Methods, devices, terminals and media for predicting noise-induced hearing loss and screening susceptible populations

    CN111584065B

  • Methods, kits, and applications for detecting polymorphisms in the 7q36.3 region associated with noise-induced hearing loss.

    CN111593108B

  • Early risk early warning system and method for high-noise exposed hearing impairment individual

    CN117153387A

  • Method for evaluating postoperative delirium possibility of surgical patient

    CN119678045A