Method and System for Assessing the Risk of Noise-Induced Hearing Loss
Through the auditory analysis model integrating genetic, environmental and physiological behavior data, extreme individuals are screened and high-risk sites are evaluated, and problems that genetic and environmental factors in the prior art are not considered are solved, achieving accurate risk assessment of noise auditory damage and personalized protection.
Patent Information
- Application Number
- CN202510237879.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-03-03
AI Technical Summary
The existing risk assessment methods for noise auditory injury fail to fully consider genetic factors and environmental exposure factors, which leads to one-sided evaluation results and it is difficult to accurately predict the risk of hearing injury in individuals.
By obtaining individual genetic data, environmental exposure data and physiological behavior data, the pre-constructed auditory analysis model was used to screen extreme individuals, conduct whole exon sequencing, identify suspected high-risk sites, and use PRS model to evaluate z-score values, divide damage risk levels and provide protective measures.
A comprehensive multi-dimensional auditory injury analysis is achieved, accurately assessing the causes of current injuries and predicting future risks, providing tailor-made protective suggestions, and improving the efficiency and effectiveness of auditory health management.
Smart Images

Figure CN119742067B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of auditory injury diagnosis, and more specifically, to a method and system for assessing the risk of noise-induced auditory injury. Background Art
[0002] Although traditional methods for assessing the risk of noise-induced auditory injury can reflect the individual's hearing damage, there are obvious limitations in aspects such as genetic susceptibility prediction and damage risk quantification.
[0003] In existing methods, for example, the Chinese patent application with the publication number CN117297596A discloses an auditory pathway assessment and analysis device and method, including: an SSAEP test signal generation component, a tinnitus simulation signal generation component, a composite signal generation component, a stimulation component, an electroencephalogram recording component, a data processing and analysis module, an amplitude determination module, and a comparison module. It is applicable to the assessment of the degree of auditory system damage in tinnitus subjects. By generating a simulated tinnitus signal that matches the tinnitus characteristics of the tinnitus subject, and associating and combining the simulated tinnitus signal with the SSAEP test signal to form a real and comprehensive composite excitation signal, thereby more accurately assessing the overall state of the auditory pathway. By generating a number of different SSAEP test signals, where the SSAEP test signal is a sine wave with different frequencies, amplitudes, or phases, different SSAEP responses can be detected in a short time. Although the above method reduces the test time and can more quickly and comprehensively evaluate the function of the auditory pathway, through research and application of the above method and the prior art, it is found that the above method and the prior art have at least the following partial defects:
[0004] Evaluating the state of the auditory pathway only based on steady-state auditory evoked potential (SSAEP) signals and simulated tinnitus signals can capture the physiological characteristics of tinnitus patients, but ignores important information such as genetic factors and environmental exposure factors that may cause or exacerbate auditory injury, thereby leading to one-sidedness of the evaluation results.
[0005] Therefore, the present invention provides a method and system for assessing the risk of noise-induced auditory injury. Summary of the Invention
[0006] In order to overcome the above-mentioned defects of the prior art, the present invention provides a method and system for assessing the risk of noise-induced auditory injury to solve the problems raised in the above background art.
[0007] To achieve the above object, the present invention provides the following technical solutions:
[0008] In the first aspect, the present invention provides a method for assessing the risk of noise-induced auditory injury, including:
[0009] Step 1: Obtain the auditory representation data of I sample individuals, where I is an integer greater than zero; the auditory representation data includes individual gene data, environmental exposure data, individual physiological data, and behavioral data;
[0010] Step 2: Input the auditory representation data of I sample individuals into a pre-constructed auditory analysis model to obtain a first screening result, where the first screening result includes Q extreme individuals;
[0011] Step 3: Obtain the Q extreme individuals output by the auditory analysis model, extract the q-th extreme individual for whole-exome sequencing to obtain a set of gene mutation sites; q = 1, 2, ……, Q;
[0012] Step 4: Perform functional annotation on the set of gene mutation sites to identify a set of suspected high-risk sites and the corresponding annotation results;
[0013] Step 5: Based on the PRS model, evaluate the set of suspected high-risk sites to obtain the z-score value of the q-th extreme individual; determine the damage risk level according to the z-score value;
[0014] Step 6: Feedback the corresponding protection measure plan to the client for display according to the damage risk level.
[0015] Furthermore, the specific training process of the auditory analysis model includes:
[0016] Pre-collect Y groups of auditory representation data, where Y is an integer greater than 1, set the corresponding screening results for the auditory representation data, mark the digital labels of the screening results as prediction labels, and convert the auditory representation data and the corresponding prediction labels into a corresponding set of feature vectors;
[0017] Use each set of feature vectors as the input of the auditory analysis model. The auditory analysis model takes a set of prediction labels corresponding to each group of auditory representation data as the output, and takes the actual prediction label corresponding to each group of auditory representation data as the prediction target. The actual prediction label is the prediction label preset corresponding to the auditory representation data; use minimizing the sum of prediction errors of all auditory representation data as the training target; train the auditory analysis model until the sum of prediction errors reaches convergence and then stop training; the auditory analysis model is a random forest, support vector machine, or deep neural network model to obtain the first screening result.
[0018] Furthermore, the method for extracting the q-th extreme individual for whole-exome sequencing to obtain a set of gene mutation sites includes:
[0019] Step a1: Extract and purify DNA from the blood sample of the q-th extreme individual using a specific reagent to obtain purified DNA;
[0020] The specific reagent is a lysis buffer, a chelating agent or a precipitating agent;
[0021] Step a2: mechanically or enzymatically decomposing the purified DNA to obtain DNA fragments, wherein the size of the DNA fragments is in the range of 100-500 bp;
[0022] Step a3: using a specific exon probe pool to capture DNA fragments in the exon region, obtaining exon sequence fragments, and removing non-exon sequences;
[0023] Step a4: connecting the adapter containing the sequencing adapter to both ends of the exon sequence fragment, and performing PCR amplification on the exon sequence fragment to obtain the exon sequence amplified fragment;
[0024] Step a5: sequencing the exon sequence amplified fragments by a sequencer connected to the sequencing platform to obtain sequencing data;
[0025] Step a6: Perform quality control on the sequencing data to obtain high-quality sequencing data, and compare the obtained high-quality sequencing data with the reference genome to determine the set of gene variation sites.
[0026] Furthermore, the method of performing quality control on the sequencing data to obtain high-quality sequencing data includes:
[0027] Step b1: using a first tool to evaluate the sequencing data, removing the sequencing data with bases below the standard quality, and obtaining first sequencing data;
[0028] Step b2: using a second tool to process the first sequencing data, remove low-quality bases at the ends of the read segments, and cut out adapter fragments to obtain second sequencing data;
[0029] Step b3: setting every F bases of the sequencer as a sliding window, performing sliding inspection on the second sequencing data, obtaining the average quality score of the bases in each sliding window, removing the window segments below the average quality score threshold, and obtaining the third sequencing data; F is an integer greater than 1;
[0030] Step b4: Use a third tool to mark the repetitive sequences generated by PCR amplification in the third sequencing data, and use a decontamination tool to clean up non-human sequences to obtain high-quality sequencing data.
[0031] Further, the method of comparing the obtained high-quality sequencing data with the reference genome to determine the set of gene variation sites includes:
[0032] Step c1: Use the BWA-MEM algorithm to compare high-quality sequencing data with the target genome to obtain a SAM file;
[0033] Step c2: Use the fifth tool to convert the SAM file into a compressed BAM format, sort it by genomic location to obtain a BAM file, and use the sixth tool to mark the duplicate sequences during the PCR process, correct the alignment offsets, and filter out unmatched or low-quality alignment fragments;
[0034] Step c3: Use the seventh tool to perform genomic recalibration on the BAM file and extract genomic variant information to generate a preliminary VCF file;
[0035] Step c4: Use the VQSR module of the seventh tool to calibrate the detected variant sites, set the variant quality threshold to screen for a set of matching gene variant sites.
[0036] Furthermore, the method for functionally annotating the set of gene variant sites to identify the set of suspected high-risk sites and the corresponding annotation results includes:
[0037] Step d1: Based on the known SNP site data, perform a preliminary screening on the set of gene variant sites to obtain a screened set of gene variant sites; the screened set of gene variant sites includes Group E gene variant sites;
[0038] Step d2: Perform an association analysis on the Group e gene variant sites, calculate the correlation coefficient of the Group e gene variant sites, where e = 1, 2,..., E;
[0039] Step d3: Preset a correlation coefficient threshold, the correlation coefficient threshold includes and where ; Compare the correlation coefficient of the Group e gene variant sites with the preset correlation coefficient threshold;
[0040] If , then mark the Group e gene variant sites as suspected high-risk sites;
[0041] If , then mark the Group e gene variant sites as medium-risk sites;
[0042] If , then mark the Group e gene variant sites as low-risk sites;
[0043] Step d4: Use a functional annotation tool to annotate the suspected high-risk sites to obtain annotation results, let e = e + 1, and return to Step d2, the functional annotation tool is Annovar or VEP; the annotation results include gene location and functional region: the gene location includes coding region, promoter or non-coding region;
[0044] Step d5: Repeat steps d2 - d4 until the loop ends when e = E, and obtain a set of suspected high-risk loci and corresponding annotation results.
[0045] Further, the loci of the PRS model include rs10783780, rs10794566, rs10794567, rs3887954, rs58011943, rs11836060, rs16973424, and rs74905175.
[0046] Further, based on the PRS model, the method for obtaining the z-score value of the q-th extreme individual by evaluating the set of suspected high-risk loci includes:
[0047] Step s1: Extract the effect value of the h-th suspected locus in the set of suspected high-risk loci based on the correlation coefficient; h = 1, 2,..., H; H is the number of suspected loci in the set of suspected high-risk loci;
[0048] Step s2: Calculate the PRS score of the q-th extreme individual;
[0049] Step s3: Repeat step s2 until the loop ends when q = Q, and obtain the PRS scores of Q extreme individuals;
[0050] Step s4: Convert the PRS score of the q-th extreme individual into a z-score value to obtain the z-score value of the q-th extreme individual.
[0051] Further, the damage risk levels include the first level, the second level, the third level, the fourth level, and the fifth level; the first level > the second level > the third level > the fourth level > the fifth level.
[0052] Further, the individual gene data includes KCNQ4, MYO6, MYO7A, MYO15A, OTOF, GJB2, SLC26A4, GRM7, TECTA, CDH23, ESRRB, HSP70, CAT, and SOD2; the environmental exposure data includes noise exposure intensity and noise exposure time; the individual physiological data includes age, gender, genetic background, health status, and history of hearing impairment; the behavioral data includes high-frequency sound sensitivity, tinnitus symptoms, and tolerance to noise.
[0053] In the second aspect, a noise-induced hearing loss risk assessment system includes:
[0054] A data acquisition module for obtaining auditory characterization data of I sample individuals, where I is an integer greater than zero; the auditory characterization data includes individual gene data, environmental exposure data, individual physiological data, and behavioral data;
[0055] An analysis module for inputting the auditory representation data of I sample individuals into a pre-constructed auditory analysis model to obtain a first screening result, where the first screening result includes Q extreme individuals;
[0056] A sequencing module for obtaining the Q extreme individuals output by the auditory analysis model, extracting the q-th extreme individual for whole exome sequencing to obtain a set of gene variation sites; q = 1, 2, ……, Q;
[0057] An identification module for performing functional annotation on the set of gene variation sites to identify a set of suspected high-risk sites and corresponding annotation results;
[0058] An evaluation module for evaluating the set of suspected high-risk sites based on the PRS model to obtain the z-score value of the q-th extreme individual; determining the damage risk level according to the z-score value;
[0059] A display module for feeding back the corresponding protection measure plan to the client for display according to the damage risk level.
[0060] The technical effects and advantages of the present invention:
[0061] 1. The present invention constructs auditory representation data by using gene data, environmental exposure data, individual physiological data and behavioral data, realizes multi-dimensional comprehensive analysis of auditory damage, reveals the multi-factor causes and interactions of auditory damage, so as to achieve accurate analysis; analyzes the gene variation sites known to be related to auditory function by using the gene data of extreme individuals, standardizes according to the PRS score (such as calculating the Z-score value), and corresponds it to the auditory damage risk level, divides into multiple evaluation levels, and formulates more strict auditory protection measures for high-risk individuals.
[0062] 2. The comprehensive method based on multi-dimensional data, accurate evaluation model and big data analysis can not only carefully analyze the causes of current auditory damage, but also predict future risks and provide customized protection suggestions for individuals, thereby improving the efficiency and effect of auditory health management at both the prevention and treatment levels. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] Figure 1 It is a flowchart of the method for evaluating the risk of noise-induced auditory damage in Example 1;
[0064] Figure 2 It is a flowchart of the method for identifying the set of suspected high-risk sites and corresponding annotation results in Example 1;
[0065] Figure 3 It is a flowchart of the method for obtaining the z-score value of the q-th extreme individual in Example 1;
[0066] Figure 4 Individual characteristic map of extremely susceptible and extremely resistant individuals in Example 1;
[0067] Figure 5 Basic information table in the verification stage of Example 1;
[0068] Figure 6 Manhattan plot of association analysis for noise-induced hearing loss risk assessment;
[0069] Figure 7 Q-Q plot of the results of whole exome association analysis;
[0070] Figure 8 PRS distribution plots for the noise-induced hearing loss group and the non-noise-induced hearing loss group;
[0071] Figure 9 Forest plot of PRS and noise-induced hearing loss risk;
[0072] Figure 10 Schematic structural diagram of the noise-induced hearing loss risk assessment system in Example 2. Detailed implementation manners
[0073] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0074] In addition, the accompanying drawings are only schematic diagrams of the present invention and are not necessarily drawn to scale. The same reference numerals in the drawings represent the same or similar parts, and thus their repeated description will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. The functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor methods and / or microcontroller methods.
[0075] It should be understood that although terms such as "first" and "second" may be used here to describe various units, these units should not be limited by these terms. These terms are only used to distinguish one unit from another. For example, without departing from the scope of the exemplary embodiments, the first unit can be called the second unit, and similarly the second unit can be called the first unit. The term "and / or" used here includes any and all combinations of one or more of the listed associated items.
[0076] Example 1
[0077] Please refer to Figure 1 As shown, this embodiment publicly provides a method for assessing the risk of noise-induced hearing loss, which is implemented based on a client and a cloud server. The client is remotely communicatively connected to the cloud server. The method includes:
[0078] Step 1: Obtain the auditory characterization data of I sample individuals, where I is an integer greater than zero;
[0079] The auditory characterization data includes individual gene data, environmental exposure data, individual physiological data, and behavioral data;
[0080] It should be understood that the sample individuals include, but are not limited to, individuals such as factory workers, construction workers, musicians, military personnel, and music lovers who like high-decibel music and are in high-noise environments for a long time.
[0081] The individual gene data includes, but is not limited to, KCNQ4, MYO6, MYO7A, MYO15A, OTOF, GJB2, SLC26A4, GRM7, TECTA, CDH23, ESRRB, HSP70, CAT, and SOD2.
[0082] It should be explained that: KCNQ4, MYO6, MYO7A, MYO15A, and OTOF are genes related to the structure and function of the cochlea; GJB2, SLC26A4, and GRM7 are genes related to neurotransmitter transmission; TECTA, CDH23, and ESRRB are genes related to auditory conduction and inner ear protection; HSP70, CAT, and SOD2 are genes for inner ear protection and repair. The individual gene data is obtained by collecting the blood or saliva of the sample individual to extract DNA and using high-throughput sequencing technology for whole-exome sequencing. The high-throughput sequencing technology is an existing technology, and any method for obtaining individual gene data can be used, and no specific limitation is made here.
[0083] The environmental exposure data includes noise exposure intensity, noise exposure time, and noise frequency; among them, the noise exposure intensity is the volume of the noise; the noise exposure time refers to the cumulative duration of the sample individual exposed in the noise environment; the noise frequency is the pitch of the noise; the noise exposure intensity, noise exposure time, and noise frequency are all collected by the sample individual carrying a noise monitoring device or a smart watch.
[0084] The individual physiological data includes age, gender, genetic background, health status, and history of hearing impairment; the individual physiological data is collected in advance through questionnaires and other means and pre-stored in the system database. Specifically, as age increases, the inner ear hair cells and auditory nerves gradually degenerate, and the risk of noise-induced hearing impairment also increases; due to their work nature (such as heavy industry, construction, military, etc.), men are more likely to be exposed to high-noise environments, so the incidence of noise-induced hearing impairment is higher than that of women; genetic genes play a greater role in noise-induced hearing impairment, and some gene mutations may make individuals more sensitive to noise. For example, genes such as GJB2 and GRM7 are related to the cell protection mechanism, signal transmission, and cell repair in the cochlea, and mutations may increase the risk of noise-induced hearing impairment. Health problems, such as hypertension, diabetes, and immune system diseases, will affect the blood circulation and cell health of the cochlea, thus increasing the risk of noise-induced hearing impairment. These chronic diseases weaken the resistance of inner ear cells to noise and make them more vulnerable to damage. People with a history of hearing impairment (such as hearing impairment caused by occupational noise exposure, explosions, or ear diseases) are more likely to have their damage aggravated due to repeated noise exposure. A history of auditory impairment causes damage to the hair cells and nerve connections in the cochlea, making the individual more sensitive to noise exposure.
[0085] The behavioral data includes high-frequency sound sensitivity, tinnitus symptoms, and noise tolerance. The high-frequency sound sensitivity refers to the auditory response of an individual to high-frequency sounds (such as above 3000 Hz). Individuals with higher high-frequency sensitivity can usually detect the signs of noise damage faster in the early stage because the hearing in the high-frequency part is often damaged first. Tinnitus symptoms are a common symptom of noise-induced hearing impairment, manifested as a continuous "buzzing" or "hissing" sound in the ear, and this phenomenon may even occur without external noise. Tinnitus is usually the result of damage to the inner ear hair cells or auditory nerves, indicating that the auditory system has been damaged; the noise tolerance refers to the ability of an individual to remain comfortable in a high-intensity or harsh noise environment. Low tolerance usually means that the individual is more sensitive to noise and feels discomfort or even pain at a lower noise intensity. The high-frequency sound sensitivity, tinnitus symptoms, and noise tolerance are obtained through hearing tests, questionnaires, and noise tolerance tests, etc.
[0086] It should be noted that: the auditory characterization data is pre-collected by those skilled in the art, and after preprocessing the data of I sample individuals, it is input into the system database through the client. The preprocessing methods are data cleaning, standardization, or normalization. The preprocessing methods are prior arts and will not be elaborated here.
[0087] Step 2: Input the auditory characterization data of I sample individuals into a pre-constructed auditory analysis model to obtain a first screening result, and the first screening result includes Q extreme individuals.
[0088] It should be noted that the extreme individuals refer to those who have severe hearing impairment under low noise exposure and no obvious impairment under noise exposure.
[0089] In implementation, the specific training process of the auditory analysis model includes:
[0090] Pre-collect Y sets of auditory representation data, where Y is an integer greater than 1. Set corresponding screening results for the auditory representation data, and collect Y sets of different auditory representation data. Those skilled in the art set corresponding screening results for the Y sets of different auditory representation data in sequence. Please refer to Figure 8 as shown.
[0091] Mark the digital labels of the screening results as prediction labels, and convert the auditory representation data and the corresponding prediction labels into a corresponding set of feature vectors;
[0092] Take each set of feature vectors as the input of the auditory analysis model. The auditory analysis model takes a set of prediction labels corresponding to each set of auditory representation data as the output, and takes the actual prediction label corresponding to each set of auditory representation data as the prediction target. The actual prediction label is the prediction label preset corresponding to the auditory representation data; take minimizing the sum of the prediction errors of all auditory representation data as the training target; where the calculation formula of the prediction error is , where is the prediction error, is the group number of the feature vector corresponding to the auditory representation data, is the prediction label corresponding to the th group of auditory representation data,
[0093] The auditory analysis model is a random forest, support vector machine or deep neural network model, which includes an input layer, a hidden layer and an output layer; each hidden layer includes multiple neurons, and there are connections between each neuron and the neurons in the next layer. The connections contain weights, which determine the importance and influence of data transmission in the neural network; an activation function is applied to each neuron between the hidden layer and the output layer. The activation function introduces non-linearity, allowing the network to learn more complex patterns and features.
[0094] This step can improve the reliability of the screening results of extreme susceptible and resistant individuals through the screening of individuals with multi-dimensional auditory representation data. The analysis of multi-level data reduces the deviation caused by misjudgment of a single factor, making the screening results more accurate, and further optimizing the reliability of the risk assessment of noise-induced hearing loss. Please refer to Figure 9Forest plot of the PRS shown and the risk of noise-induced hearing loss.
[0095] Step 3: Obtain Q extreme individuals output by the auditory analysis model, extract the q-th extreme individual for whole exome sequencing, and obtain a set of gene mutation sites; q = 1, 2, ……, Q;
[0096] In a specific embodiment, the method for extracting the q-th extreme individual for whole exome sequencing to obtain a set of gene mutation sites includes:
[0097] Step a1: Extract and purify DNA from the blood sample of the q-th extreme individual using a specific reagent to obtain purified DNA;
[0098] It should be noted that: the blood sample is obtained by a sequencing institution collecting the saliva of the extreme individual and performing pretreatment. The specific reagent is a lysis buffer, chelating agent, or precipitating agent; the method for purifying DNA includes organic solvent extraction, silica membrane column method, magnetic bead method, or salting-out method. The pretreatment method is quality detection of saliva to ensure sufficient DNA and no contamination. Any existing technology for purifying DNA can be used, and no further description is given here.
[0099] Step a2: Decompose the purified DNA by mechanical or enzymatic digestion to obtain DNA fragments, and the size of the DNA fragments is in the range of 100 - 500 bp;
[0100] Step a3: Use a specific exon probe pool to capture the DNA fragments in the exon region, obtain exon sequence fragments, and remove non-exon sequences;
[0101] The purpose is: the uncaptured non-exon sequences are removed to ensure that the sequencing data is mainly concentrated in the exon region, improving efficiency and accuracy;
[0102] Step a4: Connect adaptors containing sequencing linkers to both ends of the exon sequence fragments, and perform PCR amplification on the exon sequence fragments to obtain exon sequence amplification fragments;
[0103] Step a5: Sequence the exon sequence amplification fragments through a sequencer connected to the sequencing platform to obtain sequencing data;
[0104] It should be noted that: the sequencing platform is Illumina or BGI. The sequencer reads the amplified fragments of exon sequences using optical or electrochemical signals and generates sequencing data, and outputs the sequencing data. The sequencing data is stored in FASTQ format, which contains sequences and their corresponding quality information. The quality information is encoded in ASCII characters. Each character corresponds to a Q value, which is obtained by subtracting a specific offset (usually 33 or 64) from the ASCII value. Each character position corresponds to the quality value of the base at that position in the sequence.
[0105] Step a6: Perform quality control on the sequencing data to obtain high-quality sequencing data, and align the obtained high-quality sequencing data with the reference genome to determine the set of gene variation sites. The set of gene variation sites includes SNP, insertion, and deletion variation sites in the exon region, and the reference gene is pre-stored in the system database.
[0106] In a specific embodiment, the method for performing quality control on the sequencing data to obtain high-quality sequencing data includes:
[0107] Step b1: Use the first tool to evaluate the sequencing data, remove the sequencing data with bases below the standard quality, and obtain the first sequencing data.
[0108] It should be noted that: the first tool is the FastQC tool. FastQC is used to analyze the basic quality distribution, GC content, read length consistency, and any possible contamination and sequencing errors of the data. The standard quality base is Q30, which is specifically set by those skilled in the art and will not be specifically limited here.
[0109] Step b2: Use the second tool to process the first sequencing data, remove the low-quality bases at the ends of the reads, and cut out the adapter fragments to obtain the second sequencing data.
[0110] It should be noted that: the second tool is the Trimmomatic or Cutadapt tool. The low-quality bases at the ends of the fragments are bases below the preset quality threshold; the preset quality threshold is set by those skilled in the art according to experience and will not be limited here; the adapter fragment refers to the short sequence added to the ends of DNA fragments before sequencing, which will cause data errors and analysis biases if not removed, so adapter removal is performed when processing the sequencing data.
[0111] Step b3: Set every F bases of the sequencer as a sliding window, perform a sliding check on the second sequencing data to obtain the average quality score of the bases within each sliding window, remove the window segments with an average quality score lower than the threshold, and obtain the third sequencing data; F is an integer greater than 1; to ensure that each segment of the third sequencing data has high-quality bases.
[0112] Among them, the calculation method of the average quality score of bases within each sliding window includes:
[0113] ;
[0114] In the formula, represents the average quality score of bases within the sliding window, is the Phred quality score of the f-th base within the sliding window, represents the total number of bases in the sliding window, and P represents the probability of incorrect determination of bases.
[0115] It should be noted that: the Phred quality score represents an index of the accuracy of base calling in sequencing data, and is commonly used in the results of high-throughput sequencing. The higher the value of the Phred quality score, the higher the accuracy of the determination of that base.
[0116] Exemplarily:
[0117] = 10: The error probability is 0.1, that is, there is a 10% probability of sequencing error.
[0118] = 20: The error probability is 0.01, that is, there is a 1% probability of sequencing error.
[0119] = 30: The error probability is 0.001, that is, there is a 0.1% probability of sequencing error.
[0120] = 40: The error probability is 0.0001, that is, there is a 0.01% probability of sequencing error.
[0121] Therefore, for every 10-point increase in the Phred quality score, it means that the error probability is reduced by 10 times.
[0122] Step b4: Use a third tool to mark the repetitive sequences generated by PCR amplification in the third sequencing data, and use a decontamination tool to clean up non-human source sequences to obtain high-quality sequencing data.
[0123] It should be noted that: The third tool is the Picard tool. By comparing the sequence and position information of each read segment, duplicate read segments are marked to facilitate removal during subsequent analysis. By marking duplicate sequences, the bias introduced during the amplification process can be effectively reduced, and the accuracy of the sequencing data can be improved. The decontamination tool is the DecontaMiner tool. Using the DecontaMiner tool, the sequencing data is compared with a standard reference database (such as RefSeq) to detect and remove non-human source sequences, thereby ensuring the high quality and pollution-free of the sequencing data. Non-human source sequences refer to DNA sequences detected in the sequencing data that do not match the human genome, and these sequences are pre-stored in the standard reference database.
[0124] In a specific embodiment, the method for comparing the obtained high-quality sequencing data with a reference genome to determine a set of gene mutation sites includes:
[0125] Step c1: Use the BWA-MEM algorithm to compare the high-quality sequencing data with the target genome to obtain a SAM file.
[0126] It should be noted that: The BWA-MEM algorithm can efficiently process short read lengths and some medium read lengths of sequencing data to ensure high-precision alignment.
[0127] Step c2: Use the fifth tool to convert the SAM file into a compressed BAM format, sort it by genomic position to obtain a BAM file, and use the sixth tool to mark the duplicate sequences in the PCR process, correct the alignment offset, and filter out mismatched or low-quality alignment fragments; to ensure the accuracy of the data.
[0128] The fifth tool is the Samtools tool, and the sixth tool is the Picard tool.
[0129] Step c3: Use the seventh tool to perform genomic recalibration on the BAM file and extract genomic variation information to generate a preliminary VCF file; the seventh tool is the GATK tool.
[0130] Step c4: Use the VQSR module of the seventh tool to calibrate the detected mutation sites, set a mutation quality threshold to screen for a matching set of gene mutation sites.
[0131] It should be noted that the setting of the mutation quality threshold is set by those skilled in the art according to experience and will not be specifically limited here.
[0132] Through step-by-step filtering, calibration, and optimization processes, the final obtained set of variant sites has high reliability, reducing the interference of false positives and noise. These sets of gene variant sites with high confidence are more biologically interpretable and help identify potential functional gene variants and disease-related mutations.
[0133] Step 4: Perform functional annotation on the set of gene variant sites to identify the set of suspected high-risk sites and the corresponding annotation results;
[0134] Please refer to Figure 2 As shown, in a preferred embodiment, the method for performing functional annotation on the set of gene variant sites to identify the set of suspected high-risk sites and the corresponding annotation results includes:
[0135] Step d1: Based on the known SNP site data, perform a preliminary screening on the set of gene variant sites to obtain the screened set of gene variant sites; the screened set of gene variant sites includes Group E gene variant sites;
[0136] It should be noted that: the known SNP site data includes SNP sites known to be related to auditory function, cochlear function, nerve conduction, etc. The method of preliminary screening includes database support, functional enrichment, literature basis, and batch tool screening to effectively identify gene variant sites related to hearing.
[0137] Please refer to Figure 6 、 Figure 7 As shown, step d2: Perform association analysis on the Group e gene variant sites, and calculate the correlation coefficient of the Group e gene variant sites, where e = 1, 2,..., E;
[0138] Among them, the method for obtaining the correlation coefficient of the Group e gene variant sites includes:
[0139] ;
[0140] In the formula, represents the correlation coefficient of the Group e gene variant sites, represents the p-value, represents the effect value, represents the sample size, 、 and are preset weights, represents the logarithmic function with base e.
[0141] In the formula, the preset weights are obtained by those skilled in the art by collecting multiple sets of comprehensive parameters, setting corresponding weights for each set of comprehensive parameters, substituting the preset weights and the collected comprehensive parameters into the formula, any four formulas form a system of ternary linear equations, screening the calculated weights and taking the average value to obtain , and value.
[0142] It should be noted that: the lower the p-value, the more significant the correlation; the larger the effect value, the stronger the impact of the variant site on the risk of auditory damage; the larger the sample size, the higher the statistical credibility of the association result; the p-value, effect value and sample size are all obtained using the WES analysis software.
[0143] Step d3: Preset the correlation coefficient threshold, and the correlation coefficient threshold includes and , where ; compare the correlation coefficient of the e-th group of gene variant sites with the preset correlation coefficient threshold;
[0144] If , then mark the e-th group of gene variant sites as suspected high-risk sites, indicating SNPs that are repeatedly marked as significantly related to the risk of auditory damage in WES or annotation tools;
[0145] If , then mark the e-th group of gene variant sites as medium-risk sites, indicating that the e-th group of gene variant sites has a relatively small correlation in the WES database, with low statistical significance or effect;
[0146] If , then mark the e-th group of gene variant sites as low-risk sites, indicating no significant association or not being enriched in the auditory-related pathway.
[0147] Step d4: Use a functional annotation tool to annotate the suspected high-risk sites to obtain an annotation result, let e = e + 1, and return to step d2. The functional annotation tool is Annovar or VEP;
[0148] It should be noted that: the annotation result includes gene location and functional region: the gene location includes coding region, promoter or non-coding region.
[0149] Step d5: Repeat step d2 - step d4 until e = E to end the loop, and obtain a set of suspected high-risk sites and the corresponding annotation results.
[0150] It should be noted that: The method of identifying the suspected high-risk locus set and the corresponding annotation results can also be through literature support, WES research, or functional enrichment analysis, etc. Specifically, the literature support is to consult the published genetic studies on noise-induced hearing loss and identify the SNP loci that have been reported to be related to hearing loss. For example, specific loci of certain genes related to cochlear structure or nerve signal conduction have been mentioned many times in the literature and are significantly associated with the risk of hearing loss; capture and enrich the exon part of the genome through specific technical means, and then perform high-throughput sequencing to obtain the base sequence information of the exon region. Compared with whole-genome sequencing, its coverage range is relatively more targeted, focusing on the exon region that accounts for about 1% - 2% of the human genome but has important functions, and can more efficiently detect variations related to diseases or with important biological significance; the functional enrichment analysis is to perform functional enrichment analysis on the annotated SNP loci to determine whether they are concentrated in the known functional pathways related to the auditory system (such as cochlear development, auditory conduction pathway), so as to infer their potential association with hearing loss.
[0151] As Figure 6 shown, in the Manhattan plot of the association analysis for noise-induced hearing loss risk assessment:
[0152] The red line represents the genome-wide significance level threshold. When the -log10(p) value of a certain locus is higher than this line, it means that the association of this locus with noise-induced hearing loss has strong statistical significance, that is, this locus is very likely to be a genetic locus truly related to noise-induced hearing loss.
[0153] The blue line represents the suggestive association threshold. If the -log10(p) value of a locus exceeds the blue line but is lower than the red line, it indicates that there may be an association between this locus and noise-induced hearing loss, but its significance level has not reached the genome-wide significance level and further research and verification are needed. These thresholds help those skilled in the art quickly identify which loci on the chromosomes are meaningfully associated with the studied trait (here it is noise-induced hearing loss).
[0154] As Figure 7 shown by the Q-Q plot of the whole exome association analysis results, it can be seen that:
[0155] The observed values on the vertical axis represent the -log10-transformed values of the p-values of the actually observed association analysis, and the expected values on the horizontal axis represent the expected -log10(p) values under the null hypothesis (that is, there is no true genetic association and the observed association is caused by random factors); the Q-Q plot is used to test whether the results of the association analysis conform to the null hypothesis. If the data points are closely clustered around Figure 7On the red reference line (representing the expected distribution under the null hypothesis), it shows that the observed results are consistent with the expected results under the null hypothesis, that is, no significant genetic loci related to the studied traits are found.
[0156] Step 5: Based on the PRS model, evaluate the set of suspected high-risk loci to obtain the z-score value of the q-th extreme individual; determine the damage risk level according to the z-score value.
[0157] It should be noted that: The PRS model is a risk score constructed based on multiple genetic loci, indicating an individual's genetic susceptibility to a certain disease or specific phenotype.
[0158] Among them, the loci of the PRS model include rs10783780 (located at position 56704152 on human chromosome 12, base is A / G), rs10794566 (located at position 125521533 on human chromosome 10, base is G / A), rs10794567 (located at position 125521590 on human chromosome 10, base is G / A), rs3887954 (located at position 53186088 on human chromosome 12, base is G / C), rs58011943 (located at position 18684527 on human chromosome 17, base is G / C), rs11836060 (located at position 80750659 on human chromosome 12, base is G / A), rs16973424 (located at position 66904001 on human chromosome 17, base is A / C), rs74905175 (located at position 31815326 on human chromosome 7, base is C / T).
[0159] It should be noted that the construction of PRS is based on 7 SNP loci with p < 0.001 and consistent effect directions in the validation stage, and PRS is obtained by weighted calculation of each locus. The validation method is as follows:
[0160] The method includes two stages. The discovery stage includes 150 extremely susceptible and 150 extremely resistant individuals, and the validation stage includes 2028 non-noisy hearing loss workers and 3836 noisy hearing loss workers. In addition, this application collects basic information including the gender, age, etc. of the research subjects, please refer to Figure 4 and Figure 5 as shown. Figure 5 in which, n is the number of validation groups.
[0161] Among them, continuous variables are expressed as median (interquartile range), and categorical variables are expressed as number of people (percentage), and CNE is the cumulative noise exposure.
[0162] Whole-exome sequencing was performed on the Illumina HiSeq X10 platform. In the constructed DNA fragment library, each inserted fragment was sequenced at both ends, with 150 bp sequenced for each end. The average sequencing coverage depth was set at 50X.
[0163] After sequencing, information analysis was performed on the original sequences. First, quality control assessment of the data was required to determine whether it met the standards; if it met the standards, variant detection of the samples was performed, including SNPs and InDels, and annotation; if it did not meet the standards, additional sequencing or library reconstruction was required according to the actual situation.
[0164] Please refer to Figure 3 As shown, specifically, based on the PRS model, the method for evaluating the q-th extreme individual to obtain the z-score value of the suspected high-risk locus set includes:
[0165] Step s1: Extract the effect value of the h-th suspected locus in the suspected high-risk locus set based on the correlation coefficient; h = 1, 2,..., H;
[0166] It should be noted that: the suspected high-risk locus set includes H suspected loci, and the acquisition method of the correlation coefficient has been described above and will not be repeated here. Additionally, the magnitude of the effect value of each suspected locus is related to the risk increase multiple of the auditory impairment associated with that locus. Exemplarily, if the effect value of the suspected locus is 1.5, it means that for each increase in the allele of the suspected locus, the risk increases by 1.5 times.
[0167] Step s2: Calculate the PRS score of the q-th extreme individual, and its calculation formula is:
[0168] ;
[0169] In the formula, represents the PRS score of the q-th extreme individual, is the effect value of the h-th suspected locus, represents the genotype of the h-th suspected locus.
[0170] It should be noted that: the genotypes are coded as 0, 1, 2 (usually 0 is the non-risk allele, 1 is one risk allele, and 2 is two risk alleles).
[0171] Step s3: Repeat step s2 until the loop ends when q = Q, and obtain the PRS scores of Q extreme individuals;
[0172] Step s4: Convert the PRS score of the q-th extreme individual into a z-score value to obtain the z-score value of the q-th extreme individual;
[0173] ;
[0174] ;
[0175] In the formula, represents the z-score value of the q-th extreme individual, represents the average PRS score of Q extreme individuals, represents the standard deviation of the PRS score.
[0176] It should be noted that: compared with the complex score of the original PRS, the z-score value is more intuitive. It is difficult for those not skilled in the art to understand the meaning of the original score, but the z-score shows the risk in terms of the degree of deviation from the mean. A positive number represents a higher-than-average risk, and a negative number vice versa, which is convenient for individuals to quickly know the relative risk status of extreme individuals.
[0177] The injury risk levels include the first level, the second level, the third level, the fourth level, and the fifth level; the first level > the second level > the third level > the fourth level > the fifth level.
[0178] In implementation, based on the correspondence between the z-score value and the injury risk level, the method for determining the injury risk level of the Q-th extreme individual includes:
[0179] As shown in Table 1, set up an injury risk level table, and obtain the corresponding injury risk level according to the z-score value; the z-score value is represented by ZS.
[0180] Table 1 Injury Risk Level Table
[0181]
[0182] It should be noted that K is a preset threshold. K is collected by those skilled in the art for the corresponding z-score values during the historical auditory injury risk assessment process, calculate the mean value of the multiple collected z-score values, and use the mean value as the preset threshold.
[0183] Correspond the z-score value of the q-th extreme individual to the injury risk level table to obtain the corresponding injury risk level.
[0184] Step 6: Feed back the corresponding protection measure plan to the client for display according to the injury risk level.
[0185] In implementation, the corresponding protection measure plan includes:
[0186] Plan 1: The recommended measures for the first level include:
[0187] Completely avoid high-noise exposure (such as avoiding staying in places like concerts and high-noise factories for a long time).
[0188] Customize exclusive protective equipment, such as special noise-canceling earplugs or customized eardrum protectors.
[0189] Conduct an auditory health check every six months to monitor for signs of hearing loss or early noise damage.
[0190] If there is already a risk of damage, it is recommended to consult an otolaryngologist or audiology expert in advance.
[0191] Option 2: The recommended measures at the second level include:
[0192] Strictly follow the noise protection regulations in the workplace or in daily life.
[0193] Equip with high-quality personal auditory protection devices and ensure proper fit and regular replacement.
[0194] Avoid exposure to high-intensity noise environments (such as above 100 dB).
[0195] Strengthen health monitoring and establish a personalized auditory health record.
[0196] Option 3: The recommended measures at the third level include:
[0197] Always use protective equipment (such as earmuffs or high-efficiency noise-canceling earplugs) when working or moving in a noisy environment.
[0198] Limit the noise exposure time (such as no more than 2 hours each time).
[0199] Arrange regular auditory assessments (such as once a year).
[0200] Option 4: The recommended measures at the fourth level include:
[0201] Avoid long-term exposure to moderate noise (such as the regular machine noise in an industrial environment).
[0202] Properly use noise-canceling earplugs or earmuffs.
[0203] Increase the education on auditory health care knowledge.
[0204] Option 5: The recommended measures at the fifth level include:
[0205] No special protection is required, but healthy auditory protection habits should be developed (such as avoiding long-term exposure to noisy environments).
[0206] Regular hearing health checks (such as once every two years).
[0207] In this embodiment, auditory representation data is constructed using genetic data, environmental exposure data, individual physiological data, and behavioral data to achieve comprehensive multi-dimensional analysis of auditory impairment, reveal the multi-factor causes and interactions of auditory impairment, and thus achieve precise analysis. By using the genetic data of extreme individuals, gene mutation sites known to be related to auditory function are analyzed, standardized according to the PRS score (such as calculating the Z-score value), and corresponding to the auditory impairment risk level, multiple assessment levels are divided, and for high-risk individuals, more stringent auditory protection measures are formulated.
[0208] Based on the comprehensive method of multi-dimensional data, precise assessment model, and big data analysis, this embodiment can not only meticulously analyze the causes of current auditory impairment, but also predict future risks and provide customized protection suggestions for individuals, thereby improving the efficiency and effectiveness of auditory health management at both the prevention and treatment levels.
[0209] Embodiment 2
[0210] Please refer to Figure 10 As shown, this embodiment provides a noise-induced auditory impairment risk assessment system, including a data acquisition module, an analysis module, a sequencing module, an identification module, an assessment module, and a display module; each module is connected by wired and / or wireless means to achieve data transmission between modules; the system includes:
[0211] A data acquisition module, used to obtain auditory representation data of I sample individuals, where I is an integer greater than zero; the auditory representation data includes individual gene data, environmental exposure data, individual physiological data, and behavioral data;
[0212] An analysis module, used to input the auditory representation data of I sample individuals into a pre-constructed auditory analysis model to obtain a first screening result, where the first screening result includes Q extreme individuals;
[0213] A sequencing module, used to obtain the Q extreme individuals output by the auditory analysis model, extract the qth extreme individual for whole exome sequencing to obtain a set of gene mutation sites; q = 1, 2,..., Q;
[0214] An identification module, used to perform functional annotation on the set of gene mutation sites to identify a set of suspected high-risk sites and corresponding annotation results;
[0215] An assessment module, based on the PRS model, assesses the set of suspected high-risk sites to obtain the z-score value of the qth extreme individual; determines the impairment risk level according to the z-score value;
[0216] A display module, used to feedback the corresponding protection measure plan to the client for display according to the impairment risk level.
[0217] The formulas involved above are all calculated by removing the dimension and taking their numerical values. It is a formula obtained by collecting a large amount of data for software simulation to be closest to the actual situation. The weight factors in the formula and each preset threshold in the analysis process are set by those skilled in the art according to the actual situation or obtained through simulation of a large amount of data; the magnitude of the weight factor is a specific value obtained by quantifying each parameter for subsequent comparison. Regarding the magnitude of the weight factor, it depends on the amount of sample data and the processing coefficients initially set by those skilled in the art for each group of sample data; as long as the proportional relationship between the parameters and the quantified values is not affected.
[0218] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed in the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claimed rights.
[0219] Finally: The above is only the preferred embodiment of the present invention and is not used to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for assessing the risk of noise-induced hearing loss, characterized in that, Including: Step 1: Obtain the auditory representation data of I sample individuals, where I is an integer greater than zero; the auditory representation data includes individual gene data, environmental exposure data, individual physiological data, and behavioral data; Step 2: Input the auditory representation data of I sample individuals into a pre-constructed auditory analysis model to obtain a first screening result, where the first screening result includes Q extreme individuals; Step 3: Obtain the Q extreme individuals output by the auditory analysis model, extract the q-th extreme individual for whole-exome sequencing to obtain a set of gene mutation sites; q = 1, 2, ……, Q; The method for extracting the q-th extreme individual for whole-exome sequencing to obtain a set of gene mutation sites includes: Step a1: Extract and purify DNA from the blood sample of the q-th extreme individual using a specific reagent to obtain purified DNA; Step a2: Generate sequencing data based on the purified DNA, perform quality control on the sequencing data to obtain high-quality sequencing data, and align the obtained high-quality sequencing data with the reference genome to determine the set of gene mutation sites; The method for performing quality control on the sequencing data to obtain high-quality sequencing data includes: Step b1: Use a first tool to evaluate the sequencing data, remove the sequencing data with bases below the standard quality, and obtain first sequencing data, where the first tool is the FastQC tool; Step b2: Use a second tool to process the first sequencing data, remove the low-quality bases at the ends of the reads, and cut out the adapter fragments to obtain second sequencing data, where the second tool is the Trimmomatic or Cutadapt tool; Step b3: Set every F bases of the sequencer as a sliding window, perform a sliding check on the second sequencing data to obtain the average quality score of the bases within each sliding window, and remove the window segments with a quality score lower than the average quality score threshold to obtain third sequencing data; F is an integer greater than 1; Step b4: Use a third tool to mark the repetitive sequences generated by PCR amplification in the third sequencing data, and use a decontamination tool to clean the non-human source sequences to obtain high-quality sequencing data, where the third tool is the Picard tool; Step 4: Perform functional annotation on the set of gene mutation sites to identify a set of suspected high-risk sites and the corresponding annotation results; Step 5: Based on the PRS model, evaluate the set of suspected high-risk sites to obtain the z-score value of the q-th extreme individual; determine the damage risk level according to the z-score value; Step 6: Feedback the corresponding protection measure plan to the client for display according to the damage risk level.
2. The method for assessing the risk of noise-induced hearing loss according to claim 1, wherein The specific training process of the auditory analysis model includes: Pre-collect Y groups of auditory representation data, where Y is an integer greater than 1, set the corresponding screening results for the auditory representation data, mark the digital tags of the screening results as prediction tags, and convert the auditory representation data and the corresponding prediction tags into a corresponding set of feature vectors; Taking each set of feature vectors as the input of an auditory analysis model, the auditory analysis model outputs a set of predicted labels corresponding to each set of auditory representation data, and uses the actual predicted label corresponding to each set of auditory representation data as the prediction target, where the actual predicted label is the pre-set predicted label corresponding to the auditory representation data; taking minimizing the sum of prediction errors of all auditory representation data as the training target; training the auditory analysis model until the sum of prediction errors converges and then stopping the training; the auditory analysis model is a random forest, a support vector machine or a deep neural network model, and obtaining a first screening result.
3. The method for assessing the risk of noise-induced auditory injury according to claim 2, wherein the specific reagent is a lysis buffer, a chelating agent or a precipitating agent; Generating sequencing data based on the purified DNA includes: Decomposing the purified DNA by mechanical or enzymatic cleavage to obtain DNA fragments, the size of the DNA fragments being in the range of 100-500 bp; Using a specific exon probe pool to capture DNA fragments in the exon region, obtaining exon sequence fragments, and removing non-exon sequences; Connecting adaptors containing sequencing linkers to both ends of the exon sequence fragments, and performing PCR amplification on the exon sequence fragments to obtain exon sequence amplification fragments; Sequencing the exon sequence amplification fragments by a sequencer connected to a sequencing platform to obtain sequencing data.
4. The method for assessing the risk of noise-induced hearing loss according to claim 1, wherein The method for aligning the obtained high-quality sequencing data with a reference genome to determine a set of gene mutation sites includes: Step c1: Using the BWA-MEM algorithm to compare the high-quality sequencing data with the target genome to obtain a SAM file; Step c2: Using a fifth tool to convert the SAM file into a compressed BAM format, sorting it by genomic position to obtain a BAM file, and using a sixth tool to mark the repetitive sequences in the PCR process, correct the alignment offset, and filter out unmatched alignment fragments; Step c3: Using a seventh tool to perform genomic recalibration on the BAM file, extracting genomic variation information, and generating a preliminary VCF file; Step c4: Using the VQSR module of the seventh tool to calibrate the detected mutation sites, setting a mutation quality threshold to screen a set of matching gene mutation sites.
5. The method for assessing the risk of noise-induced hearing loss according to claim 4, wherein The method for functionally annotating a set of gene mutation sites to identify a set of suspected high-risk sites and the corresponding annotation results includes: Step d1: Based on known SNP site data, performing a preliminary screening on the set of gene mutation sites to obtain a screened set of gene mutation sites; the screened set of gene mutation sites includes a set of gene mutation sites in group E; Step d2: Performing an association analysis on the set of gene mutation sites in group e, calculating the correlation coefficient of the set of gene mutation sites in group e, where e = 1, 2,..., E; Step d3: Preset a correlation coefficient threshold, where the correlation coefficient threshold includes and , where ; Compare the correlation coefficient of the e-th group of gene mutation sites with the preset correlation coefficient threshold; If > , then mark the gene mutation sites in the e-th group as suspected high-risk sites; If , then mark the gene mutation sites in the e-th group as medium-risk sites; If , then mark the genetic variation sites in the e-th group as low-risk sites; Step d4: Using a functional annotation tool to annotate the suspected high-risk sites to obtain annotation results, setting e = e + 1, and returning to step d2; the annotation results include gene positions and functional regions: the gene positions include coding regions, promoters or non-coding regions; Step d5: Repeat steps d2 - d4 until the loop ends when e = E, and obtain a set of suspected high - risk loci and the corresponding annotation results.
6. The method for assessing the risk of noise-induced hearing loss according to claim 5, wherein The loci of the PRS model include rs10783780, rs10794566, rs10794567, rs3887954, rs58011943, rs11836060, rs16973424, and rs74905175.
7. The method for assessing the risk of noise-induced hearing loss according to claim 6, wherein Based on the PRS model, the method for obtaining the z - score value of the q - th extreme individual by evaluating the set of suspected high - risk loci includes: Step s1: Extract the effect value of the h - th suspected locus in the set of suspected high - risk loci based on the correlation coefficient; h = 1, 2, ……, H; H is the number of suspected loci in the set of suspected high - risk loci; Step s2: Calculate the PRS score of the q - th extreme individual; Step s3: Repeat step s2 until the loop ends when q = Q, and obtain the PRS scores of Q extreme individuals; Step s4: Convert the PRS score of the q - th extreme individual into a z - score value to obtain the z - score value of the q - th extreme individual.
8. The method for assessing the risk of noise-induced hearing loss according to claim 7, characterized in that, The injury risk levels include the first level, the second level, the third level, the fourth level, and the fifth level; the first level > the second level > the third level > the fourth level > the fifth level.
9. The method for assessing the risk of noise-induced hearing loss according to claim 8, wherein The individual genetic data includes KCNQ4, MYO6, MYO7A, MYO15A, OTOF, GJB2, SLC26A4, GRM7, TECTA, CDH23, ESRRB, HSP70, CAT, and SOD2; the environmental exposure data includes noise exposure intensity and noise exposure time; The individual physiological data includes age, gender, genetic background, health status, and history of hearing injury; the behavioral data includes high - frequency sound sensitivity, tinnitus symptoms, and tolerance to noise.
10. A noise-induced hearing loss risk assessment system for implementing the noise-induced hearing loss risk assessment method according to any one of claims 1-9, characterized in that Including: A data acquisition module for obtaining the auditory characterization data of I sample individuals, where I is an integer greater than zero; the auditory characterization data includes individual genetic data, environmental exposure data, individual physiological data, and behavioral data; An analysis module for inputting the auditory characterization data of I sample individuals into a pre - constructed auditory analysis model to obtain a first screening result, where the first screening result includes Q extreme individuals; A sequencing module for obtaining the Q extreme individuals output by the auditory analysis model, extracting the q - th extreme individual for whole - exome sequencing to obtain a set of gene mutation loci; q = 1, 2, ……, Q; An identification module for performing functional annotation on the set of gene mutation loci to identify a set of suspected high - risk loci and the corresponding annotation results; An evaluation module for evaluating the set of suspected high - risk loci based on the PRS model to obtain the z - score value of the q - th extreme individual; Determine the injury risk level according to the z - score value; A display module for feeding back the corresponding protection measure plan to the client for display according to the injury risk level.
Citation Information
Patent Citations
Auditory pathway evaluation and analysis device and method thereof
CN117297596A
Noisy hearing loss prediction and susceptible population screening method and apparatus, terminal and medium
CN111584065A
Mitochondrial disease prediction method and mitochondrial disease prediction system
CN117953969A
Prenatal screening data acquisition and analysis system and method
CN118588161A
Risk assessment system for chronic diseases
CN119207806A