Hearing health detection method, device and equipment and readable storage medium
By obtaining the central feature vector and reference hearing health parameters of a lossy audio set, and using feature extraction and distance calculation to determine the hearing health parameters of the target audio, the problem of inaccurate hearing health detection in existing technologies is solved, and a more objective hearing health measurement is achieved.
Patent Information
- Application Number
- CN202511110791.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-08
- Publication Date
- 2025-11-04
AI Technical Summary
Existing hearing health tests lack unified standards, and the scoring methods are highly subjective, resulting in inaccurate testing.
By obtaining the central feature vector and reference hearing health parameters of a lossy audio set, the hearing health parameters of the target audio are determined using feature extraction and distance calculation, avoiding human scoring and weight setting.
It improves the accuracy of hearing health testing and provides an objective method for measuring hearing health.
Smart Images

Figure CN120884282A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of intelligent detection, and in particular to a hearing health detection method, device and equipment and a readable storage medium. BACKGROUND
[0002] The existing hearing health score does not have a unified standard. Different scoring methods are based on various factors affecting hearing health, and the scores are weighted by quantifying these factors. However, the factors affecting hearing health cannot be completely listed, and the quantification of each factor and the selection of weights for different factors are subjective, which leads to the final score being not objective and authoritative.
[0003] Therefore, how to improve the accuracy of hearing health detection is a technical problem that needs to be solved by those skilled in the art. SUMMARY
[0004] Therefore, the purpose of the present application is to provide a hearing health detection method, device and equipment and a readable storage medium, which solves the technical problem of inaccurate hearing health detection in the prior art.
[0005] To solve the above technical problems, the present application provides a hearing health detection method, comprising:
[0006] obtaining a center feature vector corresponding to a set of lossy audios and a reference hearing health parameter corresponding to the set of lossy audios; wherein the reference audio in the set of lossy audios is an audio whose damage degree to hearing health reaches a preset condition;
[0007] performing feature extraction on the target audio to obtain an audio feature vector of the target audio;
[0008] determining a feature distance of the audio feature vector to the center feature vector;
[0009] determining a hearing health parameter of the target audio based on the feature distance and the reference hearing health parameter.
[0010] Optionally, before obtaining the center feature vector corresponding to the set of lossy audios, further comprising:
[0011] determining a target parameter for representing the damage degree; the target parameter includes a proportion of target high-frequency energy, a loudness value and a noise proportion; wherein the proportion of target high-frequency energy is the energy of acoustic signals in a set frequency range;
[0012] determining a first preset number of first reference audios with the highest proportion of target high-frequency energy based on the proportion of target high-frequency energy corresponding to each reference audio, to obtain a first set of lossy audios;
[0013] determine a second preset number of first reference audios with the maximum loudness value based on the loudness value corresponding to each reference audio, to obtain a second lossy audio set;
[0014] determine a third preset number of third reference audios with the maximum noise proportion based on the noise proportion corresponding to each reference audio, to obtain a third lossy audio set, and take the first lossy audio set, the second lossy audio set and the third lossy audio set as the lossy audio set;
[0015] Correspondingly, the center feature vector corresponding to the lossy audio set is obtained, including:
[0016] determine the overall feature vector corresponding to the lossy audio set, and take the overall feature vector as the center feature vector corresponding to the lossy audio set.
[0017] Optionally, determining the overall feature vector corresponding to the lossy audio set, and taking the overall feature vector as the center feature vector corresponding to the lossy audio set, includes:
[0018] determine the reference audio feature vector corresponding to each reference audio in the lossy audio set;
[0019] determine the feature vector mean based on all the reference audio feature vectors, and take the feature vector mean as the overall feature vector.
[0020] Optionally, determining the feature distance from the audio feature vector to the center feature vector includes:
[0021] take the feature vector mean as the center point of the lossy audio set;
[0022] take the distance between the audio feature vector and the center point as the feature distance.
[0023] Optionally, feature extraction is performed on the target audio to obtain an audio feature vector of the target audio, including:
[0024] performing Mel filtering processing on the target audio to obtain a time-frequency spectrum;
[0025] performing multi-stage residual quantization processing on the time-frequency spectrum to obtain a discrete compression codebook;
[0026] performing block processing on the time-frequency spectrum to obtain a time-frequency block, and performing block masking processing on the time-frequency block to obtain a partially visible spectrum;
[0027] processing the partially visible spectrum through a convolutional neural network to extract a target high-dimensional feature vector, the target high-dimensional feature vector representing the statistical distribution characteristics of the frequency domain structure.
[0028] Cross attention processing is performed on the discretized compression codebook and the target high-dimensional feature vector to obtain the audio feature vector.
[0029] Optionally, determining the hearing health parameter of the target audio based on the feature distance and the reference hearing health parameter comprises:
[0030] Determining the age stage of the current user, and determining the target hearing health parameter priority corresponding to the age stage; wherein different target hearing health parameter priorities correspond to different distance ranges;
[0031] Comparing the feature distance with the distance range corresponding to each priority in the target hearing health parameter priority to obtain a hearing health priority, and determining the hearing health parameter based on the hearing health priority.
[0032] Optionally, after determining the hearing health parameter of the target audio based on the feature distance and the reference hearing health parameter, the method further comprises:
[0033] When the hearing health parameter is less than a set minimum hearing health parameter, determining a target parameter corresponding to the target audio;
[0034] Based on the target parameter corresponding to the target audio, the target audio is health-corrected to obtain a health-corrected target audio.
[0035] Optionally, the preset condition comprises at least one condition of maximum damage to hearing health and minimum loss of hearing health.
[0036] Embodiments of the present application also provide a hearing health detection device, comprising:
[0037] A center feature vector determination module is configured to obtain a center feature vector corresponding to a set of lossy audios and reference hearing health parameters corresponding to the set of lossy audios; wherein a reference audio in the set of lossy audios is an audio whose damage to hearing health meets a preset condition.
[0038] An audio feature vector determination module is configured to perform feature extraction on a target audio to obtain an audio feature vector of the target audio.
[0039] A feature distance determination module is configured to determine a feature distance from the audio feature vector to the center feature vector.
[0040] A hearing health parameter determination module is configured to determine a hearing health parameter of the target audio based on the feature distance and the reference hearing health parameter.
[0041] The embodiment of the present application also provides a hearing health detection device, comprising:
[0042] a memory for storing the computer program;
[0043] a processor for executing the computer program to realize the steps of the hearing health detection method.
[0044] The embodiment of the present application also provides a readable storage medium, wherein the readable storage medium stores a computer program, and the computer program is executed by a processor to realize the steps of the hearing health detection method.
[0045] The embodiment of the present application also provides a computer program product, comprising a computer program / instruction, and the computer program / instruction is executed by a processor to realize the steps of the hearing health detection method.
[0046] It can be seen that, by acquiring the center feature vector corresponding to the lossy audio set and the reference hearing health parameter corresponding to the lossy audio set, wherein the reference audio in the lossy audio set is the audio whose damage degree to the hearing health reaches a preset condition, performing feature extraction on the target audio to obtain the audio feature vector of the target audio, determining the feature distance from the audio feature vector to the center feature vector, and determining the hearing health parameter of the target audio based on the feature distance and the reference hearing health parameter, the present application has the beneficial effects that, compared with the current hearing health scoring of each audio by experts to obtain the weighted hearing health parameter for measuring the hearing health of the audio, the present application screens the lossy audio set by the preset condition, determines the center feature vector corresponding to the lossy audio set, determines the feature distance between the audio feature vector of the target audio and the center feature vector, and measures the hearing health of the audio based on the feature distance and the reference hearing health parameter corresponding to the lossy audio set, and since the present application does not need artificial scoring or artificial setting of the weight factor, the hearing health parameter can be measured based on the objective target distance, and the accuracy of the hearing health detection is improved.
[0047] In addition, the present application also provides a hearing health detection device, equipment and readable storage medium, which also have the beneficial effects. BRIEF DESCRIPTION OF DRAWINGS
[0048] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or the prior art description. Obviously, the drawings in the following description are only the embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor on the basis of the provided drawings.
[0049] Figure 1A flowchart of a hearing health detection method provided for an embodiment of the present application;
[0050] Figure 2 A flowchart of a hearing health detection method provided for an embodiment of the present application;
[0051] Figure 3 A schematic diagram of a hearing health detection framework provided for an embodiment of the present application;
[0052] Figure 4 A schematic diagram of a farthest point distance and mean center provided for an embodiment of the present application;
[0053] Figure 5 A structural schematic diagram of a hearing health detection device provided for an embodiment of the present application;
[0054] Figure 6 A structural schematic diagram of a hearing health detection device provided for an embodiment of the present application. DETAILED DESCRIPTION
[0055] In order to make the objects, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described below in connection with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of protection of the present application.
[0056] Hearing health is often overlooked in people's daily life, and the population of people wearing earphones to listen to songs is very large. Therefore, by calculating the hearing health degree, it is an important ability of a music software to inform users whether the current listening habit will cause hearing health problems. Through the hearing health degree, users can adjust their listening habits and tastes, so that their hearing level remains healthy for a long time, and they can better and more continuously enjoy music. Therefore, the method proposed in the present case can bring more scientific and accurate calculation results for the hearing health degree, and better serve the users.
[0057] Please refer to Figure 1 , Figure 1 A flowchart of a hearing health detection method provided for an embodiment of the present application. The method can include:
[0058] S101, obtaining a center feature vector corresponding to a set of lossy audios and a reference hearing health parameter corresponding to the set of lossy audios; wherein the reference audio in the set of lossy audios is an audio whose damage to hearing health reaches a preset condition.
[0059] The execution subject of this embodiment is an electronic device. This embodiment does not limit the specific electronic device, which can be a computer, a mobile phone, etc. The lossy audio set in this embodiment is the reference audio selected from the audio. The lossy audio set is composed of reference audio. In addition to the 0-point audio, the reference audio can also be the 100-point audio (the audio with the smallest damage to hearing health), or the 0-point and 100-point audio are used at the same time (the hearing health needs to be determined by combining the first feature distance between the target audio and the 0-point audio, and the second feature distance between the target audio and the 100-point audio). This embodiment does not limit the specific preset condition. For example, the preset condition in this embodiment can be the maximum damage to hearing health (at this time, the reference audio is 0-point audio), or the preset condition in this embodiment can be the minimum damage or no damage to hearing health (at this time, the reference audio is 100-point audio). For example, this embodiment can set a preset number of target parameters, thereby determining the target parameters of each category corresponding to each reference audio in the music library, determining a set number of reference audios with the maximum target parameters corresponding to each category, and summarizing the set number of audios corresponding to each category to obtain all reference audios. This embodiment does not limit the specific target parameter. For example, the target parameter in this embodiment can be the proportion of audio high-frequency energy; or the target parameter in this embodiment can be the audio loudness; or the target parameter in this embodiment can be the noise proportion; or the target parameter in this embodiment can be the transient impulse index (TII), that is, the proportion of sound pressure change rate > 100 dB / s events per unit time. The larger the target parameter in this embodiment, the more serious the impact on hearing health. This embodiment does not limit the specific center feature vector corresponding to the reference audio. For example, when there are multiple reference audios in this embodiment, the mean of the feature vectors corresponding to the multiple reference audios can be used as the center feature vector; or this embodiment can also use the mode of the feature vectors corresponding to the multiple reference audios as the center feature vector. This embodiment does not limit the extraction method of the feature vector corresponding to the reference audio. For example, this embodiment can extract the feature vector corresponding to the reference audio based on the Encodecmae (Encodec Mask autoencoder, Encodec Mask autoencoder) model. Or this embodiment can also perform feature extraction based on the MULE (Multitask Universal Learning Encoder, Multitask Universal Learning Encoder) model; or this embodiment can also perform feature extraction based on the LAION-CLAP (Contrastive Learning-AudioPretraining, Contrastive Learning-AudioPretraining model); or it can also be based on the MERT (MusicEncoder w / Residual Transformers, MusicEncoder w / Residual Transformers) model for feature extraction.
[0060] It should be further explained that based on any of the above embodiments, in order to improve the accuracy of the center feature vector acquisition, the above can further include, before acquiring the center feature vector corresponding to the lossy audio set:
[0061] Step 1: Determine the target parameter for characterizing the degree of damage; the target parameter includes the proportion of target high-frequency energy, the loudness value and the noise proportion; wherein the proportion of target high-frequency energy is the energy of the acoustic signal in the frequency range within the set frequency range.
[0062] In this embodiment, the proportion of target high-frequency energy is determined based on frequency detection; in this embodiment, the loudness value is determined based on loudness detection; in this embodiment, the noise proportion is determined based on noise detection. The proportion of target high-frequency energy in this embodiment is a percentage parameter that measures the proportion of high-frequency energy (e.g. 4kHz-16kHz) in the total energy of the full frequency band (20Hz-20kHz); the noise proportion in this embodiment is the proportion of distortion and background noise energy in the signal to the total energy; the loudness value in this embodiment represents the acoustic energy intensity parameter that harms hearing health. The reason why this embodiment selects these three parameters is because it is considered that high-frequency energy is directly related to the fatigue of the basilar membrane cilium, loudness determines the mechanical damage threshold of hair cells, and noise components cause abnormal synapses - these three parameters exactly cover the three key damages of the auditory pathway, which can improve efficiency on the basis of ensuring accuracy.
[0063] Step 2: Determine the first preset number of first reference audios with the highest proportion of target high-frequency energy based on the proportion of target high-frequency energy corresponding to each reference audio, to obtain a first lossy audio set.
[0064] The process of determining the proportion of target high-frequency energy in this embodiment can include: measuring the proportion of target high-frequency energy of the audio based on the target high-frequency energy proportion formula, The proportion of high-frequency energy is the proportion of target high-frequency energy; wherein E(f) represents the energy at frequency f, f high_start represents the starting frequency of the high-frequency band, f high_end represents the cutoff frequency of the high-frequency band, f min represents the lowest frequency of the full frequency band, f maxThe highest frequency of the full frequency band is represented. In this embodiment, the top 100 (first preset number) songs with the highest proportion of high-frequency energy in the song library can be selected as the first lossy audio set. In this embodiment, the proportions of target high-frequency energy corresponding to each reference audio can be sorted from high to low, and the top first preset number of first reference audios are selected as the first lossy audio set. For example, the design idea of this embodiment can be to use a seed song (reference audio) as a 0-point audio reference. The reference audio is the song that has the most serious impact on hearing health. If 60-point and 80-point audios need to consider multiple factors, but defining a 0-point audio only needs to find an extreme case, so determining a 0-point reference audio is more objective and does not have subjective influence. The first reference audio in the current embodiment can also be an audio that has no impact on hearing health; or the first reference audio in this embodiment can also be an audio that has damage to hearing health that meets a preset condition that can be measured.
[0065] Step 3: Determine a second preset number of second reference audios with the largest loudness value based on the loudness value corresponding to each reference audio, and obtain a second lossy audio set.
[0066] In this embodiment, the loudness of each reference audio in the song library is detected to determine the LUFS (full-scale loudness unit) value of the audio loudness, and the second preset number of reference audios with the largest loudness are selected as the second lossy audio set. The second preset number in this embodiment can be the same as or different from the first preset number.
[0067] Step 4: Determine a third preset number of third reference audios with the largest noise proportion based on the noise proportion corresponding to each reference audio, and obtain a third lossy audio set. The first lossy audio set, the second lossy audio set, and the third lossy audio set are used as the lossy audio set.
[0068] In this embodiment, for noise detection, the third preset number (for example, 100) of audios with a noise proportion of 100% are selected as the third lossy audio set through a song attribute classification model (identifying the proportion of noise / speech / singing / acompaniment music in a song). The first preset number, the second preset number, and the third preset number in this embodiment can be the same or different, and can be set according to requirements.
[0069] Correspondingly, obtaining the center feature vector corresponding to the lossy audio set can include:
[0070] Step 5: Determine the overall feature vector corresponding to the lossy audio set, and use the overall feature vector as the center feature vector corresponding to the lossy audio set.
[0071] The embodiment does not limit the specific method for determining the overall feature vector corresponding to the lossy audio set. For example, the embodiment can perform cluster analysis on the feature vector corresponding to the lossy audio set to obtain the overall feature vector; or the embodiment can determine the mean value of the feature vector set of the lossy audio set. The embodiment gives a specific method for determining the center feature vector corresponding to the lossy audio set, and improves the accuracy of determining the center feature vector corresponding to the lossy audio set.
[0072] It needs to be further explained that, based on the above embodiment, in order to improve the accuracy of determining the overall feature vector, the above determining the overall feature vector corresponding to the lossy audio set, taking the overall feature vector as the center feature vector corresponding to the lossy audio set, can include: determining a reference audio feature vector corresponding to each reference audio in the lossy audio set; determining a feature vector mean value based on all reference audio feature vectors, and taking the feature vector mean value as the overall feature vector. The embodiment takes the mean value of the reference audio feature vectors of the reference audios in the lossy audio set as the overall feature vector, and the mean value calculation is a linear operation with low calculation complexity, which can improve the efficiency of determining the feature vector.
[0073] It needs to be further explained that, based on the above embodiment, in order to improve the accuracy of determining the feature distance, the above determining the feature distance of the audio feature vector to the center feature vector can include: taking the feature vector mean value as the center point of the lossy audio set; and taking the distance between the audio feature vector and the center point as the feature distance. In the embodiment, the feature distance is the distance between the audio feature vector and the center point, and since the center point can be understood as the most accurate reference audio, calculating the distance from the center point can improve the accuracy of calculating the feature distance.
[0074] S102, performing feature extraction on the target audio to obtain an audio feature vector of the target audio.
[0075] The embodiment does not limit the specific method for performing feature extraction on the target audio, for example, feature extraction can be performed based on a MERT (Music Encoder w / Residual Transformers, residual transformer music encoder) model; or the embodiment can perform feature extraction based on a MULE (Multitask Universal Learning Encoder, multitask universal learning encoder) model. It needs to be explained that, in order to improve the accuracy of calculating the feature distance, the method for performing feature extraction on the target audio in the embodiment and the method for performing feature extraction on the reference audio in the lossy audio set are consistent.
[0076] It needs to be further explained that, based on any of the above embodiments, in order to improve the accuracy of audio feature vector extraction, the above feature extraction on the target audio to obtain the audio feature vector of the target audio can include:
[0077] S1021, the target audio is subjected to Mel filtering processing to obtain a time-frequency spectrum.
[0078] This embodiment can generate a time-frequency spectrum based on the target audio through a Mel filter bank. In this embodiment, Mel filtering processing is performed, i.e., Mel filtering processing is performed through a Mel filter bank. This embodiment converts the target audio into a time-frequency representation consistent with the human ear's hearing characteristics, and the Mel scale simulates the non-linear frequency perception of the basilar membrane. The Mel filter bank in this embodiment is a set of band-pass filters distributed according to the Mel scale, which is used to map linear frequencies to a non-linear frequency perception space consistent with the human ear's hearing characteristics. The time-frequency spectrum in this embodiment is a three-dimensional representation of time-frequency-energy.
[0079] S1022, the time-frequency spectrum is subjected to multi-stage residual quantization processing to obtain a discretized compressed codebook.
[0080] This embodiment can input the time-frequency spectrum into a residual vector quantization encoder to generate a discretized compressed codebook. The residual vector quantization encoder in this embodiment is a multi-stage vector quantization device that compresses high-dimensional features by iteratively approximating residuals, i.e., multi-stage residual quantization processing is performed based on the residual vector quantization encoder. The discretized compressed codebook in this embodiment is a pre-trained numerical dictionary structure composed of a fixed number of prototype vectors (code words), which is used to efficiently represent the audio feature space.
[0081] S1023, the time-frequency spectrum is subjected to blocking to obtain time-frequency blocks, and the time-frequency blocks are subjected to block masking processing to obtain a partially visible spectrum.
[0082] The blocking in this embodiment refers to the process of cutting the time-frequency spectrum into sub-regions of equal size. The block masking processing in this embodiment is a structured random discarding technique that generates a partially visible and damaged spectrum by forcing the elements of the time-frequency blocks to zero.
[0083] S1024, the partially visible spectrum is processed through a convolutional neural network to extract a target high-dimensional feature vector, and the target high-dimensional feature vector represents the statistical distribution characteristics of the frequency domain structure.
[0084] The convolutional neural network in this embodiment is a biologically inspired deep learning model that extracts hierarchical features of the time-frequency spectrum through local connections and weight sharing mechanisms.
[0085] S1025, cross-attention processing is performed on the discretized compressed codebook and the target high-dimensional feature vector to obtain an audio feature vector.
[0086] The cross-attention processing in this embodiment is a multi-modal feature fusion mechanism that aligns different source features through dynamic weight distribution. Through cross-attention processing, this embodiment can fuse the discretized compressed codebook and the target high-dimensional feature vector to obtain an audio feature vector. The method of extracting the feature vector in this embodiment improves the accuracy of feature extraction, because the mel spectrum simulates the human auditory characteristics, the residual quantization preserves key information, the block masking enhances the robustness of the convolutional neural network, and finally the cross-attention realizes global optimal fusion.
[0087] S103, determining a feature distance of the audio feature vector to the center feature vector.
[0088] This embodiment does not limit the specific process of determining the feature distance of the audio feature vector to the center feature vector. For example, in this embodiment, there are multiple center feature vectors, and the mean of the center feature vectors can be used as the center to calculate the distance between the mean center and the audio feature vector; or this embodiment can cluster multiple center feature vectors to obtain a clustering center, and calculate the distance between the audio feature vector and the clustering center; or this embodiment can directly calculate the distance between the center feature vector and the audio feature vector. This embodiment does not limit the specific method of calculating the feature distance, for example, this embodiment can determine the cosine distance between the to-be-detected feature vector and the feature vector; or this embodiment can determine the Euclidean distance between the to-be-detected feature vector and the feature vector.
[0089] S104, determining the hearing health parameter of the target audio based on the feature distance and the reference hearing health parameter.
[0090] The embodiment does not limit the specific method for determining the hearing health parameter corresponding to the target audio based on the feature distance. For example, when the reference audio in the set of lossy audios is the audio with the greatest damage to hearing health, the feature distance value can be directly taken as the hearing health parameter, and the greater the target distance, the healthier the hearing. Alternatively, the embodiment can divide the feature distance into corresponding hearing health parameter priorities based on the reference hearing health parameter, and take the hearing health parameter priority as the hearing health parameter. For example, the embodiment can take the average as the center point of the 0-point song (reference audio, audio with the greatest damage to hearing health), calculate the feature distance between the audio feature vector and the center point; or the embodiment can calculate the distance between each reference audio in the set of lossy audios and the center point, and take the farthest distance as the farthest point of the 0-point song. The farthest distance is defined as d, and a 1-point song can be defined as a song within the range of d to 2d from the center point, a 2-point song can be defined as a song within the range of 2d to 3d, and so on. A song with a score of 100 or more is a 100-point song, that is, a k-point song is a song within the range of kxd to (k+1)d. By such a method, the target distance can be accurately refined into each range to obtain an accurate hearing health parameter. When 0-point and 100-point (no damage to hearing health) audios are used at the same time, the first feature distance between the target audio and the 0-point audio and the second feature distance between the target audio and the 100-point audio are combined to determine the hearing health parameter.
[0091] It needs to be further explained that based on any of the above embodiments, the determination of the hearing health parameter of the target audio based on the feature distance and the reference hearing health parameter can include: determining the age stage of the current user, and determining the target hearing health parameter priority corresponding to the age stage; wherein different target hearing health parameter priorities correspond to different distance ranges; comparing the feature distance with the distance range corresponding to each priority in the target hearing health parameter priority to obtain a hearing health priority, and determining the hearing health parameter based on the hearing health priority. The embodiment does not limit the specific division of the age stage. For example, the age stage in the embodiment can be divided into infants, children, young people, middle-aged people and old people; or the age stage can be divided into infants (0-3 years old), children (4-12 years old), children (4-12 years old), teenagers (13-25 years old) and old people (60+ years old). The embodiment can divide different priorities for different age stages based on the reference hearing health parameter corresponding to the set of lossy audios, and each priority corresponds to a different distance range. Since different age stages have different perceptions of hearing health, the embodiment divides different hearing health parameters for users of different age stages, thereby accurately determining the hearing health parameter (hearing health priority) corresponding to the user of different age stages, and improving the accuracy of the determination of the hearing health parameter.
[0092] It needs to be further explained that based on any of the above embodiments, in order to improve the experience of the user, after determining the hearing health parameter of the target audio based on the feature distance and the reference hearing health parameter, the method can further include: when the hearing health parameter is less than the set minimum hearing health parameter, determining a target parameter corresponding to the current target audio; and based on the target parameter corresponding to the target audio, performing health correction on the target audio to obtain a health-corrected target audio. In this embodiment, the greater the hearing health parameter, the healthier the current audio. When it is less than the set minimum hearing health parameter, it means that the current audio has a greater damage to the user's hearing health, so health correction is needed. The health correction in this embodiment refers to correcting the audio to reduce hearing health damage. When the target parameter is the proportion of target high-frequency energy, the loudness value and the noise proportion, for high-frequency damage correction: operation: apply a slope high-frequency attenuation filter to reduce energy above 4 kHz with a slope of 6-24 dB / octave. Biological simulation: simulate the cochlea outer hair cell active suppression mechanism to reduce the basal membrane resonance amplitude. Intensity control: use a 6 dB / oct slope when the over-standard rate is 25-30%, and increase to 24 dB / oct when it is greater than 40%. Loudness damage correction: operation: perform dynamic range expansion to restore the non-linear compression characteristics of the audio waveform. Biological simulation: reconstruct the inner ear lymphatic pressure gradient to avoid tearing of the hair cells caused by instantaneous high pressure. Intensity control: for every 5 dB increase in loudness over-standard value, the expansion index decreases by 0.1 (e.g. 85 dB → index 0.7, 90 dB → 0.6). Noise damage correction: operation: separate the fundamental harmonic components and enhance them, and suppress non-harmonic noise (gain factor λ = 0.5-1.0). Biological simulation: strengthen the harmonic locking ability of the auditory center to improve the frequency domain signal-to-noise ratio. Intensity control: when the noise over-standard rate is greater than 2%, λ = 1.0, and the harmonic structure is forced to be reconstructed. Post-correction safety guarantee. Parameter constraint: ensure that the output audio meets Fh high-frequency energy proportion ≤20%, Lh loudness value ≤80 dB, Np noise proportion ≤0.8%. Audio quality fidelity: certified by the P.862 standard (speech quality perceptual evaluation ≥3.8), and the total harmonic distortion is less than 1.2%. This embodiment gives the correction performed when the target audio damages the hearing health, thereby reducing the damage to the hearing health and improving the user's experience,
[0093] The hearing health detection method provided by the embodiment of the present application can comprise: S101, obtaining a center feature vector corresponding to a set of damaged audios and a reference hearing health parameter corresponding to the set of damaged audios; wherein the reference audio in the set of damaged audios is an audio whose damage degree to hearing health reaches a preset condition; S102, performing feature extraction on a target audio to obtain an audio feature vector of the target audio; S103, determining a feature distance from the audio feature vector to the center feature vector; S104, determining a hearing health parameter of the target audio based on the feature distance and the reference hearing health parameter. Compared with the current hearing health scoring of each audio by an expert to obtain a weighted hearing health parameter to measure the hearing health of the audio, the present application screens to obtain a set of damaged audios, determines a center feature vector corresponding to the set of damaged audios, determines a target distance between the center feature vector and an audio feature vector of a target audio, and measures the hearing health of the audio based on the target distance. Since the present application does not need artificial scoring or artificial setting of a weight factor, the hearing health parameter can be measured based on an objective target distance, thereby improving the accuracy of hearing health detection.
[0094] In order to make the present application more convenient to understand, please refer to Figure 2 , Figure 2 The flowchart of the hearing health detection method provided by the embodiment of the present application can specifically comprise:
[0095] S201, performing frequency detection, loudness detection and noise detection on audios in a music library to obtain a first set of 100 damaged audios with the highest proportion of target high-frequency energy, a second set of 100 damaged audios with the largest loudness value, and a third set of 100 damaged audios with the largest noise ratio.
[0096] The structural framework corresponding to the embodiment is shown in Figure 3 , Figure 3 The schematic diagram of the hearing health detection framework provided by the embodiment of the present application. Figure 3 The to-be-predicted song in the music library is a target audio, the to-be-predicted song representation is a feature vector corresponding to the target audio, and the song representation is a mean feature vector corresponding to 300 reference audios. The 0-minute seed song is a reference audio.
[0097] S202, combining the first set of damaged audios, the second set of damaged audios and the third set of damaged audios into a set of 300 damaged audios.
[0098] S203, performing feature extraction on the 300 audios in the set of damaged audios based on a music understanding large model to obtain a feature vector corresponding to each reference audio, and determining a center feature vector corresponding to the 300 target audios.
[0099] The embodiment can understand 300 audio input music large models, and the application extracts song characteristics by using Encodecmae. The output of Encodecmae is a (20, 1024) dimensional vector, where 20 represents the network layer, and the output of the last layer, that is, a (1, 1024) dimensional vector, is taken as the song characteristic. The mean value (mean feature vector) of the characteristics of 300 seed songs is taken as the center point of the 300 audios in the lossy audio set.
[0100] S204, determine the cosine distance between each reference audio in the lossy audio set and the center feature vector, and take the maximum distance as the farthest point distance.
[0101] For other songs to be predicted in the song library, the embodiment also inputs Encodecmae, takes the last layer characteristic, and calculates the cosine distance between the song to be predicted and the mean feature vector (center point). Please refer to Figure 4 , Figure 4 A schematic diagram of the farthest point distance and the mean center provided by the embodiment of the application is shown in the figure, where the 0-point song center point refers to the target audio center point, and the 0-point song farthest point refers to the farthest point distance.
[0102] S205, construct a distance range corresponding to each score based on the farthest point distance.
[0103] In the embodiment, the 1-point song can be defined as a song within the distance center point d-2d, the 2-point song can be defined as a song within the range of 2d-3d, the k-point audio can be defined as a song within the range of kxd to (k+1)d, and the song above 100d is a 100-point song, and d is the farthest point distance. Each score in the embodiment corresponds to a specific distance range.
[0104] S206, determine the distance between the audio feature vector of the target audio and the center feature vector, determine the score corresponding to the current distance based on the distance range corresponding to each score, and take the score corresponding to the target audio as the hearing health parameter.
[0105] The embodiment converts the distance into a score to obtain the hearing health degree of the audio.
[0106] Since the application only needs to collect a lossy audio set, and the score is obtained by characteristic distance, it is not necessary to consider all factors affecting the hearing health degree, the music understanding large model can find the inherent relationship between the reference audios, and automatically distinguish the degree of influence of the song on the hearing health, so the application is more scientific and reasonable compared with the multi-factor weighted score method.
[0107] The hearing health detection device provided by the embodiment of the application will be described below. The hearing health detection device described below can be correspondingly referred to the hearing health detection method described above.
[0108] Specifically refer to Figure 5 , Figure 5 A structure diagram of a hearing health detection device provided by an embodiment of the present application can include:
[0109] The center feature vector determination module 100 is configured to obtain a center feature vector corresponding to a set of lossy audios and a reference hearing health parameter corresponding to the set of lossy audios, wherein a reference audio in the set of lossy audios is an audio whose damage degree to hearing health reaches a preset condition.
[0110] The audio feature vector determination module 200 is configured to perform feature extraction on a target audio to obtain an audio feature vector of the target audio.
[0111] The feature distance determination module 300 is configured to determine a feature distance from the audio feature vector to the center feature vector.
[0112] The hearing health parameter determination module 400 is configured to determine a hearing health parameter of the target audio based on the feature distance and the reference hearing health parameter.
[0113] Further, based on any of the above embodiments, the hearing health detection device can further include:
[0114] A target parameter determination module is configured to determine a target parameter for characterizing the damage degree, wherein the target parameter includes a proportion of target high-frequency energy, a loudness value, and a noise proportion, and the proportion of target high-frequency energy is an acoustic signal energy in a set frequency range.
[0115] A first lossy audio set determination module is configured to determine a first preset number of first reference audios with the highest proportion of target high-frequency energy based on the proportion of target high-frequency energy corresponding to each reference audio to obtain a first lossy audio set.
[0116] A second lossy audio set determination module is configured to determine a second preset number of second reference audios with the largest loudness value based on the loudness value corresponding to each reference audio to obtain a second lossy audio set.
[0117] A lossy audio set determination module is configured to determine a third preset number of third reference audios with the largest noise proportion based on the noise proportion corresponding to each reference audio to obtain a third lossy audio set, and the first lossy audio set, the second lossy audio set, and the third lossy audio set are taken as the set of lossy audios.
[0118] Correspondingly, the center feature vector determination module 100 can include:
[0119] a center feature vector determination unit configured to determine an overall feature vector corresponding to the set of lossy audios, and determine the overall feature vector as the center feature vector corresponding to the set of lossy audios.
[0120] Further, based on any of the above embodiments, the center feature vector determination unit can comprise:
[0121] a reference audio feature vector determination unit configured to determine a reference audio feature vector corresponding to each reference audio in the set of lossy audios;
[0122] an overall feature vector determination unit configured to determine a feature vector mean based on all the reference audio feature vectors, and determine the feature vector mean as the overall feature vector.
[0123] Further, based on any of the above embodiments, the feature distance determination module 300 can comprise:
[0124] a center point determination unit configured to determine the feature vector mean as a center point of the set of lossy audios;
[0125] a feature distance determination unit configured to determine a distance between the audio feature vector and the center point as the feature distance.
[0126] Further, based on any of the above embodiments, the audio feature vector determination module 200 can comprise:
[0127] a time-frequency spectrogram generation unit configured to perform a mel filtering process on the target audio to obtain a time-frequency spectrogram;
[0128] a discretized compressed codebook generation unit configured to perform a multi-stage residual quantization process on the time-frequency spectrogram to obtain a discretized compressed codebook;
[0129] a partially visible spectrogram determination unit configured to perform a block masking process on the time-frequency spectrogram to obtain a partially visible spectrogram;
[0130] a target high-dimensional feature vector extraction unit configured to perform a convolutional neural network process on the partially visible spectrogram to extract a target high-dimensional feature vector, the target high-dimensional feature vector representing a statistical distribution characteristic of a frequency domain structure;
[0131] a feature fusion unit configured to perform a cross-attention process on the discretized compressed codebook and the target high-dimensional feature vector to obtain the audio feature vector.
[0132] Further, based on any of the above embodiments, the hearing health parameter determination module 400 can comprise:
[0133] The hearing health parameter priority determination unit is configured to determine an age stage of the current user, and determine a target hearing health parameter priority corresponding to the age stage; wherein different distance ranges correspond to different target hearing health parameter priorities;
[0134] The hearing health parameter determination unit is configured to compare the feature distance with a distance range corresponding to each of the target hearing health parameter priorities, to obtain a hearing health priority, and determine the hearing health parameter based on the hearing health priority.
[0135] Further, based on any of the above embodiments, the hearing health detection device can further include:
[0136] The target parameter determination module is configured to determine a target parameter corresponding to the target audio when the hearing health parameter is less than a set minimum hearing health parameter.
[0137] The correction module is configured to perform health correction on the target audio based on the target parameter corresponding to the target audio, to obtain a health-corrected target audio.
[0138] Further, based on any of the above embodiments, the preset condition includes at least one of a condition of maximum damage to hearing health and a condition of minimum loss of hearing health.
[0139] It should be noted that the order of the modules and units in the hearing health detection device described above can be changed without affecting the logic.
[0140] The hearing health detection device provided by the embodiment of the present application can comprise: a center feature vector determination module 100 configured to acquire a center feature vector corresponding to a set of damaged audios and a reference hearing health parameter corresponding to the set of damaged audios; wherein the reference audio in the set of damaged audios is an audio whose damage degree to hearing health reaches a preset condition; an audio feature vector determination module 200 configured to perform feature extraction on a target audio to obtain an audio feature vector of the target audio; a feature distance determination module 300 configured to determine a feature distance from the audio feature vector to the center feature vector; and a hearing health parameter determination module 400 configured to determine a hearing health parameter of the target audio based on the feature distance and the reference hearing health parameter. Compared with the current hearing health scoring of each audio by an expert to obtain a weighted hearing health parameter to measure the hearing health of the audio, the present application screens to obtain a set of damaged audios, determines a center feature vector corresponding to the set of damaged audios, determines a target distance between the center feature vector and an audio feature vector of a target audio, and measures the hearing health of the audio based on the target distance. Since the present application does not need artificial scoring or artificial setting of a weight factor, the hearing health parameter can be measured based on an objective target distance, and the accuracy of hearing health detection is improved.
[0141] The hearing health detection device described below can be referred to the hearing health detection method described above.
[0142] Please refer to Figure 6 , Figure 6 The structure diagram of the hearing health detection device provided by the embodiment of the present application can comprise:
[0143] The memory 10 is configured to store a computer program.
[0144] The processor 20 is configured to execute the computer program to implement the hearing health detection method described above.
[0145] The memory 10, the processor 20 and the communication interface 30 can communicate with each other through the communication bus 40.
[0146] In the embodiment of the present application, the memory 10 stores one or more programs, and the program can comprise program code including computer operation instructions. In the embodiment of the present application, the memory 10 can store programs for implementing the following functions:
[0147] Acquire a target feature vector corresponding to a target audio; wherein the target audio is an audio with the largest damage degree to hearing health selected from audios;
[0148] Feature extraction is performed on the to-be-detected audio to obtain a to-be-detected feature vector;
[0149] A target distance between the to-be-detected feature vector and the target feature vector is determined.
[0150] A hearing health parameter corresponding to the to-be-detected audio is determined based on the target distance.
[0151] In a possible implementation, the memory 10 can include a program storage area and a data storage area, where the program storage area can store an operating system, and application programs required by at least one function, and the like; and the data storage area can store data created during use.
[0152] In addition, the memory 10 can include a read-only memory and a random access memory, and provide instructions and data for the processor. A part of the memory can also include an NVRAM. The memory stores an operating system and operation instructions, executable modules or data structures, or a subset of them, or an extended set of them, where the operation instructions can include various operation instructions for implementing various operations. The operating system can include various system programs for implementing various basic tasks and processing hardware-based tasks.
[0153] The processor 20 can be a central processing unit (CPU), an application-specific integrated circuit, a digital signal processor, a field programmable gate array, or other programmable logic device. The processor 20 can be a microprocessor or any conventional processor, etc. The processor 20 can invoke a program stored in the memory 10.
[0154] The communication interface 30 can be an interface of a communication module, used for connecting with other devices or systems.
[0155] Of course, it should be noted that, Figure 6 The structures shown do not constitute a limitation on the hearing health detection device in the embodiments of the present application. In actual applications, the hearing health detection device can include more or fewer components than Figure 6 those shown, or combine certain components.
[0156] The computer-readable storage medium provided by the embodiments of the present application is described below. The computer-readable storage medium described below can be mutually referred to with the hearing health detection method described above.
[0157] The present application also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the steps of the hearing health detection method described above are implemented.
[0158] The computer readable storage medium can include a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media capable of storing program codes.
[0159] The embodiment of the present application also provides a computer program product, comprising computer programs / instructions, which, when executed by a processor, implement the steps of the hearing health detection method.
[0160] The embodiments in the specification are described in a progressive manner, and each embodiment focuses on the difference from other embodiments, and the same or similar parts of each embodiment can be referred to each other. For the device disclosed by the embodiments, since it corresponds to the method disclosed by the embodiments, the description is relatively simple, and the related parts can be referred to the method part.
[0161] The skilled person can further realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be realized by electronic hardware, computer software or a combination of both. In order to clearly show the interchangeability of hardware and software, the components and steps of each example have been described in the above description. Whether the functions are realized by hardware or software depends on the specific application and design constraints of the technical solution. The skilled person can use different methods to realize the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0162] Finally, it should be noted that in this paper, relationships such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between the entities or operations. Moreover, the terms "include", "contain" or any other variant are intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or device.
[0163] The above describes in detail the hearing health detection method, device, equipment and readable storage medium provided by the present application. The principle and implementation mode of the present application are described by applying specific examples in this paper. The above embodiment description is only used to help understand the method of the present application and its core idea; at the same time, for the general technical personnel in the art, according to the idea of the present application, the specific implementation mode and application range will be changed; according to the above, the content of the specification should not be understood as the limitation of the present application.
Claims
1. A method for detecting hearing health, characterized in that, include: Obtain the central feature vector corresponding to the lossy audio set and the reference hearing health parameters corresponding to the lossy audio set; wherein, the reference audio in the lossy audio set is audio that has reached a preset condition for the degree of damage to hearing health; Feature extraction is performed on the target audio to obtain the audio feature vector of the target audio; Determine the feature distance from the audio feature vector to the center feature vector; The hearing health parameters of the target audio are determined based on the feature distance and the reference hearing health parameters.
2. The hearing health test according to claim 1, characterized in that, Before obtaining the central feature vector corresponding to the lossy audio set, the following steps are also included: Determine target parameters to characterize the degree of damage; the target parameters include the proportion of target high-frequency energy, loudness value, and noise proportion; wherein, the proportion of target high-frequency energy is the acoustic signal energy within a set frequency range; Based on the proportion of target high-frequency energy corresponding to each reference audio, determine the first preset number of first reference audios with the highest proportion of target high-frequency energy, and obtain the first lossy audio set; Based on the loudness value corresponding to each reference audio, determine the second preset number of first second reference audios with the largest loudness value, and obtain the second lossy audio set; Based on the noise ratio corresponding to each reference audio, a third preset number of first third reference audios with the largest noise ratio is determined to obtain a third lossy audio set. The first lossy audio set, the second lossy audio set, and the third lossy audio set are used as the lossy audio set. Accordingly, the central feature vector corresponding to the lossy audio set is obtained, including: Determine the overall feature vector corresponding to the lossy audio set, and use the overall feature vector as the center feature vector corresponding to the lossy audio set.
3. The hearing health testing method according to claim 2, characterized in that, Determining the overall feature vector corresponding to the lossy audio set, and using the overall feature vector as the center feature vector corresponding to the lossy audio set, includes: Determine the reference audio feature vector corresponding to each reference audio in the lossy audio set; The mean of the feature vectors is determined based on all the reference audio feature vectors, and the mean of the feature vectors is used as the overall feature vector.
4. The hearing health testing method according to claim 3, characterized in that, Determining the feature distance from the audio feature vector to the center feature vector includes: The mean of the feature vectors is used as the center point of the lossy audio set; The distance between the audio feature vector and the center point is taken as the feature distance.
5. The hearing health testing method according to claim 1, characterized in that, Feature extraction is performed on the target audio to obtain the audio feature vector of the target audio, including: The target audio is subjected to Mel filtering to obtain a time-spectrum. The time-spectrum graph is subjected to multi-stage residual quantization to obtain a discretized compressed codebook; The time-frequency spectrum is divided into blocks to obtain time-frequency blocks, and the time-frequency blocks are subjected to block masking processing to obtain a partially visible spectrum. The visible spectrum is processed by a convolutional neural network to extract a high-dimensional feature vector of the target, which represents the statistical distribution characteristics of the frequency domain structure. The audio feature vector is obtained by performing cross-attention processing on the discretized compressed codebook and the target high-dimensional feature vector.
6. The hearing health testing method according to any one of claims 1 to 5, characterized in that, Determining the hearing health parameters of the target audio based on the feature distance and the reference hearing health parameters includes: Determine the current user's age group and determine the priority of the target hearing health parameters corresponding to the age group; wherein, the distance range corresponding to different priority of target hearing health parameters is different; The hearing health priority is obtained by comparing the feature distance with the distance range corresponding to each priority in the target hearing health parameter priority, and the hearing health parameter is determined based on the hearing health priority.
7. The hearing health testing method according to claim 1, characterized in that, After determining the hearing health parameters of the target audio based on the feature distance and the reference hearing health parameters, the method further includes: When the hearing health parameter is determined to be less than the set minimum hearing health parameter, the target parameter corresponding to the current target audio is determined; Based on the target parameters corresponding to the target audio, the target audio is health-corrected to obtain the health-corrected target audio.
8. The hearing health testing method according to claim 1, characterized in that, The preset conditions include at least one of the following: the condition that causes the greatest damage to hearing health and the condition that causes the least damage to hearing health.
9. A hearing health testing device, characterized in that, include: The central feature vector determination module is used to obtain the central feature vector corresponding to the lossy audio set and the reference hearing health parameters corresponding to the lossy audio set; wherein, the reference audio in the lossy audio set is audio that has reached the preset condition for the degree of damage to hearing health; The audio feature vector determination module is used to extract features from the target audio to obtain the audio feature vector of the target audio. The feature distance determination module is used to determine the feature distance from the audio feature vector to the center feature vector; A hearing health parameter determination module is used to determine the hearing health parameters of the target audio based on the feature distance and the reference hearing health parameters.
10. A hearing health testing device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the steps of the hearing health testing method as described in any one of claims 1 to 8.
11. A readable storage medium, characterized in that, The readable storage medium stores a computer program that, when executed by a processor, implements the steps of the hearing health detection method as described in any one of claims 1 to 8.
12. A computer program product, characterized in that, Includes a computer program / instruction that, when executed by a processor, implements the steps of the hearing health detection method as described in any one of claims 1 to 8.