Automatic debugging method and system, electronic equipment and storage medium
By using an automatic tuning method, the parameter adjustments of the sound-sensing device are determined based on test audio and response information. This solves the problems of expensive and complex tuning of sound-sensing devices, and enables comprehensive and personalized auditory optimization and device adaptation. It is suitable for remote and on-site tuning.
Patent Information
- Application Number
- CN202511657854.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-13
- Publication Date
- 2026-02-06
AI Technical Summary
The existing sound sensor tuning process is expensive and complicated, and users find it difficult to operate on their own. Traditional tuning equipment and services are inconvenient, especially for users in remote areas. Moreover, the tuning process focuses on pure tone testing rather than actual voice/sound/music scenarios.
An automatic tuning method is provided, which plays test audio, receives user response information, determines the matching degree between the response information and the test audio, stores mismatched audio as a benchmark for audio optimization, extracts target distinctive speech features, and adjusts the operating parameters of the sound sensing device to achieve comprehensive detection and personalized optimization.
It achieves fully automated tuning, reduces technical barriers and labor costs, is suitable for remote and on-site tuning, meets personalized listening needs, improves auditory clarity and equipment reproduction accuracy, and adapts to complex audio scenarios.
Smart Images

Figure CN121486745A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of sound sensing equipment technology, and more specifically to an automatic tuning method, system, electronic device, and storage medium. Background Technology
[0002] People with hearing impairments often need to wear sensory devices such as cochlear implants to improve or amplify their hearing. These devices need to be individually adjusted based on each patient's perceptual abilities to maximize their benefits. To adjust the device, the technician tests the user's perceptual abilities with appropriate sensory signals and adjusts the device parameters based on the user's performance in the test.
[0003] Currently, the equipment used for tuning audio devices (such as soundproof rooms, sound field testing equipment, pure tone testing equipment, middle ear analyzers, etc.) is expensive and generally inaccessible to patients / users / customers. Therefore, users typically need to go to a specialized location and have it performed by specialized personnel. The tuning process is complex and costly.
[0004] In view of this, the present invention is hereby proposed. Summary of the Invention
[0005] The present invention was proposed in view of the above-mentioned problems. According to one aspect of the present invention, an automatic tuning method is provided for a sound sensing device, the method comprising: Play test audio to the user, wherein the test audio is any one of words, syllables, music, speech, vocal music, or pure tone; Receive user response information; Determine whether the response information matches the test audio; When the response information does not match the test audio, the test audio is stored as a baseline optimized audio. Determine the target discriminative speech features corresponding to the benchmark optimized audio; Based on the correspondence between the distinctive speech features and the operating parameters of the sound sensing device, the target operating parameters corresponding to the target distinctive speech features are determined; Adjust the target operating parameters to optimize the sound sensing device.
[0006] Exemplarily, the method further includes: Play sample audio to the user; Monitor users' understanding of sample audio to obtain test data; When the test data does not match the sample audio, the sample audio is determined to be unrecognized audio. Identify the unidentified distinctive speech features corresponding to the unidentified audio; Determine the correlation between the unrecognized distinctive speech features and the operating parameters; For each of the unrecognized distinctive speech features, the operating parameters associated with the unrecognized distinctive speech feature are determined as the operating parameters corresponding to the unrecognized distinctive speech feature.
[0007] For example, determining the correlation between the unrecognized distinctive speech features and the operating parameters includes: Each of the aforementioned operating parameters is changed serially. After each change of the operating parameters, the unrecognized audio corresponding to the unrecognized distinctive speech features is played repeatedly, and user feedback data is obtained. The correlation between the changed operating parameters and the unrecognized distinctive speech features is determined based on the feedback data.
[0008] For example, the response information is audio feedback repeated by the user; Determining whether the response information matches the test audio includes: Convert the audio feedback into feedback text; Determine whether the feedback text is the same as the test text corresponding to the test audio; If they are the same, then the response information matches the test audio; Otherwise, the response information does not match the test audio.
[0009] For example, before receiving the user's response information, the method further includes: Display multiple text options, including the correct text option corresponding to the test audio; The received user response information includes: In response to the user's selection action, determine the target text option from the plurality of text options; Determining whether the response information matches the test audio includes: When the target text option is the correct text option, it is determined that the response information matches the test audio; Otherwise, it is determined that the response information does not match the test audio.
[0010] For example, the number of test audios is multiple sets; wherein, the step of adjusting the target operating parameters is performed each time it is determined that the response information does not match the test audio; or, the step of adjusting the target operating parameters is performed after all multiple sets of test audios have been played.
[0011] For example, determining the target discriminative speech features corresponding to the benchmark optimized audio includes: The target discriminative speech features corresponding to the benchmark optimized audio are determined using the confusion error matrix.
[0012] According to another aspect of the present invention, an automatic tuning system is provided for a sound sensing device, the system comprising: An audio playback module is used to play test audio to the user, wherein the test audio is any one of words, syllables, music, speech, or vocal music; The monitoring module is used to receive user response information; The comparison module is used to determine whether the response information matches the test audio; when the response information does not match the test audio, the test audio is stored as a benchmark optimized audio; and the target discriminative speech features corresponding to the benchmark optimized audio are determined. The control module is used to determine the target operating parameters corresponding to the target distinctive speech features based on the correspondence between the distinctive speech features and the operating parameters of the sound sensing device; and to adjust the target operating parameters to optimize the sound sensing device.
[0013] According to another aspect of the present invention, an electronic device is provided, including a processor and a memory, wherein the memory stores a computer program, and the processor is used to execute the computer program to implement the method as described above.
[0014] According to another aspect of the present invention, a computer-readable storage medium is provided, which stores a computer program / instructions that, when executed by a processor, implement the method described above.
[0015] In the aforementioned technical solution, by playing diverse test audio such as words, syllables, music, speech, vocal music, and pure tones, the feature dimensions of different types of audio can be comprehensively covered, ensuring a comprehensive and thorough detection of the performance of the audio-sensing device in various auditory scenarios, including daily communication, music appreciation, and professional listening. When the user's response does not match the test audio, that audio is locked as the benchmark optimization audio, and its target distinctive speech features are further extracted precisely. Then, the parameters are adjusted according to the correspondence between the features and the device's operating parameters. This not only achieves a closed-loop optimization of "problem location - feature extraction - precise parameter tuning," but also enables customized optimization for personalized listening scenarios such as dialects, special accents, and specific music styles through targeted adaptation of distinctive speech features. This significantly improves the audio-sensing device's accuracy in reproducing complex audio, its ability to present details, and the clarity of the user's hearing. Meanwhile, the entire tuning process is automated, requiring no technical personnel assistance, significantly reducing the technical threshold and labor costs associated with tuning. The operation is simple and intuitive, requiring no complex auxiliary equipment. It is compatible with traditional on-site tuning scenarios and can flexibly adapt to remote manual or fully automated remote tuning needs, thus broadening its applicability. Furthermore, the solution focuses on the specific speech features causing current user recognition errors. Parameter adjustments avoid the blindness of generalized debugging, offering greater individual targeting. Personalized process optimization can be performed based on differences in auditory sensitivity and usage habits among different users, fully meeting the customized auditory needs of various users.
[0016] The above description is merely an overview of the technical solution of the present invention. In order to better understand the technical means of the present invention and to implement it in accordance with the contents of the specification, and in order to make the above and other objects, features and advantages of the present invention more apparent and understandable, specific embodiments of the present invention are described below. Attached Figure Description
[0017] The above and other objects, features, and advantages of the present invention will become more apparent from the more detailed description of the embodiments of the invention in conjunction with the accompanying drawings. The drawings are provided to further illustrate the embodiments of the invention and form part of the specification. They are used together with the embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings, the same reference numerals generally represent the same parts or steps.
[0018] Figure 1 A schematic flowchart of an automatic machine adjustment method according to an embodiment of the present invention is shown; Figure 2 A schematic block diagram of an automatic machine adjustment system according to an embodiment of the present invention is shown; Figure 3 A schematic block diagram of an electronic device according to an embodiment of the present invention is shown. Detailed Implementation
[0019] To make the objectives, technical solutions, and advantages of the present invention more apparent, exemplary embodiments according to the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are merely a part of the embodiments of the present invention, and not all of the embodiments of the present invention. It should be understood that the present invention is not limited to the exemplary embodiments described herein. Based on the embodiments of the present invention described herein, all other embodiments obtained by those skilled in the art without inventive effort should fall within the protection scope of the present invention.
[0020] A typical example of a sound-sensing device is the multi-channel cochlearimplant (CI) system. These CI systems consist of an external earpiece with a microphone and transmitter, a battery-powered in-ear or behind-the-ear speech processor, and an internal receiver and electrode array. The microphone detects sound information and sends it to the speech processor, which encodes the sound information into a digital signal. This information is then sent to the earpiece so that the transmitter can transmit electrical signals via radio frequency waves through the skin to the internal receiver located in the implant's mastoid bone. The receiver sends electrical impulses to the electrodes implanted in the cochlea, thereby stimulating the auditory nerve and enabling the implant recipient to receive the sensation of sound.
[0021] Multichannel CI systems use multiple sensors or electrodes. Each sensor is associated with a corresponding channel that carries a signal within a specific frequency range. Therefore, the sensitivity or gain perceived by the implantee can be varied for each channel independently of the other channels.
[0022] In recent years, hearing enhancement systems (CI) have made significant progress in improving the quality of life for people with severe hearing loss. CI systems have evolved from providing a minimum level of pitch response to allowing patients with implanted CI systems to recognize over 80% of words in test scenarios. Much of this improvement is based on advancements in speech coding technology. For example, the introduction of Chinese speech coding (MTone), continuous interleaved sampling (CIS), and HiResolution has helped improve the performance of CI systems and other digital hearing enhancement systems employing multichannel and / or speech processing technologies. Once a CI system is implanted, or if the user is wearing another type of digital hearing enhancement mechanism, appropriate speech coding and mapping strategies must be selected to enhance the performance of the CI system for daily operation. Mapping strategies (MAPs) refer to the adjustment of parameters corresponding to one or more independent channels of a multichannel CI system or other hearing enhancement system. The selection of each of these strategies typically occurs during an introductory period of approximately six or seven weeks, during which the hearing enhancement system is tuned. During this tuning period, users of such systems are asked to provide feedback on their perception of the device's performance. However, the tuning process is not specific to any particular user; rather, it is designed for "general users." More specifically, to create a mapping for the voice processor, the tuner first determines the electrical dynamic range of each electrode or sensor used. The programming system supplies current to each electrode via the CI system to obtain measurements of the electrical threshold (Level T) and maximum comfort level (Level C) defined by the device manufacturer. Level T, or the lowest stimulation level, is the gentlest current that allows the user to hear comfortably for 100% of the time. Level C is the maximum signal level that the user can comfortably hear for an extended period. The voice processor is then programmed or “mapped” using one of several coding strategies so that the current delivered to the implant will fall within this measured dynamic range, i.e., between Level T and Level C. After establishing the Level T and Level C and creating the mapping, the microphone is activated so that the patient can hear speech and sounds in the environment. From this point on, the tuning process continues as a traditional hearing test, where the hearing enhancement device user is asked to listen to tones at different frequencies and volumes in order to further adjust the gain of each channel within the established threshold range so that the patient can better hear a variety of tones at different volumes and frequencies. Therefore, current tuning practices focus on allowing the user to adapt to the signals generated by the hearing device.
[0023] The aforementioned tuning techniques meet the common needs of ordinary users. This approach is favored because designing an optimal mapping strategy (MAP) for an individual user involves an enormous amount of time and a vast number of potential variables. For example, when a user attempts to add subjective input to the tuning process of a hearing enhancement system, additional complexity arises because each change in the system's mapping requires the user to adjust to the new signal. Therefore, after a mapping change, the user may perceive an enhancement in their hearing, when in reality, they haven't adapted to the new mapping. By the time the user adapts, their hearing may have actually deteriorated.
[0024] Generally, the systems described above for tuning the sensory device are extremely expensive and typically unavailable to most patients / users / clients. Furthermore, users (such as audiologists) need to know how to connect the various device components for satisfactory operation. However, many users lack the technical expertise to set up and operate the device. Since the sensory device only requires occasional adjustments, the cost of acquiring the necessary equipment can be prohibitive. Therefore, the device is often placed in locations serving multiple users, such as the offices of doctors, audiologists, or other clinicians. However, traveling to these facilities for testing can be inconvenient or impossible (for example, some tests, such as the EABR test, are only available in a few locations with the necessary equipment and qualified personnel; if equipment malfunctions, these tests cannot be performed, rendering them impossible). Another drawback is that current tuning processes focus on pure-tone testing, while hearing-impaired individuals are actually exposed to speech / sound / music. Finally, the current manual tuning process remains cumbersome, making tuning extremely inconvenient for implant recipients in remote areas who require a combination of short-term automated tuning and long-term on-site tuning services. In view of this, the present invention provides an automated device setup method, system, electronic device, and storage medium. This method can effectively simplify the device setup process, reduce setup costs, and achieve automated, remote device setup. Furthermore, this method can be optimized for individual users, meeting the needs of different users. The method, system, electronic device, and storage medium are described in detail below.
[0025] According to one aspect of the present invention, an automatic tuning method is provided. The method is used for a sound-sensing device, which includes, but is not limited to, cochlear implants, hearing aids, etc.
[0026] Figure 1 A schematic flowchart of an automatic machine adjustment method according to an embodiment of the present invention is shown. Figure 1As shown, the method may include steps S110, S120, S130, S140, S150, S160 and S170.
[0027] In step S110, test audio is played to the user. The test audio can be any one of words, syllables, music, speech, vocal music, or pure tone.
[0028] In this example, the test audio can be played via a specific audio playback device or a user terminal device. For example, the test audio can be played using a user's mobile phone, tablet, etc., which allows the user to adjust the device in different locations without being limited by the device.
[0029] In this example, the test audio can be any of the following: words, syllables, music, speech, or vocal music. That is, it can perform traditional pure tone tests, or it can use words, syllables, music, speech, or vocal music for testing. This makes the test more similar to the user's actual usage environment, and the device can be tuned more accurately based on this test.
[0030] In step S120, the user's response information is received.
[0031] In this example, after each test audio playback, the user can repeat what they hear; this feedback is the response information. User response information includes, but is not limited to, voice and text, which will not be elaborated upon further.
[0032] In step S130, it is determined whether the response information matches the test audio.
[0033] In this example, text matching can be used to determine if the response information matches the content of the test audio; if they match, it's a match. Of course, a pre-trained neural network model can also be used to automatically compare the response information and the test audio, which will not be elaborated upon here.
[0034] In step S140, if the response information does not match the test audio, the test audio is stored as the baseline optimized audio.
[0035] It's understandable that when the response information matches the test audio, it indicates that the user's acquisition of the test audio using the sound-sensing device is effective. Conversely, when the response information does not match the test audio, it indicates a deviation in the user's acquisition of the test audio using the sound-sensing device. In this example, the test audio is directly stored as the baseline optimization audio, serving as the basis for tuning and optimization.
[0036] In step S150, the target discriminative speech features corresponding to the benchmark optimized audio are determined.
[0037] In this example, we consider automated tuning based on the distinctive features of speech. Two distinct feature sets have been proposed by the academic community through the study of distinctive features of speech. The first set, proposed by Chompsky and Halle (1968), is based on articulation location, while the other set, proposed by Jakobson, Fan, and Halle (1963), is based on the acoustic properties of various speech sounds. These attributes describe a small set of contrasting acoustic properties that are perceptually relevant to speech discrimination. More specifically, the different distinctive features and their potential acoustic correlations can be broadly categorized into three types: primary source features, secondary consonant source features, and resonance properties. Primary source features can be further distinguished based on whether the speech is audible or silent. Aaudible speech corresponds to sounds associated with vowels; therefore, such speech corresponds to a single periodic source, and the start of the speech is not abrupt. Silent speech is the opposite of audible speech. Primary source features can also be classified based on whether the speech is consonant or non-consonant. Consonant speech corresponds to sounds associated with consonants. Such speech is characterized by the presence of zeros in the associated sound spectrum. Secondary consonant features can be further distinguished based on whether the speech is interrupted or continuous. Continuous speech is also called semivowels because they have similar timbre. Continuous speech has little or no friction when air flows freely from the speaker's mouth. Continuous speech is produced when the vocal tract is not fully closed. In contrast, interrupted speech ends abruptly. Secondary consonant features can also be distinguished based on whether the speech is blocked or unblocked. Blocked speech is characterized by abrupt termination rather than gradual decay, while unblocked speech is characterized by gradual decay. In addition, secondary consonant features can be characterized as harsh or rounded. The former usually has an irregular waveform, while the latter usually has a smooth waveform. Secondary consonant features characterized as rounded also have a broader autocorrelation function relative to the corresponding normalized harsh features. Secondary consonant features can also be classified based on whether the sound is voiced or unvoiced. Resonance features can be further distinguished based on whether the speech is compact or diffuse. Compact features are associated with a sound having a relative dominance of a central format region, while diffuse features mean a sound with one or more non-central formats. Resonance features can also be characterized as deep or sharp. Speech characterized by a deep, resonant tone is predominantly low-frequency, while speech characterized by a sharp, piercing tone is predominantly high-frequency. Furthermore, resonance features can be characterized as flat or ordinary, depending on the presence of downward shifts in certain or all formats, typically related to vowel reduction and reduction of the speaker's lip and mouth. Resonance features can also be further characterized as sharp or flat, the latter representing the rising tone in the second and / or higher formats. Additionally, resonance features can be characterized as tense or relaxed, depending on the amount and duration of sound energy. Resonance features can also be classified based on whether the speech has nasality or nasal murmurs.
[0038] The aforementioned distinguishing features of speech and their potential acoustic correlations are merely examples of the many different distinguishing features of speech. Based on the description herein, the present invention can determine relationships with one or more adjustable operating parameters of a sound-sensing device. Therefore, regardless of the specific distinguishing features of speech in a particular context, the present invention can determine the relationship between distinguishing speech features and adjustable operating parameters of a sound-sensing device to enhance the capabilities of a particular sound-sensing device for a particular user.
[0039] In some implementations of this example, determining the target discriminative speech features corresponding to the benchmark optimized audio includes: using a confusion error matrix (CEM) to determine the target discriminative speech features corresponding to the benchmark optimized audio. Specifically, the CEM can record logs of the test audio played to the user and the user's response information; that is, it stores the test audio and the user's response information (which can be the response information itself or a text representation of the response information). After determining the benchmark optimized audio, the CEM matrix can be used to determine the target discriminative speech features corresponding to that benchmark optimized audio. Alternatively, the discriminative speech features corresponding to each test audio can be predetermined. When any test audio is determined as the benchmark optimized audio, the discriminative speech features corresponding to that benchmark optimized audio can be determined based on the predetermined correspondence.
[0040] It is understandable that various test audio recordings can be associated with distinctive speech features. Therefore, by comparing the played test audio with the user's response information, it can be determined whether the user can accurately perceive each distinctive speech feature. User error feedback (i.e., response information that does not match the test audio) indicates that, despite the use of a sound-sensing device, one or more distinctive speech features represented by incorrectly pronounced words or syllables were not correctly perceived. In practical scenarios, the use of distinctive speech features is flexible; any one of the various speech features used in this invention can be used, or a set of one or more feature sets can be used. Furthermore, the distinctive features of various languages (English / Chinese) and dialects (Mandarin / Cantonese) have both commonalities and unique characteristics. This also facilitates the optimization of dialect listening comprehension.
[0041] In this paper, distinctive phonetic features include, but are not limited to: interrupted vs. continuous, compact vs. diffuse, low vs. high, tense vs. relaxed, harsh vs. rounded, stressed vs. rhythmic, retroflex vs. non-retroflex, nasal vs. lateral, front nasal vs. back nasal, the four tones of Mandarin, male vs. female voices, and phonetic features inherent in local dialects.
[0042] In step S160, the target operating parameters corresponding to the target distinctive speech features are determined based on the correspondence between the distinctive speech features and the operating parameters of the sound sensing device.
[0043] It is understood that the operating parameters of the sound sensing device in this article refer to the adjustable parameters of the sound sensing device. Taking the CI system as an example, the operating parameters in the CI system are shown in Table 1 below.
[0044] Table 1 CI System Operating Parameters
[0045] Generally, for CI systems, stimulation strategies mainly include the MTone strategy based on frequency resolution, the L-CIS strategy based on time resolution, or a combination of both. Before measuring the T / C value, the stimulation rate and pulse width must be specified, and the stimulation rate × number of stimulation channels must be ≤ 30 kHz. After the T / C value is obtained, the following parameters can be adjusted: 1. Number of active channels and number of stimulation channels. 2. Spectrum allocation. 3. Global / single channel gain parameters (corresponding to channel gain in the table). 4. Global T / C value. 5. Current mapping curve parameters (corresponding to Q value). 6. Baseline parameter, i.e., the minimum sound pressure level that can produce stimulation. Decreasing this parameter is equivalent to increasing the input dynamic range of the CI system, thus making weaker sounds audible; increasing this parameter has the opposite effect, that is, even weaker sounds are ignored. 7. Percentage of random variation in pulse frequency (corresponding to Jitter value). 8. Channel stimulation order (corresponding to Order value) to accommodate abnormal frequency response cases.
[0046] As mentioned above, the perception of distinctive speech features is related to speech discrimination. That is, by adjusting the operating parameters corresponding to the distinctive speech features, the user's auditory discrimination of the speech corresponding to the distinctive speech features can be improved, thereby more accurately improving the user's hearing.
[0047] In step S170, the target operating parameters are adjusted to optimize the sound sensing device.
[0048] In the aforementioned technical solution, by playing diverse test audio such as words, syllables, music, speech, vocal music, and pure tones, the feature dimensions of different types of audio can be comprehensively covered, ensuring a comprehensive and thorough detection of the performance of the audio-sensing device in various auditory scenarios, including daily communication, music appreciation, and professional listening. When the user's response does not match the test audio, that audio is locked as the benchmark optimization audio, and its target distinctive speech features are further extracted precisely. Then, the parameters are adjusted according to the correspondence between the features and the device's operating parameters. This not only achieves a closed-loop optimization of "problem location - feature extraction - precise parameter tuning," but also enables customized optimization for personalized listening scenarios such as dialects, special accents, and specific music styles through targeted adaptation of distinctive speech features. This significantly improves the audio-sensing device's accuracy in reproducing complex audio, its ability to present details, and the clarity of the user's hearing. Meanwhile, the entire tuning process is automated, requiring no technical personnel assistance, significantly reducing the technical threshold and labor costs associated with tuning. The operation is simple and intuitive, requiring no complex auxiliary equipment. It is compatible with traditional on-site tuning scenarios and can flexibly adapt to remote manual or fully automated remote tuning needs, thus broadening its applicability. Furthermore, the solution focuses on the specific speech features causing current user recognition errors. Parameter adjustments avoid the blindness of generalized debugging, offering greater individual targeting. Personalized process optimization can be performed based on differences in auditory sensitivity and usage habits among different users, fully meeting the customized auditory needs of various users.
[0049] For example, the method further includes: playing sample audio to a user; monitoring the user's understanding of the sample audio to obtain test data; determining the sample audio as unrecognized audio when the test data does not match the sample audio; determining unrecognized distinctive speech features corresponding to the unrecognized audio; determining the correlation between the unrecognized distinctive speech features and operating parameters; and for each unrecognized distinctive speech feature, determining the operating parameters associated with that unrecognized distinctive speech feature as the operating parameters corresponding to that unrecognized distinctive speech feature. This example can be referred to as the steps of determining the relationship between distinctive speech features and operating parameters.
[0050] In this example, the steps of playing sample audio to the user and monitoring the user's understanding of the sample audio to obtain test data are implemented in a manner similar to steps S110 and S120. When the test data does not match the sample audio, the sample audio is determined to be unrecognized audio. The specific implementation of determining the unrecognized distinctive speech features corresponding to the unrecognized audio is similar to steps S140 and S150, and will not be elaborated further. The sample audio may include, but is not limited to, the content of the test audio, and will not be elaborated further.
[0051] After identifying unrecognized distinctive speech features, the correlation between these features and operating parameters can be further determined. This correlation can be determined empirically or iteratively. For example, each parameter can be changed sequentially, and the correlation between each parameter and the distinctive speech feature can be determined based on whether the user's perception improves. Taking a low-pitched voice feature as an example, since the low-pitched voice feature is dominated by energy in the low-frequency range of speech, the channel gain parameter responsible for low frequencies can be changed to test the user's perception of the low-pitched voice. The perception result can then be used to determine whether the low-pitched voice feature is related to the channel gain parameter. Of course, field theory modeling (MFT), genetic algorithms, neural networks, fuzzy logic, etc., can also be used to determine the correlation between parameters and distinctive speech features, which will not be elaborated here.
[0052] It is understandable that when determining the relationship between distinctive speech features and operating parameters, testing can be conducted simultaneously by multiple users. This facilitates the rapid identification of the relationship between different distinctive speech features and operating parameters corresponding to different sample audio files, thereby enabling the rapid and accurate establishment of a relationship table between distinctive speech features and operating parameters. This relationship table can provide an accurate basis for determining the target operating parameters during actual sound tuning.
[0053] The aforementioned technical solution plays sample audio and monitors user comprehension, identifying mismatched sample audio as unrecognized audio. It then further mines the correlation between the unrecognized distinctive speech features and device operating parameters, selecting relevant parameters as associated parameters for the corresponding features. This helps establish a more comprehensive feature-parameter correspondence library, providing a more accurate basis for subsequent sound tuning optimization, reducing the randomness of parameter adjustments, and improving the ability to recognize and reproduce different audio in different scenarios.
[0054] For example, determining the correlation between unrecognized distinctive speech features and operating parameters includes: changing each operating parameter serially; after each change of operating parameters, repeatedly playing the unrecognized audio corresponding to the unrecognized distinctive speech features and obtaining user feedback data; and determining the correlation between the changed operating parameters and the unrecognized distinctive speech features based on the feedback data.
[0055] User feedback data refers to the content repeated by the user after hearing the unrecognized audio again. This content can be either speech or text, and this article does not impose any restrictions on it.
[0056] In this example, after receiving feedback data, it's possible to determine whether the user's hearing has improved. For instance, if the sample audio is "sam" and the user's first test result is "sham," then "sham" can be identified as unrecognized audio. After changing the operating parameters, "sam" can be played again. If the user's feedback result is "sam" this time, it indicates that the user's hearing has improved. In this case, it can be determined that the changed operating parameters are related to the unrecognized distinctive speech features.
[0057] The above technical solution determines the correlation between each operating parameter and unrecognized distinctive speech features by sequentially changing each parameter and combining it with user feedback data. This allows for the individual examination of the impact of each parameter on specific speech features. This approach clearly quantifies the degree of correlation between a single parameter and a feature, avoids interference caused by adjusting multiple parameters simultaneously, ensures the accuracy of correlation judgment, and provides reliable basic data for subsequent targeted parameter optimization.
[0058] For example, the response information is audio feedback repeated by the user; determining whether the response information matches the test audio includes: converting the audio feedback into feedback text; determining whether the feedback text is the same as the test text corresponding to the test audio; if they are the same, the response information matches the test audio; otherwise, the response information does not match the test audio.
[0059] In this example, we consider directly receiving the user's repeated speech as the response information. That is, the user can directly repeat the perceived audio aloud. This clearly reflects the user's auditory perception. In subsequent steps, we consider directly determining whether the response information matches the test audio through text comparison. This transforms the auditory response into intuitive textual differences, objectively quantifying the matching result. This approach reduces subjective judgment errors, improves the accuracy and consistency of the match judgment between the response and the test audio, and provides a reliable trigger for subsequent sound tuning optimization.
[0060] For example, before receiving the user's response information, the method further includes: displaying multiple text options, wherein the multiple text options include a correct text option corresponding to the test audio; receiving the user's response information, including: in response to the user's selection operation, determining a target text option among the multiple text options; determining whether the response information matches the test audio, including: if the target text option is a correct text option, determining that the response information matches the test audio; otherwise, determining that the response information does not match the test audio.
[0061] Multiple text options can be displayed on the human-computer interaction page. This page may include, but is not limited to, mobile phone screens, computer screens, and screens of specific equipment in professional sound mixing environments, etc., which will not be elaborated upon further. The multiple text options include both correct and incorrect options. For example, if the test audio being played is "Sam," possible choices could include the correct option "Sam" and the incorrect option "sham."
[0062] The above solution simplifies user feedback and lowers the barrier to entry by displaying multiple text options containing the correct choice, using the user's selected target text option as the response information and determining the match. This is particularly suitable for users who cannot express themselves verbally. Furthermore, the text-based matching judgment is intuitive and efficient, quickly determining the user's recognition of the test audio and providing timely trigger signals for sound device tuning optimization, especially suitable for scenarios where audio repetition is inconvenient.
[0063] For example, the number of test audios is multiple sets; wherein the step of adjusting the target operating parameters is performed each time it is determined that the response information does not match the test audio; or, the step of adjusting the target operating parameters is performed after all multiple sets of test audios have been played.
[0064] In some implementations of this example, the operating parameters can be adjusted after each playback of the test audio if the test audio is not correctly detected. This can quickly improve the ability to recognize specific distinctive speech features.
[0065] In other implementations of this example, after all test audio has been played, the target discriminative speech features corresponding to each audio element can be optimized based on all established benchmarks to determine the target operating parameters that need adjustment. This unified adjustment approach integrates multiple sets of test data, avoiding the bias of a single adjustment and achieving more comprehensive and balanced parameter optimization. Furthermore, it improves tuning efficiency compared to word-by-word adjustments. In one specific implementation, the operating parameters corresponding to each target discriminative speech feature can be output to a professional who can determine the target operating parameters to be adjusted. Of course, this process can also be performed using a pre-trained neural network model, i.e., specifying the parameter adjustment strategy through the neural network model, which will not be elaborated upon here.
[0066] In some embodiments, after adjustment is complete, the baseline optimized audio can be played repeatedly, and the tuning effect can be determined based on user feedback. If the user feedback is exactly the same as the baseline optimized audio, the tuning ends; otherwise, the operating parameters continue to be adjusted based on the baseline optimized audio corresponding to the error feedback.
[0067] In this example, the parameter adjustment strategy can specify one or more operating parameters of the hearing device to be changed to correct perceived hearing loss. It is worth noting that the implementation of the strategy can be limited to situations where the user misidentifies a test word or syllable. For example, if a test word with a low-pitched sound characteristic is misidentified, a strategy aimed at correcting this misperception can be determined. Since low-pitched sound characteristics are dominated by energy in the low-frequency range of speech, the implemented strategy can include adjusting parameters of the hearing device that affect low-frequency processing. For example, the strategy could specify that the mapping should be updated to increase the channel gain parameter responsible for low frequencies. Additionally, the frequency range of each channel of the hearing device can be varied. It is understood that, due to the non-linear nature of user hearing, more than one parameter can be adjusted during parameter adjustment, and the adjustment of one parameter can be offset by adjusting another parameter (i.e., increasing or decreasing it), which will not be elaborated further.
[0068] For example, before storing the test audio as the benchmark optimized audio, the method further includes: repeatedly playing the test audio and re-acquiring the response information; determining whether the re-acquiring response information matches the test audio; and if the re-acquiring response information does not match the test audio, performing the step of storing the test audio as the benchmark optimized audio. This solution, by repeatedly playing the test audio and re-acquiring the user's response information before storing it as the benchmark optimized audio, effectively eliminates misjudgments caused by user subjective errors (such as distraction, slips of the tongue, or misoperation) or accidental environmental interference (such as transient noise) in a single test. This design ensures that the locked benchmark optimized audio is audio that users cannot accurately identify due to limitations in the performance of the sound-sensing device, avoiding including mismatches caused by non-device factors in the optimization scope, and significantly improving the accuracy and reliability of benchmark optimized audio selection. Based on the benchmark optimization of the selected audio, the target distinctive speech features are extracted and the operating parameters are adjusted. This allows subsequent tuning actions to better match the actual performance defects of the equipment, reduce invalid parameter adjustments, improve the overall tuning efficiency and optimization effect, and at the same time avoid unreasonable adjustments to equipment parameters due to misjudgment, ensuring that the original good performance of the sound sensing equipment is not affected.
[0069] According to another aspect of the present invention, an automatic adjustment system is provided. This system is used in sound-sensing devices. Figure 2 A schematic block diagram of an automatic machine adjustment system according to an embodiment of the present invention is shown. Figure 2 As shown, the system includes: an audio playback module 210, a monitoring module 220, a comparison module 230, and a control module 240.
[0070] The audio playback module 210 is used to play test audio to the user. The test audio can be any of the following: words, syllables, music, speech, or vocal music.
[0071] The monitoring module 220 is used to receive user response information.
[0072] The comparison module 230 is used to determine whether the response information matches the test audio; when the response information does not match the test audio, the test audio is stored as the benchmark optimized audio; and the target distinctive speech features corresponding to the benchmark optimized audio are determined.
[0073] The control module 240 is used to determine the target operating parameters corresponding to the target distinctive speech features based on the correspondence between the distinctive speech features and the operating parameters of the sound sensing device; and to adjust the target operating parameters to optimize the sound sensing device.
[0074] In some embodiments, the system may further include a knowledge base 250 for storing the correspondence between distinctive speech features and operating parameters of the sound sensing device.
[0075] In this example, the audio playback module can be any of various analog or digital sound playback systems, a computer system storing digitized audio, or a text-to-speech (TTS) system capable of generating synthesized speech from input or stored text. The audio playback module can either simply play recorded and / or generated audio to the user, or it can be connected to the hearing device being tested via some communication link. For example, in the case of a selected digital hearing aid and / or cochlear implant system, an A / C input jack may be included in the hearing device, through which the audio playback module connects directly to the user's hearing device to play audio without generating sound through an acoustic transducer. The audio playback module can be configured to play any of a variety of different test words and / or syllables (test audio) to the user, or it can play media such as tapes or CDs, or load test audio into a computer system for playback, or it can automatically generate synthesized speech that mimics the test speech.
[0076] The monitoring module can be a person who records various test words / syllables provided to the user and the user's responses, or it can be a speech recognition system configured to recognize user responses by speech or convert user responses into text. It can also be implemented using the microphone of the user's terminal device.
[0077] The comparison module can be based on the CEM principle. This module can acquire the played test audio and the user's response information.
[0078] The control module can be a computer or an information processing system that coordinates the operations of various system components. The control module can access the confusion error matrix (CEM) and mapping parameter knowledge base to determine the optimal mapping strategy for the hearing device. More specifically, the control module can correctly determine the operating parameters of the hearing device based on the user's response to the test audio. In addition to initializing and controlling the operation of the system components, the control module can communicate with the user's hearing device. The corresponding control system provides an operating interface through which the parameters of the user's hearing device are automatically adjusted. The control logic is based on the CEM matrix and mapping parameter knowledge base obtained during the testing process. Finally, the optimal mapping strategy determined by the control system can be implemented in the user's hearing device. The modules in this example can be included in one or more computer systems, thus offering the advantage of flexible deployment.
[0079] In some embodiments, some modules of the system (such as the comparison module and the control module) can be deployed on a cloud server, which can reduce the weight of local devices and lower equipment costs. At the same time, deploying them in the cloud can acquire valuable data from different users over a long period of time. When using neural network models such as field theory modeling to determine the relationship between parameters and distinctive speech features, this data helps to quickly establish and develop effective neural network models. The more effective data acquired, the more mature the neural network model, and the shorter the system setup time.
[0080] According to another aspect of the present invention, an electronic device is also provided. Figure 3 A schematic block diagram of an electronic device according to an embodiment of the present invention is shown. Figure 3 As shown, the electronic device 300 includes a processor 310 and a memory 320. The memory 320 stores a computer program, which the processor 310 executes to implement the method described above.
[0081] According to another aspect of the present invention, a computer-readable storage medium is also provided. The storage medium stores a computer program / instructions that, when executed by a processor, implement the method described above. The storage medium may, for example, include a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a portable compact disc read-only memory (CD-ROM), a USB memory, or any combination of the above storage media. The computer-readable storage medium may be any combination of one or more computer-readable storage media.
[0082] Those skilled in the art will readily understand the implementation structure, working principle, and beneficial effects of the system, electronic device, and computer-readable storage medium by reading the above methods. For the sake of brevity, further details will not be elaborated upon here.
[0083] Although exemplary embodiments have been described herein with reference to the accompanying drawings, it should be understood that the above exemplary embodiments are merely illustrative and are not intended to limit the scope of the invention. Various changes and modifications can be made therein by those skilled in the art without departing from the scope and spirit of the invention. All such changes and modifications are intended to be included within the scope of the invention as claimed in the appended claims.
[0084] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0085] In the several embodiments provided by this invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed.
[0086] Numerous specific details are set forth in the specification provided herein. However, it will be understood that embodiments of the invention may be practiced without these specific details. In some instances, well-known methods, structures, and techniques have not been shown in detail so as not to obscure the understanding of this specification.
[0087] Similarly, it should be understood that, in order to streamline the invention and aid in understanding one or more of the various aspects of the invention, features of the invention are sometimes grouped together in a single embodiment, figure, or description thereof in the description of exemplary embodiments of the invention. However, this approach should not be construed as reflecting an intention that the claimed invention requires more features than are expressly recited in each claim. Rather, as reflected in the corresponding claims, its inventive point lies in solving the corresponding technical problem with fewer features than all of those in a single disclosed embodiment. Therefore, the claims following the detailed description are hereby expressly incorporated into that detailed description, wherein each claim itself is a separate embodiment of the invention.
[0088] Those skilled in the art will understand that, apart from the mutual exclusion of features, all features disclosed in this specification (including the accompanying claims, abstract, and drawings) and all processes or elements of any method or apparatus so disclosed may be combined in any combination. Unless otherwise expressly stated, each feature disclosed in this specification (including the accompanying claims, abstract, and drawings) may be replaced by an alternative feature that serves the same, equivalent, or similar purpose.
[0089] Furthermore, those skilled in the art will understand that although some embodiments described herein include certain features but not others included in other embodiments, combinations of features from different embodiments are intended to be within the scope of the invention and form different embodiments. For example, in the claims, any of the claimed embodiments can be used in any combination.
[0090] The various component embodiments of the present invention can be implemented in hardware, or as software modules running on one or more processors, or a combination thereof. Those skilled in the art will understand that microprocessors or digital signal processors (DSPs) can be used in practice to implement some or all of the functions of some modules in the electronic device according to embodiments of the present invention. The present invention can also be implemented as an apparatus program (e.g., a computer program and computer program product) for performing some or all of the methods described herein. Such programs implementing the present invention can be stored on a computer-readable medium or can be in the form of one or more signals. Such signals can be downloaded from an Internet website, provided on a carrier signal, or provided in any other form.
[0091] It should be noted that the above embodiments are illustrative of the invention and not restrictive, and that those skilled in the art can devise alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses should not be construed as limiting the claims. The word "comprising" does not exclude the presence of elements or steps not listed in the claims. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The invention can be implemented by means of hardware comprising several different elements and by means of a suitably programmed computer. In the unit claims enumerating several means, several of these means may be embodied by the same item of hardware. The use of the words first, second, and third, etc., does not indicate any order. These words can be interpreted as names.
[0092] The above description is merely a specific embodiment of the present invention or an explanation of that embodiment. The scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. The scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. An automatic machine adjustment method, characterized in that, For use in a sound-sensing device, the method includes: Play test audio to the user, wherein the test audio is any one of words, syllables, music, speech, vocal music, or pure tone; Receive user response information; Determine whether the response information matches the test audio; When the response information does not match the test audio, the test audio is stored as a baseline optimized audio. Determine the target discriminative speech features corresponding to the benchmark optimized audio; Based on the correspondence between the distinctive speech features and the operating parameters of the sound sensing device, the target operating parameters corresponding to the target distinctive speech features are determined; Adjust the target operating parameters to optimize the sound sensing device.
2. The method according to claim 1, characterized in that, The method further includes: Play sample audio to the user; Monitor users' understanding of sample audio to obtain test data; When the test data does not match the sample audio, the sample audio is determined to be unrecognized audio. Identify the unidentified distinctive speech features corresponding to the unidentified audio; Determine the correlation between the unrecognized distinctive speech features and the operating parameters; For each of the unrecognized distinctive speech features, the operating parameters associated with the unrecognized distinctive speech feature are determined as the operating parameters corresponding to the unrecognized distinctive speech feature.
3. The method according to claim 2, characterized in that, Determining the correlation between the unrecognized distinctive speech features and the operating parameters includes: Each of the aforementioned operating parameters is changed serially. After each change of the operating parameters, the unrecognized audio corresponding to the unrecognized distinctive speech features is played repeatedly, and user feedback data is obtained. The correlation between the changed operating parameters and the unrecognized distinctive speech features is determined based on the feedback data.
4. The method according to claim 1, characterized in that, The response information is audio feedback repeated by the user; Determining whether the response information matches the test audio includes: Convert the audio feedback into feedback text; Determine whether the feedback text is the same as the test text corresponding to the test audio; If they are the same, then the response information matches the test audio; Otherwise, the response information does not match the test audio.
5. The method according to claim 1, characterized in that, Before receiving the user's response information, the method further includes: Display multiple text options, including the correct text option corresponding to the test audio; The received user response information includes: In response to the user's selection action, determine the target text option from the plurality of text options; Determining whether the response information matches the test audio includes: When the target text option is the correct text option, it is determined that the response information matches the test audio; Otherwise, it is determined that the response information does not match the test audio.
6. The method according to claim 1, characterized in that, The number of test audios is multiple sets; wherein, the step of adjusting the target operating parameters is performed each time it is determined that the response information does not match the test audio; or, the step of adjusting the target operating parameters is performed after all multiple sets of test audios have been played.
7. The method according to any one of claims 1-6, characterized in that, Determining the target discriminative speech features corresponding to the benchmark optimized audio includes: The target discriminative speech features corresponding to the benchmark optimized audio are determined using the confusion error matrix.
8. An automatic machine adjustment system, characterized in that, For a sound sensing device, the system includes: An audio playback module is used to play test audio to the user, wherein the test audio is any one of words, syllables, music, speech, or vocal music; The monitoring module is used to receive user response information; The comparison module is used to determine whether the response information matches the test audio; when the response information does not match the test audio, the test audio is stored as a benchmark optimized audio; and the target discriminative speech features corresponding to the benchmark optimized audio are determined. The control module is used to determine the target operating parameters corresponding to the target distinctive speech features based on the correspondence between the distinctive speech features and the operating parameters of the sound sensing device; and to adjust the target operating parameters to optimize the sound sensing device.
9. An electronic device, characterized in that, It includes a processor and a memory, wherein the memory stores a computer program, and the processor is used to execute the computer program to implement the method as claimed in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The system stores a computer program / instructions that, when executed by a processor, implement the method as described in any one of claims 1-7.