An audio device adaptive control method and apparatus

By analyzing the background sound and speech loudness characteristics of the microphone audio signal during silent periods, and combining the stability of speech features and link transmission characteristics, the audio processing parameters are dynamically adjusted, solving the problem of repeated microphone signal processing in online education systems and improving audio quality and teaching interaction.

CN120972573BActive Publication Date: 2026-04-17深圳市辰益兴电子有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
深圳市辰益兴电子有限公司
Filing Date
2025-09-10
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing online education systems cannot identify whether microphone audio signals have been externally processed, leading to repeated processing that causes sound distortion and affects the quality of teaching interaction.

Method used

By analyzing the background sound characteristics and speech loudness characteristics of the microphone audio signal during silent periods, and combining the stability of speech characteristics with the transmission characteristics of the audio signal link, it is possible to determine whether the audio signal has undergone external processing and dynamically adjust the audio processing parameters.

Benefits of technology

It effectively avoids the audio quality degradation caused by repeated processing, improves the clarity of audio communication and user experience, and enhances teaching effectiveness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120972573B_ABST
    Figure CN120972573B_ABST
Patent Text Reader

Abstract

This application provides an adaptive control method and apparatus for audio devices, relating to the field of audio signal processing technology. The key technical points are: acquiring a microphone audio signal and identifying silent periods from the microphone audio signal; analyzing the background sound features and speech loudness features of the microphone audio signal during the silent periods; determining whether the microphone audio signal has undergone external processing based on the background sound features and the speech loudness features; and adjusting the audio processing parameters of the microphone audio signal based on the determination result. The adaptive control method and apparatus for audio devices provided by this application has the advantage of avoiding audio quality degradation caused by repeated processing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of audio signal processing technology, and more specifically, to an adaptive control method and apparatus for audio devices. Background Technology

[0002] In modern online education scenarios, intelligent audio systems need to adaptively control audio devices based on personalized teaching syllabus data. However, when users preprocess audio signals using third-party software, existing systems cannot detect this external processing, leading to repeated processing and sound distortion. Specifically, after receiving audio signals that have been denoised or amplified by third-party software, the system still performs secondary processing according to the original signal characteristics, resulting in excessive attenuation of vocal details or abnormal signal levels, producing hollow sound effects, digital clipping, and other phenomena. This superimposed processing not only destroys the naturalness of the audio but also hinders communication between teachers and students, severely impacting the quality of teaching interaction. The lack of a detection mechanism for the processing status of the audio signal link in existing technologies, and the inability to dynamically adjust processing strategies, has become a key bottleneck restricting the effectiveness of intelligent audio systems.

[0003] To address the aforementioned problems, this application proposes a new solution. Summary of the Invention

[0004] The purpose of this application is to provide an adaptive control method and apparatus for audio devices, which has the advantage of being able to effectively identify whether the microphone audio signal has been externally processed and dynamically adjust the audio processing parameters according to the identification result, thereby avoiding repeated processing that leads to a decrease in audio quality.

[0005] This application provides a first aspect: an adaptive control method for audio devices, the technical solution of which is as follows:

[0006] Acquire microphone audio signals and identify silent periods from the microphone audio signals;

[0007] During the silent period, the background sound features and speech loudness features of the microphone audio signal are analyzed;

[0008] Based on the background sound characteristics and the speech loudness characteristics, determine whether the microphone audio signal has undergone external processing;

[0009] Based on the judgment result, adjust the audio processing parameters of the microphone audio signal.

[0010] Furthermore, this application also proposes that the step of determining whether the microphone audio signal has undergone external processing based on the background sound features and the speech loudness features includes:

[0011] Evaluate the stability of the speech features of the microphone audio signal;

[0012] During the silent period, the transmission characteristics of the microphone audio signal link are measured by injecting a reference signal and receiving its return.

[0013] Based on the background sound features, the speech loudness features, the stability of the speech features, and the transmission characteristics of the microphone audio signal link, it is determined whether the microphone audio signal has undergone external processing.

[0014] Furthermore, this application also proposes a step of measuring the transmission characteristics of the microphone audio signal link during the silent period by injecting a reference signal and receiving its return, comprising:

[0015] During the silent period, a low-level reference signal with spectral characteristics that change non-periodically over time is generated to avoid being identified and suppressed by external preprocessing applications.

[0016] Inject the reference signal into the microphone audio signal link;

[0017] Receive the return of the reference signal;

[0018] The transmission characteristics of the microphone audio signal link are measured based on the spectral distortion and energy attenuation of the returned signal.

[0019] Furthermore, this application also proposes a step for determining whether the microphone audio signal has undergone external processing based on the background sound features, the speech loudness features, the stability of the speech features, and the transmission characteristics of the microphone audio signal link, including:

[0020] The transmission characteristics of the microphone audio signal link are evaluated to form a preliminary indication as to whether the microphone audio signal has undergone external processing;

[0021] Based on the preliminary indication and the stability of the speech features, the preliminary indication is verified to form a verified judgment.

[0022] Continuously monitor the background sound features and the speech loudness features to obtain auxiliary judgment information;

[0023] Based on the verified judgment, the auxiliary judgment information, and the transmission characteristics of the microphone audio signal link, it is confirmed whether the microphone audio signal has undergone external processing.

[0024] Furthermore, this application also proposes that, during the silent period, the step of analyzing the background sound characteristics of the microphone audio signal includes:

[0025] During the silent period, the background sound characteristics of the microphone audio signal are analyzed, and the analysis includes:

[0026] Analyze the spectral variation patterns of the background sound to identify unnatural spectral evolution;

[0027] Analyze the energy fluctuation patterns of the background sound to identify unnatural energy fluctuations;

[0028] Analyze and identify the instantaneous events of the background sound to distinguish its source.

[0029] Furthermore, this application also proposes that the step of analyzing the speech loudness characteristics of the microphone audio signal includes:

[0030] The loudness of the speech portion of the microphone audio signal is measured to obtain the loudness value of the speech portion;

[0031] Based on the loudness value, determine the loudness variation range of the speech portion;

[0032] Analyze the distribution of the loudness variation range.

[0033] Furthermore, this application also proposes that the step of analyzing the distribution of the loudness variation range includes:

[0034] The distribution of the loudness variation range is analyzed, and the analysis includes:

[0035] Based on the evolution characteristics of the loudness variation range over time, evaluate the smoothness or abruptness of the loudness variation range;

[0036] Based on the smoothness or abruptness, identify the unnatural evolution of the distribution of the loudness variation range.

[0037] Furthermore, this application also proposes that the step of adjusting the audio processing parameters of the microphone audio signal based on the judgment result includes:

[0038] When the judgment result indicates the presence of external processing, the audio processing parameters of the microphone audio signal are determined based on the characteristics of the external processing reflected by the judgment result.

[0039] Continuously acquire the external processing and judgment results of the microphone audio signal;

[0040] Based on the continuously acquired external processing judgment results, the updated values ​​of the audio processing parameters of the microphone audio signal are calculated;

[0041] The audio processing parameters of the microphone audio signal are gradually adjusted to adapt to the characteristics of external processing that change over time.

[0042] Furthermore, this application also proposes a step of progressively adjusting the audio processing parameters of the microphone audio signal, including:

[0043] Acquire the voice activity status of the microphone audio signal;

[0044] Obtain the background sound level of the microphone audio signal;

[0045] The adjustment step size of the audio processing parameters is determined based on the voice activity state or the background sound level.

[0046] The adjustment cycle of the audio processing parameters is determined based on the voice activity state or the background sound level; and

[0047] The audio processing parameters of the microphone audio signal are gradually adjusted according to the adjustment step size and the adjustment period.

[0048] Furthermore, this application also proposes an audio device adaptive control device, which includes:

[0049] The acquisition module is used to acquire microphone audio signals and identify silent periods from the microphone audio signals;

[0050] The analysis module is used to analyze the background sound features and speech loudness features of the microphone audio signal during the silent period.

[0051] The judgment module is used to determine whether the microphone audio signal has undergone external processing based on the background sound characteristics and the voice loudness characteristics.

[0052] An adjustment module is used to adjust the audio processing parameters of the microphone audio signal based on the judgment result.

[0053] As can be seen from the above, the audio device adaptive control method and apparatus provided in this application analyzes the background sound characteristics and speech loudness characteristics during silent periods to determine whether the audio signal has undergone external processing, and dynamically adjusts the audio processing parameters according to the determination results. This effectively solves the problem of audio quality degradation caused by the inability of existing systems to detect external processing behavior, and has significant advantages in improving the intelligence level of audio processing, enhancing the naturalness of audio, and improving user experience. Attached Figure Description

[0054] Figure 1 This is a schematic diagram of an adaptive control method for an audio device provided in this application.

[0055] Figure 2 This is a schematic diagram of an adaptive control system for an audio device provided in this application.

[0056] In the diagram: 210, Acquisition module; 220, Analysis module; 230, Judgment module; 240, Adjustment module. Detailed Implementation

[0057] The technical solutions of this application will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this application, and not all embodiments. The components of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.

[0058] It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. Furthermore, in the description of this application, terms such as "first," "second," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0059] Traditional online education systems typically adjust audio processing parameters, such as noise reduction or voice enhancement, based on pre-defined syllabi or scenario requirements, when adaptively controlling audio devices. However, when students introduce third-party auxiliary tools (such as speech-to-text software) to improve learning efficiency or meet specific needs, these tools often pre-process the microphone audio signal at the operating system level, performing tasks like intelligent noise reduction and voice enhancement. Existing systems lack the ability to detect this external pre-processing, causing them to repeat the processing on already processed signals. This leads to sound distortion or quality degradation, severely impacting teaching effectiveness and user experience.

[0060] For example, suppose in an online language learning course, the system is configured to activate a "clear dialogue mode" when a student speaks, to enhance the voice and suppress background noise. However, if the student is simultaneously using professional speech-to-text software, this software will perform preliminary noise reduction and gain processing on the audio signal through its system-level plugins before the system receives the microphone audio stream. At this point, the signal received by the online education system is no longer the original signal, but because it cannot recognize this fact, it will still perform a second round of noise reduction and gain processing according to its preset logic. This repeated processing may cause the voice to sound hollow, unnatural, or even distorted or clipped, making it difficult for teachers and classmates to clearly understand the student's speech, seriously hindering normal teaching interaction.

[0061] In response, this application proposes an adaptive control method for audio devices, comprising:

[0062] Acquire microphone audio signals and identify silent periods from the microphone audio signals;

[0063] During this silent period, analyze the background sound characteristics and speech loudness characteristics of the microphone audio signal;

[0064] Based on the background sound characteristics and the speech loudness characteristics, determine whether the microphone audio signal has undergone external processing;

[0065] Based on the judgment result, adjust the audio processing parameters of the microphone audio signal.

[0066] This application introduces an intelligent judgment mechanism to determine whether the microphone audio signal has undergone external processing, and adaptively adjusts the audio processing parameters according to the judgment result, thereby effectively avoiding audio distortion and quality degradation caused by repeated processing, and significantly improving the clarity of audio communication and user experience.

[0067] To make the technical solution of this application easier and clearer to understand, some key terms and implementation environments involved are explained below.

[0068] Microphone audio signals refer to the acoustic information stream captured by a microphone and converted into electrical or digital signals, which carries various sound components such as speech and background noise.

[0069] A silent period is a time in which the microphone audio signal does not contain primary speech activity; it is typically characterized by background noise or ambient sound rather than human voice. The purpose of identifying silent periods is to analyze background sound features and speech loudness characteristics without interfering with normal speech communication, as performing such analysis during speech activity could introduce additional complexity or interference.

[0070] Background sound characteristics refer to the acoustic properties of the non-speech components of a microphone audio signal, such as the spectral distribution, energy level, and temporal fluctuation patterns of noise.

[0071] Speech loudness characteristics refer to the perceived loudness attribute of the speech portion in a microphone audio signal, reflecting the volume and dynamic range of the speech.

[0072] External processing refers to any form of preprocessing performed on the microphone audio signal by other systems, software, or hardware devices before it reaches the adaptive control system of this audio device, such as noise reduction, echo cancellation, automatic gain control, equalization, etc.

[0073] Audio processing parameters refer to the various configurations within the adaptive control system of this audio device used to adjust the audio signal processing effect, such as noise reduction intensity, gain, equalization curve, compression ratio, etc.

[0074] The implementation environment of this application is mainly the audio processing module in an online education system, but its principle is also applicable to other scenarios that require intelligent adaptive processing of microphone audio signals, such as video conferencing and voice assistants.

[0075] The method of this application first requires acquiring a microphone audio signal and identifying silent periods from that signal. Acquiring the microphone audio signal can be achieved in various ways, such as converting the analog microphone signal to a digital signal using an analog-to-digital converter, or directly obtaining a digital audio stream from the operating system or audio driver interface. After acquiring the audio signal, it is necessary to identify the silent periods within it. One implementation method is to use a simple energy threshold detection; when the energy of the audio signal is below a preset threshold and persists for a period of time, it is determined to be a silent period.

[0076] During this silent period, the background noise characteristics and speech loudness characteristics of the microphone audio signal are analyzed. For the analysis of background noise characteristics, the root mean square energy value of the signal during the silent period can be calculated to obtain the overall energy level of the background noise. For the analysis of speech loudness characteristics, the peak level or average RMS energy of the speech portion can be measured when speech activity is detected to obtain the speech loudness value.

[0077] Subsequently, based on the analyzed background sound features and speech loudness features, the system determines whether the microphone audio signal has undergone external processing. For example, if the background sound RMS energy during a silent period is below a certain extremely low threshold, or the speech loudness value is above a certain abnormally high threshold, it is preliminarily determined that external noise reduction or gain processing may have occurred.

[0078] Finally, based on the judgment result, the system adjusts the audio processing parameters of the microphone audio signal. For example, if the judgment result indicates the presence of external noise reduction processing, the system can correspondingly reduce or turn off the intensity of its own noise reduction module. If the judgment result indicates the presence of external gain processing, the system can correspondingly reduce its own gain to avoid signal overload. This adjustment can take effect immediately or in stages.

[0079] The core innovation of this application lies in its ability to intelligently detect whether the microphone audio signal has been preprocessed by other system-level or third-party software. Traditional audio processing systems typically assume that the received microphone signal is raw and apply fixed or simple environment-aware processing strategies accordingly. However, when the signal has been externally preprocessed, this blind processing can lead to problems such as repeated noise reduction and secondary gain, resulting in negative listening experiences such as hollowness, distortion, or clipping.

[0080] This application analyzes background sound features and speech loudness features during silent periods to identify unnatural acoustic characteristics introduced by external preprocessing. For example, a signal that has undergone external noise reduction may have abnormally low background noise energy during silent periods, even below the minimum noise level in a natural environment; while a speech signal that has undergone external gain may have abnormally high loudness or unnaturally compressed dynamic range. By identifying these features, this application can accurately determine the presence and type of external processing and adaptively adjust its own audio processing parameters accordingly.

[0081] In some embodiments described above, this application proposes a method for determining whether a microphone audio signal has undergone external processing based on background sound features and speech loudness features. However, in practical applications, relying solely on background sound features and speech loudness features for judgment may have limitations in terms of accuracy or robustness when faced with certain complex or covert external processing. For example, some advanced audio preprocessing algorithms may subtly adjust the dynamic range or spectral details of the signal without significantly altering the background sound or overall speech loudness, making accurate identification difficult based solely on the aforementioned features.

[0082] In this regard, this application further proposes that the steps for determining whether the microphone audio signal has undergone external processing include:

[0083] Evaluate the stability of the speech features of the microphone audio signal;

[0084] During this silent period, the transmission characteristics of the microphone audio signal link are measured by injecting a reference signal and receiving its return.

[0085] Based on background sound characteristics, speech loudness characteristics, the stability of speech characteristics, and the transmission characteristics of the microphone audio signal link, it can be determined whether the microphone audio signal has undergone external processing.

[0086] Specifically, assessing the stability of speech features in microphone audio signals involves analyzing the variation patterns of intrinsic properties of the speech signal, such as fundamental frequency, formant frequencies, energy envelope, and spectral tilt, over a period of speech activity. For example, the variance, standard deviation, or rate of change of these features over a period can be calculated. If external processing (such as dynamic range compression, noise thresholding, or certain speech enhancement algorithms) unnaturally interferes with the speech signal, it may cause abnormal changes in the stability of these speech features. For example, naturally occurring fluctuations may be excessively smoothed out, or unnatural abrupt changes may occur in some cases. Monitoring the stability of these features provides an additional dimension for judging external processing.

[0087] In this method, during a silent period, the transmission characteristics of the microphone audio signal link are measured by injecting a reference signal and receiving its return. This can be understood as a probe-like detection method for the entire audio signal path. Specifically, a reference signal with known characteristics is generated and injected into a certain point in the microphone audio signal link, for example, at the operating system level or the driver level.

[0088] The returned reference signal is then received from the microphone output. By comparing the differences between the injected reference signal and the received returned signal—such as spectral distortion, energy attenuation, and phase changes—the transmission characteristics of the signal link can be inferred. If external processing modules not controlled by this system exist in the signal link, these modules will have specific effects on the reference signal, thereby altering its transmission characteristics and causing the measurement results to deviate from the expected "clean" link characteristics. For example, an external noise reduction unit might suppress specific frequency components in the reference signal, while an external gain unit might amplify the signal. This measurement provides direct, physical evidence of whether the signal has undergone external processing.

[0089] In practical applications, determining whether a microphone audio signal has undergone external processing involves comprehensively considering various factors, including background sound characteristics, speech loudness characteristics, the stability of speech features, and the transmission characteristics of the microphone audio signal link. Multiple judgment rules can be set, and external processing is only confirmed when multiple features simultaneously meet specific conditions.

[0090] This application's solution effectively overcomes the limitations of relying solely on background noise and speech loudness by introducing an assessment of the stability of speech features and a measurement of the microphone audio signal link transmission characteristics. Background noise and speech loudness alone may not be sufficient to identify external processing that subtly modifies the signal. By assessing the stability of speech features, unnatural dynamic changes can be captured, thus revealing the presence of external processing.

[0091] More importantly, by injecting a reference signal during silent periods and measuring the link's transmission characteristics, this application provides a more direct and objective detection method. External processing modules, whether software or hardware, will have specific and measurable effects on the signals passing through them. By analyzing the changes in the reference signal after passing through the entire link, such as spectral distortion and energy attenuation, it is possible to directly infer whether there are any unexpected processing steps in the link. This method does not rely on subjective judgment of the signal content but is based on objective measurement of the signal's physical characteristics. Therefore, it can more accurately and robustly identify various types of external processing, including covert processing that has little impact on background noise and speech loudness.

[0092] In some embodiments described above in this application, a method for measuring the link transmission characteristics of a microphone audio signal is proposed by injecting a reference signal and receiving its return. However, in practical applications, if the injected reference signal has predictable or periodic characteristics, some advanced external preprocessing applications, such as intelligent noise reduction algorithms or voice activity detectors, may identify and suppress the reference signal, thereby causing the measurement results to be distorted or unable to accurately reflect the true link transmission characteristics, thus affecting the accuracy of the judgment of external processing.

[0093] In this regard, this application further proposes the following steps for measuring the transmission characteristics of the microphone audio signal link by injecting a reference signal and receiving its return during a silent period:

[0094] During this quiet period, a low-level reference signal with spectral characteristics that change non-periodically over time is generated to avoid being identified and suppressed by external preprocessing applications.

[0095] Inject the reference signal into the microphone audio signal link;

[0096] Receive the return of the reference signal;

[0097] The transmission characteristics of the microphone audio signal link are measured based on the spectral distortion and energy attenuation of the returned signal.

[0098] Specifically, generating a low-level reference signal with a non-periodic spectral characteristic over time aims to make this reference signal difficult for external preprocessing applications to identify and suppress as noise or a specific signal. The non-periodic spectral characteristic means that the frequency components or energy distribution of the signal exhibit irregular and unpredictable patterns of change over time. This non-periodicity makes it difficult for external processing applications to separate and eliminate it from the background using fixed filters or pattern recognition algorithms. "Low level" means that the amplitude of the reference signal is controlled at a very low level, typically below the energy of normal speech or significant background noise, and possibly even below the threshold perceptible to the human ear. This ensures that the reference signal does not interfere with the user experience during injection and does not trigger more aggressive signal processing strategies from external processing applications.

[0099] Injecting the reference signal into the microphone audio signal link means superimposing or replacing the current audio stream with the generated reference signal at a point before the microphone audio signal enters the system's processing module, using software or hardware. This injection can be performed at the operating system level in the virtual audio device driver or at the audio interface card driver level. Receiving the return of the reference signal means that the system captures the reference signal from the output of the microphone audio signal link, after it has been transmitted through the link and possibly processed externally.

[0100] Measuring the transmission characteristics of a microphone audio signal link based on the spectral distortion and energy attenuation of the returned signal involves inferring the link's characteristics by comparing the injected original reference signal with the received returned signal. Spectral distortion refers to changes in the frequency components of the returned signal compared to the original reference signal; for example, some frequencies may be attenuated or amplified, or new harmonic components may appear. Energy attenuation refers to a decrease in the overall energy level of the returned signal relative to the original reference signal.

[0101] By analyzing these distortions and attenuations, the transfer function or frequency response curve of the microphone audio signal link can be constructed, thereby revealing whether there is external processing in the link and its specific processing characteristics, such as whether there is high-pass filtering, low-pass filtering, equalization of specific frequency bands, or dynamic range compression.

[0102] This application's solution effectively addresses the limitation in existing technologies where reference signals can be identified and suppressed by external preprocessing applications by designing a reference signal. Traditional reference signals, such as simple sine waves or white noise, have relatively fixed or easily predictable spectral characteristics. When these signals are injected into an audio link, if external processing applications such as intelligent noise reduction or speech enhancement are present, these predictable reference signals may be misidentified as noise or non-speech components and filtered or suppressed. Once the reference signal is suppressed, the returned signal received by this system will not accurately reflect the true link transmission characteristics, leading to misjudgments of external processing.

[0103] This application generates a low-level reference signal with spectral characteristics that vary non-periodically over time, making the signal highly random and unpredictable in both the spectral and time domains. This characteristic makes it difficult for external preprocessing applications to distinguish it from normal background noise or ambient sound using their internal pattern recognition or fixed filtering algorithms, thus avoiding misidentification and suppression of the reference signal. Simultaneously, the low-level design ensures that the signal will not trigger a more aggressive response from external processing applications, nor will it cause auditory interference to the user. Therefore, this application ensures that the injected reference signal can penetrate external processing applications and accurately reflect the actual impact of external processing on the signal, resulting in more accurate and reliable measurements of the microphone audio signal link transmission characteristics.

[0104] Through the above technical solution, this application significantly improves the accuracy and reliability of measuring the link transmission characteristics of microphone audio signals. By generating a low-level reference signal with spectral characteristics that change non-periodically over time, this signal can effectively circumvent the identification and suppression of external preprocessing applications, thereby ensuring that the returned signal can truly reflect the actual impact of external processing on the signal. This enables the system to obtain more accurate link transmission characteristic data, thus providing a more solid and reliable basis for determining whether the microphone audio signal has undergone external processing, greatly improving the overall accuracy and robustness of the judgment, and avoiding judgment errors caused by misprocessing of the reference signal.

[0105] In some embodiments described above, this application proposes determining whether a microphone audio signal has undergone external processing based on background sound features, speech loudness features, the stability of speech features, and the transmission characteristics of the microphone audio signal link. However, while this approach provides multi-dimensional criteria for judgment, in practical applications, simply combining these features may lead to misjudgments due to occasional fluctuations in a particular feature or the complexity of external processing. This is especially true when dealing with subtle or dynamically changing external processing, where the reliability of the judgment may be affected.

[0106] In response, this application further proposes the following steps for determining whether a microphone audio signal has undergone external processing based on background sound features, speech loudness features, the stability of speech features, and the transmission characteristics of the microphone audio signal link:

[0107] Evaluate the transmission characteristics of the microphone audio signal link to form a preliminary indication of whether the microphone audio signal has undergone external processing;

[0108] Based on this preliminary indication, and in conjunction with the stability of the speech features, the preliminary indication is verified to form a verified judgment.

[0109] Continuously monitor the background sound features and the speech loudness features to obtain auxiliary judgment information;

[0110] Based on the verified judgment, the auxiliary judgment information, and the transmission characteristics of the microphone audio signal link, it is confirmed whether the microphone audio signal has undergone external processing.

[0111] Specifically, assessing the transmission characteristics of the microphone audio signal link to form a preliminary indication of whether the microphone audio signal has undergone external processing refers to using direct evidence obtained from measurements of the microphone audio signal link's transmission characteristics as the primary basis for determining the presence or absence of external processing. For example, if the measured link frequency response curve deviates significantly from the expected "clean" link, or if the phase response exhibits a non-linear change, a preliminary, physical-based indication of external processing can be generated. This preliminary indication typically has a high degree of confidence because it directly reflects the physical alteration of the signal path.

[0112] Based on this initial indication, and in conjunction with the stability of speech features, the initial indication is verified to form a verified judgment. This means cross-validating the initial judgment based on link characteristics with the analysis of the speech content itself. For example, if the initial indication strongly suggests the presence of external processing, and at the same time the stability analysis of speech features also shows unnatural smoothing or dynamic range compression, then these two independent pieces of evidence corroborate each other, forming a more reliable and convincing verified judgment. This verification mechanism can effectively reduce the risk of misjudgment that may arise from a single piece of evidence.

[0113] Continuous monitoring of background sound features and speech loudness features to obtain auxiliary judgment information refers to the system's uninterrupted real-time analysis of the microphone audio signal's background sound features (such as noise level and spectral distribution) and speech loudness features (such as speech peaks and average loudness) throughout the entire audio communication process. This continuously acquired information, as auxiliary judgment information, can provide dynamic, context-sensitive clues. For example, if external processing is intermittent, or its intensity changes over time, continuous monitoring can capture these changes, providing real-time updated background information for the final judgment.

[0114] Finally, based on the verified judgment, the auxiliary judgment information, and the transmission characteristics of the microphone audio signal link, it is confirmed whether the microphone audio signal has undergone external processing.

[0115] This application first utilizes the most direct and objective evidence—the transmission characteristics of the microphone audio signal link—to form a preliminary indication, providing a solid foundation for judgment. Subsequently, by combining the stability of speech features to verify this preliminary indication, an analysis of the signal content itself is introduced. This makes the judgment no longer single-dimensional, but rather increases confidence through mutual corroboration of different types of evidence, thereby effectively reducing the risk of false positives or false negatives. Furthermore, continuous monitoring of background sound features and speech loudness features provides dynamic and real-time auxiliary information for judgment, enabling the system to adapt to dynamic changes in external processing and avoiding judgment lag or failure caused by changes in external processing characteristics.

[0116] In some preferred embodiments, it is assumed that the student is using a smart headset with advanced adaptive noise cancellation and voice enhancement features that dynamically adjust based on the environment. Relying solely on background sound characteristics and voice loudness characteristics may not be sufficient to accurately determine the real-time state of the headset's internal processing.

[0117] According to the scheme of this application, the system first evaluates the transmission characteristics of the microphone audio signal link. For example, by injecting a reference signal and analyzing the returned signal, the system finds that the frequency response of the link has slight dynamic changes in certain frequency bands, which forms a preliminary indication of the presence of external processing.

[0118] Next, the system verifies the initial indication by combining it with an assessment of the stability of the student's speech features. For example, if the initial indication shows the presence of some form of dynamic processing, and the system observes that the instantaneous loudness changes or spectral details of the student's speech are abnormally smoothed, this further verifies the existence of external processing and forms a verified judgment.

[0119] Meanwhile, the system continuously monitors background sound characteristics and speech loudness characteristics. This information serves as auxiliary judgment information, helping the system understand the current intensity and pattern of external processing. For example, if background noise suddenly increases, but the background noise level monitored by the system remains unchanged, this may indicate that external noise reduction is actively working.

[0120] Ultimately, based on the verified judgment, continuously acquired auxiliary judgment information, and the real-time transmission characteristics of the microphone audio signal link, the system confirms whether the microphone audio signal has undergone external processing and identifies its dynamic characteristics. For example, the system may confirm that the headphones are performing moderate-intensity adaptive noise cancellation and slight speech compression. Based on this accurate confirmation, the system can adjust its own audio processing parameters accordingly, such as reducing its own noise cancellation intensity and speech compression ratio, thereby avoiding superimposition with the headphone's internal processing and ensuring that the final output audio quality is natural and clear, avoiding distortion caused by over-processing.

[0121] Specifically, the steps for analyzing the background sound characteristics of the microphone audio signal during the silent period include:

[0122] During this silent period, the background sound characteristics of the microphone audio signal are analyzed, including:

[0123] Analyze the spectral variation patterns of the background sound to identify unnatural spectral evolution;

[0124] Analyze the energy fluctuation patterns of the background sound to identify unnatural energy fluctuations;

[0125] Analyze and identify the instantaneous events of the background sound to distinguish its source.

[0126] Analyzing the spectral variation patterns of the background sound to identify unnatural spectral evolution refers to meticulously observing the changes in the frequency components of the microphone audio signal over time during silent periods.

[0127] Analyzing the energy fluctuation patterns of the background sound to identify unnatural energy fluctuations refers to monitoring the changes in the overall energy or specific frequency band energy of the microphone audio signal over time during silent periods.

[0128] Analyzing and identifying transient events in the background sound to distinguish its source refers to detecting and identifying non-speech sound events with short durations and rapid energy changes during silent periods. For example, transient detection algorithms can be used to identify typing sounds, keyboard clicks, mouse clicks, etc.

[0129] By conducting multi-dimensional analysis of the spectral variation patterns, energy fluctuation patterns, and instantaneous events of background sound, this application can identify unnatural spectral evolution, energy fluctuations, and digital artifacts introduced by external processing. This enables the system to more accurately determine whether the microphone audio signal has undergone external processing, especially when faced with external processing that does not significantly affect the background sound but alters its intrinsic characteristics. It provides a more reliable basis for judgment, thereby further improving the overall accuracy and robustness of the judgment.

[0130] Specifically, the steps for analyzing the speech loudness characteristics of the microphone audio signal include:

[0131] The loudness of the speech portion of the microphone audio signal is measured to obtain the loudness value of that speech portion;

[0132] Based on this loudness value, determine the range of loudness variation for this speech segment;

[0133] Analyze the distribution of the loudness variation range.

[0134] The loudness measurement of the speech portion of the microphone audio signal, to obtain its loudness value, refers to quantifying the perceived loudness of speech segments after identifying the speech activity periods in the microphone audio signal. Loudness measurement differs from simple peak level or root mean square energy measurement; it more closely approximates the human ear's perception of sound volume.

[0135] Determining the loudness range of a speech segment based on the loudness value involves calculating the difference between the maximum and minimum loudness values ​​within a certain time window after acquiring a series of speech loudness values, or statistically analyzing their dynamic range. This range reflects the dynamic range of speech changes from the softest to the loudest.

[0136] Analyzing the distribution of loudness variation range refers to performing statistical or pattern recognition analysis on the determined loudness variation range. For example, a histogram of the loudness variation range can be plotted to observe its statistical characteristics such as peak value, skewness, and kurtosis. The trend of loudness variation range over time can also be analyzed; for example, whether it remains within a relatively stable range or exhibits anomalous compression or expansion. By analyzing this distribution, unnatural loudness characteristics introduced by external processing can be identified.

[0137] This application's solution effectively overcomes the limitations of judging external processing solely based on simple speech energy or peak level by conducting a deeper and more comprehensive analysis of speech loudness characteristics. While traditional speech energy or peak level measurements can reflect volume, they often fail to capture the human ear's perception of dynamic changes in sound and struggle to identify subtle loudness variations caused by complex external processing (such as dynamic range compression or multi-segment compression). For example, an advanced dynamic range compressor can significantly compress the dynamic range of speech loudness without substantially altering the average speech energy, making the speech sound lifeless.

[0138] By measuring the loudness of the speech component, this application obtains loudness values ​​that more closely match human auditory perception, providing a more accurate basis for subsequent analysis. More importantly, by determining the range of loudness variation in the speech component and analyzing its distribution, this application can directly reveal whether the dynamic characteristics of the speech signal have been unnaturally altered by external processing. For example, if the loudness variation range is abnormally narrow, or its distribution exhibits an unnatural concentration trend, this strongly indicates the presence of dynamic range compression. This detailed analysis of the dynamic characteristics of loudness enables the system to identify external processing that subtly adjusts the speech loudness, thus providing more accurate and reliable evidence for judging external processing.

[0139] Specifically, the steps for analyzing the distribution of loudness variation range mentioned above include:

[0140] The distribution of the loudness variation range was analyzed, including:

[0141] Based on the evolution characteristics of the loudness variation range over time, assess the smoothness or abruptness of the loudness variation range;

[0142] Based on this smoothness or abruptness, identify the unnatural evolution of the distribution of the loudness variation range.

[0143] Assessing the smoothness or abruptness of the loudness variation range based on its evolution over time involves analyzing the trends of these range values ​​across consecutive time frames or speech segments after obtaining the loudness variation range. If the loudness variation range fluctuates drastically within a short period, or if the expected natural dynamic changes are excessively smoothed, this may indicate the presence of external processing. For example, an aggressive dynamic range compressor or limiter might force the loudness variation range into a narrow interval, making its changes over time appear abnormally smooth; while a noise threshold or gain control with inappropriate release time might cause abrupt jumps in the loudness variation range at the beginning or end of the speech.

[0144] Identifying unnatural evolutions in the distribution of loudness variation range based on smoothness or abruptness involves comparing the assessed smoothness or abruptness with a pre-defined distribution model of loudness variation range in normal speech. If the smoothness or abruptness of the currently analyzed loudness variation range exceeds the normal range or exhibits a pattern inconsistent with natural speech, it can be determined as an unnatural evolution of the distribution of loudness variation range, thus indicating the presence of external processing.

[0145] This application captures such unnatural dynamic evolution by evaluating the smoothness or abruptness of loudness variation range. Excessive smoothing of the loudness variation range indicates the presence of dynamic range compression or limiters; unnatural abrupt changes may indicate inappropriate intervention by noise thresholding or automatic gain control. This fine-grained analysis of loudness dynamics allows the system to identify external processing that subtly adjusts speech loudness, providing more accurate and reliable evidence for judging external processing. This method reveals the impact of external processing on speech naturalness, even if this impact is not apparent in static loudness metrics, it can be identified through dynamic evolution characteristics.

[0146] In some embodiments described above, the audio processing parameters of the microphone audio signal are adjusted based on the judgment result. However, if the characteristics of the external processing are dynamically changing (e.g., the user may turn third-party software on or off at any time, or its processing intensity may change with the environment), simply adjusting based on a single judgment result may not be able to continuously adapt to such changes, resulting in poor adjustment effects or the recurrence of audio quality problems. Furthermore, abrupt parameter adjustments may also cause discomfort to the user.

[0147] In this regard, this application further proposes that the steps of adjusting the audio processing parameters of the microphone audio signal based on the judgment result include:

[0148] When the judgment result indicates the presence of external processing, the audio processing parameters of the microphone audio signal are determined based on the characteristics of the external processing reflected by the judgment result.

[0149] Continuously acquire the external processing and judgment results of the microphone audio signal;

[0150] Based on the continuously acquired external processing judgment results, the updated values ​​of the audio processing parameters of the microphone audio signal are calculated;

[0151] The audio processing parameters of the microphone audio signal are gradually adjusted to adapt to the changing characteristics of external processing over time.

[0152] Specifically, when the judgment result indicates the presence of external processing, the audio processing parameters of the microphone audio signal are determined based on the characteristics of the external processing reflected in the judgment result. This means that when the system first or again confirms that the microphone audio signal has undergone external processing, it calculates and sets the initial parameters of the system's internal audio processing module based on the specific type and intensity of the external processing identified by the judgment module. For example, if the judgment result indicates that strong noise reduction has been applied externally, the system will correspondingly set its own noise reduction parameters to a lower value or directly disable the noise reduction function.

[0153] Continuously acquiring the external processing judgment results of the microphone audio signal means that the judgment process is not stopped after a one-time judgment, but is periodically or re-executed under specific conditions. This ensures that the system can detect any changes in the external processing status in real time, such as when the user turns on or off third-party software, or when the intensity of the external processing is adjusted.

[0154] Based on the continuously acquired external processing judgment results, the updated values ​​of the audio processing parameters of the microphone audio signal are calculated. This means that when the continuous judgment results indicate a change in the external processing characteristics, the system will recalculate the ideal values ​​of its own audio processing parameters according to the new judgment results. For example, if the external noise reduction intensity weakens, the system may calculate an updated value that requires a moderate increase in its own noise reduction intensity.

[0155] Gradual adjustment of the microphone's audio signal processing parameters to adapt to the changing characteristics of external processing over time means that after calculating the updated parameter values, the system does not immediately jump to the new values, but rather gradually approaches the target values ​​in small steps over a certain period of time. This gradual adjustment can be achieved using smoothing functions, such as linear interpolation, exponential decay, or S-curves, ensuring a smooth parameter change process and avoiding sudden changes in audio quality or user discomfort caused by abrupt parameter changes. For example, if the noise reduction intensity needs to be adjusted from 0dB to -10dB, the system may adjust it by 0.5dB every tens of milliseconds over several seconds, rather than adjusting it by 10dB all at once.

[0156] Through the above technical solution, this application significantly improves the adaptability of audio processing parameter adjustment and user experience. By continuously acquiring external processing judgment results and calculating parameter update values, this application ensures that the system can respond to dynamic changes in external processing in real time. By gradually adjusting audio processing parameters, this application avoids sudden changes in audio quality or user discomfort that may be caused by abrupt parameter changes, making the adjustment process smooth and imperceptible. This enables the system to provide high-quality audio output stably over a long period of time, effectively adapting to the characteristics of external processing changing over time, thereby further optimizing audio quality and user experience.

[0157] In some embodiments described above, a gradual adjustment of the audio processing parameters of the microphone audio signal is proposed to adapt to the time-varying characteristics of external processing. However, if the step size and period of such gradual adjustment are fixed, or if it is based solely on the judgment results of external processing, it may not adequately take into account subtle changes in the current audio environment and user state. For example, in situations where the user is speaking or there is significant background noise, overly aggressive or overly slow adjustments may negatively impact the user experience, resulting in an unsmooth or inefficient adjustment process.

[0158] In this regard, this application further proposes the following steps for progressively adjusting the audio processing parameters of the microphone audio signal:

[0159] Obtain the voice activity status of the microphone audio signal;

[0160] Obtain the background sound level of the microphone audio signal;

[0161] The adjustment step size of the audio processing parameter is determined based on the state of the voice activity or the level of the background sound.

[0162] The adjustment cycle of the audio processing parameters is determined based on the state of the speech activity or the background sound level; and

[0163] The audio processing parameters of the microphone audio signal are gradually adjusted according to the adjustment step size and the adjustment period.

[0164] Specifically, acquiring the speech activity status of the microphone audio signal means that the system detects in real time whether there is human speech activity in the current microphone audio signal. This is usually achieved through a speech activity detection algorithm, which can distinguish between speech segments and non-speech segments.

[0165] Obtaining the background sound level of the microphone audio signal refers to the system measuring the background noise energy or loudness in the current microphone signal in real time. This can be done during silent periods or when speech is present using a noise estimation algorithm.

[0166] The adjustment step size of the audio processing parameter is determined based on the state of the speech activity or the background noise level. For example, when speech activity is detected, to avoid perceptible interference with the speech, the adjustment step size can be set to be smaller, such as 0.1 dB each time, making the parameter change smoother and less noticeable. If it is a silent period and the background noise level is low, a slightly larger step size can be used, such as 0.5 dB each time, to speed up the adjustment and improve response efficiency.

[0167] The adjustment cycle for the audio processing parameter is determined based on the voice activity status or background noise level. This means the system dynamically selects the time interval between two parameter adjustments based on the current environment and user status. During voice activity, the adjustment cycle can be set longer to reduce potential impact on speech and ensure speech continuity. During silent periods or when background noise levels are high, a shorter adjustment cycle can be set to adapt to environmental changes more quickly and reach optimal processing status rapidly.

[0168] Finally, based on the adjustment step size and the adjustment period, the audio processing parameters of the microphone audio signal are gradually adjusted. This means that the system applies the dynamically determined step size and period to the parameter update process, gradually updating the audio processing parameters in a smooth manner, such as through linear interpolation, exponential smoothing, or S-curve algorithms, to ensure the smoothness of the adjustment process and the user's insensitivity.

[0169] Through the above technical solution, this application significantly improves the intelligence and user-friendliness of the gradual adjustment of audio processing parameters. By dynamically determining the adjustment step size and period based on the voice activity state and background sound level, this application enables the system to complete parameter adaptation more efficiently and smoothly without affecting user perception. This avoids auditory discomfort caused by sudden parameter changes during voice activity, and also avoids adaptation lag caused by slow adjustment during silence. Ultimately, it significantly improves the adaptive control capability of audio devices and user experience, ensuring smooth and high-quality audio communication in various complex dynamic environments.

[0170] Secondly, referring to Figure 2 This application proposes an adaptive control device for an audio device, the device comprising:

[0171] Acquisition module 210 is used to acquire microphone audio signals and identify silent periods from the microphone audio signals;

[0172] Analysis module 220 is used to analyze the background sound features and speech loudness features of the microphone audio signal during the silent period;

[0173] The judgment module 230 is used to determine whether the microphone audio signal has undergone external processing based on the background sound characteristics and the voice loudness characteristics;

[0174] The adjustment module 240 is used to adjust the audio processing parameters of the microphone audio signal according to the judgment result.

[0175] By analyzing background sound features and speech loudness features during silent periods, it can determine whether the audio signal has undergone external processing and dynamically adjust audio processing parameters based on the judgment results. This effectively solves the problem of audio quality degradation caused by the inability of existing systems to detect external processing behavior, and has significant advantages in improving the intelligence level of audio processing, enhancing the naturalness of audio, and improving user experience.

[0176] Furthermore, in some preferred embodiments, the audio device adaptive control device proposed in this application can perform any of the steps in the above methods.

[0177] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.

Claims

1. An adaptive control method for an audio device, characterized in that, include: Acquire microphone audio signals and identify silent periods from the microphone audio signals; During the silent period, the background sound features and speech loudness features of the microphone audio signal are analyzed; Based on the background sound characteristics and the speech loudness characteristics, determine whether the microphone audio signal has undergone external processing; Based on the judgment result, adjust the audio processing parameters of the microphone audio signal; The step of determining whether the microphone audio signal has undergone external processing based on the background sound features and the speech loudness features includes: Evaluate the stability of the speech features of the microphone audio signal; During the silent period, the transmission characteristics of the microphone audio signal link are measured by injecting a reference signal and receiving its return. Based on the background sound features, the speech loudness features, the stability of the speech features, and the transmission characteristics of the microphone audio signal link, it is determined whether the microphone audio signal has undergone external processing; The step of adjusting the audio processing parameters of the microphone audio signal based on the judgment result includes: When the judgment result indicates the presence of external processing, the audio processing parameters of the microphone audio signal are determined based on the characteristics of the external processing reflected by the judgment result. Continuously acquire the external processing and judgment results of the microphone audio signal; Based on the continuously acquired external processing judgment results, the updated values ​​of the audio processing parameters of the microphone audio signal are calculated; The audio processing parameters of the microphone audio signal are gradually adjusted to adapt to the characteristics of external processing that change over time. The step of progressively adjusting the audio processing parameters of the microphone audio signal includes: Acquire the voice activity status of the microphone audio signal; Obtain the background sound level of the microphone audio signal; The adjustment step size of the audio processing parameters is determined based on the voice activity state or the background sound level. The adjustment cycle of the audio processing parameters is determined based on the voice activity state or the background sound level; and The audio processing parameters of the microphone audio signal are gradually adjusted according to the adjustment step size and the adjustment period.

2. The adaptive control method for an audio device according to claim 1, characterized in that, The step of measuring the transmission characteristics of the microphone audio signal link by injecting a reference signal and receiving its return during the silent period includes: During the silent period, a low-level reference signal with spectral characteristics that change non-periodically over time is generated to avoid being identified and suppressed by external preprocessing applications. Inject the reference signal into the microphone audio signal link; Receive the return of the reference signal; The transmission characteristics of the microphone audio signal link are measured based on the spectral distortion and energy attenuation of the returned signal.

3. The adaptive control method for an audio device according to claim 1, characterized in that, The step of determining whether the microphone audio signal has undergone external processing based on the background sound features, the speech loudness features, the stability of the speech features, and the transmission characteristics of the microphone audio signal link includes: The transmission characteristics of the microphone audio signal link are evaluated to form a preliminary indication as to whether the microphone audio signal has undergone external processing; Based on the preliminary indication and the stability of the speech features, the preliminary indication is verified to form a verified judgment. Continuously monitor the background sound features and the speech loudness features to obtain auxiliary judgment information; Based on the verified judgment, the auxiliary judgment information, and the transmission characteristics of the microphone audio signal link, it is confirmed whether the microphone audio signal has undergone external processing.

4. The adaptive control method for an audio device according to claim 1, characterized in that, The step of analyzing the background sound characteristics of the microphone audio signal during the silent period includes: During the silent period, the background sound characteristics of the microphone audio signal are analyzed, and the analysis includes: Analyze the spectral variation patterns of background sounds to identify unnatural spectral evolution; Analyze the energy fluctuation patterns of background sounds to identify unnatural energy fluctuations; Analyze and identify instantaneous events in the background sound to distinguish their sources.

5. The adaptive control method for an audio device according to claim 1, characterized in that, The steps for analyzing the speech loudness characteristics of the microphone audio signal include: The loudness of the speech portion of the microphone audio signal is measured to obtain the loudness value of the speech portion; Based on the loudness value, determine the loudness variation range of the speech portion; Analyze the distribution of the loudness variation range.

6. The adaptive control method for an audio device according to claim 5, characterized in that, The steps for analyzing the distribution of the loudness variation range include: The distribution of the loudness variation range is analyzed, and the analysis includes: Based on the evolution characteristics of the loudness variation range over time, evaluate the smoothness or abruptness of the loudness variation range; Based on the smoothness or abruptness, identify the unnatural evolution of the distribution of the loudness variation range.

7. An audio device adaptive control apparatus for performing the method according to any one of claims 1 to 6, characterized in that, The device includes: The acquisition module is used to acquire microphone audio signals and identify silent periods from the microphone audio signals; The analysis module is used to analyze the background sound features and speech loudness features of the microphone audio signal during the silent period. The judgment module is used to determine whether the microphone audio signal has undergone external processing based on the background sound characteristics and the voice loudness characteristics. The adjustment module is used to adjust the audio processing parameters of the microphone audio signal based on the judgment result.

Citation Information

Patent Citations

  • Method and system for automatically identifying voice recording equipment source

    CN102394062A

  • Apparatus and method for improving the audibility of specific sounds to a user

    CN105075289A