Adaptive noise reduction method and system for multi-mode audio SoC main control chip

By employing time-frequency domain dual deconstruction and dynamic modeling techniques in a multimodal audio SoC main control chip, the separation and adaptability issues of traditional audio noise reduction methods in complex noise environments are solved, achieving efficient and stable audio signal processing in complex environments.

CN120998219APending Publication Date: 2025-11-21HANK ELECTRONICS
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202511119500.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-11
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Traditional audio noise reduction methods struggle to effectively separate noise from speech signals in complex noise environments, leading to failures or inefficiencies in speech recognition and sound processing tasks, and they cannot flexibly adapt to dynamic noise scenarios.

Method used

A multimodal audio SoC main control chip is used to perform dual time-frequency domain deconstruction, construct a full-frequency domain noise interference topology table, separate multimodal interference factors and model the real-time noise environment, and realize neural coding gain and audio boundary reconstruction through dynamic noise reduction response learning optimization, and finally build an iterative noise reduction optimization engine.

Benefits of technology

Accurately identify noise sources in complex noisy environments, enhance speech clarity, preserve audio details, ensure the stability and adaptability of noise reduction effects, reduce human distortion, and improve speech recognition capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120998219A_ABST
    Figure CN120998219A_ABST
Patent Text Reader

Abstract

The invention relates to the field of adaptive noise reduction, in particular to an adaptive noise reduction method and system for a multi-mode audio SoC main control chip. The method comprises the following steps: acquiring an original audio input signal according to an SoC main control chip, performing time-frequency domain dual deconstruction and cross-frequency domain noise interference structure analysis, and constructing a full-frequency domain noise interference topology table; carrying out multi-mode interference factor separation on the full-frequency-domain noise interference topology table, carrying out multi-noise environment modeling, and constructing a real-time noise scene model; time sequence noise slope fluctuation modeling is carried out on the real-time noise scene model, dynamic noise reduction response learning optimization is carried out, and a dynamic noise reduction optimization strategy is constructed; and performing neural coding gain and audio boundary line reconstruction according to the original audio input signal to obtain a coding gain key audio signal and a weak audio optimization signal. According to the invention, by flexibly adjusting the noise reduction intensity of the audio signal, the real-time scene noise reduction performance is optimized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of adaptive noise reduction, and more particularly to an adaptive noise reduction method and system for a multimodal audio SoC main control chip. Background Technology

[0002] With the rapid development of artificial intelligence, the Internet of Things, and smart devices, audio processing technology has been widely applied in various devices, especially in smart homes, voice assistants, autonomous driving, augmented reality (AR), and virtual reality (VR). As one of the primary means of human interaction, the clarity and accuracy of audio signals directly impact user experience and device performance. Especially in noisy environments such as public places, vehicles, and industrial sites, audio signals are often interfered with by various noise sources, leading to failures or inefficiencies in tasks such as speech recognition and sound processing. Therefore, how to effectively perform audio noise reduction, especially adaptive noise reduction in complex environments, has become a crucial problem that urgently needs to be solved in audio technology.

[0003] Traditional noise reduction methods typically rely on hardware-level noise filtering, linear noise suppression algorithms, or simple frequency domain filtering techniques. While these methods can suppress environmental noise to some extent, they often suffer from drawbacks such as incomplete separation of noise and speech signals and insufficient processing effectiveness. In situations with multiple noise sources (such as background music, crowd noise, and traffic noise), traditional noise reduction techniques may remove some useful speech signals, causing speech distortion or information loss. Furthermore, existing noise reduction methods often depend on fixed noise models and cannot flexibly adapt to dynamic changes in different noise scenarios. Therefore, their effectiveness in dynamic noise environments is often insufficient to meet practical needs. Summary of the Invention

[0004] To address the aforementioned technical problems, this invention proposes an adaptive noise reduction method and system for a multimodal audio SoC main control chip, thereby resolving at least one of the aforementioned technical problems.

[0005] To achieve the above objectives, this invention provides an adaptive noise reduction method for a multimodal audio SoC main control chip, comprising the following steps: Step S1: Collect the original audio input signal from the SoC main control chip, perform time-frequency domain dual deconstruction and cross-frequency domain noise interference structure analysis, and construct a full-frequency domain noise interference topology table; Step S2: Separate multimodal interference factors from the full-frequency domain noise interference topology table, model multiple noise environments, and construct a real-time noise scene model; Step S3: Model the temporal noise slope fluctuation of the real-time noise scene model, and perform dynamic noise reduction response learning optimization to construct a dynamic noise reduction optimization strategy; Step S4: Perform neural coding gain and audio boundary reconstruction based on the original audio input signal to obtain the key audio signal of coding gain and weak audio optimization signal; Step S5: Perform effective audio reconstruction on the key audio signal of the encoded gain and the weak audio optimization signal to construct the noise-reduced and optimized audio input signal; Step S6: Perform real-time audio noise reduction offset evaluation on the noise reduction optimized audio input signal, and perform iterative noise reduction migration optimization on the dynamic noise reduction optimization strategy to build an iterative noise reduction optimization engine.

[0006] This invention effectively analyzes the time and frequency characteristics of audio signals by performing a dual time-frequency domain deconstruction on the original audio input signal. This dual deconstruction allows for the simultaneous identification of noise and useful audio signals in both the time and frequency domains, thus more accurately determining the location and frequency distribution of noise sources. By analyzing the spread of noise interference across multiple frequency domains, noise sources with broad frequency coverage spanning multiple frequency bands can be identified. This analysis helps to build a more comprehensive noise interference topology table, providing strong support for subsequent noise reduction processing. The constructed topology table details the noise characteristics (such as intensity and frequency range) at different frequency bands, providing a precise data foundation for subsequent interference separation, modeling, and optimization. With the fusion of multimodal signals, different types of noise sources (such as traffic noise, wind noise, and voice interference) can be effectively separated. Wind noise may be more pronounced in the low-frequency band, while human voice is mainly in the mid-to-high frequency band; this separation allows for more accurate subsequent noise reduction. By modeling the interaction between different noise sources in real time, a more realistic and dynamic noise environment model can be established. This model not only works effectively in static noise environments but also handles dynamically changing noise scenarios (such as sudden changes in background noise), providing more flexible adaptability for denoising algorithms. Through real-time noise scenario models, the state and characteristics of noise sources can be continuously updated, ensuring that the denoising strategy remains optimal at all times. This real-time modeling capability greatly enhances its responsiveness and adaptability. By modeling the temporal fluctuations of noise scenarios, dynamic changes in noise intensity and frequency can be captured. Traffic noise may fluctuate with changes in traffic flow, or environmental noise may change due to wind speed variations. Accurately capturing these changes allows for rapid noise reduction responses. By learning and optimizing noise responses, intelligent adjustments can be made for different noise fluctuation patterns. During sudden noise increases, it can quickly switch to a high-efficiency denoising mode, while reverting to a low-power, low-latency mode during calm periods. As the noise environment changes, the denoising strategy can be adaptively adjusted to ensure continuous and stable denoising effects, especially when multiple noises coexist, enabling real-time strategy optimization. Through neural coding gain technology, the quality of audio signals can be enhanced without introducing excessive distortion. This means that not only can noise be removed, but the clarity of speech or audio can also be enhanced, making the denoised audio signal more natural and intelligible. By reconstructing the boundaries of the audio signal, the boundaries and details of the sound can be accurately restored, avoiding the "blurring" or "hollowing" phenomenon that occurs during the noise reduction process. This is crucial for the fidelity of speech or music signals, preserving more speech details or musical nuances. This technology allows low-volume or weak speech signals to be accurately amplified, improving the ability to recognize human voices, especially in noisy environments, ensuring that the speech signal is clearly discernible. After noise reduction processing, the useful audio signal can be effectively reconstructed.By preserving key audio signals and optimizing weak audio signals, a balance between signal clarity and detail can be achieved, making the audio output more realistic. The reconstructed audio signal allows for more balanced noise reduction, especially in complex, noisy environments, maximizing the restoration of the original sound and reducing human distortion. This stage allows for flexible adjustment of the noise reduction intensity, avoiding audio "compression" caused by excessive noise reduction and maintaining a natural audio presentation. Real-time offset evaluation of the noise reduction effect allows for immediate detection of deviations in the noise reduction strategy and necessary adjustments. This process ensures that the noise reduction effect remains optimal, especially under constantly changing environmental noise conditions. The iterative noise reduction optimization strategy continuously learns and updates the noise reduction model. In new noise environments, it can intelligently adapt through historical feedback, reducing manual intervention and adjustments. By establishing an iterative learning mechanism for the noise reduction optimization engine, it ensures that the noise reduction processing remains efficient during long-term operation, reducing the risk of failure in complex environments.

[0007] This specification provides an adaptive noise reduction system for a multimodal audio SoC main control chip, used to perform the adaptive noise reduction method for a multimodal audio SoC main control chip as described above, including: The noise structure analysis module is used to collect the original audio input signal from the SoC main control chip, perform time-frequency domain dual deconstruction and cross-frequency domain noise interference structure analysis, and construct a full-frequency domain noise interference topology table. The noise environment modeling module is used to separate multimodal interference factors from the full-frequency domain noise interference topology table, and to model multiple noise environments to build a real-time noise scene model. The dynamic noise reduction module is used to model the temporal noise slope fluctuation of the real-time noise scene model, and to learn and optimize the dynamic noise reduction response to build a dynamic noise reduction optimization strategy. An effective audio optimization module is used to perform neural coding gain and audio boundary reconstruction based on the original audio input signal to obtain the key audio signal of coding gain and weak audio optimization signal; The effective audio reconstruction module is used to effectively reconstruct the key audio signal of the encoded gain and the weak audio optimization signal to build a noise-reduced and optimized audio input signal. The iterative noise reduction optimization module is used to perform real-time audio noise reduction offset evaluation on the noise reduction optimization audio input signal, and to perform iterative noise reduction migration optimization on the dynamic noise reduction optimization strategy, thus building an iterative noise reduction optimization engine.

[0008] This invention employs a dual time-frequency domain deconstruction approach, enabling the system to simultaneously perform detailed analysis of audio signals in both time and frequency dimensions, thereby accurately distinguishing useful signals from noise components. It is particularly effective at identifying noise sources with a wide frequency range. By analyzing the frequency domain structure and mutual interference relationships of noise, the system can construct a full-frequency domain noise interference topology table, detailing the characteristics of different noise sources, such as intensity, frequency range, and interference modes. This topology table provides a reliable data foundation for subsequent noise source separation and noise reduction strategy design. The system can handle complex noise environments, even when multiple noise sources coexist, maintaining accurate noise analysis and management to ensure that no important interference signals are missed during noise reduction. By effectively separating different types of noise in the audio signal (such as mechanical noise, human voice, and environmental noise), the system can formulate different noise reduction strategies for each noise source, thus processing each noise source more precisely. The construction of a real-time noise scene model allows the system to dynamically identify changes in noise sources and the environment. Changes in traffic noise intensity, fluctuations in wind noise, etc., are reflected in the noise scene model in real time, ensuring that noise reduction strategies are updated continuously. By modeling the temporal slope fluctuations of noise, the system can predict future trends in noise sources. Noise intensity may suddenly increase or decrease, and this change can be predicted in advance through temporal modeling. Based on the temporal fluctuations of noise, the system can optimize the response of the noise reduction algorithm in real time. When noise fluctuations intensify, the system can quickly switch to a stronger noise reduction mode; when noise weakens, the system can return to an energy-saving mode, thus ensuring efficient noise reduction while reducing power consumption. Through neural coding gain, the system can intelligently enhance audio signals, especially at low volumes or under noise masking, effectively improving audio clarity and intelligibility. By reconstructing the boundaries of audio signals, the system can more accurately preserve the temporal structure and frequency characteristics of audio signals. This is crucial for the fidelity of speech or music signals, avoiding problems such as blurring and loss of audio details during noise reduction. Through effective reconstruction, the system can seamlessly fuse the key audio signals of the coded gain with the optimized weak audio signals, thereby restoring the closest possible audio signal to reality. This ensures the naturalness and clarity of the audio, and the sound quality will not degrade due to the noise reduction process. During noise reduction, the system effectively preserves the original characteristics of the audio and avoids excessive "modification" of the audio signal. This ensures that the audio quality after noise reduction is as close as possible to the original quality, which is crucial, especially in applications such as speech recognition. Through iterative optimization, the system can learn from feedback on noise reduction effects in different environments, continuously improving the accuracy and adaptability of the noise reduction strategy. This allows the system to continuously optimize performance over long periods of operation, avoiding the limitations of fixed algorithms. Dynamically optimizing based on real-time changing noise patterns, the system provides optimal noise reduction results in both high-noise and quiet environments. Attached Figure Description

[0009] Figure 1 This is a flowchart illustrating the steps of an adaptive noise reduction method for a multimodal audio SoC main control chip according to the present invention. Figure 2 This is a detailed flowchart illustrating the implementation steps of step S1. Figure 3 This is a detailed flowchart illustrating the implementation steps of step S2; Figure 4 This is a flowchart illustrating the detailed implementation steps of step S3. Detailed Implementation

[0010] It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.

[0011] This application provides an adaptive noise reduction method and system for a multimodal audio SoC main control chip. The execution entities of the adaptive noise reduction method and system for the multimodal audio SoC main control chip include, but are not limited to, mechanical equipment, data processing platforms, cloud server nodes, network upload devices, etc., which can be considered as general computing nodes in this application. The data processing platform includes, but is not limited to, at least one of an audio image management system, an information management system, and a cloud data management system.

[0012] Please see Figures 1 to 4 This invention provides an adaptive noise reduction method for a multimodal audio SoC main control chip, which includes the following steps: Step S1: Collect the original audio input signal from the SoC main control chip, perform time-frequency domain dual deconstruction and cross-frequency domain noise interference structure analysis, and construct a full-frequency domain noise interference topology table; Step S2: Separate multimodal interference factors from the full-frequency domain noise interference topology table, model multiple noise environments, and construct a real-time noise scene model; Step S3: Model the temporal noise slope fluctuation of the real-time noise scene model, and perform dynamic noise reduction response learning optimization to construct a dynamic noise reduction optimization strategy; Step S4: Perform neural coding gain and audio boundary reconstruction based on the original audio input signal to obtain the key audio signal of coding gain and weak audio optimization signal; Step S5: Perform effective audio reconstruction on the key audio signal of the encoded gain and the weak audio optimization signal to construct the noise-reduced and optimized audio input signal; Step S6: Perform real-time audio noise reduction offset evaluation on the noise reduction optimized audio input signal, and perform iterative noise reduction migration optimization on the dynamic noise reduction optimization strategy to build an iterative noise reduction optimization engine.

[0013] This invention effectively analyzes the time and frequency characteristics of audio signals by performing a dual time-frequency domain deconstruction on the original audio input signal. This dual deconstruction allows for the simultaneous identification of noise and useful audio signals in both the time and frequency domains, thus more accurately determining the location and frequency distribution of noise sources. By analyzing the spread of noise interference across multiple frequency domains, noise sources with broad frequency coverage spanning multiple frequency bands can be identified. This analysis helps to build a more comprehensive noise interference topology table, providing strong support for subsequent noise reduction processing. The constructed topology table details the noise characteristics (such as intensity and frequency range) at different frequency bands, providing a precise data foundation for subsequent interference separation, modeling, and optimization. With the fusion of multimodal signals, different types of noise sources (such as traffic noise, wind noise, and voice interference) can be effectively separated. Wind noise may be more pronounced in the low-frequency band, while human voice is mainly in the mid-to-high frequency band; this separation allows for more accurate subsequent noise reduction. By modeling the interaction between different noise sources in real time, a more realistic and dynamic noise environment model can be established. This model not only works effectively in static noise environments but also handles dynamically changing noise scenarios (such as sudden changes in background noise), providing more flexible adaptability for denoising algorithms. Through real-time noise scenario models, the state and characteristics of noise sources can be continuously updated, ensuring that the denoising strategy remains optimal at all times. This real-time modeling capability greatly enhances its responsiveness and adaptability. By modeling the temporal fluctuations of noise scenarios, dynamic changes in noise intensity and frequency can be captured. Traffic noise may fluctuate with changes in traffic flow, or environmental noise may change due to wind speed variations. Accurately capturing these changes allows for rapid noise reduction responses. By learning and optimizing noise responses, intelligent adjustments can be made for different noise fluctuation patterns. During sudden noise increases, it can quickly switch to a high-efficiency denoising mode, while reverting to a low-power, low-latency mode during calm periods. As the noise environment changes, the denoising strategy can be adaptively adjusted to ensure continuous and stable denoising effects, especially when multiple noises coexist, enabling real-time strategy optimization. Through neural coding gain technology, the quality of audio signals can be enhanced without introducing excessive distortion. This means that not only can noise be removed, but the clarity of speech or audio can also be enhanced, making the denoised audio signal more natural and intelligible. By reconstructing the boundaries of the audio signal, the boundaries and details of the sound can be accurately restored, avoiding the "blurring" or "hollowing" phenomenon that occurs during the noise reduction process. This is crucial for the fidelity of speech or music signals, preserving more speech details or musical nuances. This technology allows low-volume or weak speech signals to be accurately amplified, improving the ability to recognize human voices, especially in noisy environments, ensuring that the speech signal is clearly discernible. After noise reduction processing, the useful audio signal can be effectively reconstructed.By preserving key audio signals and optimizing weak audio signals, a balance between signal clarity and detail can be achieved, making the audio output more realistic. The reconstructed audio signal allows for more balanced noise reduction, especially in complex, noisy environments, maximizing the restoration of the original sound and reducing human distortion. This stage allows for flexible adjustment of the noise reduction intensity, avoiding audio "compression" caused by excessive noise reduction and maintaining a natural audio presentation. Real-time offset evaluation of the noise reduction effect allows for immediate detection of deviations in the noise reduction strategy and necessary adjustments. This process ensures that the noise reduction effect remains optimal, especially under constantly changing environmental noise conditions. The iterative noise reduction optimization strategy continuously learns and updates the noise reduction model. In new noise environments, it can intelligently adapt through historical feedback, reducing manual intervention and adjustments. By establishing an iterative learning mechanism for the noise reduction optimization engine, it ensures that the noise reduction processing remains efficient during long-term operation, reducing the risk of failure in complex environments.

[0014] In the embodiments of the present invention, see Figure 1 This is a flowchart illustrating the steps of an adaptive noise reduction method for a multimodal audio SoC main control chip according to the present invention. In this example, the steps of the adaptive noise reduction method for the multimodal audio SoC main control chip include: Step S1: Collect the original audio input signal from the SoC main control chip, perform time-frequency domain dual deconstruction and cross-frequency domain noise interference structure analysis, and construct a full-frequency domain noise interference topology table; In this embodiment, the SoC chip integrates a MEMS microphone array input path (which can be single-channel or multi-array, with a sampling rate of 16kHz / 24kHz / 48kHz and a bit depth of 16bit / 24bit). Real-time sampling is achieved through an on-chip ADC module, forming a time-continuous raw audio waveform data stream. It is recommended to use an on-chip buffered FIFO + DMA transfer mechanism for non-blocking data scheduling, ensuring a processing latency of no more than 5ms. The raw audio input signal is input to the joint deconstruction engine, where a short-time Fourier transform (STFT) is first used for spectral decomposition. Using a 256-point FFT window (with a sliding step of 128 points), the spectrum of one time window is extracted for each frame. Then, a continuous wavelet transform (CWT) is applied for high-resolution temporal feature capture, particularly useful for detecting rapidly flashing impulse noise. In the deconstruction result, the system obtains a three-dimensional spectral tensor (Frame × Frequency × Power), which represents the trajectory of noise energy variation throughout the entire time-frequency space. To facilitate subsequent semantic modeling, all spectral tensors are normalized after extraction and band-packed according to frequency layers (e.g., 0–500Hz, 500Hz–2kHz, 2kHz–8kHz). Based on the dual-channel output of STFT and CWT, the system further performs statistical analysis on the background noise energy envelope and modulation rate characteristics of each frequency band. For periodic noise sources (e.g., fans, mechanical motors), the system identifies amplitude modulation frequencies in their spectra between 5–50Hz and constructs noise modulation envelope curves. The system also distinguishes between sudden noise (e.g., collision sounds) and constant background noise by calculating the signal's "zero-crossing rate (ZCR)" and "spectral rolloff," thereby forming a multi-type sound field interference vector library. Finally, the aforementioned spectral tensors and interference vectors are projected onto a "noise interference topology table," which uses frequency bands as the main axis and time windows as the horizontal axis, recording the noise type, energy intensity, modulation mode, and spatial propagation trend of each frequency band within each time window. Its structure is similar to a heat matrix or a relational network diagram, which can be used to reference the "identity" and "dynamic characteristics" of noise when quickly calling noise reduction strategies.

[0015] Step S2: Separate multimodal interference factors from the full-frequency domain noise interference topology table, model multiple noise environments, and construct a real-time noise scene model; In this embodiment, noise interference feature data, including frequency range, energy level, and temporal characteristics, is extracted from a full-frequency domain noise interference topology table. This ensures the data reflects the characteristics of various noise interference sources. Different types of noise features, such as ambient noise, traffic noise, and human voice interference, are extracted, and their frequency range, energy intensity, and occurrence time are recorded to ensure data integrity. Suitable signal separation techniques, such as Independent Component Analysis (ICA) or Non-negative Matrix Factorization (NMF), are selected to separate multimodal interference factors. These methods can effectively extract independent noise components from composite signals. Using ICA, the mixed audio signal is decomposed into multiple independent source signals, each corresponding to a different type of noise interference source. The selected separation algorithm is applied to process the extracted noise features. This process should ensure accurate separation of various noise interference sources and record the characteristic parameters of each source. During processing, if ICA successfully divides the signal into three categories, the feature information of ambient noise (e.g., low-frequency band), traffic noise (e.g., mid-frequency band), and human voice interference (e.g., high-frequency band) is recorded respectively. Environmental characteristic analysis is performed on the separated multimodal interference factors to identify the main noise sources affecting audio signal quality. This process should consider the intensity, frequency, and impact of the noise sources on the signal. The intensity variation of traffic noise over specific time periods is analyzed, and its peak periods and frequency characteristics are recorded for subsequent modeling. Appropriate modeling methods are selected, such as stochastic process models (e.g., Gaussian processes) or physics-based models (e.g., sound propagation models), to construct a multi-noise environment model. These models can simulate the behavior of different noise sources and their interactions. Using Gaussian process models can effectively capture the dynamic changes of noise sources in time and space. Based on the analysis results, a real-time noise scene model is constructed to ensure that the model can dynamically reflect the changing characteristics of noise sources. The model should include the location, intensity, frequency, and variation patterns of the noise sources. The model should record the distribution of various noise sources at different time periods, forming a dynamic model containing time, frequency, and spatial information.

[0016] Step S3: Model the temporal noise slope fluctuation of the real-time noise scene model, and perform dynamic noise reduction response learning optimization to construct a dynamic noise reduction optimization strategy; In this embodiment, energy slope modeling is performed on each activated noise mode (e.g., low-frequency rolling, high-frequency motor whistling, overlapping human voices, etc.) in the noise scene model along the time series dimension. The first derivative (ΔE / Δt) of the energy envelope of each frequency band is calculated within a time window (typically 25ms frame length, 10ms frame shift). The time slope data of multiple frequency bands are fused into a two-dimensional tensor representation. A sliding window smoothing technique (window width 3-5 frames) is used to handle local anomalous abrupt changes, forming a continuous fluctuation spectrum that can be used as model input. The noise mode vector + slope sequence + real-time SNR evaluation are used as state inputs; actions are different combinations of noise reduction parameters, such as time window width, frequency band filtering intensity, and speech enhancement gain. A composite reward mechanism is formulated by combining short-term speech clarity improvement (using PESQ / ESTOI metrics) and noise reduction gain (SNR improvement). Training employs Dueling DQN (a two-layer network structure that estimates state values ​​and the dominance function separately), which has fast convergence speed and is suitable for online inference environments with limited SoC chip resources. During training, various typical working scenarios were simulated, such as an open factory (reverberation + low-frequency rolling), a closed container environment (high-density human voice interference), and mechanical start-up noise scenarios, to train and fine-tune the policy network. After approximately 20,000 training iterations, the average SNR improved by about 4.8 dB and the PESQ improved by 0.62 on the simulation set, indicating that the policy has good generalization ability. After the above learning and optimization process, the system generates a multi-policy mapping table to match and cache different noise slope curves with their corresponding optimal action sequences, facilitating quick recall during online runtime. The following optimization mechanisms are adopted: for newly identified noise slope patterns, the initial response is initiated using the nearest neighbor policy matching method; when a sudden change in the current noise slope is detected (Δ slope is greater than a set threshold, such as ±2.1 dB / ms), the policy response window is quickly switched; when the policy execution effect is poor (PESQ decreases for two consecutive frames), it automatically rolls back to the previous optimal state.

[0017] Step S4: Perform neural coding gain and audio boundary reconstruction based on the original audio input signal to obtain the key audio signal of coding gain and weak audio optimization signal; In this embodiment, a lightweight temporal neural network compression coding model (TCN-Encoder) is used based on the raw, unprocessed audio stream acquired by the front end. This model can extract key audio segments with high semantic weights without destroying the temporal dependency structure. The raw audio signal is segmented into frames in 20ms increments, with each frame containing 320-512 sampling points (sampling rate 16kHz). After segmentation, overlapping processing (50% frame shift) is performed. A three-layer stacked one-dimensional convolutional module is used, with a receptive field set to 20-60ms. Variable convolutional kernel lengths (3 / 5 / 7) are used to capture semantic information at different speech rates. Weight scoring is performed based on the energy distribution of the encoded signal channels in each frame, and a gain factor (usually set in the range of 1.2-1.8) is applied to high-expression-density regions to achieve dynamic gain coding of high-semantic-energy segments. Speech often exhibits blurred boundaries or fragmented segments in complex background environments, especially during the transition phases from silence to speech, which are easily misidentified as background noise by filters and lost. To address this, this section introduces a **Boundary Learning-Based Speech Recovery Network (B-BNet)** to reconstruct audio boundary lines. Possible speech abrupt changes are identified by calculating the first derivative (ΔE) of the energy change in each frame, forming a candidate boundary set. A dual-channel Bi-GRU network (64 hidden units × 2) is used, inputting five frames of context energy information, to predict whether the current frame marks the start or end of real speech. For identified speech fragments, speech compensation and reconstruction are performed using semantic vectors from adjacent segments (extracted using the Transformer semantic compression module), with the interpolation length controlled within the 5-25ms range. The output of the boundary reconstruction is the weak speech optimization signal. These audio segments are often suppressed or deleted in traditional noise reduction processes, but are effectively preserved and enhanced in this scheme, thereby improving the overall structural coherence and information restoration accuracy of the speech.

[0018] Step S5: Perform effective audio reconstruction on the key audio signal of the encoded gain and the weak audio optimization signal to construct the noise-reduced and optimized audio input signal; In this embodiment, the key audio signal for coded gain and the weak audio optimization signal are extracted, ensuring that both signals are of good quality and have undergone preprocessing. The waveform and spectral characteristics of the signals should be checked to confirm they are within the dynamic range. The peak value of the coded gain signal is recorded as -10 dB, and the peak value of the weak audio optimization signal as -20 dB, ensuring they can be effectively blended together. Feature analysis is performed on these two signals to extract their main frequency components, time-domain characteristics, and instantaneous amplitude. This step helps to understand the signal characteristics and provides a basis for the subsequent reconstruction process. The coded gain signal is analyzed using FFT to obtain its frequency component distribution, and the main frequency peaks (such as 500 Hz and 1500 Hz) are recorded for consideration during subsequent reconstruction. A suitable audio reconstruction method is selected, typically using Inverse Fourier Transform (IFFT) or signal synthesis, to reconstruct the coded gain signal and weak audio optimization signal into a high-quality audio signal. The IFFT window size is set to 2048 points to ensure sufficient frequency resolution and time-domain detail during reconstruction. A synthesis strategy is determined, including the signal superposition method. Methods such as linear superposition and weighted superposition can be used to ensure the preservation of the characteristics of different signals. The weight of the coded gain signal is set to 0.7, and the weight of the weak audio optimization signal is set to 0.3. The final output signal is generated through weighted synthesis. The selected reconstruction method is applied to synthesize the coded gain key audio signal and the weak audio optimization signal to generate a noise-reduced optimized audio input signal. The synthesis process is ensured to effectively preserve the semantic information and dynamic range of the audio. The two signals are superimposed using weighted superposition to generate a new audio signal, and the amplitude and frequency characteristics of the final signal are recorded. The reconstructed audio signal is verified to ensure that it meets the sound quality requirements. Evaluation can be performed through subjective listening tests and objective indicators (such as signal-to-noise ratio and clarity index). Listening tests are conducted to ensure that the reconstructed signal has improved clarity and intelligibility, and the test results are recorded for subsequent analysis. The reconstructed noise-reduced optimized audio input signal is evaluated to check its performance in different environments, ensuring that it can effectively resist noise interference. The signal-to-noise ratio changes of the signal under different background noise conditions are recorded to ensure that the final signal's SNR reaches at least 15 dB.

[0019] Step S6: Perform real-time audio noise reduction offset evaluation on the noise reduction optimized audio input signal, and perform iterative noise reduction migration optimization on the dynamic noise reduction optimization strategy to build an iterative noise reduction optimization engine.

[0020] In this embodiment, based on the ITU-T P.862 PESQ model, and combining the short-time energy method and the spectral smoothness index, the perceptual difference (ΔPESQ) between the denoised speech signal and the noise-free reference speech is calculated. The evaluation window is set to 256ms and the step size is 64ms. Frames with ΔPESQ greater than 0.45 (indicating severe noise residue) or spectral fluctuation residuals greater than ±3dB are designated as outliers. The standard deviation of residual fluctuation (σ_res) within a 5-second sliding time window is introduced to determine whether a transfer learning optimization update is needed. Once a denoising offset is detected, the system enters the "transfer learning denoising optimization" process. The core idea of ​​this mechanism is to transfer and optimize the static noise reduction strategy obtained from the original training to the current sound field environment, forming a task-scene-specific model. Based on the interference features extracted from the noise scene model (such as periodic frequencies and background harmonic components), the target frequency band attenuation function of the current noise reduction strategy is compared. If the peak offset of the two exceeds ±80Hz or the energy attenuation inconsistency exceeds 20%, it is judged as a strategy mismatch. The strategy optimization network with LSTM structure is used to generate a new noise reduction weight update path based on the residual fluctuation trend. In the experiment, the network consists of two hidden layers (128 units per layer). Incremental learning is completed in local frame segments through a lightweight backpropagation module (only the last two layer parameters are activated when running in the chip). Usually, a strategy fine-tuning cycle is controlled within 600ms, which will not have a significant impact on the chip response. Using a BERT-based sound field semantic compression encoder, a 5-second noise-reducing input signal is compressed into a 32-dimensional vector representing the current acoustic environment. The closest encoded feature from the historical policy library (Euclidean distance threshold < 0.25) is matched and used as the initial noise reduction value, significantly accelerating system convergence. A low-power monitoring module in the SoC performs a soft update evaluation of the policy effect every 10 seconds, marking the optimal policy as the current preferred weight. In experiments, the engine demonstrated a 0.43 improvement in mean PESQ and a 93.5% improvement in speech recognition accuracy in three complex dynamic environments (factory machinery area, densely populated area, and noisy traffic area), proving its strong robustness and adaptability to sound field drift and audio variability. In this embodiment, see Figure 2 The diagram below illustrates the detailed implementation steps of step S1. In this embodiment, the detailed implementation steps of step S1 include: The original audio input signal is acquired based on the SoC main control chip; The original audio input signal is decomposed into multiple frequencies to obtain multiple audio frequency layered time windows; Multiple audio frequencies are layered into time windows and then deconstructed in both the time and frequency domains to obtain the spectral distribution tensor and the sound field interference vector. Dual-modal deep audio semantic mining is performed on the spectral distribution tensor and sound field interference vector to extract semantically valid audio vectors and audio noise feature vectors. Based on the audio noise feature vector, cross-frequency domain noise interference structure analysis is performed, and a full-frequency domain noise interference topology table is constructed.

[0021] In this embodiment, the SoC main control chip first completes the audio acquisition task of the external sound field environment. The audio front-end module integrated in the SoC chip typically includes a high-sensitivity analog-to-digital converter (ADC) and an audio interface controller (such as PDM, I²S, TDM interfaces), which can connect to a multi-channel microphone array (such as a 24-channel linear array or beamforming array) to obtain rich sound source information. Sampling parameters often use a 16-bit quantization depth and a 16kHz-48kHz sampling rate, which can be dynamically adjusted according to the environment. To ensure the integrity and robustness of the acquired signal, dynamic gain control (AGC) and DC offset calibration modules are typically used for pre-amplifier signal optimization. During signal acquisition, to reduce CPU load and ensure low-latency transmission, the SoC chip automatically moves the audio data to the SRAM buffer via a DMA controller and caches it in frames in the Ring Buffer. Each frame is typically 1020ms (160320 samples), corresponding to the response requirements of tasks such as AGV voice control. The system periodically triggers interrupts (such as a 1ms interrupt cycle) for data upload and processing. Choose a suitable frequency decomposition algorithm, such as Short-Time Fourier Transform (STFT) or Wavelet Transform, to perform multi-frequency decomposition on the original audio signal. These algorithms can provide effective decomposition results in both the time and frequency domains. Using Wavelet Transform, the instantaneous frequency features of the signal can be effectively extracted, better handling non-stationary signals. Decompose the original audio signal to obtain multiple audio frequency-layered time windows with different frequency ranges. Each time window should cover a different frequency range of the audio signal for subsequent analysis. Set the frequency range as low frequency (0-200Hz), mid frequency (200-2000Hz), and high frequency (2000-20000Hz), with each time window lasting 50ms to ensure effective capture of audio features. Organize the multiple audio frequency-layered time windows obtained from the decomposition and store them as structured data for subsequent time-frequency domain analysis. Create a three-dimensional array, where the first dimension is the frequency layer, the second dimension is the time window, and the third dimension is the signal amplitude, to facilitate subsequent calculations and processing. Choose an appropriate time-frequency domain analysis method, such as spectrogram generation, to simultaneously analyze the distribution characteristics of the audio signal in time and frequency. Use tools such as matplotlib to generate time-frequency plots to visualize the energy distribution of each frequency layer within different time windows. Calculate the spectral distribution for each audio frequency layer within its time windows, forming a spectral distribution tensor. This tensor should contain the energy distribution of each frequency layer under different time windows. For each frequency layer, calculate its energy spectral density (PSD) and organize the results into a multidimensional array representing frequency, time, and energy. Based on the time-frequency domain analysis, extract the sound field interference vector to describe the interference characteristics of the audio signal. This process should consider factors such as the sound source location and environmental reflections.By analyzing the spectral distribution, abnormal peaks in the spectrum are identified and labeled as sound field interference. Their specific frequency and amplitude information are recorded to form a sound field interference vector. Deep mining is performed on the spectral distribution tensor and the sound field interference vector to extract semantically valid audio vectors and audio noise feature vectors. Audio feature extraction is performed using a deep learning model (such as a convolutional neural network). A convolutional neural network is trained to learn features from the spectrogram, extracting semantic and noise features of the sound. The extracted audio features are organized into vector format, representing semantic validity and noise features respectively. Each vector should contain numerical and descriptive information of the relevant features for subsequent analysis. A vector containing 10 features is generated to describe the pitch, loudness, rhythm, etc. of the audio, and a noise feature vector is generated to describe the frequency and intensity of the interference. The effectiveness and accuracy of the extracted audio feature vectors are verified by analysis, ensuring that the extracted features can truly reflect the semantic and noise characteristics of the audio signal. A k-fold cross-validation method is used to evaluate the performance of the extracted features in an audio classification task, ensuring that they can effectively distinguish different types of audio signals. Choose appropriate noise feature analysis methods, such as Principal Component Analysis (PCA) or Independent Component Analysis (ICA), to conduct in-depth analysis of audio noise feature vectors and identify the structural characteristics of noise interference. Use PCA to reduce the dimensionality of audio noise features and extract the principal components to help identify the main sources of noise interference. Analyze the noise interference features in different frequency ranges and identify their impact on the overall audio signal. This step should consider the correlation between frequency domains and construct a cross-frequency domain structure of noise interference. During the analysis, record the impact of low-frequency noise on mid- and high-frequency audio and establish an interference model between frequency domains. Based on the analysis results, construct a full-frequency domain noise interference topology table to describe the noise interference relationships between each frequency band. The topology table should be able to visually display the mutual influence and interference intensity of noise in different frequency bands. Generate a two-dimensional table to record the noise characteristics of each frequency band and the interference coefficients of other frequency bands, forming a comprehensive noise interference topology structure.

[0022] In this embodiment, see Figure 3 The diagram below illustrates the detailed implementation steps of step S2. In this embodiment, the detailed implementation steps of step S2 include: Multimodal interference factors are separated from the full-frequency domain noise interference topology table to generate multiple noise interference nodes; Noise envelope feature analysis is performed on multiple noise interference nodes to obtain the envelope feature curve of each interference node; For each interference node, the envelope characteristic curve is analyzed in the main interference noise frequency band to generate the main interference noise frequency band characteristics. Scene noise pattern analysis is performed based on the characteristics of the main interference noise frequency band to generate dynamic noise patterns for the real-time scene. Multi-noise environment modeling is performed on the dynamic noise patterns of real-time scenes to construct a real-time noise scene model.

[0023] In this embodiment, the full-frequency domain noise interference topology table is organized to ensure data integrity and accuracy. Each noise interference node should include its corresponding frequency range, interference intensity, and related characteristics. A structure containing 10 noise interference nodes is constructed, with each node recording its frequency range (e.g., 0-200Hz, 200-400Hz, etc.) and its corresponding interference value (e.g., dBSPL). Multimodal separation techniques, such as Independent Component Analysis (ICA) or Non-negative Matrix Factorization (NMF), are used to separate the noise interference nodes. These methods can effectively extract independent interference components from composite signals. The ICA algorithm is used to analyze the noise interference signal to identify each independent interference source, ensuring that each interference node accurately reflects its original signal characteristics. Based on the separation results, multiple noise interference nodes are generated, each representing an independent noise interference source. The frequency characteristics, time characteristics, and intensity information of each node are recorded for subsequent analysis. If the separation results show three main interference sources, three independent noise interference nodes are generated, and their characteristics, including frequency peaks and corresponding interference intensities, are recorded respectively. Choose a suitable envelope feature extraction method, such as Hilbert transform or envelope detection algorithm, to analyze the envelope features of each noise interference node. Envelope features reflect the overall trend of noise signal changes. Use Hilbert transform to process the signal of each noise interference node, extracting its envelope curve to represent the amplitude change of the signal. Generate an envelope feature curve for each noise interference node. Each curve should clearly show the periodic changes and instantaneous amplitude of the noise signal, facilitating subsequent analysis of the main interference noise frequency band. The analyzed envelope curves can show the intensity changes of the noise signal within a specific time period, forming a visualized graph, recording the maximum, minimum, and their positions. Record the envelope feature curves of each interference node in a database and display them through visualization tools to facilitate operator analysis and understanding of noise characteristics. Generate a set of charts containing the envelope features of all interference nodes, displaying the envelope curve and characteristic parameters of each node to aid in subsequent analysis. Use spectral analysis methods (such as FFT or STFT) to analyze the noise envelope feature curves for the main interference noise frequency band to identify the main interference frequency components. The envelope feature curves are converted to the frequency domain using Fast Fourier Transform (FFT) to identify the main peaks in the spectrum. The main interference frequency band features are extracted from the spectrum, and the center frequency, bandwidth, and interference intensity of each band are recorded. These features will be used for subsequent noise pattern analysis. If the spectrum of a noise interference node shows a significant peak at 500Hz, the center frequency of that band is recorded as 500Hz, and the interference intensity as 85dB. The main interference noise frequency band features are compiled and recorded in a database for subsequent scene noise pattern analysis. A table containing all the main interference frequency bands is generated, recording the corresponding feature parameters for subsequent analysis and model building.Based on the frequency characteristics of the main interference noise band, a noise pattern analysis framework is established to analyze noise characteristics under different scenarios. This framework should be able to correlate different interference frequency bands with scenario characteristics. Scenario classification criteria are set, such as "traffic noise," "industrial noise," and "environmental noise," and classification is performed based on frequency band characteristics. Data mining techniques are used to statistically analyze noise characteristics under different scenarios, generating dynamic noise patterns for real-time scenarios. These patterns should reflect real-time environmental noise changes. Noise data within a specific time period is analyzed to generate dynamic noise pattern diagrams, showing the changes in noise intensity under different scenarios. Based on the dynamic noise patterns, a real-time noise scenario model framework is designed. This model should be able to integrate various noise characteristics and reflect changes in environmental noise in real time. The model structure is defined, including an input layer (noise characteristics), a processing layer (noise pattern analysis), and an output layer (noise scenario model). Model parameters are determined and optimized to improve the model's adaptability and accuracy. The impact of scenario changes on noise characteristics should be considered to ensure the model can be updated in real time. Model parameters are adjusted based on historical data to more accurately reflect the current environmental noise state. The constructed real-time noise scenario model was validated to ensure its ability to effectively predict and reflect noise changes in the actual environment. Model calibration was performed using actual noise data. Noise data collected from real-time environmental monitoring equipment was used to validate the model and confirm its accuracy and effectiveness under different scenarios.

[0024] In this embodiment, the specific steps for modeling the dynamic noise patterns of the real-time scene using multiple noise environments and constructing a real-time noise scene model are as follows: A deep time-frequency domain variation analysis was performed on the spectral distribution tensor and the sound field interference vector to identify the dynamic change characteristics of the noise source and its movement path. The dynamic change characteristics are subjected to noise frequency trend evolution to generate a noise frequency change trend; Calculate the noise source locations at multiple time points based on the movement path; Based on the noise source locations at the multiple time points, spatial location situation prediction is performed to obtain noise source location prediction data; Based on the noise frequency variation trend and noise source location prediction data, the dynamic change law of the noise source is evolved, thereby generating a dynamic change logic diagram of the noise source. Based on the dynamic change logic diagram of noise sources, a multi-noise environment model is constructed for the dynamic noise pattern of real-time scenes to build a real-time noise scene model.

[0025] In this embodiment, data on the spectral distribution tensor and the acoustic interference vector are collected. This data should include frequency distribution information and interference characteristics within different time windows for subsequent in-depth analysis. The spectral distribution tensor can be a three-dimensional array, where the first dimension is the frequency layer, the second dimension is the time window, and the third dimension is the signal strength. The acoustic interference vector records the interference intensity at each frequency layer. Time-frequency analysis techniques, such as Continuous Wavelet Transform (CWT) or Short-Time Fourier Transform (STFT), are used to perform in-depth change analysis on the spectral distribution tensor. These methods can effectively capture the signal's time and frequency variation characteristics. CWT is used to analyze the dynamic changes in different frequency bands, obtaining the temporal local features of the spectrum to identify the dynamic changes of noise sources. Through the above analysis, the dynamic change characteristics of noise sources are extracted, such as the rate of frequency change, the rate of amplitude change, and the duration. These features will be used for subsequent noise frequency trend evolution analysis. If a frequency band is found to increase rapidly within a specific time period, the rate of frequency change of that band can be recorded as part of the dynamic change characteristics. Choose an appropriate trend analysis method, such as linear regression, time series analysis, or exponential smoothing, to analyze the extracted dynamic change characteristics and generate a noise frequency change trend. Use linear regression to analyze the noise frequency change trend over time and identify patterns of frequency increase or decrease. Based on the extracted dynamic change characteristics, generate a noise frequency change trend graph to show the evolution of noise frequency over different time periods. This graph should clearly reflect the frequency change pattern. Generate a line graph containing time (X-axis) and frequency (Y-axis) to indicate the rise and fall of frequency over a specific time period. Based on the dynamic change characteristics from the previous step, integrate the noise source's movement path data. This data should include the noise source's position coordinates and their change characteristics at different time points. Record the noise source's position change every second for subsequent calculations and analysis. Calculate the noise source's position at multiple time points based on the movement path. The noise source's speed and direction should be considered in this process to accurately predict its position at a specific time point. If the noise source's speed during the process is 1 m / s, and at a certain time point it is (10, 20), then the position at subsequent time points can be calculated using a formula. Choose an appropriate prediction model, such as a Kalman filter or an autoregressive moving average (ARMA) model, to predict the spatial location and situation of the noise source. The Kalman filter effectively handles dynamic changes and uncertainties to predict the future location of the noise source in real time. Based on the previously calculated time-point locations, use the selected prediction model to perform state estimation, obtaining the noise source's predicted location data. This data should include the estimated location of the noise source at future time points. Using the Kalman filter, predict the noise source's position change within the next 5 seconds based on its velocity and current state. Employ data mining techniques, such as cluster analysis or association rule mining, to analyze the dynamic changes of the noise source. These methods help identify patterns and regularities in the noise source's behavior.Cluster analysis is used to classify the characteristics of noise sources across different time periods to uncover potential patterns. Based on this analysis, a dynamic change logic diagram of noise sources is generated, illustrating the behavioral relationships of noise sources at different times and locations. The logic diagram should clearly reflect the changing patterns of noise sources. A network diagram containing the states and changing relationships of different noise sources is generated, demonstrating the dynamic changes of noise sources under specific conditions. Based on the dynamic change logic diagram of noise sources, a multi-noise environment modeling framework is designed. This framework should be able to integrate the dynamic characteristics of different noise sources and reflect their impact on the environment. A model structure is designed, including an input layer (dynamic change characteristics), a processing layer (noise source behavior analysis), and an output layer (noise environment model). Model parameters are determined and optimized to improve the model's adaptability and accuracy. The impact of environmental changes on noise characteristics should be considered to ensure the model can be updated in real time. Model parameters are adjusted based on historical data to more accurately reflect the noise state of the current environment. The constructed real-time noise scene model is validated to ensure it can effectively predict and reflect noise changes in the actual environment. Model calibration is performed using actual noise data. Noise data collected using real-time environmental monitoring equipment was used to validate the model and confirm its accuracy and effectiveness in different scenarios.

[0026] In this embodiment, see Figure 4 The diagram below illustrates the detailed implementation steps of step S3. In this embodiment, the detailed implementation steps of step S3 include: Noise topology analysis is performed on a real-time noise scene model to extract unstructured noise vectors and structured repetitive noise. Calculate the noise intensity and noise periodic variation characteristics of the structured repetitive noise; Based on the noise intensity and noise periodic variation characteristics, a time-series noise slope fluctuation model is performed to construct a periodic disturbance trajectory diagram of structural noise; The long-term noise interference frequency range is calculated by analyzing the periodic disturbance trajectory diagram to obtain the long-term noise periodic frequency range. Adaptive noise reduction threshold analysis is performed on the long-term noise cycle frequency range to obtain adaptive noise reduction parameters for structural noise; High-frequency filtering and denoising are performed on unstructured noise vectors to obtain high-frequency filtering and denoising parameters for unstructured noise. Based on the adaptive noise reduction parameters and the high-frequency filtering noise reduction parameters, dynamic noise reduction response learning and optimization are performed to construct a dynamic noise reduction optimization strategy.

[0027] In this embodiment, noise data is extracted from a real-time noise scene model, ensuring that the data includes different types of noise features, such as unstructured noise vectors and structured repetitive noise. This data should cover a certain time range for comprehensive analysis. Noise data from the past 30 minutes is collected, recording the noise intensity and frequency characteristics per second, forming a matrix containing multi-dimensional data. Signal processing techniques are used to classify the noise signals. Methods such as wavelet transform or Fourier transform can be used to decompose the noise signals into different frequency components, thereby identifying unstructured noise vectors and structured repetitive noise. Wavelet transform is used to analyze the signal, extracting periodic structured noise and highly random unstructured noise, and recording their characteristic parameters. For the extracted structured repetitive noise, its noise intensity is calculated. Root mean square (RMS) or sound pressure level (SPL) indicators can be used to quantitatively describe the noise intensity. If the average SPL of a certain structured noise over 1 minute is 85 dB, this value is recorded, and its impact on the environment is analyzed. Periodic analysis is performed on the structured repetitive noise to extract its periodic variation characteristics. Noise cycles can be identified through autocorrelation function or periodogram analysis. Calculate the autocorrelation function of the noise signal to find the time interval of the repeating cycle and record its period value (e.g., a period of 2 seconds). Choose an appropriate modeling method, such as linear regression or time series analysis, to capture the slope fluctuation characteristics of noise intensity and cycle changes. Use time series analysis tools to model the change of noise intensity over time and identify the slope fluctuations of noise intensity. Based on the above modeling results, generate a periodic disturbance trajectory diagram of the structured noise, showing the trend of noise intensity change over time and its slope fluctuations. Plot a line graph of noise intensity change over time, marking the intensity values ​​and slope changes at different time points to help analyze the dynamic characteristics of the noise. Select a spectrum analysis method to calculate the long-term noise interference frequency range of the periodic disturbance trajectory diagram. Methods such as FFT or STFT can be used to analyze the frequency domain characteristics of the noise signal. Use FFT to analyze the noise signal and identify the main frequency components of long-term noise. Based on the spectrum analysis results, extract the frequency range of long-term noise and record its center frequency and bandwidth. This process should consider the spectral characteristics of the noise signal and the interference effects. The center frequency of long-term noise is recorded as 500Hz, and the bandwidth as 50Hz, forming frequency range data. A suitable analysis method, such as statistical analysis or machine learning algorithms, is selected to perform adaptive noise reduction threshold analysis for the long-term noise's periodic frequency range. An adaptive noise reduction model is trained using a machine learning model based on the data characteristics of the long-term noise. Based on the analysis results, adaptive noise reduction parameters for structural noise are generated. These parameters should be adjustable in real time to adapt to changes in different noise environments. If the analysis results show that the average noise intensity within the frequency range is 80dB, then the corresponding noise reduction threshold is set to 75dB for effective noise reduction.Choose a suitable high-frequency filtering method, such as a Butterworth filter or a Chebyshev filter, to denoise unstructured noise vectors. Using a Butterworth filter, set the cutoff frequency to 1000Hz to filter out noise above this frequency. Apply the high-frequency filter to the unstructured noise vector to obtain the denoised signal. Ensure the filtering process does not affect the effective information of the signal. After applying the filter, record the signal strength and frequency characteristics after denoising to evaluate the denoising effect. Choose a suitable learning algorithm, such as reinforcement learning or adaptive filtering algorithm, to achieve dynamic denoising response learning optimization. Use the adaptive filtering algorithm to adjust the denoising parameters in real time to adapt to changes in noise in the environment. Based on the aforementioned analysis results, combine the adaptive denoising parameters and the high-frequency filtering denoising parameters to construct a dynamic denoising optimization strategy. This strategy should be able to respond to environmental changes in real time and adjust the denoising strategy accordingly. Set the system to automatically increase the denoising intensity when the noise intensity exceeds a set threshold and update the denoising parameters in real time. Record the constructed dynamic denoising optimization strategy in a database and display it through visualization tools to help operators monitor the denoising effect. Generate a dynamic noise reduction strategy diagram to show the noise reduction response under different environments, which facilitates real-time adjustment and optimization.

[0028] In this embodiment, step S4 includes the following steps: Dynamic amplitude spectrum recognition is performed on semantically valid audio vectors to extract dynamic amplitude spectrum characteristics; Based on dynamic amplitude spectrum characteristics, audio semantic characteristics are analyzed to extract key audio activity blocks and weak audio semantic blocks. Neural coding gain is applied to key audio activity blocks to obtain the coded gain key audio signal; Audio boundary lines are reconstructed from weak audio semantic blocks to obtain audio semantic boundary lines; Morphological restoration is performed based on audio semantic boundary lines to generate weak audio optimized signals.

[0029] In this embodiment, a Fast Fourier Transform (FFT) is used to transform the audio vector in the frequency domain to calculate its dynamic amplitude spectrum. The amplitude spectrum represents the energy distribution of the signal at various frequencies and reflects the characteristics of the audio. With a sampling frequency of 44.1 kHz, the amplitude spectrum of the audio signal is calculated using FFT, obtaining amplitude values ​​at different frequencies and forming an amplitude spectrum matrix. Dynamic amplitude spectrum characteristics, including instantaneous amplitude changes and frequency component changes, are extracted from the calculated amplitude spectrum. These characteristics will be used for subsequent audio semantic feature analysis. The changes in the amplitude spectrum within a specific time window are analyzed to identify the main energy change points, and amplitude peaks and their corresponding time positions are recorded. Appropriate analysis methods, such as cluster analysis or pattern recognition, are selected to perform audio semantic feature analysis based on the dynamic amplitude spectrum characteristics. This step aims to identify key and weak activity blocks in the audio. Using the K-means clustering algorithm, the amplitude spectrum characteristics are divided into multiple categories to identify different audio activity blocks. Based on the analysis results, key audio activity blocks are extracted. These blocks typically have significant semantic features, such as prominent pitch changes or obvious volume increases. If an audio segment exhibits significant energy concentration at 2000Hz, it can be identified as a key audio activity block, and its start and end times recorded. Simultaneously, weak audio semantic blocks are extracted; these blocks typically have low energy and are difficult to identify. A similar method is used to analyze the amplitude spectrum to identify audio segments with lower energy. If the amplitude of an audio segment is below 50dB, it can be labeled as a weak audio semantic block. A neural network model, such as a convolutional neural network (CNN), is used to perform neural coding gain processing on the key audio activity blocks. This step aims to enhance important audio information and improve its signal-to-noise ratio. A CNN model is designed, taking the amplitude spectrum of the key audio activity block as input and the enhanced audio signal as output. The key audio activity block is processed through the neural network model to obtain the coded gain key audio signal. It is ensured that the processed signal remains semantically valid. The processed signal may increase in volume by 10dB, thereby enhancing its intelligibility in mixed audio environments. The coded gain key audio signal is recorded in a database, and its features are displayed using visualization tools to help understand the processing effect. A suitable audio boundary reconstruction method, such as an edge detection algorithm, is selected to analyze the boundary features of the weak audio semantic blocks. This process aims to clearly define the start and end positions of the audio. The Canny edge detection algorithm is used to identify the boundaries of weak audio segments. Audio semantic boundary lines are obtained through boundary line reconstruction. These boundary lines should accurately reflect the start and end positions of the weak audio blocks. The start time of the weak audio block is recorded as 3 seconds, and the end time as 5 seconds. Morphological processing methods, such as dilation and erosion, are selected to perform morphological repair processing on the audio semantic boundary lines. This step aims to optimize the signal morphology of the weak audio blocks and reduce the impact of noise. Morphological dilation is used to enhance the signal strength of the weak audio blocks, making them more prominent.morphological restoration is used to generate optimized weak audio signals, ensuring a significant improvement in signal quality. The optimized signal exhibits higher signal strength in spectral analysis, improving its intelligibility within the overall audio. A suitable learning algorithm, such as reinforcement learning or adaptive filtering, is selected to optimize the dynamic noise reduction response. This step aims to continuously adjust the noise reduction strategy based on environmental changes. Adaptive filtering algorithms are used to adjust noise reduction parameters in real time to adapt to different noise environments. Based on the aforementioned analysis results, a dynamic noise reduction optimization strategy is constructed by combining adaptive noise reduction parameters and high-frequency filtering noise reduction parameters. This strategy should be able to respond to environmental changes in real time to ensure optimal noise reduction performance. The noise reduction intensity is automatically adjusted when the noise intensity exceeds a set threshold to achieve the best noise reduction effect. The constructed dynamic noise reduction optimization strategy is recorded in a database and displayed using visualization tools to help operators monitor the noise reduction effect. A dynamic noise reduction strategy graph is generated to show the noise reduction response under different environments, facilitating real-time adjustment and optimization.

[0030] In this embodiment, step S5 includes the following steps: Effective audio reconstruction is performed on the key audio signal with coded gain and the weak audio optimization signal to construct an effective gain reconstructed audio signal. Adaptive noise reduction processing is performed on the audio noise feature vector based on the dynamic noise reduction optimization strategy to extract adaptive noise-reduced audio. Cross-modal audio fusion is performed on the adaptive noise-reduced audio signal and the effective gain reconstructed audio signal to construct a noise-reduced optimized audio input signal.

[0031] In this embodiment, a key audio signal for coded gain and a weak audio optimization signal are collected. These signals should have undergone the previous processing steps and contain enhanced audio information and optimized weak audio segments. The signal after coded gain can be an enhanced audio file with a signal strength increased by 10dB compared to the original signal, while the weak audio optimization signal is a morphologically repaired signal. A suitable audio reconstruction method, such as Inverse Fourier Transform (IFFT) or synthesis, is selected to effectively reconstruct the coded gain signal and the weak audio signal. This process aims to merge the two signals into a complete audio signal. IFFT is used to convert the frequency domain signal back to the time domain, ensuring that the reconstructed signal retains the overall characteristics and semantic information of the audio. The key audio signal for coded gain is combined with the weak audio optimization signal to generate an effectively gained reconstructed audio signal. During reconstruction, the phase and amplitude information of the audio signal must be preserved to maintain sound quality. The amplitudes of the two signals are added by linear superposition to form a new audio signal, ensuring good performance in the spectrum. Audio noise feature vectors are collected; these vectors should contain the frequency, intensity, and other relevant features of the noise signal for adaptive noise reduction processing. The noise intensity in the noise feature vector is recorded as 80dB, with a frequency range of 300Hz to 3000Hz, ensuring data accuracy. Adaptive noise reduction processing is applied to the audio noise feature vector according to a preset dynamic noise reduction optimization strategy. This strategy should be able to adjust the noise reduction parameters in real time according to environmental changes. If the ambient noise intensity exceeds a set threshold, the noise reduction intensity is automatically increased, and the noise reduction parameters are adjusted to achieve the best noise reduction effect. After applying the dynamic noise reduction optimization strategy, the denoised audio signal is extracted, ensuring that the influence of background noise is reduced while retaining effective audio information. The denoised audio signal can reduce the noise intensity to 60dB while maintaining the integrity of key speech information. A suitable audio fusion method, such as a weighted average or a deep learning model, is selected to perform cross-modal fusion of the adaptively denoised noise audio and the effective gain reconstructed audio signal. Using a weighted average, weights are assigned according to the energy and importance of the signals, and the two signals are merged. The adaptively denoised noise audio and the effective gain reconstructed audio signal are fused to generate a new noise-optimized audio input signal. The fusion process ensures that the semantic information and quality of the audio are preserved to the greatest extent possible. The weight of the effective gain signal is set to 0.7, and the weight of the noise reduction signal is set to 0.3. The final audio signal is generated by weighted averaging.

[0032] In this embodiment, step S6 includes the following steps: Calculate the audio signal-to-noise ratio, speech intelligibility index, and semantic fidelity of the noise-reduced audio input signal; Audio distortion quantization analysis is performed based on semantic fidelity to generate audio distortion quantization values; Real-time audio noise reduction offset evaluation is performed on the audio signal-to-noise ratio, speech intelligibility index, and audio distortion quantization value to obtain the audio noise reduction offset evaluation value. Based on the audio noise reduction offset evaluation value, noise reduction feedback analysis is performed to generate real-time audio noise reduction feedback information; Based on real-time audio noise reduction feedback, the dynamic noise reduction optimization strategy is iteratively migrated and optimized to build an iterative noise reduction optimization engine.

[0033] In this embodiment, noise-reduced and optimized audio input signals are collected to ensure they contain effective speech information and background noise. The signal should have undergone the previous processing steps and contain clear speech with corresponding noise. The noise-reduced audio signal is recorded to ensure that its signal-to-noise ratio (SNR) and intelligibility index can be effectively evaluated over multiple time periods. The signal-to-noise ratio (SNR) is calculated using the formula: SNR = 10 * log10(P_signal / P_noise), where P_signal is the effective signal power and P_noise is the noise power. The power of each signal is calculated by analyzing the time-domain waveform or frequency-domain characteristics of the audio signal. If the calculated effective signal power is 0.5 W and the noise power is 0.05 W, then the SNR is 10 * log10(0.5 / 0.05) = 10 dB, and this value is recorded for subsequent analysis. Intelligibility is evaluated using speech intelligibility metrics (such as CSIG or COVL). The CSIG calculation formula typically involves the spectral and time-domain characteristics of the signal, using a weighted average method to derive the intelligibility index. If the CSIG (Critical Speech Intelligence Index) is calculated to be 4.2 (out of 5), this value is recorded, reflecting the clarity and intelligibility of the audio signal. Distortion quantization analysis is then performed on the audio signal based on semantic fidelity. Semantic fidelity can be assessed by its similarity to the original signal, typically quantified using mean squared error (MSE) or signal correlation coefficient (CC). By comparing the similarity between the denoised signal and the reference signal, an MSE of 0.02 is calculated and recorded as the audio distortion quantization value. A comprehensive evaluation method is chosen to calculate the audio denoising offset evaluation value. This evaluation typically combines signal-to-noise ratio (SNR), CSIG, and audio distortion quantization value for a comprehensive analysis. The evaluation formula is set as: Offset Evaluation Value = α * SNR + β * CSIG - γ * Distortion, where α, β, and γ are the weights of each indicator. The denoising offset evaluation value is calculated in real-time based on the calculated SNR, CSIG, and distortion quantization value. The calculation formula is ensured to accurately reflect the overall quality of the audio signal. If the SNR is 10 dB, CSIG is 4.2, and the distortion quantization value is 0.02, the offset evaluation value is calculated to be 8.5 after weighting. This value is recorded for subsequent analysis. Based on the audio noise reduction offset evaluation value, noise reduction feedback analysis is performed to generate real-time audio noise reduction feedback information. This feedback information should reflect the effectiveness and shortcomings of the current noise reduction strategy. If the offset evaluation value is lower than the set threshold, feedback information is issued indicating that the noise reduction intensity needs to be increased. The generated real-time audio noise reduction feedback information is recorded in the database for subsequent analysis and strategy adjustment. The feedback information content, including the offset evaluation value and suggested adjustment measures, is recorded to support subsequent optimization decisions. Based on the real-time audio noise reduction feedback information, the dynamic noise reduction strategy is iteratively optimized. This ensures that the strategy can adaptively adjust according to the real-time evaluation value and feedback information to improve the noise reduction effect.If feedback indicates poor noise reduction performance, increase noise reduction efforts in the high-frequency bands and adjust noise reduction parameters in real time. Build an iterative noise reduction optimization engine to ensure it can respond to environmental changes and feedback in real time, continuously optimizing the noise reduction strategy. This engine should incorporate machine learning techniques, enabling it to learn and improve based on historical data. Train the noise reduction optimization model using reinforcement learning algorithms, updating noise reduction parameters in real time to adapt to different noise environments. Record the iterative noise reduction optimization strategies in a database and visualize their effects to help operators monitor and evaluate the optimization process. Generate a dynamic noise reduction optimization strategy change graph, showing noise reduction parameter adjustments and effect evaluations over different time periods, facilitating subsequent real-time adjustments and optimizations.

[0034] In this embodiment, an adaptive noise reduction system for a multimodal audio SoC main control chip is provided, used to execute the adaptive noise reduction method for a multimodal audio SoC main control chip as described above, including: The noise structure analysis module is used to collect the original audio input signal from the SoC main control chip, perform time-frequency domain dual deconstruction and cross-frequency domain noise interference structure analysis, and construct a full-frequency domain noise interference topology table. The noise environment modeling module is used to separate multimodal interference factors from the full-frequency domain noise interference topology table, and to model multiple noise environments to build a real-time noise scene model. The dynamic noise reduction module is used to model the temporal noise slope fluctuation of the real-time noise scene model, and to learn and optimize the dynamic noise reduction response to build a dynamic noise reduction optimization strategy. An effective audio optimization module is used to perform neural coding gain and audio boundary reconstruction based on the original audio input signal to obtain the key audio signal of coding gain and weak audio optimization signal; The effective audio reconstruction module is used to effectively reconstruct the key audio signal of the encoded gain and the weak audio optimization signal to build a noise-reduced and optimized audio input signal. The iterative noise reduction optimization module is used to perform real-time audio noise reduction offset evaluation on the noise reduction optimization audio input signal, and to perform iterative noise reduction migration optimization on the dynamic noise reduction optimization strategy, thus building an iterative noise reduction optimization engine.

[0035] This invention employs a dual time-frequency domain deconstruction approach, enabling the system to simultaneously perform detailed analysis of audio signals in both time and frequency dimensions, thereby accurately distinguishing useful signals from noise components. It is particularly effective at identifying noise sources with a wide frequency range. By analyzing the frequency domain structure and mutual interference relationships of noise, the system can construct a full-frequency domain noise interference topology table, detailing the characteristics of different noise sources, such as intensity, frequency range, and interference modes. This topology table provides a reliable data foundation for subsequent noise source separation and noise reduction strategy design. The system can handle complex noise environments, even when multiple noise sources coexist, maintaining accurate noise analysis and management to ensure that no important interference signals are missed during noise reduction. By effectively separating different types of noise in the audio signal (such as mechanical noise, human voice, and environmental noise), the system can formulate different noise reduction strategies for each noise source, thus processing each noise source more precisely. The construction of a real-time noise scene model allows the system to dynamically identify changes in noise sources and the environment. Changes in traffic noise intensity, fluctuations in wind noise, etc., are reflected in the noise scene model in real time, ensuring that noise reduction strategies are updated continuously. By modeling the temporal slope fluctuations of noise, the system can predict future trends in noise sources. Noise intensity may suddenly increase or decrease, and this change can be predicted in advance through temporal modeling. Based on the temporal fluctuations of noise, the system can optimize the response of the noise reduction algorithm in real time. When noise fluctuations intensify, the system can quickly switch to a stronger noise reduction mode; when noise weakens, the system can return to an energy-saving mode, thus ensuring efficient noise reduction while reducing power consumption. Through neural coding gain, the system can intelligently enhance audio signals, especially at low volumes or under noise masking, effectively improving audio clarity and intelligibility. By reconstructing the boundaries of audio signals, the system can more accurately preserve the temporal structure and frequency characteristics of audio signals. This is crucial for the fidelity of speech or music signals, avoiding problems such as blurring and loss of audio details during noise reduction. Through effective reconstruction, the system can seamlessly fuse the key audio signals of the coded gain with the optimized weak audio signals, thereby restoring the closest possible audio signal to reality. This ensures the naturalness and clarity of the audio, and the sound quality will not degrade due to the noise reduction process. During noise reduction, the system effectively preserves the original characteristics of the audio and avoids excessive "modification" of the audio signal. This ensures that the audio quality after noise reduction is as close as possible to the original quality, which is crucial, especially in applications such as speech recognition. Through iterative optimization, the system can learn from feedback on noise reduction effects in different environments, continuously improving the accuracy and adaptability of the noise reduction strategy. This allows the system to continuously optimize performance over long periods of operation, avoiding the limitations of fixed algorithms. Dynamically optimizing based on real-time changing noise patterns, the system provides optimal noise reduction results in both high-noise and quiet environments.

[0036] Therefore, the embodiments should be considered as exemplary and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of the equivalents of the application are intended to be included within the invention.

[0037] The above description is merely a specific embodiment of the present invention, enabling those skilled in the art to understand or implement it. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the present invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features of the invention herein.

Claims

1. An adaptive noise reduction method for a multimodal audio SoC main control chip, characterized in that, Includes the following steps: Step S1: Collect the original audio input signal from the SoC main control chip, perform time-frequency domain dual deconstruction and cross-frequency domain noise interference structure analysis, and construct a full-frequency domain noise interference topology table; Step S2: Separate multimodal interference factors from the full-frequency domain noise interference topology table, model multiple noise environments, and construct a real-time noise scene model; Step S3: Model the temporal noise slope fluctuation of the real-time noise scene model, and perform dynamic noise reduction response learning optimization to construct a dynamic noise reduction optimization strategy; Step S4: Perform neural coding gain and audio boundary reconstruction based on the original audio input signal to obtain the key audio signal of coding gain and weak audio optimization signal; Step S5: Perform effective audio reconstruction on the key audio signal of the encoded gain and the weak audio optimization signal to construct the noise-reduced and optimized audio input signal; Step S6: Perform real-time audio noise reduction offset evaluation on the noise reduction optimized audio input signal, and perform iterative noise reduction migration optimization on the dynamic noise reduction optimization strategy to build an iterative noise reduction optimization engine.

2. The adaptive noise reduction method for a multimodal audio SoC main control chip according to claim 1, characterized in that, The specific steps of step S1 are as follows: The original audio input signal is acquired based on the SoC main control chip; The original audio input signal is decomposed into multiple frequencies to obtain multiple audio frequency layered time windows; Multiple audio frequencies are layered into time windows and then deconstructed in both the time and frequency domains to obtain the spectral distribution tensor and the sound field interference vector. Dual-modal deep audio semantic mining is performed on the spectral distribution tensor and sound field interference vector to extract semantically valid audio vectors and audio noise feature vectors. Based on the audio noise feature vector, cross-frequency domain noise interference structure analysis is performed, and a full-frequency domain noise interference topology table is constructed.

3. The adaptive noise reduction method for a multimodal audio SoC main control chip according to claim 1, characterized in that, The specific steps of step S2 are as follows: Multimodal interference factors are separated from the full-frequency domain noise interference topology table to generate multiple noise interference nodes; Noise envelope feature analysis is performed on multiple noise interference nodes to obtain the envelope feature curve of each interference node; For each interference node, the envelope characteristic curve is analyzed in the main interference noise frequency band to generate the main interference noise frequency band characteristics. Scene noise pattern analysis is performed based on the characteristics of the main interference noise frequency band to generate dynamic noise patterns for the real-time scene. Multi-noise environment modeling is performed on the dynamic noise patterns of real-time scenes to construct a real-time noise scene model.

4. The adaptive noise reduction method for a multimodal audio SoC main control chip according to claim 3, characterized in that, The specific steps for modeling the dynamic noise patterns of real-time scenes and constructing a real-time noise scene model are as follows: A deep time-frequency domain variation analysis was performed on the spectral distribution tensor and the sound field interference vector to identify the dynamic change characteristics of the noise source and its movement path. The dynamic change characteristics are subjected to noise frequency trend evolution to generate a noise frequency change trend; Calculate the noise source locations at multiple time points based on the movement path; Based on the noise source locations at the multiple time points, spatial location situation prediction is performed to obtain noise source location prediction data; Based on the noise frequency variation trend and noise source location prediction data, the dynamic change law of the noise source is evolved, thereby generating a dynamic change logic diagram of the noise source. Based on the dynamic change logic diagram of noise sources, a multi-noise environment model is constructed for the dynamic noise pattern of real-time scenes to build a real-time noise scene model.

5. The adaptive noise reduction method for a multimodal audio SoC main control chip according to claim 1, characterized in that, Step S3 is as follows: Noise topology analysis is performed on a real-time noise scene model to extract unstructured noise vectors and structured repetitive noise. Calculate the noise intensity and noise periodic variation characteristics of the structured repetitive noise; Based on the noise intensity and noise periodic variation characteristics, a time-series noise slope fluctuation model is performed to construct a periodic disturbance trajectory diagram of structural noise; The long-term noise interference frequency range is calculated by analyzing the periodic disturbance trajectory diagram to obtain the long-term noise periodic frequency range. Adaptive noise reduction threshold analysis is performed on the long-term noise cycle frequency range to obtain adaptive noise reduction parameters for structural noise; High-frequency filtering and denoising are performed on unstructured noise vectors to obtain high-frequency filtering and denoising parameters for unstructured noise. Based on the adaptive noise reduction parameters and the high-frequency filtering noise reduction parameters, dynamic noise reduction response learning and optimization are performed to construct a dynamic noise reduction optimization strategy.

6. The adaptive noise reduction method for a multimodal audio SoC main control chip according to claim 1, characterized in that, The specific steps of step S4 are as follows: Dynamic amplitude spectrum recognition is performed on semantically valid audio vectors to extract dynamic amplitude spectrum characteristics; Based on dynamic amplitude spectrum characteristics, audio semantic characteristics are analyzed to extract key audio activity blocks and weak audio semantic blocks. Neural coding gain is applied to key audio activity blocks to obtain the coded gain key audio signal; Audio boundary lines are reconstructed from weak audio semantic blocks to obtain audio semantic boundary lines; Morphological restoration is performed based on audio semantic boundary lines to generate weak audio optimized signals.

7. The adaptive noise reduction method for a multimodal audio SoC main control chip according to claim 1, characterized in that, The specific steps of step S5 are as follows: Effective audio reconstruction is performed on the key audio signal with coded gain and the weak audio optimization signal to construct an effective gain reconstructed audio signal. Adaptive noise reduction processing is performed on the audio noise feature vector based on the dynamic noise reduction optimization strategy to extract adaptive noise-reduced audio. Cross-modal audio fusion is performed on the adaptive noise-reduced audio signal and the effective gain reconstructed audio signal to construct a noise-reduced optimized audio input signal.

8. The adaptive noise reduction method for a multimodal audio SoC main control chip according to claim 1, characterized in that, The specific steps of step S6 are as follows: Calculate the audio signal-to-noise ratio, speech intelligibility index, and semantic fidelity of the noise-reduced audio input signal; Audio distortion quantization analysis is performed based on semantic fidelity to generate audio distortion quantization values; Real-time audio noise reduction offset evaluation is performed on the audio signal-to-noise ratio, speech intelligibility index, and audio distortion quantization value to obtain the audio noise reduction offset evaluation value. Based on the audio noise reduction offset evaluation value, noise reduction feedback analysis is performed to generate real-time audio noise reduction feedback information; Based on real-time audio noise reduction feedback, the dynamic noise reduction optimization strategy is iteratively migrated and optimized to build an iterative noise reduction optimization engine.

9. An adaptive noise reduction system for a multimodal audio SoC main control chip, characterized in that, The method for performing the adaptive noise reduction method for a multimodal audio SoC main control chip as described in claim 1 includes: The noise structure analysis module is used to collect the original audio input signal from the SoC main control chip, perform time-frequency domain dual deconstruction and cross-frequency domain noise interference structure analysis, and construct a full-frequency domain noise interference topology table. The noise environment modeling module is used to separate multimodal interference factors from the full-frequency domain noise interference topology table, and to model multiple noise environments to build a real-time noise scene model. The dynamic noise reduction module is used to model the temporal noise slope fluctuation of the real-time noise scene model, and to learn and optimize the dynamic noise reduction response to build a dynamic noise reduction optimization strategy. An effective audio optimization module is used to perform neural coding gain and audio boundary reconstruction based on the original audio input signal to obtain the key audio signal of coding gain and weak audio optimization signal; The effective audio reconstruction module is used to effectively reconstruct the key audio signal of the encoded gain and the weak audio optimization signal to build a noise-reduced and optimized audio input signal. The iterative noise reduction optimization module is used to perform real-time audio noise reduction offset evaluation on the noise reduction optimization audio input signal, and to perform iterative noise reduction migration optimization on the dynamic noise reduction optimization strategy, thus building an iterative noise reduction optimization engine.

Citation Information

Cited By

  • Intelligent voice interaction system and method

    CN121583247A

  • Sensor data noise reduction processing method under electromagnetic interference environment of flight control system

    CN122309943A

  • Sensor data denoising processing method in electromagnetic interference environment of flight control system

    CN122309943B