Internet of Things equipment real-time noise reduction method and system based on edge computing

By introducing spatial enhancement processing and spectral structure feature extraction at the edge of IoT devices, fixed-directional background noise is identified and classified, solving the problem of background noise misprocessing in existing technologies and improving speech recognition accuracy and device stability.

CN121884845APending Publication Date: 2026-04-17SHENZHEN ZUNTE DIGITAL CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHENZHEN ZUNTE DIGITAL CO LTD
Filing Date
2026-02-14
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing edge computing-based real-time noise reduction technologies for IoT devices cannot effectively identify the spatial frequency characteristics differences between fixed-direction background noise and real speech signals, resulting in incorrect processing of background noise and affecting speech recognition accuracy and device stability.

Method used

By introducing spatial enhancement processing at the edge of IoT devices, a directional sound source mapping structure is generated, fixed-directional background noise is identified and its spectral structure features are extracted, a spatial frequency clustering index is constructed, and frequency clustering parameters are set to achieve classification processing of enhanced sound signals.

Benefits of technology

It effectively solves the problem that beamforming methods cannot identify background noise in a fixed direction, improves speech recognition accuracy and device stability, reduces data transmission pressure and computational complexity, and enhances usability in voice interaction environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121884845A_ABST
    Figure CN121884845A_ABST
Patent Text Reader

Abstract

The invention discloses an Internet of Things equipment real-time noise reduction method and system based on edge computing, and relates to the technical field of Internet of Things equipment noise reduction, and the method specifically comprises the following steps: recognizing a direction number with a fixed direction background noise condition based on a direction sound source mapping structure generated in real time; the identified direction number is calibrated as a background aggregation direction number; performing spectrum structure feature extraction on the enhanced sound signal corresponding to the background aggregation direction number, analyzing the spatial frequency aggregation degree of the enhanced sound signal, and setting a parameter value of the frequency aggregation degree according to an analysis result; and based on a set frequency aggregation degree, judging whether the enhanced sound signal corresponding to the background aggregation direction number should be reserved as voice content, and performing classification processing according to a judgment result. According to the invention, the problem that the background noise in the fixed direction in the Internet of Things equipment is easy to be mistakenly recognized as voice is solved, accurate discrimination and classification processing based on the frequency aggregation characteristics are realized, and the accuracy of voice recognition and the system stability are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of noise reduction technology for Internet of Things (IoT) devices, and more specifically to a real-time noise reduction method and system for IoT devices based on edge computing. Background Technology

[0002] Real-time noise reduction for IoT devices based on edge computing refers to deploying noise reduction algorithms on edge nodes (such as sensors, embedded devices, and edge gateways) close to the data source in an IoT system. This allows for rapid local processing of collected environmental data (such as sound, images, and vibration signals), instantly filtering out interference noise, extracting useful information, and completing data purification and optimization without uploading the original data to the cloud. The core value of this approach lies in leveraging the low latency, high efficiency, and local data processing advantages of edge computing to solve the problems of high latency, large bandwidth consumption, and privacy leaks caused by traditional noise reduction methods that rely on cloud computing. It is particularly suitable for scenarios with high real-time requirements, such as industrial equipment status monitoring, intelligent voice recognition, remote medical equipment, or security systems. Furthermore, with the surge in the number of IoT devices, centralized computing models struggle to handle the massive data traffic. Edge computing-based real-time noise reduction methods not only improve the overall system response speed and robustness but also reduce network load and energy consumption, providing crucial support for building efficient, intelligent, and reliable IoT systems.

[0003] Existing edge computing-based real-time noise reduction technologies for IoT devices typically deploy lightweight signal processing modules and intelligent algorithm models on the device side or edge nodes to perform real-time noise reduction on raw signals (such as audio, electromagnetic, image, vibration, etc.) collected by sensors. The implementation of this technology generally includes the following key steps: First, a signal preprocessing module is integrated at the IoT terminal or edge gateway to perform preliminary filtering, noise reduction, and formatting on the collected signals; second, optimized noise reduction algorithms (such as adaptive filtering, Fourier transform, convolutional neural networks, time-frequency feature extraction, etc.) are run on the edge computing nodes to model and suppress interference noise, extracting clear and effective target signals; then, the noise-reduced data is further encoded and compressed, and local decisions are made or uploaded to the cloud according to actual application needs (such as control response, data analysis, alarm notification, etc.); in some high-level systems, a dynamic parameter tuning function based on a feedback mechanism is also included, which can optimize the noise reduction strategy in real time according to changes in environmental noise to improve robustness and adaptability. The entire process fully leverages the advantages of edge computing in "nearby processing and distributed collaboration," achieving not only low-latency and high-efficiency data purification but also effectively reducing data transmission pressure and reliance on cloud computing, while enhancing the system's real-time performance, stability, and data privacy protection capabilities.

[0004] The existing technology has the following shortcomings: When spatial filtering is used in IoT edge terminals to improve the quality of forward speech signals, if there are low-frequency background noise sources with a fixed direction in the device's operating environment, such as motors or air conditioners, whose sound source direction is singular and whose spectrum distribution is relatively stable, the spatial characteristics of this background noise are often similar to the directional distribution of the target speech. This noise may be incorrectly identified as a foreground speech signal by the current beamforming logic and processed with excessive attention. Because this type of noise has a stable spatial structure and lacks drastic spectral jumps, it easily overlaps with the target speech in both direction and frequency. This leads the enhancement algorithm to directly and mistakenly preserve it as speech without first checking the speech spectrum structure. Existing edge computing-based real-time noise reduction technologies for IoT devices cannot determine whether a signal should be retained as speech content based on the spatial frequency concentration in the presence of background noise in a fixed direction. The main reason is that current beamforming logic relies too heavily on directional information and has not established a rejection mechanism between the speech spectrum structure and background noise. This results in background noise being repeatedly amplified and processed, which in turn continuously suppresses the real speech signal. This not only leads to a decrease in speech recognition accuracy but may also cause serious consequences such as lost voice commands, false trigger responses, and failures to wake up from the far field. Ultimately, this reduces the availability and stability of edge devices in actual IoT voice interaction scenarios.

[0005] The information disclosed in the background section is only intended to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention

[0006] The purpose of this invention is to provide a real-time noise reduction method and system for IoT devices based on edge computing, so as to solve the problems in the background art mentioned above.

[0007] To achieve the above objectives, the present invention provides the following technical solution: a real-time noise reduction method for IoT devices based on edge computing, specifically including the following steps: Spatial enhancement processing is performed on the input sound signal corresponding to the preset direction number in the IoT device environment to generate a directional sound source mapping structure containing each direction number and its corresponding enhanced sound signal in real time. Based on the real-time generated directional sound source mapping structure, the directional numbers of cases with fixed directional background noise are identified, and the identified directional numbers are labeled as background aggregation directional numbers. The spectral structure features of the enhanced acoustic signals corresponding to the background aggregation direction numbers are extracted, their spatial frequency aggregation degree is analyzed, and the parameter value of the frequency aggregation degree is set according to the analysis results. Based on the set frequency concentration, it is determined whether the enhanced sound signal corresponding to the background concentration direction number should be retained as speech content, and the classification process is performed according to the determination result. The enhanced audio signal, which is preserved as voice content, is subjected to noise reduction processing at the edge of the IoT device to improve the voice signal quality and complete closed-loop processing.

[0008] Preferably, based on the real-time generated directional sound source mapping structure, the directional numbers of cases with fixed directional background noise are identified, and the identified directional numbers are labeled as background aggregation directional numbers, specifically: The frequency distribution range, frequency interval energy distribution, and spectral structure change trend of the enhanced sound signal corresponding to each direction number in the real-time generated directional sound source mapping structure are obtained. It is determined whether each direction number has the characteristics of constant frequency distribution range, concentrated frequency interval energy, and spectral structure change amplitude below a preset threshold during continuous change. The direction number that meets all the judgment conditions is labeled as the background aggregation direction number.

[0009] Preferably, the enhanced acoustic signal corresponding to the background aggregation direction number is subjected to spectral structure feature extraction, its spatial frequency aggregation degree is analyzed, and the parameter value of frequency aggregation degree is set according to the analysis results. Specifically, the following steps are included: Extract the three structural feature data of the enhanced sound signal corresponding to the background aggregation direction number in a number of preset frequency intervals: mean amplitude, variance amplitude, and number of times the main frequency position changes. Generate the first spectrum factor, the second spectrum factor, and the third spectrum factor respectively. Based on the first, second, and third spectral factors of the mapping, a spatial frequency clustering index is generated. Determine the pre-set spatial frequency aggregation index threshold range and compare it with the generated spatial frequency aggregation index. Evaluate the degree of spatial frequency aggregation of the enhanced sound signal corresponding to the background aggregation direction number based on the comparison results. Based on the evaluation results, the parameter value for frequency clustering is set.

[0010] Preferably, the specific acquisition logics for the first spectral factor, the second spectral factor, and the third spectral factor are as follows: Extract the mean amplitude, variance amplitude, and number of fundamental frequency position changes of the enhanced sound signal corresponding to the background focusing direction number within several pre-defined frequency intervals, and label them as follows: , and , This indicates that the enhanced sound signal corresponding to the background focusing direction number is in the preset number of the [missing information]. The average amplitude within each frequency range This indicates that the enhanced sound signal corresponding to the background focusing direction number is in the preset number of the [missing information]. The amplitude variance within each frequency range This indicates that the enhanced sound signal corresponding to the background focusing direction number is in the preset number of the [missing information]. The number of times the dominant frequency position changes within a frequency range. , It is a positive integer; Constructing a local gain function ,in For the pre-set trend amplification factor, The result is used as the first spectral factor ; The model is trained using historical samples to determine a pre-set threshold for amplitude variance segmentation. The mapping is performed using a piecewise logarithmic function, which is: In the formula, and All of these are preset compression ratio adjustment factors. The enhanced sound signal corresponding to the background focusing direction number is in the pre-set number of the first... The amplitude variance within each frequency interval is mapped using a piecewise logarithmic function. The result is used as the second spectral factor ; The formula for calculating the third spectral factor is as follows: In the formula, The third spectral factor, For the pre-set number Non-zero weighting coefficients for each frequency interval This is the average number of times the dominant frequency position changes across all frequency ranges. This represents the standard deviation of the number of times the dominant frequency position changes across all frequency ranges.

[0011] Preferably, the first spectral factor based on the mapping Second spectral factor and the third spectral factor Spatial frequency clustering index is generated by weighted summation. The specific calculation formula is as follows: In the formula, The spatial frequency clustering index. , and These are the non-zero weight coefficients corresponding to the first, second, and third spectral factors, respectively. .

[0012] Preferably, a pre-defined spatial frequency clustering index threshold range is determined. And after being determined, it is combined with the generated spatial frequency clustering index. A comparison was performed, and the degree of spatial frequency clustering of the enhanced acoustic signal corresponding to the background clustering direction number was evaluated based on the comparison results. The specific comparison analysis is as follows: like The spatial frequency concentration of the enhanced acoustic signal corresponding to the background concentration direction number is low. like The spatial frequency concentration of the enhanced acoustic signal corresponding to the background focusing direction number is of medium degree; like The spatial frequency concentration of the enhanced acoustic signal corresponding to the background focusing direction number is high.

[0013] Preferably, based on the evaluation results, the parameter value for frequency clustering is set as follows: When the evaluation result indicates that the spatial frequency clustering degree is low, the parameter value of the frequency clustering degree is set to the first preset value, which is the pre-calibrated minimum clustering degree value. When the evaluation result indicates that the spatial frequency concentration is moderate, the parameter for frequency concentration is set to a second preset value, which is an intermediate value greater than the first preset value. When the evaluation result indicates that the spatial frequency concentration is high, the parameter for frequency concentration is set to a third preset value, which is the maximum value greater than the second preset value.

[0014] Preferably, based on a set frequency concentration level, it is determined whether the enhanced acoustic signal corresponding to the background concentration direction number should be retained as speech content, and the classification process is performed according to the determination result, specifically as follows: When the set frequency concentration parameter is the first preset value, the enhanced sound signal corresponding to the background concentration direction number is directly determined to be the speech content and retained for subsequent speech processing. When the set frequency concentration parameter is set to the second preset value, the short-time energy envelope variation range of the enhanced sound signal corresponding to the background concentration direction number and the root mean square error of the main frequency position between adjacent frames are extracted and compared with the preset speech structure reference threshold. If both meet the retention conditions, the enhanced sound signal is labeled as speech content; otherwise, it is classified as background noise signal. When the set frequency concentration parameter is the third preset value, the frequency distribution density fluctuation rate of the enhanced sound signal corresponding to the background concentration direction number is extracted in the target frequency range and compared with the preset rejection criteria. If the rejection condition is met, it is classified as a background noise signal; otherwise, it is temporarily stored as a signal to be judged and enters the subsequent buffer pool.

[0015] Preferably, the real-time noise reduction system for IoT devices based on edge computing includes a spatial enhancement construction module, a background aggregation recognition module, a spectrum aggregation analysis module, a speech discrimination and classification module, and an edge noise reduction closed-loop module; The spatial enhancement module performs spatial enhancement processing on the input sound signals corresponding to preset direction numbers in the IoT device environment, and generates a directional sound source mapping structure containing each direction number and its corresponding enhanced sound signal in real time. The background aggregation identification module identifies the direction number of the background noise with a fixed direction based on the real-time generated directional sound source mapping structure, and marks the identified direction number as the background aggregation direction number. The spectrum aggregation analysis module extracts the spectral structure features of the enhanced acoustic signal corresponding to the background aggregation direction number, analyzes its spatial frequency aggregation degree, and sets the parameter value of the frequency aggregation degree based on the analysis results. The speech discrimination and classification module, based on the set frequency clustering degree, determines whether the enhanced sound signal corresponding to the background clustering direction number should be retained as speech content, and performs classification processing according to the judgment result; The edge noise reduction closed-loop module performs noise reduction processing on the edge side of the IoT device for the enhanced audio signal that is retained as voice content, in order to improve the voice signal quality and complete the closed-loop processing.

[0016] The technical effects and advantages provided by the present invention in the above technical solution are as follows: 1. This invention introduces an "enhanced acoustic signal spatial frequency aggregation discrimination mechanism" at the edge of IoT devices, effectively solving the problem that existing beamforming methods cannot identify the spatial frequency feature differences between fixed-direction background noise and real speech signals. The solution extracts spectral structure features from the enhanced acoustic signals corresponding to each direction number, quantifies and constructs three-dimensional spectral factors (amplitude trend, energy stability, frequency jump characteristics), and fuses them into a spatial frequency aggregation index. This index is then compared with a preset threshold range for judgment, enabling the system to perform "spectral spatial dual-dimensional" modeling and separation identification of potential fixed noise directions. This avoids problems such as decreased recognition rate or false triggering of responses due to incorrect enhancement of non-speech content.

[0017] 2. This invention achieves a structured expression and parametric analysis of the spatial frequency clustering characteristics of enhanced acoustic signals by constructing multidimensional spectral factors, adjusting mapping weights, and classifying frequency clustering levels. In particular, the calculation of the three spectral factors introduces amplitude gradient functions, piecewise logarithmic functions, and standardized frequency perturbation functions, which not only improves the ability to discriminate steady-state low-frequency noise in complex backgrounds, but also significantly enhances the system's tolerance to dynamic speech structures and the flexibility of its preservation strategy, thereby improving the robustness and accuracy of the overall speech recognition system.

[0018] 3. This invention executes at the edge of IoT devices, relying on lightweight feature extraction and structural modeling logic to effectively reduce data transmission pressure and computational complexity, while supporting embedded deployment scenarios with high real-time performance and low resource consumption. Without introducing external deep learning models or training annotations, the solution can achieve intelligent discrimination and differentiated processing of enhanced signals, constructing a closed-loop architecture of "identification-classification-processing," thus improving the stability, availability, and user experience of edge terminals in actual IoT voice interaction environments. Attached Figure Description

[0019] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this invention. For those skilled in the art, other drawings can be obtained based on these drawings.

[0020] Figure 1 This is a flowchart illustrating the real-time noise reduction method and system for IoT devices based on edge computing, as described in this invention.

[0021] Figure 2 This is a schematic diagram of the modules of the real-time noise reduction method and system for IoT devices based on edge computing of the present invention. Detailed Implementation

[0022] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, they are provided so that the description of this disclosure will be more complete and fully convey the concept of the exemplary embodiments to those skilled in the art.

[0023] This invention provides, for example Figure 1 The real-time noise reduction method for IoT devices based on edge computing, as shown, specifically includes the following steps: Spatial enhancement processing is performed on the input sound signal corresponding to the preset direction number in the IoT device environment to generate a directional sound source mapping structure containing each direction number and its corresponding enhanced sound signal in real time. To achieve spatial enhancement of input acoustic signals corresponding to preset direction numbers in an IoT device environment, a directional filtering algorithm can be invoked in software at the edge to process the acquired raw acoustic signals by direction. Specifically, the system constructs a corresponding directional pointing template based on each preset direction number. For example, by setting a delay model or spatial reference vector corresponding to the sound source direction, and combining classic spatial filtering techniques such as delay-summation, minimum variance distortionless response (MVDR), or generalized sidelobe cancellation (GSC), the system enhances the direction of the raw acoustic signal. During processing, the system weights and superimposes data collected from multiple sensor channels to maximize the energy of the acoustic signal from the target direction while suppressing interference noise from other directions. Each direction number corresponds to one spatial enhancement process, thereby obtaining an enhanced acoustic signal matching that number, effectively extracting signal content with strong energy characteristics in that direction.

[0024] After completing the spatial enhancement processing corresponding to each directional number, the system constructs a directional sound source mapping structure in software. Specifically, this is achieved by establishing a mapping table from directional number to enhanced acoustic signal. The system first allocates a mapping unit for each directional number. This unit stores the enhanced acoustic signal data and adds relevant metadata, including the enhancement timestamp, the spatial filtering model used, and a signal feature summary (such as energy spectrum distribution and frequency domain envelope). This structure can be implemented using arrays, hash tables, or key-value mapping dictionaries, with the directional number as a unique key and the corresponding enhanced acoustic signal as the value, thus achieving structured organization of sound source spatial information in the software. This mapping structure not only records the processing results of signals from each direction but also provides an efficient and readily available data organization foundation for subsequent periodic comparisons, frequency clustering assessments, and noise discrimination analysis.

[0025] The core purpose of employing the above-described spatial enhancement processing and generating a directional sound source mapping structure is to construct a real-time, multi-directional, structured acoustic signal analysis mechanism suitable for edge computing environments. In complex IoT application environments, background noise and speech signals may highly overlap spatially, especially when the noise source direction is fixed and the spectrum is stable, making it easy for existing noise reduction algorithms to mistakenly identify it as target speech. By independently enhancing the input acoustic signal according to its direction number and storing it in a structured manner, the system can independently identify and track signals from each direction, thereby avoiding a one-size-fits-all approach to processing the entire signal. This structuring process not only improves the accuracy of noise detection and direction recognition but also significantly reduces the frequency of uploading the original signal to the cloud for processing, achieving efficient computing and low-latency response at the edge, laying the foundation for subsequent spectrum aggregation assessment and speech signal preservation judgment.

[0026] Based on the real-time generated directional sound source mapping structure, the directional numbers of cases with fixed directional background noise are identified, and the identified directional numbers are labeled as background aggregation directional numbers. In this embodiment, based on the real-time generated directional sound source mapping structure, the directional numbers of cases with fixed-directional background noise are identified, and the identified directional numbers are labeled as background aggregation directional numbers. Specifically: The frequency distribution range, frequency interval energy distribution, and spectral structure change trend of the enhanced sound signal corresponding to each direction number in the real-time generated directional sound source mapping structure are obtained. It is determined whether each direction number has the characteristics of constant frequency distribution range, concentrated frequency interval energy, and spectral structure change amplitude below a preset threshold during continuous change. The direction number that meets all the judgment conditions is labeled as the background aggregation direction number.

[0027] To obtain the frequency distribution range, frequency interval energy distribution, and spectral structure change trend of the enhanced acoustic signal corresponding to each direction number in the real-time generated directional sound source mapping structure, this can be achieved by performing frame segmentation, windowing, and Fourier transform processing on the spatially enhanced acoustic signal within the built-in digital signal processing flow of the edge computing node. Specifically, the system first divides the enhanced acoustic signal corresponding to each direction number into multiple short-time analysis frames in the time domain, applies a window function to each frame to suppress spectral leakage, and then performs a Fast Fourier Transform (FFT) on each frame signal to obtain the frequency domain amplitude spectrum. By statistically analyzing the amplitude change of each frequency component in consecutive frames, the frequency distribution range of that direction number can be constructed, i.e., identifying the frequency intervals where the main energy is concentrated in the signal frequency domain. Furthermore, by normalizing the energy integral value within each frequency interval, the frequency interval energy distribution information can be obtained, reflecting the concentration of energy in the spectrum. Simultaneously, by performing differential analysis on the rate of change of energy at each frequency point between several consecutive frames, the spectral structure change trend can be obtained, quantifying the stability or volatility of the enhanced acoustic signal in the frequency dimension. These processing steps can all be executed in real time in a streaming manner on edge devices via software, generating structured feature data that can be used for subsequent identification and judgment.

[0028] To determine whether each directional number exhibits characteristics of constant frequency distribution range, concentrated energy within frequency intervals, and spectral structure variation amplitude below a preset threshold during continuous change, software can be used on an edge computing platform to perform time-series modeling and feature statistical analysis of the spectral characteristics of each enhanced acoustic signal. Specifically, the system first extracts the frequency distribution range of each directional number within a set number of consecutive time frames. By statistically analyzing the positions of significant frequency components in each frame, if the frequency concentration interval remains within a fixed or minimal offset range within consecutive frames, it is considered to have a "constant frequency distribution range." Next, the energy distribution of the frequency interval in each frame is integrated and normalized. If the main energy peak is stably distributed in a few frequency intervals, and the energy proportion of that region is higher than a set concentration threshold, then the directional number is considered to have the characteristic of "concentrated energy within frequency intervals." Simultaneously, the system calculates the difference in the spectrum of each frame along the frequency axis, extracts the spectral variation amplitude, and quantifies the stability of these difference sequences using methods such as standard deviation and moving average. If the statistical measure of its rate of change is lower than a preset threshold, then its spectral structure variation amplitude is considered small. Of the three features, the system will only label the direction number as having fixed background noise if all the judgment conditions are met simultaneously, thus completing the subsequent aggregation direction identification. This method can be implemented in real time on edge devices using embedded signal processing software, exhibiting high responsiveness and scalability.

[0029] Directional numbers that satisfy all feature criteria are designated as background aggregation directional numbers. This can be achieved by constructing a directional number management logic based on feature-based judgment rules within the edge computing device. Specifically, after determining the frequency distribution range, frequency interval energy distribution, and spectral structure change trend of the enhanced acoustic signal corresponding to each directional number, the system records whether each directional number satisfies each feature condition in the current period in a state matrix using Boolean values. Subsequently, the software system performs joint judgment on this matrix, filtering out directional numbers where all features are "satisfied," and uses these as a candidate set of background aggregation directional numbers. To facilitate subsequent identification, classification, and tracking, the system assigns a clear logical identifier to these directional numbers that satisfy all feature conditions. For example, a state label field can be set, labeled as "background aggregation," and recorded in the directional number mapping structure. Alternatively, an independent data structure can be constructed within the system to store information about these identified directional numbers, such as a list of background aggregation directional numbers. The calibration process requires no manual intervention and is entirely automated by the software logic within the edge computing device, exhibiting excellent real-time performance and suitability for multi-directional parallel identification. In this way, not only can we effectively isolate noise directions with high stability and concentrated energy, but we can also provide a clear basis for subsequent classification and speech preservation decisions.

[0030] The spectral structure features of the enhanced acoustic signals corresponding to the background aggregation direction numbers are extracted, their spatial frequency aggregation degree is analyzed, and the parameter value of the frequency aggregation degree is set according to the analysis results. In this embodiment, the spectral structure features of the enhanced acoustic signal corresponding to the background aggregation direction number are extracted, its spatial frequency aggregation degree is analyzed, and the parameter value of the frequency aggregation degree is set according to the analysis results. Specifically, the following steps are included: Extract the three structural feature data of the enhanced sound signal corresponding to the background aggregation direction number in a number of preset frequency intervals: mean amplitude, variance amplitude, and number of times the main frequency position changes. Generate the first spectrum factor, the second spectrum factor, and the third spectrum factor respectively. In software implementation, a time-spectrum map can be generated by performing a short-time Fourier transform (STFT) on the enhanced acoustic signal. Then, structural feature extraction is performed on the spectrum data of each frame according to several pre-divided frequency intervals (e.g., 100Hz–300Hz, 300Hz–500Hz, etc.). First, for each frequency interval, the amplitude distribution of that interval over several consecutive frames is statistically analyzed, and its average amplitude value (i.e., amplitude mean) and amplitude variance are calculated, reflecting the intensity level and fluctuation degree of the signal energy within that frequency interval, respectively. Then, the dominant frequency position (i.e., the frequency point with the largest amplitude) within that frequency interval is identified in each frame. The changes in this dominant frequency position between adjacent frames are compared. If the dominant frequency position changes, the dominant frequency change count is accumulated, ultimately obtaining the number of times the dominant frequency position of that frequency interval changes throughout the entire frame sequence. To ensure accuracy, the system typically smooths the amplitude data to eliminate spurious changes caused by noise fluctuations, and windows the spectrum at the boundaries of the frequency intervals to avoid edge interference. These three characteristic data (mean amplitude, variance amplitude, and number of times the dominant frequency position changes) can all be automatically extracted locally on the edge device by the software using an efficient vector calculation method and updated in real time, serving as the input basis for subsequent spectrum factor calculation and spatial clustering analysis.

[0031] Based on the first, second, and third spectral factors of the mapping, a spatial frequency clustering index is generated. Determine the pre-set spatial frequency aggregation index threshold range and compare it with the generated spatial frequency aggregation index. Evaluate the degree of spatial frequency aggregation of the enhanced sound signal corresponding to the background aggregation direction number based on the comparison results. Determining a pre-defined spatial frequency clustering index threshold range can be achieved through statistical modeling based on historical sample data during model initialization or adaptive learning phases in software. Specifically, before deployment or in the initial stages of operation, the system can collect a large amount of enhanced acoustic signal sample data from typical scenarios. It then extracts spectral structure features from these data, identifying two types of signals: "speech content" and "background noise," calculating their corresponding first, second, and third spectral factors. Following predetermined calculation rules, it generates a series of spatial frequency clustering index sample values. Subsequently, the software models the clustering index distributions corresponding to the two types of signals using statistical methods such as probability density estimation or cluster analysis. It extracts boundary values, the interval with the maximum average difference, or the point with the maximum inter-class variance as the upper and lower limits of the discrimination threshold range, ultimately forming a stable "spatial frequency clustering index threshold range." In actual operation, this threshold range can be stored as an adjustable parameter structure, supporting periodic correction through an adaptive feedback mechanism from edge devices, further improving adaptability and distinguishability in different environments. This method can be implemented entirely in software through an edge computing platform, without relying on external servers, making it suitable for deployment in resource-constrained IoT terminals for online modeling and updates.

[0032] Based on the evaluation results, the parameter value for frequency clustering is set.

[0033] In this embodiment, the specific acquisition logic of the first spectrum factor, the second spectrum factor, and the third spectrum factor is as follows: Extract the mean amplitude, variance amplitude, and number of fundamental frequency position changes of the enhanced sound signal corresponding to the background focusing direction number within several pre-defined frequency intervals, and label them as follows: , and , This indicates that the enhanced sound signal corresponding to the background focusing direction number is in the preset number of the [missing information]. The average amplitude within each frequency range This indicates that the enhanced sound signal corresponding to the background focusing direction number is in the preset number of the [missing information]. The amplitude variance within each frequency range This indicates that the enhanced sound signal corresponding to the background focusing direction number is in the preset number of the [missing information]. The number of times the dominant frequency position changes within a frequency range. , It is a positive integer; Constructing a local gain function ,in For the pre-set trend amplification factor, The result is used as the first spectral factor ; This section constructs a local gain function and generates a first spectral factor to more accurately reflect the amplitude distribution trend of the enhanced acoustic signal in various frequency intervals, thereby characterizing potential energy concentration features and serving as a key basis for assessing the degree of spatial frequency clustering. Specifically, the amplitude mean of the enhanced acoustic signal corresponding to the background clustering direction number is first extracted in each preset frequency interval. These averages reflect the energy intensity of each frequency band. Then, the energy variation between adjacent frequency bands is calculated. And through trend amplification coefficient Constructing a local gain function The core design of this function lies in the fact that the greater the energy change between adjacent frequency bands, the higher the local gain value, thus giving that frequency band greater attention in the overall calculation and reflecting its contribution to the overall spectral structure change. Finally, the mean amplitude of each frequency band is multiplied by its corresponding gain value, i.e. The first spectral factor is obtained by summing over all frequency bands. This factor not only integrates the energy intensity itself, but also dynamically integrates the trend information between frequency bands, making the description of spatial frequency concentration more sensitive and directional. It is especially suitable for identifying background noise signals with the characteristics of "low-frequency broadband concentration and stable non-jumping", providing a more refined data foundation for subsequent differentiation from speech content.

[0034] The model is trained using historical samples to determine a pre-set threshold for amplitude variance segmentation. The mapping is performed using a piecewise logarithmic function, which is: In the formula, and All of these are preset compression ratio adjustment factors. The enhanced sound signal corresponding to the background focusing direction number is in the pre-set number of the first... The amplitude variance within each frequency interval is mapped using a piecewise logarithmic function. The result is used as the second spectral factor ; This section uses a piecewise logarithmic function to map the amplitude variance. The core purpose is to transform the amplitude fluctuation of the enhanced acoustic signal across different frequency ranges into a quantifiable and comparable evaluation metric to reflect the energy dispersion of the signal in that direction. Specifically, amplitude variance... The smaller the value, the more concentrated and cohesive the signal energy; conversely, the larger the value, the greater the amplitude fluctuation, indicating that the signal in that frequency band is unstable and has a discrete structure. To enhance this differentiated representation, the system first trains the model using historical samples, extracts the distribution boundaries of typical background noise and speech signals in terms of spectral variance, and then sets two levels of threshold intervals. The magnitude variance is divided into three intervals. Then, different logarithmic function mapping strategies are designed for each interval: when... This indicates that the energy in that frequency band is highly concentrated, and a smoother frequency response should be used. The function performs a weak linear amplification; when Located in the middle section Then, a compression ratio adjustment factor is introduced. ,pass Appropriate amplification should be applied to moderate fluctuations; while when This indicates that the frequency band changes drastically and lacks speech stability characteristics; therefore, a higher modulation factor is used. ,make It produces a stronger suppression effect on highly fluctuating signals. Finally, it applies to all frequency bands. Summing up the results yields the second spectral factor. This serves as an important metric for measuring the discreteness of signal energy in that direction. This segmentation, compression, and logarithmic mapping method retains sensitivity to the clustering features of real speech while effectively suppressing the interference of abnormal noise fluctuations on the overall judgment, thus enhancing the system's ability to distinguish between speech and background noise.

[0035] The formula for calculating the third spectral factor is as follows: In the formula, The third spectral factor, For the pre-set number Non-zero weighting coefficients for each frequency interval This is the average number of times the dominant frequency position changes across all frequency ranges. This represents the standard deviation of the number of times the dominant frequency position changes across all frequency ranges.

[0036] This section calculates the number of times the main frequency position changes. The third spectral factor is constructed using the normalized deviation in different frequency ranges. This study aims to quantify the instability of enhanced acoustic signals in terms of frequency structure. Specifically, the number of fundamental frequency position changes represents the frequency jumps of the dominant frequency in that frequency range per unit time. The larger the value, the more unstable the fundamental frequency in that range is, and the more likely it is to belong to a non-speech signal. First, the number of fundamental frequency changes is extracted for each frequency range. Its mean over all intervals and standard deviation Normalization is performed, i.e., the standard deviation offset is calculated. This value represents the relative degree of anomalousness of frequency jumps within a specific frequency range compared to the overall trend. To further adjust the influence of different frequency ranges in the overall calculation, a preset non-zero weighting coefficient is used. Weighted summation is performed to generate standardized contribution values ​​for each frequency band. Finally, the weighted summation of the contribution values ​​from all frequency bands yields the third spectral factor. The larger the value of this factor, the more drastic the frequency jumps of the enhanced acoustic signal in that direction, indicating an unstable time-frequency structure and a higher likelihood of it being background noise. This method enhances the model's ability to identify spectral jump anomalies through standard deviation normalization and achieves a fine characterization of complex signal structures by combining a frequency band weighting mechanism, thereby improving the accuracy and robustness of the system in distinguishing between speech and non-speech signals.

[0037] Trend amplification factor This is primarily used to amplify the influence weight of the amplitude mean difference between adjacent frequency intervals in calculating the first spectral factor. It can be achieved through empirical regression optimization using a training set constructed based on speech samples and background noise samples. In the software implementation, a batch of training samples containing real speech and pseudo-speech (such as fan or motor noise) is first selected. The amplitude mean variation difference of all samples in each frequency band is calculated, and binary classification training is performed with their real annotations (speech / noise). The range of difference thresholds with the strongest discriminative power is trained through logistic regression or decision tree analysis. Then, this difference range is linearly fitted with the target gain adjustment amplitude to obtain the optimal trend amplification coefficient. The range of values ​​to be determined is ultimately set. This represents the optimal robustness value within this interval. This process can be performed offline before device deployment, with parameters written to the model configuration file. Compression ratio adjustment factor. and This method is used to perform nonlinear piecewise mapping of amplitude variance values ​​across different intervals during the calculation of the second spectral factor, to avoid mapping distortion caused by drastic changes in noise amplitude within a certain frequency band. In implementation, based on a large sample of speech and non-speech data, the amplitude variance of frequency intervals is extracted. Clustering methods (such as K-means or DBSCAN) are used to divide its distribution into three regions: "low fluctuation," "medium fluctuation," and "high fluctuation." Statistical analysis is then performed on the distribution density of amplitude variance values ​​within each region. The intermediate dividing point is selected as... Then, by adjusting the compression magnitude of different piecewise functions through minimum mean square error fitting or KL divergence comparison, the final value is determined. and The value of is chosen such that its mapping curve in each segment has both non-linear compression capability and avoids feature loss due to excessively low values. Non-zero weighting coefficients for each frequency range This is used to measure the importance of a frequency band to the assessment of frequency structure stability. Its setting can employ a speech signal statistical weighting model or a task-oriented (e.g., speech recognition) inverse optimization model. In a speech recognition-oriented approach, firstly, speech effectiveness is scored across all frequency intervals (e.g., based on speech signal-to-noise ratio, speech clarity index), obtaining a ranking of the impact of each frequency band on recognition accuracy; then, the softmax function is used to normalize this impact value, generating a non-zero positive real weight for each frequency interval. , so that all Satisfying the constraint that the weighted normalization is 1 (i.e.) This weight can be dynamically updated periodically based on the speech recognition results of the device, or it can be fixedly written into the edge model configuration table after initial training for evaluating the instability of the main frequency in noise reduction tasks.

[0038] By constructing a first, second, and third spectral factor, the amplitude variation trend, energy distribution stability, and dominant frequency fluctuation of the enhanced acoustic signal in various frequency intervals are comprehensively quantified and evaluated to assess the degree of spatial frequency clustering. The first spectral factor is constructed using a local gain function, which amplifies the amplitude differences between adjacent frequency intervals and multiplies them by the average amplitude of the current interval to form a weighted amplitude, used to capture whether significant energy concentrations occur in certain intervals of the spectrum. A large first spectral factor value indicates significant local energy concentration, reflecting certain spectral clustering characteristics. The second spectral factor performs piecewise logarithmic mapping on the amplitude variance, judging the fluctuation within the frequency band based on the magnitude of the amplitude variance and a set threshold interval. A smaller value indicates more concentrated energy and less fluctuation, helping to identify stable spectral regions. The third spectral factor calculates the standardized weighted deviation of the number of dominant frequency position changes and combines it with frequency band weighting coefficients to characterize whether the dominant frequency exhibits stable concentration in multiple intervals. If the overall deviation is small, it indicates a concentrated frequency center distribution, i.e., strong spatial frequency clustering. Therefore, when the first spectral factor is high, the second spectral factor is low, and the third spectral factor is also low, it can be inferred that the enhanced acoustic signal has a significant frequency clustering structure and is very likely to belong to background noise in a stable direction, thus providing a strong mathematical basis for subsequent judgment.

[0039] In this embodiment, the first spectral factor based on the mapping Second spectral factor and the third spectral factor Spatial frequency clustering index is generated by weighted summation. The specific calculation formula is as follows: In the formula, The spatial frequency clustering index. , and These are the non-zero weight coefficients corresponding to the first, second, and third spectral factors, respectively. .

[0040] To achieve spatial frequency concentration index The generation of the spectrum factor (PSF), based on the acquired first spectral factor (PSF), second spectral factor (SSF), and third spectral factor (TSF), is performed by weighted summation. Each factor represents a different dimension of spectral structure features. Specifically, the system combines the PSF with its corresponding non-zero weight coefficients. Direct multiplication is used, but for SSF and TSF, since smaller values ​​usually represent a more concentrated spectrum, their reciprocals are added to a constant 1 (i.e., ...). , We construct corresponding weighting terms so that smaller values ​​contribute more to the final SFC. Three weighting coefficients. , and All three coefficients are preset non-zero real values ​​used to adjust the relative importance of the three types of factors in different noise reduction tasks. Their values ​​can be optimized through training based on the statistical performance of historical speech samples in actual noise scenarios, or dynamically assigned by the strategy model to ensure that the weight distribution of the three structural features in the overall evaluation is reasonable and physically interpretable. Simultaneously, to ensure that the SFC output after weighted summation can be compared with the standardized evaluation interval in subsequent calculations, these three coefficients must satisfy… This results in a normalized linear combination expression. This processing method can be implemented through a pre-deployed speech processing model in an edge computing device, dynamically generating evaluation results while performing real-time spectral feature extraction. It does not rely on high-performance cloud resources and has good practicality and scalability.

[0041] In this embodiment, a pre-set spatial frequency clustering index threshold range is determined. And after being determined, it is combined with the generated spatial frequency clustering index. A comparison was performed, and the degree of spatial frequency clustering of the enhanced acoustic signal corresponding to the background clustering direction number was evaluated based on the comparison results. The specific comparison analysis is as follows: like The spatial frequency concentration of the enhanced acoustic signal corresponding to the background concentration direction number is low. This indicates that the spectral distribution of the enhanced sound signal is highly discrete across multiple frequency ranges, lacking significant amplitude continuity, frequency concentration, or consistent dominant frequency variation. Such characteristics typically appear in natural speech segments or unstable noise in the environment, exhibiting complex dynamics, large frequency fluctuations, and a lack of sustainable enhancement properties. Therefore, this enhanced sound signal may contain target speech information. If it is mistakenly identified as background noise and removed, it will lead to a decrease in speech recognition accuracy or loss of semantic information. In such cases, the signal should generally be retained and incorporated into subsequent speech processing.

[0042] like The spatial frequency concentration of the enhanced acoustic signal corresponding to the background focusing direction number is of medium degree; This indicates that the enhanced acoustic signal exhibits a certain degree of concentration and stability in its spectral structure, but has not yet reached a significant aggregation characteristic. This state may originate from a speech signal partially interfered with by noise, or from certain weak noise sources that possess speech characteristics but are stably present. Such signals have a certain degree of uncertainty, and in practical IoT voice interaction applications, processing them directly without discrimination may lead to misidentification or incorrect response by the system in noise-sensitive conditions. Therefore, secondary verification is usually required, taking into account context or other redundant features, to determine whether to preserve or reduce noise.

[0043] like The spatial frequency concentration of the enhanced acoustic signal corresponding to the background focusing direction number is high.

[0044] This indicates that the signal exhibits a highly consistent average amplitude, frequency distribution, and dominant frequency trend within the preset frequency range, demonstrating typical high-frequency clustering. This type of spectral pattern is commonly seen in persistent, fixed-directional background noise (such as motor noise, air conditioner noise, etc.), typically possessing stable energy output and directional consistency, fundamentally different from the transient, abrupt, and highly fluctuating structure of speech signals. Therefore, it can be determined that the signal is highly likely pseudo-speech or non-speech background noise. Further enhancement would mislead subsequent recognition models; thus, it should be prioritized as background noise masking or attenuation to improve the robustness and accuracy of the system's speech processing.

[0045] In this embodiment, based on the evaluation results, the parameter value of frequency clustering is set as follows: When the evaluation result indicates that the spatial frequency clustering degree is low, the parameter value of the frequency clustering degree is set to the first preset value. The first preset value is the minimum clustering degree value that is pre-calibrated, which is used to indicate the situation where the frequency characteristics are dispersed and do not have a stable clustering trend. When the evaluation result indicates that the spatial frequency clustering degree is moderate, the parameter value of the frequency clustering degree is set to the second preset value. The second preset value is the middle value that is greater than the first preset value, which is used to indicate that the frequency characteristics have a certain clustering trend but have not yet formed an obvious concentration state. When the evaluation result indicates that the spatial frequency concentration is high, the parameter for frequency concentration is set to a third preset value. The third preset value is the maximum value greater than the second preset value, which is used to indicate that the frequency features are concentrated and stably aggregated in a specific frequency range.

[0046] In practical implementation, the frequency clustering parameter can be set by constructing a parameter mapping table or setting a set of conditional judgment logic chains. The system first determines the current clustering level (low, medium, high) of the enhanced acoustic signal based on the comparison results of the spatial frequency clustering index. Then, in the built-in parameter matching module, it automatically retrieves the preset parameter value corresponding to that level. For example, if the clustering level is "low," the first preset value is matched, which can be determined by the statistical characteristics of the most dispersed spectral structure in historical noise samples; if it is "medium," an intermediate transition value between the first and third values ​​is set to accommodate the transitional spectral distribution; if it is "high," the maximum parameter value is matched, indicating that the current frequency clustering level has reached a highly concentrated state. The setting process for each parameter value can determine the initial value by classifying and analyzing the training data using a clustering algorithm, and then fine-tuning it through model feedback during actual deployment, thereby achieving dynamic configuration and long-term stability. The entire setting process can be embedded in the system's software inference logic and automatically invoked through conditional statements, ensuring real-time response to signal state changes without manual intervention.

[0047] The reason for introducing this hierarchical parameter setting mechanism for frequency clustering is that signals with different clustering levels have significantly different impacts on subsequent speech judgment logic. Without discriminative parameter settings, low-cluster background noise can easily be misjudged as speech content, or highly clustered interference signals can escape processing thresholds. By adopting a hierarchical parameter setting mechanism, the system's discrimination sensitivity and robustness in different complex sound fields can be effectively enhanced: when the spectrum is relatively dispersed, a lower frequency clustering setting value avoids misidentification; when the clustering trend is strong, a higher parameter value activates the speech structure screening mechanism. This not only improves the intelligent adaptability of the entire noise reduction solution but also ensures its universality and portability in different device deployment environments, truly achieving accurate separation and processing of speech and noise based on software logic.

[0048] Based on the set frequency concentration, it is determined whether the enhanced sound signal corresponding to the background concentration direction number should be retained as speech content, and the classification process is performed according to the determination result. In this embodiment, based on a set frequency concentration level, it is determined whether the enhanced sound signal corresponding to the background concentration direction number should be retained as speech content, and the classification process is performed according to the determination result, specifically: When the set frequency concentration parameter is the first preset value, the enhanced sound signal corresponding to the background concentration direction number is directly determined to be the speech content and retained for subsequent speech processing. When the set frequency clustering parameter is set to the first preset value, the enhanced sound signal corresponding to the background clustering direction number is judged as a signal with low spatial frequency clustering. Such signals typically exhibit dispersed spectral features, unconcentrated frequency distribution, frequent changes in the dominant frequency, and no periodic or directional clustering trend. At the software implementation level, the spectral structure feature data of the enhanced sound signal (such as mean amplitude, amplitude variance, number of dominant frequency position changes, etc.) can be aggregated and analyzed. The calculated spatial frequency clustering index (SFC) is then compared with a preset SFC threshold. Once the SFC is less than the threshold, automatic judgment logic is triggered. Under this logic, the system no longer performs further structural complexity analysis or speech rejection verification on the signal. Instead, it directly marks the signal as "reserved" and stores it in the device's local buffer queue or directly transmits it to the subsequent speech recognition module. The reason for this is that low-clustering frequency features are more common in real speech, especially in natural language dialogue or far-field speech scenarios, where speech fragments composed of multiple frequencies frequently appear. Over-screening could lead to misjudging real speech as noise. Therefore, by simplifying the processing path and quickly retaining such signals, the system's adaptability to diverse speech structures and recognition completeness can be effectively improved, thereby enhancing the practicality and stability of edge devices in voice interaction tasks.

[0049] When the set frequency concentration parameter is set to the second preset value, the short-time energy envelope variation range of the enhanced sound signal corresponding to the background concentration direction number and the root mean square error of the main frequency position between adjacent frames are extracted and compared with the preset speech structure reference threshold. If both meet the retention conditions, the enhanced sound signal is labeled as speech content; otherwise, it is classified as background noise signal. When the set frequency clustering parameter is set to the second preset value, it indicates that the enhanced sound signal corresponding to the current background clustering direction number has a moderate degree of spatial frequency clustering characteristics. That is, its spectral characteristics show a certain degree of concentration trend, but have not yet reached the threshold for judging high-confidence speech or obvious background noise. To further determine whether the signal should be retained as speech content, the short-time energy envelope variation range and the root mean square error of the dominant frequency position between adjacent frames can be extracted for analysis. In terms of software implementation, the enhanced sound signal is first processed by short-time frame segmentation. The energy envelope value of each frame is calculated, and the variation range between the maximum and minimum envelope values ​​within the entire analysis period is statistically analyzed to reflect the degree of speech amplitude fluctuation. Then, the frequency of the dominant frequency in each frame is extracted, and the difference between the dominant frequencies of adjacent frames is calculated. The root mean square error of the difference sequence is then calculated to characterize the stability of frequency changes over time. The system compares the two structural features mentioned above with a pre-established speech structure reference threshold model. If the energy variation range and dominant frequency fluctuation of the signal are both within the statistical range of typical speech, it indicates that the signal has certain speech regularity, and the system determines it as speech content and retains it. Conversely, if both indicators deviate from the normal speech structure range, it is inferred to be a pseudo-speech signal or incorrectly amplified background noise, classified as noise data, and removed in subsequent processing. This joint discrimination mechanism can effectively alleviate the ambiguity problem of signals in the middle region being difficult to classify directly, and enhance the system's ability to recognize and adapt to complex sound structures in the edge sound field.

[0050] When the set frequency concentration parameter is the third preset value, the frequency distribution density fluctuation rate of the enhanced sound signal corresponding to the background concentration direction number is extracted in the target frequency range and compared with the preset rejection criteria. If the rejection condition is met, it is classified as a background noise signal; otherwise, it is temporarily stored as a signal to be judged and enters the subsequent buffer pool.

[0051] When the set frequency concentration parameter is the third preset value, it indicates that the enhanced sound signal corresponding to the current background concentration direction number has a highly concentrated spectral structure. This type of signal usually has stable directivity in space and high frequency concentration, and is very likely a long-term stable background noise source (such as low-frequency motor noise, constant-speed air conditioner noise, etc.). To further confirm whether this type of signal is indeed background noise, the system can extract the frequency distribution density fluctuation rate of the signal in the target frequency range for comparative analysis. The specific implementation method is as follows: First, the enhanced sound signal is divided into several sub-frequency bands in the preset target frequency range. The energy distribution density of each sub-frequency band in consecutive frames is statistically analyzed to form a frequency density time series. Then, the standard deviation or rate of change analysis is performed on the series to quantify the degree of fluctuation of its energy density over time, i.e., the frequency distribution density fluctuation rate. If the volatility is below the preset rejection threshold (i.e., the energy density is essentially constant over time), the enhanced signal's spectral structure is considered too stable and lacks the time-varying characteristics of speech, thus it can be identified as background noise and directly rejected. If the volatility is above the threshold but has not yet reached the typical characteristic range of speech, the signal is temporarily stored in a buffer pool for further comprehensive judgment after introducing more structural features. The significance of this strategy is that, while ensuring the accuracy of rejecting high-frequency clustered signals, it retains signals that may have special speech structures but cannot be accurately identified at present, reducing the accidental deletion of speech segments due to structural judgment errors, and improving the overall noise reduction robustness and fidelity of the system.

[0052] The enhanced audio signal, which is preserved as voice content, is subjected to noise reduction processing at the edge of the IoT device to improve the voice signal quality and complete closed-loop processing.

[0053] Performing noise reduction processing on the enhanced audio signal, which is preserved as speech content, at the edge of IoT devices is primarily aimed at further improving the clarity and recognizability of the signal after spatial enhancement and speech discrimination filtering. This ensures that the final speech data used for speech recognition or command parsing possesses high quality and robustness. Specific implementation methods include: firstly, selecting an appropriate edge noise reduction algorithm based on the structural characteristics of the preserved signal (such as frequency distribution, energy spectral density, envelope changes, etc.), such as spectral subtraction, adaptive Wiener filtering, or neural network speech enhancement models. A lightweight model is then invoked in the local embedded processing unit to perform spectral domain processing on the signal, suppressing remaining background noise components. In the software implementation, a frame-level dynamic threshold adjustment strategy can be adopted to optimize noise reduction parameters in real time based on the signal-to-noise ratio of the input signal, while ensuring that speech details are not excessively attenuated. After cleaning in the frequency domain, the signal is reconstructed back to the time domain and subjected to short-term smoothing processing to improve speech coherence and perceptible quality.

[0054] The reason for setting the noise reduction process at the edge is primarily based on the real-time response requirements of IoT devices and the advantages of edge computing deployment: On the one hand, the edge has preliminary computing resources, enabling local signal noise reduction processing without uploading to the cloud, significantly reducing latency; on the other hand, edge processing can also dynamically update the noise reduction model weights according to the local device environment, making the noise reduction strategy more adaptable to the specific deployment scenario. Furthermore, by performing the noise reduction step after enhancement discrimination, misprocessing of non-speech data can be avoided, improving resource utilization efficiency and ensuring the accuracy and stability of the overall voice interaction of the system, thus constructing a closed-loop local voice preprocessing workflow.

[0055] like Figure 2 The real-time noise reduction system for IoT devices based on edge computing shown includes a spatial enhancement construction module, a background aggregation recognition module, a spectrum aggregation analysis module, a speech discrimination and classification module, and an edge noise reduction closed-loop module. The spatial enhancement module performs spatial enhancement processing on the input sound signals corresponding to preset direction numbers in the IoT device environment, and generates a directional sound source mapping structure containing each direction number and its corresponding enhanced sound signal in real time. The background aggregation identification module identifies the direction number of the background noise with a fixed direction based on the real-time generated directional sound source mapping structure, and marks the identified direction number as the background aggregation direction number. The spectrum aggregation analysis module extracts the spectral structure features of the enhanced acoustic signal corresponding to the background aggregation direction number, analyzes its spatial frequency aggregation degree, and sets the parameter value of the frequency aggregation degree based on the analysis results. The speech discrimination and classification module, based on the set frequency clustering degree, determines whether the enhanced sound signal corresponding to the background clustering direction number should be retained as speech content, and performs classification processing according to the judgment result; The edge noise reduction closed-loop module performs noise reduction processing on the edge side of the IoT device for the enhanced audio signal that is retained as voice content, in order to improve the voice signal quality and complete the closed-loop processing.

[0056] The above formulas are all dimensionless calculations. The formulas are derived from software simulations based on a large amount of collected data to obtain the most recent real-world results. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.

[0057] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more sets of available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.

[0058] It should be understood that in the various embodiments of this application, the order of the above-mentioned processes does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0059] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0060] In the several embodiments provided in this application, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection through some interfaces, devices, or units, and may be electrical, mechanical, or other forms.

[0061] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0062] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0063] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A real-time noise reduction method for IoT devices based on edge computing, characterized in that, Specifically, the following steps are included: Spatial enhancement processing is performed on the input sound signal corresponding to the preset direction number in the IoT device environment to generate a directional sound source mapping structure containing each direction number and its corresponding enhanced sound signal in real time. Based on the real-time generated directional sound source mapping structure, the directional numbers of cases with fixed directional background noise are identified, and the identified directional numbers are labeled as background aggregation directional numbers. The spectral structure features of the enhanced acoustic signals corresponding to the background aggregation direction numbers are extracted, their spatial frequency aggregation degree is analyzed, and the parameter value of the frequency aggregation degree is set according to the analysis results. Based on the set frequency concentration, it is determined whether the enhanced sound signal corresponding to the background concentration direction number should be retained as speech content, and the classification process is performed according to the determination result. The enhanced audio signal, which is preserved as voice content, is subjected to noise reduction processing at the edge of the IoT device to improve the voice signal quality and complete closed-loop processing.

2. The real-time noise reduction method for IoT devices based on edge computing according to claim 1, characterized in that, Based on the real-time generated directional sound source mapping structure, the directional numbers of cases with fixed-directional background noise are identified, and the identified directional numbers are labeled as background aggregation directional numbers. Specifically: The frequency distribution range, frequency interval energy distribution, and spectral structure change trend of the enhanced sound signal corresponding to each direction number in the real-time generated directional sound source mapping structure are obtained. It is determined whether each direction number has the characteristics of constant frequency distribution range, concentrated frequency interval energy, and spectral structure change amplitude below a preset threshold during continuous change. The direction number that meets all the judgment conditions is labeled as the background aggregation direction number.

3. The real-time noise reduction method for IoT devices based on edge computing according to claim 2, characterized in that, The spectral structure features of the enhanced acoustic signals corresponding to the background aggregation direction numbers are extracted, their spatial frequency aggregation degree is analyzed, and the parameter value of the frequency aggregation degree is set according to the analysis results. The specific steps include: Extract the three structural feature data of the enhanced sound signal corresponding to the background aggregation direction number in a number of preset frequency intervals: mean amplitude, variance amplitude, and number of times the main frequency position changes. Generate the first spectrum factor, the second spectrum factor, and the third spectrum factor respectively. Based on the first, second, and third spectral factors of the mapping, a spatial frequency clustering index is generated. Determine the pre-set spatial frequency aggregation index threshold range and compare it with the generated spatial frequency aggregation index. Evaluate the degree of spatial frequency aggregation of the enhanced sound signal corresponding to the background aggregation direction number based on the comparison results. Based on the evaluation results, the parameter value for frequency clustering is set.

4. The real-time noise reduction method for IoT devices based on edge computing according to claim 3, characterized in that, The specific acquisition logics for the first spectral factor, the second spectral factor, and the third spectral factor are as follows: Extract the mean amplitude, variance amplitude, and number of fundamental frequency position changes of the enhanced sound signal corresponding to the background focusing direction number within several pre-defined frequency intervals, and label them as follows: , and , This indicates that the enhanced sound signal corresponding to the background focusing direction number is in the preset number of the [missing information]. The average amplitude within each frequency range This indicates that the enhanced sound signal corresponding to the background focusing direction number is in the preset number of the [missing information]. The amplitude variance within each frequency range This indicates that the enhanced sound signal corresponding to the background focusing direction number is in the preset number of the [missing information]. The number of times the dominant frequency position changes within a frequency range. , It is a positive integer; Constructing a local gain function ,in For the pre-set trend amplification factor, The result is used as the first spectral factor ; The model is trained using historical samples to determine a pre-set threshold for amplitude variance segmentation. The mapping is performed using a piecewise logarithmic function, which is: In the formula, and All of these are preset compression ratio adjustment factors. The enhanced sound signal corresponding to the background focusing direction number is in the pre-set number of the first... The amplitude variance within each frequency interval is mapped using a piecewise logarithmic function. The result is used as the second spectral factor ; The formula for calculating the third spectral factor is as follows: In the formula, The third spectral factor, For the pre-set number Non-zero weighting coefficients for each frequency interval This is the average number of times the dominant frequency position changes across all frequency ranges. This represents the standard deviation of the number of times the dominant frequency position changes across all frequency ranges.

5. The real-time noise reduction method for IoT devices based on edge computing according to claim 4, characterized in that, First spectral factor based on mapping Second spectral factor and the third spectral factor Spatial frequency clustering index is generated by weighted summation. The specific calculation formula is as follows: In the formula, The spatial frequency clustering index. , and These are the non-zero weight coefficients corresponding to the first, second, and third spectral factors, respectively. .

6. The real-time noise reduction method for IoT devices based on edge computing according to claim 5, characterized in that, Determine the pre-set spatial frequency aggregation index threshold range And after being determined, it is combined with the generated spatial frequency clustering index. A comparison was performed, and the degree of spatial frequency clustering of the enhanced acoustic signal corresponding to the background clustering direction number was evaluated based on the comparison results. The specific comparison analysis is as follows: like The spatial frequency concentration of the enhanced acoustic signal corresponding to the background concentration direction number is low. like The spatial frequency concentration of the enhanced acoustic signal corresponding to the background focusing direction number is of medium degree; like The spatial frequency concentration of the enhanced acoustic signal corresponding to the background focusing direction number is high.

7. The real-time noise reduction method for IoT devices based on edge computing according to claim 6, characterized in that, Based on the evaluation results, the parameter values ​​for frequency clustering are set as follows: When the evaluation result indicates that the spatial frequency clustering degree is low, the parameter value of the frequency clustering degree is set to the first preset value, which is the pre-calibrated minimum clustering degree value. When the evaluation result indicates that the spatial frequency concentration is moderate, the parameter for frequency concentration is set to a second preset value, which is an intermediate value greater than the first preset value. When the evaluation result indicates that the spatial frequency concentration is high, the parameter for frequency concentration is set to a third preset value, which is the maximum value greater than the second preset value.

8. The real-time noise reduction method for IoT devices based on edge computing according to claim 7, characterized in that, Based on the set frequency concentration, it is determined whether the enhanced sound signal corresponding to the background concentration direction number should be retained as speech content, and the classification process is performed according to the determination result, specifically: When the set frequency concentration parameter is the first preset value, the enhanced sound signal corresponding to the background concentration direction number is directly determined to be the speech content and retained for subsequent speech processing. When the set frequency concentration parameter is set to the second preset value, the short-time energy envelope variation range of the enhanced sound signal corresponding to the background concentration direction number and the root mean square error of the main frequency position between adjacent frames are extracted and compared with the preset speech structure reference threshold. If both meet the retention conditions, the enhanced sound signal is labeled as speech content; otherwise, it is classified as background noise signal. When the set frequency concentration parameter is the third preset value, the frequency distribution density fluctuation rate of the enhanced sound signal corresponding to the background concentration direction number is extracted in the target frequency range and compared with the preset rejection criteria. If the rejection condition is met, it is classified as a background noise signal; otherwise, it is temporarily stored as a signal to be judged and enters the subsequent buffer pool.

9. A real-time noise reduction system for IoT devices based on edge computing, used to implement the real-time noise reduction method for IoT devices based on edge computing as described in any one of claims 1-8, characterized in that, It includes a spatial enhancement module, a background aggregation recognition module, a spectrum aggregation analysis module, a speech discrimination and classification module, and an edge noise reduction closed-loop module; The spatial enhancement module performs spatial enhancement processing on the input sound signals corresponding to preset direction numbers in the IoT device environment, and generates a directional sound source mapping structure containing each direction number and its corresponding enhanced sound signal in real time. The background aggregation identification module identifies the direction number of the background noise with a fixed direction based on the real-time generated directional sound source mapping structure, and marks the identified direction number as the background aggregation direction number. The spectrum aggregation analysis module extracts the spectral structure features of the enhanced acoustic signal corresponding to the background aggregation direction number, analyzes its spatial frequency aggregation degree, and sets the parameter value of the frequency aggregation degree based on the analysis results. The speech discrimination and classification module, based on the set frequency clustering degree, determines whether the enhanced sound signal corresponding to the background clustering direction number should be retained as speech content, and performs classification processing according to the judgment result; The edge noise reduction closed-loop module performs noise reduction processing on the edge side of the IoT device for the enhanced audio signal that is retained as voice content, in order to improve the voice signal quality and complete the closed-loop processing.