AI intelligent sound equipment signal processing method and system, sound equipment device and storage medium
By combining adaptive filtering algorithms and edge computing with cloud analytics, the smart speaker system monitors environmental characteristics in real time and dynamically adjusts filtering parameters, solving the noise suppression problem of smart speakers in complex environments and improving the accuracy of speech recognition and user interaction experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-22
- Publication Date
- 2026-04-03
AI Technical Summary
Existing noise suppression methods for smart speakers are ill-suited to complex and ever-changing usage scenarios, affecting the accuracy of speech recognition.
An adaptive filtering algorithm is used to monitor environmental sound characteristics in real time and dynamically adjust filtering parameters. Combined with edge computing and cloud analysis, noise suppression and voice signal enhancement are achieved.
It achieves accurate identification and suppression of noise in different scenarios, improves the voice interaction performance of smart speakers in complex environments, and ensures voice signal quality and recognition accuracy.
Smart Images

Figure CN121789667A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of signal processing technology, and in particular to an AI smart speaker signal processing method, system, speaker device, and storage medium. Background Technology
[0002] With the increasing popularity of smart homes, smart speakers, as an important human-computer interaction interface, are playing an increasingly important role in daily life. Smart speakers need to accurately recognize user voice commands in complex environmental noise, which places higher demands on voice signal processing, especially noise suppression in different scenarios.
[0003] Existing smart speakers typically use filters with fixed parameters for noise suppression, or switch between multiple preset filter parameters for different noise environments. This noise suppression method based on preset parameters is quite common in practical applications.
[0004] However, due to the variability and unpredictability of environmental noise, fixed-parameter filtering schemes are difficult to adapt to complex and ever-changing usage scenarios, which can easily lead to the loss of effective speech signals or poor noise suppression, affecting the accuracy of speech recognition. This situation needs further improvement. Summary of the Invention
[0005] To address the problem that existing audio noise suppression methods are ill-suited to complex and ever-changing usage scenarios, thus affecting the accuracy of speech recognition, this application provides an AI-powered smart speaker signal processing method, system, speaker device, and storage medium, employing the following technical solution: In a first aspect, this application provides an AI smart speaker signal processing method, comprising the following steps: The audio acquisition module receives audio signals from the environment in real time to obtain raw audio data. Based on the original audio data, an adaptive filtering algorithm is used to suppress background noise and obtain an enhanced speech signal; Based on the enhanced speech signal, speech features are extracted and matched with speech models in the database to obtain speech command recognition results; Based on the voice command recognition result, corresponding feedback information is generated and output through the speaker of the audio device; The adaptive filtering algorithm includes: Based on the raw audio data, environmental sound characteristics are monitored in real time to determine the current scene type; Based on the scene type, a preset noise suppression mode is automatically selected, and the filtering algorithm parameters are dynamically adjusted according to the selected noise suppression mode.
[0006] Optionally, the method further includes: The audio signal is preprocessed locally by the edge computing unit, and the preprocessing results are sent to the cloud server for in-depth analysis. Preliminary feedback is generated based on local preprocessing results, and the information in the preliminary feedback is updated after receiving the analysis results from the cloud server.
[0007] Optionally, the preset noise suppression modes include quiet mode, noisy city mode, and music mode.
[0008] Optionally, the filtering algorithm parameters can be dynamically adjusted according to the selected noise suppression mode, specifically including the following steps: Real-time calculation of the spectral characteristics of current environmental noise to obtain noise spectral distribution data; Based on the selected noise suppression mode, the corresponding frequency band noise threshold is determined, and the noise evaluation standard is obtained; Based on the noise spectrum distribution data and the noise assessment criteria, calculate the noise suppression target value for each frequency band; Based on the noise suppression target value and the preset parameters of the noise suppression mode, the frequency response coefficient of the filter is dynamically adjusted to obtain the updated filter parameters; The updated filtering parameters are used to perform adaptive filtering on the audio signal.
[0009] Optionally, the frequency response coefficient of the dynamically adjusted filter specifically includes: In quiet mode, a first noise threshold is set to suppress ambient noise with a signal strength lower than the first noise threshold; In the noisy city mode, a second noise threshold is set to suppress noise signals within a preset frequency range, wherein the second noise threshold is greater than the first noise threshold. In music mode, the signal of the preset music frequency band is retained, and the noise signal outside the preset music frequency band is attenuated by a preset intensity.
[0010] Optionally, determine the current scene type, which includes the following steps: Extract time-domain and frequency-domain features from the original audio data; Calculate the statistical parameters of the time-domain and frequency-domain features to construct a scene feature vector; The scene feature vector is input into a pre-trained scene classification model to obtain the scene type probability distribution; Based on the probability distribution of the scene types, the type with the highest confidence level is selected as the current scene type.
[0011] Optionally, extracting time-domain and frequency-domain features includes the following steps: The short-time energy, zero-crossing rate, and duration of the audio signal are calculated to obtain its time-domain characteristics. Spectral analysis is performed on the audio signal to extract the frequency band energy ratio and spectral centroid, thereby obtaining frequency domain characteristics; Feature templates are established based on the time-domain and frequency-domain features for scene type identification and matching.
[0012] Secondly, this application provides an AI intelligent audio signal processing system, comprising: The audio acquisition module is used to receive audio signals from the environment in real time and obtain raw audio data; The noise suppression module is used to suppress background noise using an adaptive filtering algorithm based on the original audio data to obtain an enhanced speech signal; The speech recognition module is used to extract speech features based on the enhanced speech signal and match them with speech models in the database to obtain speech command recognition results; The feedback control module is used to generate corresponding feedback information based on the voice command recognition result and output it through the speaker of the audio device; The noise suppression module includes: The scene recognition unit is used to monitor environmental sound characteristics in real time based on the original audio data and determine the current scene type; The parameter adjustment unit is used to automatically select a preset noise suppression mode according to the scene type, and dynamically adjust the filtering algorithm parameters according to the selected noise suppression mode.
[0013] Thirdly, this application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described AI smart speaker signal processing method.
[0014] Fourthly, this application provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of the above-described AI smart speaker signal processing method.
[0015] In summary, this application includes at least one of the following beneficial technical effects: This application proposes an adaptive filtering scheme based on scene recognition. After acquiring environmental audio signals through an audio acquisition module, the system performs real-time analysis of the raw audio data, monitors its acoustic characteristics, and determines the specific scene type. After identifying the scene type, the system automatically selects the most suitable processing method from preset noise suppression modes and dynamically adjusts the parameter configuration of the filtering algorithm according to the mode. The voice signal after adaptive filtering is further feature-extracted and matched with voice models in the database to obtain accurate voice command recognition results. Finally, the system generates corresponding feedback information based on the recognition results and outputs it to the user through the speaker. This achieves accurate identification and suppression of noise in different scenes, improving the voice interaction performance of smart speakers in complex environments. This application first performs frequency domain analysis on environmental noise, obtaining the spectral distribution characteristics of the noise in real time to understand the noise energy distribution in each frequency band. After obtaining the spectral data, the system determines the corresponding frequency band noise threshold based on the currently selected noise suppression mode, establishing a benchmark standard for noise assessment. The actual obtained spectral distribution data is compared and analyzed with the assessment standard to calculate the specific suppression target value required for each frequency band. Based on these suppression target values and combined with preset parameters in different modes, the system can adaptively adjust the frequency response characteristics of the filter to generate filter parameter configurations. Finally, the updated filter parameters are used to process the original audio signal, achieving precise noise suppression. This not only adapts to various complex noise environments but also achieves optimal noise suppression while ensuring the quality of the speech signal. This application designs three differentiated filtering parameter adjustment schemes for different scenarios. In quiet environments, the system sets a lower first noise threshold to specifically process weak environmental noise and ensure clear recognition of voice commands. In noisy urban environments, considering the significant increase in noise intensity, the system adopts a higher second noise threshold and focuses on suppressing noise within a specific frequency range to cope with complex urban environmental noise. In music playback scenarios, the system adopts a more flexible processing strategy. By pre-dividing the main frequency bands of the music signal, it selectively attenuates noise in other frequency bands while preserving music characteristics. This achieves precise control of noise suppression, ensuring the accuracy of voice interaction while avoiding over-processing of effective signals. Attached Figure Description
[0016] Figure 1 This is a flowchart illustrating an AI smart speaker signal processing method according to an embodiment of this application; Figure 2 This is a schematic diagram of a module of an AI intelligent audio signal processing system according to an embodiment of this application; Figure 3This is an internal structural diagram of an electronic device according to an embodiment of this application. Detailed Implementation
[0017] The terminology used in the following embodiments of this application is for the purpose of describing particular embodiments only and is not intended to be limiting of this application. As used in the specification and appended claims of this application, the singular expressions “a,” “an,” “the,” “the,” “the,” and “this” are intended to include the plural expressions as well, unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in this application refers to any or all possible combinations including one or more of the listed items.
[0018] Hereinafter, the terms "first" and "second" are used for descriptive purposes only and should not be construed as implying or suggesting relative importance or implicitly indicating the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature, and in the description of the embodiments of this application, unless otherwise stated, "multiple" means two or more.
[0019] The embodiments of this application will now be described in further detail with reference to the accompanying drawings.
[0020] Firstly, this application provides an AI smart speaker signal processing method, referring to... Figure 1 It includes the following steps: S100: The audio acquisition module receives audio signals from the environment in real time to obtain raw audio data.
[0021] The audio acquisition module may include a microphone array for acquiring sound signals from the environment and converting the acquired analog signals into digital audio data.
[0022] S200. Based on the original audio data, an adaptive filtering algorithm is used to suppress background noise and obtain an enhanced speech signal.
[0023] The adaptive filtering algorithm first monitors the environmental sound characteristics in real time based on the original audio data, and determines the current scene type by analyzing the time domain and frequency domain characteristics of the audio signal.
[0024] Specifically, the system calculates the short-time energy, zero-crossing rate, and duration of the signal in the time domain, and simultaneously extracts frequency domain features such as the band energy ratio and spectral centroid after obtaining the spectrum through FFT transformation, thus constructing a scene feature vector. This feature vector is then input into a pre-trained scene classification model to obtain the scene type probability distribution, and the type with the highest confidence is selected as the current scene.
[0025] Understandably, by simultaneously analyzing time-domain and frequency-domain features and combining them with a pre-trained scene classification model, the system can accurately identify the current scene type, providing a reliable basis for subsequent noise suppression and improving scene adaptability.
[0026] Optionally, after determining the scene type, the system automatically selects the most suitable processing method from preset noise suppression modes. In this embodiment, the preset noise suppression modes include quiet mode, noisy city mode, and music mode. In quiet mode, a first noise threshold (e.g., -40dB) is set to suppress environmental noise below this threshold. In noisy city mode, a higher second noise threshold (e.g., -20dB) is set to focus on suppressing noise signals within a preset frequency range. In music mode, the system retains preset music frequency band signals while attenuating noise in other frequency bands at preset intensities. By calculating the spectral characteristics of environmental noise in real time, noise spectral distribution data is obtained, and the corresponding frequency band noise threshold is determined according to the selected mode, thereby obtaining a noise evaluation standard. Based on these data, the noise suppression target value for each frequency band is calculated, the frequency response coefficient of the filter is dynamically adjusted, and finally, updated filtering parameters are obtained and applied to signal processing.
[0027] Understandably, by setting different noise suppression modes and dynamically adjusting parameters, the system can adopt corresponding processing strategies for the noise characteristics of different scenarios, which not only ensures effective noise suppression but also avoids over-processing of effective signals, thus improving the practicality of the system.
[0028] Optionally, the system also performs local preprocessing of the audio signal through an edge computing unit, while simultaneously sending the preprocessing results to a cloud server for in-depth analysis. Preliminary feedback can be quickly generated based on the local preprocessing results, and the feedback information is updated after receiving the analysis results from the cloud server, achieving efficient collaboration between local processing and cloud analysis.
[0029] Understandably, the architecture of edge computing and cloud collaboration not only ensures the system's real-time response capability, but also improves the accuracy of processing through in-depth cloud analysis, achieving a good balance between response speed and processing precision.
[0030] S300: Based on the enhanced speech signal, extract speech features and match them with the speech model in the database to obtain the speech command recognition result.
[0031] Specifically, after obtaining the enhanced speech signal, the system extracts its speech features and matches these features with pre-stored speech models in a database to obtain accurate speech command recognition results. Speech feature extraction includes calculating and analyzing characteristic parameters such as the signal's fundamental frequency, formants, and Mel-frequency cepstral coefficients. The system employs a deep neural network model for feature matching to improve the accuracy of command recognition.
[0032] Understandably, through in-depth analysis of speech features and the application of neural network models, the system can accurately recognize users' voice commands, improving the accuracy and naturalness of human-computer interaction.
[0033] S400 generates corresponding feedback information based on the voice command recognition results and outputs it through the speaker of the audio device.
[0034] Finally, based on the voice command recognition results, the system generates corresponding feedback information and outputs it to the user through the smart speaker. This feedback process can include various forms such as voice replies and prompts, providing users with a clear interactive experience. The generation of feedback information takes into account the characteristics of the current scene, appropriately adjusting the output volume and timbre to ensure the feedback effect. Through a scene-adaptive feedback mechanism, the system can adjust output parameters according to the characteristics of the current environment, ensuring that users can clearly receive feedback information and improving the user experience.
[0035] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0036] Secondly, this application provides an AI intelligent audio signal processing system. The AI intelligent audio signal processing system of this application will be described below in conjunction with the above-mentioned AI intelligent audio signal processing method.
[0037] Reference Figure 2 An AI-powered intelligent audio signal processing system includes: The audio acquisition module is used to receive audio signals from the environment in real time and obtain raw audio data; The noise suppression module is used to suppress background noise based on the original audio data using an adaptive filtering algorithm to obtain an enhanced speech signal; The speech recognition module is used to extract speech features based on the enhanced speech signal and match them with speech models in the database to obtain speech command recognition results. The feedback control module is used to generate corresponding feedback information based on the voice command recognition results and output it through the speaker of the audio device; The noise suppression module includes: The scene recognition unit is used to monitor environmental sound characteristics in real time based on raw audio data and determine the current scene type; The parameter adjustment unit is used to automatically select a preset noise suppression mode according to the scene type, and dynamically adjust the filtering algorithm parameters according to the selected noise suppression mode.
[0038] In one embodiment, the system is configured to further include: The edge computing module includes: The preprocessing unit is used to perform local preprocessing on the audio signal; The data transmission unit is used to send the preprocessing results to the cloud server for in-depth analysis; The feedback update module includes: The preliminary feedback unit is used to generate preliminary feedback based on the local preprocessing results; The information update unit is used to update the initial feedback information after receiving the analysis results from the cloud server.
[0039] Optional preset noise suppression modes include Quiet Mode, Noise Mode, and Music Mode.
[0040] In another embodiment, the parameter adjustment module specifically includes: The spectrum analysis unit is used to calculate the spectrum characteristics of the current environmental noise in real time and obtain noise spectrum distribution data. The threshold determination unit is used to determine the corresponding frequency band noise threshold based on the selected noise suppression mode, and obtain the noise evaluation standard. The target calculation unit is used to calculate the noise suppression target value for each frequency band based on noise spectrum distribution data and noise assessment standards. The parameter update unit is used to dynamically adjust the frequency response coefficient of the filter based on the noise suppression target value and the preset parameters of the noise suppression mode, so as to obtain the updated filter parameters. The filtering unit is used to perform adaptive filtering of the audio signal using updated filtering parameters.
[0041] Specifically, the parameter update unit includes: The quiet mode subunit is used to set a first noise threshold and suppress ambient noise with a signal strength lower than the first noise threshold. The noisy city mode subunit is used to set a second noise threshold to suppress noise signals within a preset frequency range, wherein the second noise threshold is greater than the first noise threshold. The music mode subunit is used to retain the signal of the preset music frequency band and perform preset intensity attenuation processing on the noise signal outside the preset music frequency band.
[0042] Optionally, the scene recognition module includes: The feature extraction unit is used to extract time-domain and frequency-domain features from the raw audio data; The feature calculation unit is used to calculate the statistical parameters of time-domain and frequency-domain features and construct scene feature vectors; The scene classification unit is used to input scene feature vectors into a pre-trained scene classification model to obtain the scene type probability distribution. The type determination unit is used to select the type with the highest confidence as the current scene type based on the probability distribution of scene types.
[0043] The feature extraction unit specifically includes: The time-domain feature subunit is used to calculate the short-time energy, zero-crossing rate, and duration of the audio signal to obtain the time-domain features; The frequency domain feature subunit is used to perform spectral analysis on audio signals, extract the frequency band energy ratio and spectral centroid, and obtain frequency domain features; The feature template subunit is used to build feature templates based on time-domain and frequency-domain features for scene type recognition and matching.
[0044] In one embodiment, this application provides an electronic device, which may be a server, and its internal structure diagram may be as follows: Figure 3 As shown, the electronic device includes a processor, memory, and a network interface connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage medium. The database stores data. The network interface communicates with external terminals via a network connection. When the computer program is executed by the processor, it implements an AI-powered intelligent audio signal processing method.
[0045] Those skilled in the art will understand that Figure 3 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the electronic device to which the present application is applied. The specific electronic device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.
[0046] In one embodiment, an electronic device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.
[0047] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical storage, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0048] The above are all preferred embodiments of this application, and are not intended to limit the scope of protection of this application. Therefore, all equivalent changes made in accordance with the structure, shape and principle of this application should be covered within the scope of protection of this application.
Claims
1. An AI-powered intelligent speaker signal processing method, characterized in that, Includes the following steps: The audio acquisition module receives audio signals from the environment in real time to obtain raw audio data. Based on the original audio data, an adaptive filtering algorithm is used to suppress background noise and obtain an enhanced speech signal; Based on the enhanced speech signal, speech features are extracted and matched with speech models in the database to obtain speech command recognition results; Based on the voice command recognition result, corresponding feedback information is generated and output through the speaker of the audio device; The adaptive filtering algorithm includes: Based on the raw audio data, environmental sound characteristics are monitored in real time to determine the current scene type; Based on the scene type, a preset noise suppression mode is automatically selected, and the filtering algorithm parameters are dynamically adjusted according to the selected noise suppression mode.
2. The AI intelligent speaker signal processing method according to claim 1, characterized in that, The method further includes: The audio signal is preprocessed locally by the edge computing unit, and the preprocessing results are sent to the cloud server for in-depth analysis. Preliminary feedback is generated based on local preprocessing results, and the information in the preliminary feedback is updated after receiving the analysis results from the cloud server.
3. The AI intelligent speaker signal processing method according to claim 1, characterized in that, The preset noise suppression modes include quiet mode, noisy city mode, and music mode.
4. The AI intelligent speaker signal processing method according to claim 3, characterized in that, The filtering algorithm parameters are dynamically adjusted based on the selected noise suppression mode, specifically including the following steps: Real-time calculation of the spectral characteristics of current environmental noise to obtain noise spectral distribution data; Based on the selected noise suppression mode, the corresponding frequency band noise threshold is determined, and the noise evaluation standard is obtained; Based on the noise spectrum distribution data and the noise assessment criteria, calculate the noise suppression target value for each frequency band; Based on the noise suppression target value and the preset parameters of the noise suppression mode, the frequency response coefficient of the filter is dynamically adjusted to obtain the updated filter parameters; The updated filtering parameters are used to perform adaptive filtering on the audio signal.
5. The AI intelligent speaker signal processing method according to claim 4, characterized in that, The frequency response coefficient of the dynamically adjusted filter specifically includes: In quiet mode, a first noise threshold is set to suppress ambient noise with a signal strength lower than the first noise threshold; In the noisy city mode, a second noise threshold is set to suppress noise signals within a preset frequency range, wherein the second noise threshold is greater than the first noise threshold. In music mode, the signal of the preset music frequency band is retained, and the noise signal outside the preset music frequency band is attenuated by a preset intensity.
6. The AI intelligent speaker signal processing method according to claim 1, characterized in that, Determining the current scene type involves the following steps: Extract time-domain and frequency-domain features from the original audio data; Calculate the statistical parameters of the time-domain and frequency-domain features to construct a scene feature vector; The scene feature vector is input into a pre-trained scene classification model to obtain the scene type probability distribution; Based on the probability distribution of the scene types, the type with the highest confidence level is selected as the current scene type.
7. The AI intelligent speaker signal processing method according to claim 6, characterized in that, Extract time-domain and frequency-domain features. Includes the following steps: The short-time energy, zero-crossing rate, and duration of the audio signal are calculated to obtain its time-domain characteristics. Spectral analysis is performed on the audio signal to extract the frequency band energy ratio and spectral centroid, thereby obtaining frequency domain characteristics; Feature templates are established based on the time-domain and frequency-domain features for scene type identification and matching.
8. An AI-powered intelligent audio signal processing system, characterized in that, include: The audio acquisition module is used to receive audio signals from the environment in real time and obtain raw audio data; The noise suppression module is used to suppress background noise using an adaptive filtering algorithm based on the original audio data to obtain an enhanced speech signal; The speech recognition module is used to extract speech features based on the enhanced speech signal and match them with speech models in the database to obtain speech command recognition results; The feedback control module is used to generate corresponding feedback information based on the voice command recognition result and output it through the speaker of the audio device; The noise suppression module includes: The scene recognition unit is used to monitor environmental sound characteristics in real time based on the original audio data and determine the current scene type; The parameter adjustment unit is used to automatically select a preset noise suppression mode according to the scene type, and dynamically adjust the filtering algorithm parameters according to the selected noise suppression mode.
9. An electronic device, characterized in that, The method includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the AI smart audio signal processing method according to any one of claims 1-7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the AI intelligent audio signal processing method according to any one of claims 1-7.
Citation Information
Patent Citations
Long-distance ad hoc network full duplex voice communication system for fire fighting
CN119296559A
Processing method, device and system based on voice instruction
CN120183410A
Intelligent voice interaction and analysis system and method based on large model
CN120388563A
Self-contained environment sensing artificial intelligence tuning method, tuning system and terminal equipment
CN120636424A
Interphone audio processing method, system and device based on edge AI and medium
CN120748384A