A multi-form sibilance dynamic perception and adaptive suppression method

By combining spectrum analysis and FIR frequency dividers, refined dynamic perception and adaptive suppression of different sibilance patterns are achieved, solving the problems of coarse mid-frequency band division and insufficient adaptability in existing technologies, and improving the quality and comfort of speech signals.

CN121281547BActive Publication Date: 2026-05-19BEIJING FANGWEI ZHILIAN TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING FANGWEI ZHILIAN TECHNOLOGY CO LTD
Filing Date
2025-09-24
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Existing sibilance suppression methods have overly coarse frequency band divisions, making it impossible to accurately identify different sibilance patterns. Furthermore, they lack adaptability in dynamic environments, affecting speech quality and comfort.

Method used

The sibilance is divided into multiple sub-bands by spectrum analysis and FIR crossover. Dynamic suppression is achieved by using adaptive threshold and linear gain adjustment. Frequency division and suppression are performed for different phoneme categories. The amplitude and ratio of frequency points are calculated by using the Gosser algorithm to achieve fine dynamic perception and suppression.

Benefits of technology

It effectively suppresses the harshness of sibilance while preserving the natural brightness and clarity of speech, adapting to speech signal processing in different environments and reducing the impact on other frequency bands.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121281547B_ABST
    Figure CN121281547B_ABST
Patent Text Reader

Abstract

The application discloses a multi-form sibilance dynamic perception and self-adaptive suppression method, and relates to the field of speech signal processing; specifically, first, input real-time audio or speech stream speech signals, and perform preprocessing according to frames; then, perform spectrum analysis to obtain corresponding high-frequency and low-frequency sibilance perception frequency bands of each frame, and determine frequency division points of each frequency band in each frame of speech signals. Next, a FIR frequency divider performs frequency division processing on audio signals of each frame according to the frequency division points of the high-frequency and low-frequency sibilance perception frequency bands, to obtain five subbands; perform root mean square (RMS) calculation on two subbands of Hz to Hz and Hz to Hz, to obtain instantaneous loudness of each subband, and perform judgment on the sibilance self-adaptive threshold value; when the RMS value of a certain subband exceeds the threshold value, compression processing is triggered; finally, perform weighted synthesis on the subband audio signals subjected to compression processing and the remaining subband audio signals, to obtain balanced output speech of each frame. The application can reduce the harshness of sibilance while retaining the natural brightness and clarity of speech.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of speech signal processing, and in particular to a method for dynamic perception and adaptive suppression of multimorphic dental sounds. Background Technology

[0002] Sibilants are a common type of prominent noise component in human speech. Phonemes such as "s," "z," and "ch" produce high-frequency, concentrated sound due to friction of airflow against the alveoli, tongue, or hard palate during pronunciation. Excessive sibilants can sound harsh and shrill to the listener, even affecting speech comprehension and auditory comfort. The prominence of sibilants is particularly pronounced in recording, communication, and broadcasting scenarios, especially when using high-sensitivity microphones, near-field pickup, or dynamically compressed audio signals.

[0003] In existing technologies, methods for suppressing dental sounds mainly fall into two categories:

[0004] One method involves reducing the energy of specific high-frequency bands using an equalizer. This method is static and crude, and it easily weakens the high-frequency components in speech, making the speech lack brightness and airiness.

[0005] Another type uses a dynamic compressor to attenuate the signal energy in a specific frequency band when it exceeds a threshold. Although this is more flexible, the frequency band settings are usually wider, making it difficult to distinguish the characteristics of different phonemes. This may result in some sibilants not being fully suppressed, or the overall timbre being affected when over-suppressed.

[0006] Since the frequency distribution of different dentate sounds is not the same, most existing methods identify and process dentate sounds as a single frequency band, failing to make full use of this difference. Therefore, they have natural limitations in terms of sound quality fidelity and suppression effect.

[0007] In addition, the sibilant characteristics of speech signals are greatly affected by the speaker's timbre, the recording environment, and the frequency response curve of the equipment. Static or single-band suppression schemes cannot adapt to the ever-changing real-world scenarios.

[0008] Therefore, it is necessary to propose a more refined dynamic recognition and multi-class modeling dynamic suppression algorithm for deaberration. Summary of the Invention

[0009] To address the issues of overly coarse frequency band segmentation, inaccurate detection, and insufficient dynamic adaptability in existing sibilant processing methods, this invention proposes a multi-morphological sibilant dynamic perception and adaptive suppression method. This method can perform frequency segmentation and suppression for different phoneme categories and can adaptively adjust in dynamic environments, thereby reducing the harshness of sibilants while preserving the natural brightness and clarity of speech.

[0010] The method for dynamic perception and adaptive suppression of multimorphic dental sounds includes the following steps:

[0011] Step 1: Input real-time audio or speech stream signals and preprocess them frame by frame;

[0012] The preprocessing process is as follows: standard speech is sampled in mono at a sampling rate of 16 kHz or 48 kHz, and stereo is processed in the form of two mono audio streams; then the speech signal is divided into fixed frame lengths, each frame typically 10–30 ms, and the speech data of each frame is stored in a buffer.

[0013] Step 2: Perform spectral analysis on each frame of preprocessed speech signal to obtain the low-frequency sibilance perception band and high-frequency sibilance perception band corresponding to each frame, and determine the frequency division point of each band in each frame of speech signal.

[0014] The specific steps are as follows:

[0015] Step 201: Divide the sibilants in the speech signal of the current frame into two categories, and determine the main distribution frequency bands and the core frequency bands that affect the strength of sibilants in each category:

[0016] The sibilant sounds “z”, “c”, “s”, “j”, “x” and “t” are mainly distributed in the 4kHz to 15kHz frequency band, and the core influence on the strength of sibilant sounds is in the 6kHz to 12kHz frequency band.

[0017] The sibilant sounds “zh”, “ch”, and “sh” are mainly distributed in the frequency range of 1.5kHz to 12kHz, with the core influence on the strength of sibilant sounds being in the frequency range of 1.5kHz to 4kHz.

[0018] Step 202: Divide the frequency bands that affect the strength of sibilance in both categories into frequency points and calculate the ratio of each frequency point in the two categories of frequency bands.

[0019] First, the frequency points in the 6kHz to 12kHz and 1.5kHz to 4kHz frequency bands are divided into 250Hz intervals. For all the divided frequency points, the Gossel algorithm is used to detect them and calculate the amplitude of each frequency point.

[0020] Then, the amplitudes at all frequency points are statistically analyzed to obtain an overall mean M:

[0021]

[0022] For the first The amplitude at each frequency point.

[0023] Finally, the amplitude at each frequency point is compared with the mean to obtain the ratio of each frequency point;

[0024] Step 203: For each type of frequency band, calculate the ratio of every five adjacent frequency points, and select the frequency band covered by the group of five frequency points with the highest ratio as the sibilance perception frequency band.

[0025] Step 204: Divide the two types of dentition perception frequency bands into low-frequency dentition perception frequency band and high-frequency dentition perception frequency band, and further determine their respective frequency division points;

[0026] The low-frequency sibilance perception band corresponding to the 1.5kHz to 4kHz frequency range, with the first of the five determined frequency points being the starting frequency of the band. The fifth frequency point is the end frequency of the frequency band. ;frequency and frequency This is the frequency division point.

[0027] The high-frequency sibilance perception band corresponding to the 6kHz to 12kHz frequency range, the first of the five frequency points finally determined is the starting frequency of the frequency band. The fifth frequency point is the end frequency of the frequency band. ;frequency and frequency This is the frequency division point.

[0028] Step 3: The FIR frequency divider divides the audio signal of each frame into five sub-bands based on the frequency division points of the high and low frequency sibilance sensing bands.

[0029] The low-frequency sibilance perception frequency range is Hz to Hz, the range of high-frequency sibilance perception frequency band is up to Hz to Hz divides the audio signal into five sub-bands: 0 to 100 Hz. Hz, Hz to Hz, Hz to Hz, Hz to Hz and Hz – upper limit of signal frequency; each frequency band is separated by independent FIR bandpass and high- and low-pass filters to ensure that signals of different frequency bands do not interfere with each other.

[0030] Step 4, for Hz to Hz, Hz to The instantaneous loudness of each sub-band is obtained by performing root mean square (RMS) calculations on the two sub-bands of Hz. and adaptive threshold for dental sounds If the RMS value of a sub-band exceeds the threshold, compression processing is triggered.

[0031] Adaptive threshold for sibilance It is three times the RMS mean of the audio of the five preceding historical frames of the current frame, and the current frame starts from the 6th frame.

[0032] Compression The formula is as follows:

[0033]

[0034] The ratio is a preset compression ratio parameter used to control the compression intensity.

[0035] The RMS value LD' of the compressed audio is calculated as follows:

[0036]

[0037] During compression, a linear gain adjustment method is used to attenuate the loudness of the corresponding frequency band. Specifically, the amplitude of the sub-band signal is converted into a linear scaling factor according to the compression amount and then multiplied to achieve dynamic suppression.

[0038] Step 5: The compressed sub-band audio signal is weighted and synthesized with the audio signals of the remaining sub-bands to obtain balanced output speech for each frame.

[0039] The advantages of this invention are:

[0040] This invention analyzes and processes dentate sounds, adding dynamic perception and dynamic suppression to achieve dynamic perception and adaptive suppression of multi-morphological dentate sounds. It effectively suppresses dentate sounds in high-sampling-rate audio signals during the sound pickup process, while minimizing damage to other parts of the audio signal. Attached Figure Description

[0041] Figure 1 This is a flowchart of the steps of a multi-morphological dental sound dynamic perception and adaptive suppression method according to the present invention;

[0042] Figure 2 This is a schematic diagram of the principle of a multi-morphological dental sound dynamic perception and adaptive suppression method of the present invention. Detailed Implementation

[0043] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the described embodiments of the present invention are within the scope of protection of the present invention.

[0044] This invention provides a method for dynamic perception and adaptive suppression of multimorphic sibilance, which can be widely applied in voice communication, voice enhancement, speech recognition preprocessing, speech synthesis postprocessing, broadcast and television recording, audio and video post-production, mobile terminal voice interaction, and acoustic devices such as hearing aids, conference systems, headphones, and smart speakers.

[0045] This invention leverages the audio input / output interface and processor computing power of an embedded platform to automate the detection of audio acquisition and playback paths without relying on external auxiliary equipment. It is widely applicable to production testing and troubleshooting scenarios for devices including, but not limited to, VoIP terminals, smart access control systems, intercom devices, video conferencing terminals, and industrial embedded systems.

[0046] like Figure 1 As shown, the specific steps are as follows:

[0047] Step 1: Input real-time audio or speech stream signals and preprocess them frame by frame;

[0048] The preprocessing process is as follows: standard speech is sampled in mono at a sampling rate of 16 kHz or 48 kHz, and stereo is processed in the form of two mono audio streams; then the speech signal is divided into fixed frame lengths, each frame typically 10–30 ms, and the speech data of each frame is stored in a buffer.

[0049] Step 2: Perform spectral analysis on each frame of preprocessed speech signal to obtain the low-frequency sibilance perception band and high-frequency sibilance perception band corresponding to each frame, and determine the frequency division point of each band in each frame of speech signal.

[0050] The specific steps are as follows:

[0051] Step 201: Divide the sibilants in the speech signal of the current frame into two categories, and determine the main distribution frequency bands and the core frequency bands that affect the strength of sibilants in each category:

[0052] The sibilant sounds “z”, “c”, “s”, “j”, “x” and “t” are mainly distributed in the 4kHz to 15kHz frequency band, and the core influence on the strength of sibilant sounds is in the 6kHz to 12kHz frequency band.

[0053] The sibilant sounds “zh”, “ch”, and “sh” are mainly distributed in the frequency range of 1.5kHz to 12kHz, with the core influence on the strength of sibilant sounds being in the frequency range of 1.5kHz to 4kHz.

[0054] Step 202: Divide the frequency bands that affect the strength of sibilance in both categories into frequency points and calculate the ratio of each frequency point in the two categories of frequency bands.

[0055] First, the frequency points in the 6kHz to 12kHz and 1.5kHz to 4kHz frequency bands are divided into 250Hz intervals. For all the divided frequency points, the Gossel algorithm is used to detect them and calculate the amplitude of each frequency point.

[0056] Then, the amplitudes at all frequency points are statistically analyzed to obtain an overall mean M:

[0057]

[0058] For the first The amplitude at each frequency point.

[0059] Finally, the amplitude at each frequency point is compared with the mean to obtain the ratio of each frequency point;

[0060] Step 203: For each type of frequency band, calculate the ratio of every five adjacent frequency points, and select the frequency band covered by the group of five frequency points with the highest ratio as the sibilance perception frequency band.

[0061] Step 204: Divide the two types of dentition perception frequency bands into low-frequency dentition perception frequency band and high-frequency dentition perception frequency band, and further determine their respective frequency division points;

[0062] The low-frequency sibilance perception band corresponding to the 1.5kHz to 4kHz frequency range, with the first of the five determined frequency points being the starting frequency of the band. The fifth frequency point is the end frequency of the frequency band. ;frequency and frequency This is the frequency division point.

[0063] The high-frequency sibilance perception band corresponding to the 6kHz to 12kHz frequency range, the first of the five frequency points finally determined is the starting frequency of the frequency band. The fifth frequency point is the end frequency of the frequency band. ;frequency and frequency This is the frequency division point.

[0064] Step 3: The FIR frequency divider divides the audio signal of each frame into five sub-bands based on the frequency division points of the high and low frequency sibilance sensing bands.

[0065] The low-frequency sibilance perception frequency range is Hz to Hz, the range of high-frequency sibilance perception frequency band is up to Hz to Hz divides the audio signal into five sub-bands: 0 to 100 Hz. Hz, Hz to Hz, Hz to Hz, Hz to Hz and Hz – upper limit of signal frequency; each frequency band is separated by independent FIR bandpass and high- and low-pass filters to ensure that signals of different frequency bands do not interfere with each other.

[0066] Step 4, for Hz to Hz, Hz to The instantaneous loudness of each sub-band is obtained by performing root mean square (RMS) calculations on the two sub-bands of Hz. and adaptive threshold for dental sounds If the RMS value of a sub-band exceeds the threshold, compression processing is triggered.

[0067] Adaptive threshold for sibilance It is three times the RMS mean of the audio of the five preceding historical frames of the current frame, and the current frame starts from the 6th frame.

[0068] Compression The formula is as follows:

[0069]

[0070] The ratio is a preset compression ratio parameter used to control the compression intensity.

[0071] The RMS value LD' of the compressed audio is calculated as follows:

[0072]

[0073] During compression, a linear gain adjustment method is used to attenuate the loudness of the corresponding frequency band. Specifically, the amplitude of the sub-band signal is converted into a linear scaling factor according to the compression amount and then multiplied to achieve dynamic suppression.

[0074] Step 5: The compressed sub-band audio signal is weighted and synthesized with the audio signals of the remaining sub-bands to obtain balanced output speech for each frame.

[0075] This invention includes the following three processes:

[0076] 1. Multimorphic dental sound modeling and classification:

[0077] Different dentate sounds are mapped and categorized according to their corresponding frequency ranges. First, spectral features are extracted from the speech data to reveal the distribution characteristics of dentate components in the time-frequency domain. Then, based on the frequency concentration areas and energy distribution patterns, a correspondence between dentate sound morphology and characteristic frequency bands is established by modeling different categories of dentate sounds, and a dentate sound recognition representation model is constructed on this basis. The focus of dynamic dentate sound perception is concentrated in two frequency bands: 1.5kHz to 4kHz and 6kHz to 12kHz.

[0078] 2. Dynamic perception of dental sounds:

[0079] Analysis was conducted on two frequency bands, 1.5kHz to 4kHz and 6kHz to 12kHz, to determine the specific range of sibilance impact.

[0080] First, the two frequency bands are divided into frequency points at 250Hz intervals, resulting in frequency points of 1.5kHz, 1.75kHz, and 2.0kHz, as well as frequency points of 6.0kHz, 6.25kHz, and 6.5kHz. For all the divided frequency points, the amplitude and mean of each point are calculated to obtain the ratio of each frequency point. The relative energy distribution of each frequency point is obtained from the ratio.

[0081] After the ratio calculation is completed, the frequency points are scanned point by point, and the ratio of every five adjacent points is calculated. Among all the combined results, the group with the highest ratio sum is found, and the frequency band covered by the five frequency points in this group is determined to be the sibilance perception band. The sub-frequency points of each sub-frequency point are then further determined.

[0082] The core of this method lies in point-by-point calculation, ratio normalization, and joint determination of adjacent points. In this way, the main concentration region of sibilance can be accurately determined within a defined frequency range. The final output sibilance band serves as the basis for subsequent processing, used for further suppression or optimization.

[0083] 3. Dynamic suppression of sibilant sounds:

[0084] After the sibilance perception frequency band is determined, the entire audio signal is divided using an FIR frequency divider.

[0085] Taking 48kHz sampling rate audio as an example, when the perceived low-frequency sibilance frequency range is... Hz to Hz, the range of high-frequency sibilance perception frequency band is up to Hz to Hz divides the audio signal into five sub-bands: 0 to 100 Hz. Hz, Hz to Hz, Hz to Hz, Hz to Hz and Hz–24kHz. Each sub-band is separated by independent FIR bandpass and high-pass / low-pass filters to ensure that signals from different sub-bands do not interfere with each other.

[0086] After frequency division, the root mean square (RMS) value is calculated only for the high and low dentate sound perception bands to obtain the instantaneous loudness of the sub-band. Subsequently, the calculated RMS value LD is compared with the dentate sound adaptive threshold T and compression processing is performed.

[0087] During compression, a linear gain adjustment method is used to attenuate the loudness of the corresponding frequency band. Specifically, the amplitude of the sub-band signal is scaled by multiplying the compression amount by a linear scaling factor, thereby achieving dynamic suppression. This process is executed continuously in each frame to ensure that the energy of the sibilant frequency band is effectively controlled while keeping the original energy of the non-sibilant frequency band unaffected. This method dynamically suppresses the sibilant frequency band while leaving other frequency bands unaffected, providing accurate input for subsequent signal reconstruction.

[0088] Example:

[0089] like Figure 2 As shown, the principle and process are as follows:

[0090] Step 1: Speech Signal Preprocessing. To ensure the algorithm's compatibility under different audio sampling conditions, the system first obtains the sampling rate and number of channels of the input audio or speech stream. For standard speech processing, 16 kHz or 48 kHz mono is typically used. If the input is stereo, the stereo is processed separately as two mono speech streams. The audio is then segmented into fixed frame lengths, typically 10 ms each, and the speech data is stored in a buffer.

[0091] Step 2: Dental consonant morphology modeling. By analyzing the spectral characteristics of the speech signal, a mapping relationship is established between different types of dentate consonants and their corresponding frequency bands.

[0092] In this process, the preprocessed signal is first subjected to a short-time Fourier transform to extract the amplitude and energy distribution of each frame in the frequency domain. Then, based on phonetic experiments and dentine frequency spectrum statistics, dentine sounds are divided into two categories: 6kHz to 12kHz as the core influence frequency band; and 1.5kHz to 4kHz as the core influence frequency band. By establishing a morphological category-frequency range feature table for each type of dentine sound, dentine sounds can be quickly identified and located in subsequent processing stages. Simultaneously, based on the modeling results, the key frequency bands for dynamic sensing are determined to be 1.5kHz to 4kHz and 6kHz to 12kHz, thus providing a basis for precise suppression.

[0093] Step 3: Historical speech loudness monitoring.

[0094] The algorithm calculates the root mean square (RMS) value of each frame of signal and continuously monitors the loudness changes of speech over five historical frames, providing data support for adaptive thresholding. If the number of historical frames is less than five, only the existing historical frames are used. The average RMS value of several past frames is calculated, and three times this average is used as the dynamic threshold to adapt to changes in speech intensity while avoiding false triggering of non-sibilant components. The threshold is updated every frame to ensure that the algorithm can dynamically adjust with changes in speech energy, providing stable and accurate parameters for subsequent dynamic sibilant suppression processing.

[0095] Step 4: Dynamic perception of dental sounds.

[0096] This stage aims to accurately identify the main concentrated frequency bands of sibilance. First, the key frequency bands of 1.5kHz to 4kHz and 6kHz to 12kHz are divided into frequency points at 250Hz intervals. The amplitude of each frequency point is calculated using the Gosser algorithm to achieve real-time and efficient calculations on a small number of frequency points. Then, the amplitude of each frequency point is compared with the overall mean to obtain the ratio of each point, thus obtaining the relative energy distribution. Based on this, the frequency bands in both frequency ranges are scanned point-by-point in windows of five adjacent frequency points each, and the ratios are statistically summed. The window with the highest ratio sum is selected as the concentrated frequency band of sibilance. The final output sibilance band serves as the basis for subsequent frequency division and dynamic suppression.

[0097] Step 5: Frequency Division. In the frequency division stage, an FIR crossover is used to divide the audio signal into multiple sub-bands. Utilizing the filtering characteristics of a linear time-invariant system, the input signal is decomposed into several non-overlapping frequency bands in the frequency domain, each sub-band containing signal components within a specific frequency range. Due to its linear phase characteristics, the FIR filter ensures that the waveform of the sub-band signal remains undistorted in the time domain, thus maintaining the naturalness and clarity of the speech signal. The crossover design typically includes a combination of low-pass, band-pass, and high-pass filters. The boundaries of each sub-band are determined based on the distribution characteristics of the serrated frequency band, generating multiple non-interfering sub-band signals.

[0098] Step Six: Dynamic Adaptive Suppression of Sibilance. This stage is triggered by comparing the sub-band RMS value with an adaptive threshold. When the RMS value of a sub-band in a frame exceeds the historical threshold, the compression amplitude is calculated, and a linear gain scaling is applied to the sub-band signal to adjust the sub-band amplitude, achieving dynamic attenuation. This process continues every frame, ensuring that the energy of the sibilance band is effectively controlled, while the non-sibilance band maintains its original amplitude unaffected, ensuring the naturalness and integrity of the speech. In practical applications, the compression ratio parameter can be preset according to different devices or scenarios to flexibly adjust the suppression intensity, thereby achieving adaptive control.

[0099] Step 7: Reconstruction of multi-channel subband signals.

[0100] The sub-band signals, after dynamic suppression processing, are weighted and superimposed with the residual signals to generate the final output speech. During reconstruction, weighted coefficients are used to superimpose the sub-band signals, and temporal smoothing is performed using overlapping windowing or cross-fading methods to avoid phase differences or amplitude abrupt changes caused by sub-band splicing. The final output signal not only effectively suppresses sibilance but also preserves the naturalness and clarity of other speech frequency bands to the greatest extent possible. Simultaneously, the output signal can be used to update historical RMS and adaptive thresholds, achieving closed-loop feedback and ensuring that the suppression effect remains stable and efficient in continuous speech streams.

[0101] The method described in this invention operates within an embedded device. By performing spectral analysis, subband frequency division, and dynamic compression on the input speech signal, it achieves dynamic recognition and suppression of different types of sibilance. The algorithm includes the following steps:

[0102] (1) Speech signal preprocessing;

[0103] (2) Multimorphic dental sound modeling and classification, including classifying dental sounds into different categories and determining the characteristic frequency bands of each category;

[0104] (3) Historical speech loudness monitoring, used to calculate dynamic adaptive threshold;

[0105] (4) Dynamic perception of dentine sounds: the concentrated frequency band of dentine sounds is determined by point-by-point frequency analysis and joint determination of adjacent frequency points;

[0106] (5) Frequency division processing: The audio signal is divided into multiple sub-bands using an FIR frequency divider;

[0107] (6) Dynamic adaptive suppression of sibilance, which is performed by compression based on the subband RMS value and adaptive threshold;

[0108] (7) Multi-channel subband signal reconstruction: The suppressed subband signal and the residual signal are weighted and superimposed to generate the output speech.

[0109] The method described in this invention supports automated execution on embedded platforms and can upload detection or processing results to a host computer or management platform via serial port, network, or shared memory interface. Configuration parameters include sibilance category classification, frequency point interval, historical frame count, dynamic threshold multiplier, and compression ratio, to adapt to the detection and suppression requirements of different devices and speech scenarios. The method is based on an audio processing system, including: a signal acquisition module for acquiring input speech signals; a signal processing module; an output module for generating processed output speech signals; and a communication module for uploading processing results and interacting with a host computer or management platform.

[0110] This invention implements a complete processing flow from signal preprocessing, sibilance modeling, historical loudness monitoring, dynamic sensing, frequency division processing, dynamic adaptive suppression, to multi-channel subband signal reconstruction. This method can perform refined dynamic suppression of different types of sibilance, effectively reducing interference from sibilance on high-sampling-rate audio while maintaining speech naturalness. It is suitable for audio detection and troubleshooting scenarios in embedded phones, VoIP terminals, video conferencing equipment, and industrial embedded systems. Through continuous frame processing and adaptive threshold adjustment, precise control and stable output of sibilance are achieved, ensuring good performance of high-quality speech signals in various application environments.

Claims

1. A method for dynamic perception and adaptive suppression of multimorphic dental sounds, characterized in that, The specific steps include: Step 1: Input real-time audio or speech stream signals and preprocess them frame by frame; Step 2: Perform spectrum analysis on each frame of preprocessed speech signal to obtain the low-frequency sibilance perception band and high-frequency sibilance perception band corresponding to each frame, and determine the frequency division point of each band in each frame of speech signal. The specific steps are as follows: Step 201: Divide the sibilants in the speech signal of the current frame into two categories, and determine the distribution frequency bands of each category and the frequency bands that corely affect the strength of the sibilants: The sibilant sounds "z", "c", "s", "j", "x", and "t" are distributed in the frequency band from 4kHz to 15kHz, with the core influence on the strength of sibilant sounds in the frequency band from 6kHz to 12kHz. The dentate sounds "zh", "ch", and "sh" are distributed in the frequency band from 1.5kHz to 12kHz, with the core influence on the strength of dentate sounds in the frequency band from 1.5kHz to 4kHz. Step 202: Divide the frequency bands that affect the strength of sibilance in both categories into frequency points and calculate the ratio of each frequency point in the two categories of frequency bands. Step 203: For each type of frequency band, calculate the ratio of every five adjacent frequency points, and select the frequency band covered by the group of five frequency points with the highest ratio as the sibilance perception frequency band. Step 204: Divide the two types of dentition perception frequency bands into low-frequency dentition perception frequency band and high-frequency dentition perception frequency band, and further determine their respective frequency division points; Step 3: The FIR frequency divider divides the audio signal of each frame into five sub-bands based on the frequency division points of the high and low frequency sibilance sensing bands. Step 4, for Hz to Hz, Hz to The instantaneous loudness of each sub-band is obtained by performing root mean square (RMS) calculations on the two sub-bands of Hz. and adaptive threshold for dental sounds If the RMS value of a sub-band exceeds the threshold, compression processing is triggered. Compression The formula is as follows: Where ratio is a preset compression ratio parameter used to control compression intensity; The RMS value LD' of the compressed audio is calculated as follows: During the compression process, a linear gain adjustment method is used to attenuate the loudness of the corresponding frequency band; specifically, the amplitude of the sub-band signal is converted into a linear proportional coefficient according to the compression amount and then multiplied to achieve dynamic suppression. Step 5: The compressed sub-band audio signal is weighted and synthesized with the audio signals of the remaining sub-bands to obtain balanced output speech for each frame.

2. The method as described in claim 1, characterized in that, The preprocessing in step one is as follows: standard speech is in mono with a sampling rate of 16 kHz or 48 kHz, and stereo is in the form of two mono audio streams; then the speech signal is divided into fixed frame lengths, each frame being 10–30 ms, and the speech data of each frame is stored in a buffer.

3. The method as described in claim 1, characterized in that, Step 202 specifically involves: First, dividing the frequency bands from 6kHz to 12kHz and from 1.5kHz to 4kHz into frequency points at intervals of 250Hz, and then using the Gosser algorithm to detect each of the divided frequency points and calculate the amplitude of each frequency point. Then, the amplitudes at all frequency points are statistically analyzed to obtain an overall mean M: For the first The amplitude at each frequency point; Finally, the amplitude at each frequency point is compared with the mean to obtain the ratio of each frequency point.

4. The method as described in claim 1, characterized in that, Step 204 specifically involves: identifying the low-frequency sibilance perception frequency band corresponding to the 1.5kHz to 4kHz frequency band, and determining the first of the five frequency points as the starting frequency of the frequency band. The fifth frequency point is the end frequency of the frequency band. ;frequency and frequency This is the frequency division point; The high-frequency sibilance perception band corresponding to the 6kHz to 12kHz frequency range, the first of the five frequency points finally determined is the starting frequency of the frequency band. The fifth frequency point is the end frequency of the frequency band. ;frequency and frequency This is the frequency division point.

5. The method as described in claim 1, characterized in that, Step three specifically involves: the low-frequency sibilance perception frequency band range is... Hz to Hz, the range of high-frequency sibilance perception frequency band is up to Hz to Hz divides the audio signal into five sub-bands: 0 to 100 Hz. Hz, Hz to Hz, Hz to Hz, Hz to Hz and Hz – upper frequency limit; each frequency band is separated by independent FIR bandpass and high-pass / low-pass filters to ensure that signals in different frequency bands do not interfere with each other.

6. The method as described in claim 1, characterized in that, In step four, the sibilance adaptive threshold It is three times the RMS mean of the audio of the five preceding historical frames of the current frame, and the current frame starts from the 6th frame.