Speech feature processing method and device, equipment and medium

Through adaptive frequency resolution and time resolution adjustment, combined with nonlinear transformation and perceptual weighting processing of auditory perception models, a Meer spectrum representation is generated, which solves the problem of balance between anti-noise performance and key information retention of speech features in a noisy environment, and improves the quality and accuracy of speech recognition and synthesis.

CN120340477APending Publication Date: 2025-07-18PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 10 Cited by

Patent Information

Application Number
CN202510692019.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing speech feature extraction technology is difficult to balance frequency resolution, time resolution and auditory perception characteristics in a noisy environment, which makes it difficult to achieve an effective balance between speech features’s anti-noise performance and key information retention, affecting the accuracy and quality of speech recognition.

Method used

The fused Mel band energy is generated through adaptive frequency resolution adjustment, combined with time resolution analysis and auditory perception model, nonlinear transformation and perceptual weighting are performed, and Mel spectrum representation is generated.

Benefits of technology

Effectively reduce noise interference, enhance the key information retention ability of voice signals, improve speech recognition accuracy and speech synthesis quality, and achieve coordinated optimization of noise robustness and auditory perception adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120340477A_ABST
    Figure CN120340477A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of voice processing, can be applied to business scenes of financial science and technology, medical health and the like, and discloses a voice feature processing method, device, equipment and medium. Performing time resolution analysis based on the fused Mel band energy to generate a multi-scale Mel spectrum amplitude value, and performing nonlinear transformation on the multi-scale Mel spectrum amplitude value according to the noise intensity parameter to generate a noise suppression Mel component; and generating a perception weighting coefficient according to an auditory perception model, and executing frequency domain energy adjustment on the noise suppression Mel component to generate Mel spectrum representation. On the basis of frequency resolution self-adaption, time resolution dynamic adjustment and auditory perception modeling, nonlinear transformation and perception weighting processing are applied to the multi-scale Mel spectrum amplitude value, the influence of noise interference on voice features can be effectively reduced, and the key information retention capacity of voice signals is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of speech processing, and in particular, to a speech feature processing method, apparatus, device, and storage medium. Background Art

[0002] In the field of speech processing, Mel Frequency Cepstral Coefficients (MFCC) and Mel spectrograms have long been widely used as core features in speech recognition and speech synthesis systems. However, the existing technologies still have obvious deficiencies in aspects such as noise environment adaptability, matching with human auditory characteristics, and balance between speech quality and recognition performance, which limit the performance of the system in complex application scenarios.

[0003] In the field of fintech business, applications such as intelligent customer service, remote identity verification, and voice risk warning rely heavily on speech recognition systems. However, the existing feature extraction methods based on traditional MFCC or Mel spectrograms are extremely sensitive to background noise. Especially in telephone channels, counter recordings, or unstructured dialogue environments, slight noise interference can lead to a significant decrease in recognition accuracy.

[0004] In the field of healthcare business, speech interaction is gradually popularized in scenarios such as remote medical consultations, intelligent assisted diagnosis and treatment, and patient health management. However, the traditional calculation method of Mel spectrograms does not fully consider the sensitivity changes and time resolution characteristics of the human auditory system in different frequency ranges, resulting in insufficient capture of key information in patient speech. For example, during the process of a doctor collecting medical history or a patient describing symptoms, subtle intonation changes or short acoustic features may contain important clinical significance, but the existing feature processing mechanism is difficult to accurately retain such key information, thus affecting subsequent speech understanding and medical decision-making support.

[0005] In the field of general speech processing, there is also a problem that it is difficult to balance speech quality and speech recognition performance. Improvements in traditional methods to enhance speech clarity or reduce noise often lead to changes in the distribution of speech features, thereby causing a decrease in the matching of speech recognition models and an increase in the recognition error rate. Conversely, if only the recognition performance is optimized, the speech quality may be ignored, resulting in a decline in user experience. In actual deployment, especially in application scenarios with dynamic environmental changes, large device heterogeneity, and complex noise characteristics of speech data, the existing feature extraction technologies are difficult to achieve an ideal balance between speech quality and recognition effect, resulting in limited system performance.

[0006] In summary, the current speech feature extraction technologies have problems that need to be urgently improved in terms of robustness in a noise environment, adaptability to human auditory perception characteristics, and balance between speech quality and recognition accuracy, and cannot fully meet the comprehensive requirements of system performance in the fields of fintech, healthcare, and other high-reliability speech applications. Summary of the Invention

[0007] The main objective of the present invention is to provide a method, device, equipment, and storage medium for processing speech features, aiming to solve the technical problem that the prior art fails to balance frequency resolution, time resolution, and auditory perception characteristics in a noisy environment, making it difficult to achieve an effective balance between the anti-noise performance of speech features and the retention of key information.

[0008] To achieve the above objective, the present invention provides a method for processing speech features, including:

[0009] Obtain the original audio signal;

[0010] Perform adaptive frequency resolution adjustment on the original audio signal to generate fused Mel-band energy;

[0011] Perform time resolution analysis based on the fused Mel-band energy to generate multi-scale Mel-spectrum amplitude values;

[0012] Perform non-linear transformation on the multi-scale Mel-spectrum amplitude values according to the noise intensity parameter to generate noise-suppressed Mel components;

[0013] Generate a perceptual weighting coefficient according to the auditory perception model;

[0014] Perform frequency-domain energy adjustment on the noise-suppressed Mel components based on the perceptual weighting coefficient to generate a Mel-spectrum diagram representation.

[0015] Furthermore, to achieve the above objective, the present invention provides a device for processing speech features, including:

[0016] An audio acquisition module for obtaining the original audio signal;

[0017] A frequency resolution adjustment module for performing adaptive frequency resolution adjustment on the original audio signal to generate fused Mel-band energy;

[0018] A time resolution analysis module for performing time resolution analysis based on the fused Mel-band energy to generate multi-scale Mel-spectrum amplitude values;

[0019] A noise suppression module for performing non-linear transformation on the multi-scale Mel-spectrum amplitude values according to the noise intensity parameter to generate noise-suppressed Mel components;

[0020] A perceptual weighting generation module for generating a perceptual weighting coefficient according to the auditory perception model;

[0021] A frequency-domain energy adjustment module for performing frequency-domain energy adjustment on the noise-suppressed Mel components based on the perceptual weighting coefficient to generate a Mel-spectrum diagram representation.

[0022] Further, to achieve the above object, the present invention further provides a computer device, which includes a memory, a processor, and a voice feature processing program stored on the memory and executable on the processor. When the voice feature processing program is executed by the processor, the steps of the voice feature processing method as described above are implemented.

[0023] Further, to achieve the above object, the present invention further provides a computer-readable storage medium, on which a voice feature processing program is stored. When the voice feature processing program is executed by a processor, the steps of the voice feature processing method as described above are implemented.

[0024] Beneficial effects: The present invention relates to the technical field of voice processing and can be applied to business scenarios such as fintech and healthcare. It discloses a voice feature processing method, device, equipment, and medium, including: obtaining an original audio signal, performing adaptive frequency resolution adjustment on the original audio signal to generate fused mel-band energy, performing time resolution analysis based on the fused mel-band energy to generate multi-scale mel-spectrum amplitude values, performing non-linear transformation on the multi-scale mel-spectrum amplitude values according to a noise intensity parameter to generate noise-suppressed mel components, generating a perceptual weighting coefficient according to an auditory perception model, and performing frequency-domain energy adjustment on the noise-suppressed mel components based on the perceptual weighting coefficient to generate a mel-spectrum diagram representation. By applying noise adaptive non-linear transformation and perceptual weighting processing to multi-scale mel-spectrum amplitude values based on frequency resolution adaption, time resolution dynamic adjustment, and auditory perception modeling, the present invention can effectively reduce the influence of noise interference on voice features, enhance the ability to retain key information of voice signals, and at the same time improve the voice recognition accuracy and voice synthesis quality, realizing the collaborative optimization of noise robustness and auditory perception adaptability. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] The following will further illustrate the present invention with reference to the drawings. In the drawings:

[0026] Figure 1 is a schematic diagram of an application environment of the voice feature processing method in an embodiment of the present invention;

[0027] Figure 2 is a schematic flowchart of an embodiment of the voice feature processing method of the present invention;

[0028] Figure 3 is a schematic diagram of functional modules of a preferred embodiment of the voice feature processing device of the present invention;

[0029] Figure 4 is a schematic diagram of a structure of a computer device in an embodiment of the present invention;

[0030] Figure 5 is another schematic diagram of a structure of a computer device in an embodiment of the present invention. Detailed implementation manners

[0031] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0032] The voice feature processing method provided by the embodiments of the present invention can be applied to application environments such as Figure 1 wherein the client communicates with the server through a network. The server can obtain the original audio signal through the client, perform adaptive frequency resolution adjustment on the original audio signal to generate fused mel-band energy, perform time resolution analysis on the fused mel-band energy to generate multi-scale mel-spectrum amplitude values, perform non-linear transformation on the multi-scale mel-spectrum amplitude values according to the noise intensity parameter to generate noise-suppressed mel components, generate perceptual weighting coefficients according to the auditory perception model, and perform frequency-domain energy adjustment on the noise-suppressed mel components based on the perceptual weighting coefficients to generate a mel-spectrum map representation. Based on frequency resolution adaption, dynamic adjustment of time resolution, and auditory perception modeling, the present invention applies noise adaptive non-linear transformation and perceptual weighting processing to the multi-scale mel-spectrum amplitude values, which can effectively reduce the influence of noise interference on voice features, enhance the ability to retain key information of voice signals, improve the voice recognition accuracy and voice synthesis quality at the same time, and achieve the collaborative optimization of noise robustness and auditory perception adaptability. Among them, the client can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, and portable wearable devices. The server can be implemented by an independent server or a server cluster composed of multiple servers. The present invention will be described in detail below through specific embodiments.

[0033] Please refer to Figure 2 , Figure 2 which is a schematic flowchart of an embodiment of the voice feature processing method provided by the present invention. It should be noted that although the logical order is shown in the flowchart, in some cases, the steps shown or described herein can be executed in a different order.

[0034] As Figure 2 shown, the voice feature processing method proposed by the present invention includes the following steps:

[0035] S10. Obtain the original audio signal;

[0036] In this embodiment, the purpose of obtaining the original audio signal is to provide accurate and complete input data for subsequent frequency resolution adjustment and time resolution analysis. The original audio signal generally refers to audio data that has not undergone any frequency adjustment, time compression, or energy normalization processing. Its sources can be voice recordings in natural environments, real-time acquisitions by audio sensors, or standard voice data sets stored historically. When obtaining the original audio signal, it is necessary to ensure the time continuity and sampling consistency of the signal to avoid time-frequency feature distortion caused by problems such as sampling rate changes, frame loss, and truncation. The audio signal can adopt various sampling formats, such as 16kHz 16bit single-channel linear PCM encoding, or 48kHz floating-point encoding format, and different audio resolutions are selected according to different application requirements. To ensure the stability of subsequent frequency processing, preliminary abnormal segment elimination processing also needs to be synchronously performed when obtaining the original audio signal, such as detecting obvious silent segments, strong noise segments, or signal saturation distortion segments, and marking them for subsequent analysis reference. The original audio signal can also be appended with acquisition metadata, such as acquisition time, device type, environmental noise level, acquisition scene label, etc., for subsequent adaptive parameter adjustment. In the actual acquisition process, a dedicated microphone array, a single high-directivity microphone, or an embedded system such as a mobile terminal, a wearable device, or a medical monitoring device can be used to complete signal acquisition, and the number and type of input channels are flexibly configured according to different scenarios.

[0037] The actual implementation methods for obtaining the original audio signal include real-time access and buffering of data through an audio acquisition interface. For example, underlying sound card APIs such as ALSA, CoreAudio, and WASAPI are called in the operating system to continuously receive audio samples in a streaming manner, and a circular buffer is established to ensure data continuity; it is also possible to batch-read standard format audio files stored locally, such as WAV, FLAC, and PCM files, and perform sampling rate correction and channel number standardization processing uniformly after reading. To enhance robustness, the signal energy can also be monitored in real time during the acquisition stage, and the sampling gain can be dynamically adjusted or a sampling interruption protection mechanism can be triggered to avoid signal saturation distortion.

[0038] In the field of medical and health, a high-fidelity audio acquisition module can be integrated into implantable or wearable physiological monitoring devices to collect natural sound signals such as the patient's breathing, coughing, and voice communication. The sampling rate is usually set to 16kHz to balance data volume and retention of voice details, and the input gain is dynamically adjusted in combination with a real-time noise level evaluation module.

[0039] In the fintech business scenario, voice instructions or identity verification voices of customers can be collected through mobile terminals or intelligent teller machines. The sampling frequency can be set to 24kHz to improve the recognition accuracy, and the quality of the sampled signal is judged in real time in combination with an environmental background sound detection module, and a re-recording prompt is given when the noise background exceeds the standard.

[0040] For intelligent customer service and intelligent voice interaction systems, signal preprocessing can be performed in a far-field microphone array, including beamforming and spatial denoising, to enhance the signal quality of the target speaker. For different hardware platforms, floating-point or fixed-point sampling formats can be adapted, and the sampling buffer strategy can be adjusted according to the device memory and computing power. For example, a segmented double-buffer sampling mechanism can be adopted on resource-constrained devices to ensure sampling continuity and reduce latency.

[0041] In this embodiment, by obtaining raw audio signals that are standardized, highly continuous, and stable in quality, a data basis can be provided for subsequent processing. Using signal inputs with a unified sampling rate and a unified encoding format helps to improve the consistency and stability of the feature processing module, reduce feature extraction errors caused by abnormal input signals, further enhance the accuracy and robustness of the subsequent generated Mel spectrogram representation, and thus effectively support speech recognition, speech synthesis, and speech enhancement applications in complex noise environments.

[0042] S20, perform adaptive frequency resolution adjustment on the raw audio signal to generate fused Mel band energies;

[0043] In this embodiment, performing adaptive frequency resolution adjustment on the raw audio signal to generate fused Mel band energies aims to solve the problem that traditional Mel filtering processing cannot flexibly extract differential features for different frequency bands. The core idea of adaptive frequency resolution adjustment is to dynamically set the bandwidth of the frequency filter bank according to the frequency distribution characteristics of the raw audio signal, so as to maintain a higher frequency resolution in the low-frequency region and appropriately reduce the frequency resolution in the high-frequency region to match the characteristics that the human auditory system is more sensitive to low-frequency details and less sensitive to high-frequency details. This processing method can not only improve the ability of speech features to capture key information, but also effectively compress redundant features and reduce the complexity of subsequent feature processing.

[0044] In the specific process of implementing adaptive frequency resolution adjustment, first perform a short-time Fourier transform on the raw audio signal to extract the amplitude spectrum in the time-frequency domain. Subsequently, based on a preset frequency division rule, the entire frequency axis is divided into a low-frequency band and a high-frequency band. The low-frequency band is usually defined as the range between 0 Hz and 1000 Hz or 1500 Hz, and the high-frequency band is the frequency range above that. For low-frequency band signals, a densely set Mel filter bank is used, that is, a narrower bandwidth is set in the low-frequency band region to retain more detailed low-frequency change information. For high-frequency band signals, a sparsely set Mel filter bank is used, that is, a wider bandwidth is set in the high-frequency band region to highlight the main energy distribution and suppress tiny high-frequency noise perturbations.

[0045] To achieve a continuous frequency band energy distribution, after the division is completed, Mel filtering is performed on the low-frequency band and the high-frequency band respectively, and the first Mel frequency band energy distribution and the second Mel frequency band energy distribution are calculated separately. The first Mel frequency band energy distribution of the low-frequency band emphasizes the extraction of key information such as fundamental frequency, formants, and low-frequency resonances, while the second Mel frequency band energy distribution of the high-frequency band focuses on capturing the overall energy of high-frequency features such as fricatives and plosives. After obtaining these two parts of the Mel frequency band energy distributions, the first and second Mel frequency band energy distributions are fused by means of frequency band energy splicing or weighted transition to generate a fused Mel frequency band energy. This fusion not only maintains the integrity of the low-frequency details but also reasonably compresses and generalizes the high-frequency features, forming a feature representation that conforms to the characteristics of auditory perception and is suitable for subsequent processing.

[0046] During the adaptive frequency resolution adjustment process, to further improve flexibility, the frequency division threshold can also be dynamically fine-tuned or the density of the low-frequency and high-frequency filter banks can be adjusted based on the real-time analysis of the energy distribution characteristics of the original audio signal. For example, when it is detected that the energy of the input signal is mainly concentrated in the low-frequency region, the low-frequency band range can be appropriately expanded to further enhance the ability to extract low-frequency details; when it is detected that the high-frequency energy of the signal increases, the low-frequency band range can be compressed to improve the retention strength of high-frequency features. This dynamic adjustment mechanism can adaptively optimize the frequency resolution setting according to the actual requirements of different application scenarios, further improving the feature extraction effect of the system in complex environments.

[0047] In the field of medical and health, considering the characteristics that the physiological signal features such as breath sounds and heart sounds are mainly concentrated in the low-frequency band, the frequency division threshold can be set to 1200 Hz. 50 dense Mel filters are used for high-resolution filtering in the range of 0 to 1200 Hz, and 20 sparse Mel filters are used for low-resolution filtering above 1200 Hz to ensure sufficient capture of key low-frequency details while avoiding excessive interference from high-frequency noise.

[0048] In the fintech business scenario, such as identity verification voice collection or customer service voice recognition, since the acoustic features such as vowels and consonants in the voice commands cover a wide frequency band, the frequency division threshold can be set to 1000 Hz, and the density of the low-frequency filter is dynamically adjusted according to the real-time environmental noise level. For example, the number of low-frequency filters is increased in a noisy environment to enhance robustness.

[0049] In an intelligent interaction system, a frequency energy distribution analysis module can be integrated in the signal acquisition stage. According to parameters such as the detected speaker's gender and speaking intensity, different frequency division schemes can be adaptively selected. For example, the frequency division threshold is appropriately increased for female speakers to optimize the high-frequency feature extraction effect.

[0050] Different implementation methods can also be optimized and adjusted based on hardware resources. For example, in an embedded device, to reduce the consumption of computing resources, a fixed-frequency partition and pre-generated Mel filter bank can be adopted. In a server-side environment with sufficient computing resources, adaptive frequency partitioning and dynamic filter configuration can be calculated in real time to achieve a more refined feature extraction process.

[0051] In this embodiment, by performing adaptive frequency resolution adjustment on the original audio signal, the problems of low-frequency information loss and high-frequency redundancy in the traditional Mel spectrum feature extraction process can be effectively solved. By adopting the adaptive strategy of high resolution in the low frequency and low resolution in the high frequency, the feature extraction process can be dynamically optimized according to the actual characteristics of the audio signal, improving the sensitivity to the detailed changes of speech, while suppressing the interference of irrelevant noise, thereby enhancing the robustness and expressiveness of speech features in a complex environment and laying a solid foundation for subsequent time resolution analysis, noise suppression, and perceptual weighting processing.

[0052] S30. Perform time resolution analysis based on the fused Mel band energy to generate multi-scale Mel spectrum amplitude values;

[0053] In this embodiment, performing time resolution analysis based on the fused Mel band energy to generate multi-scale Mel spectrum amplitude values aims to solve the problem that the traditional single time-frequency analysis window cannot simultaneously take into account the extraction of fast-changing speech features and stable speech features. Time resolution analysis refers to localizing the fused Mel band energy in the time domain and extracting feature patterns at different time scales, so as to capture the dual characteristics of both fast-changing, transient features and stable, continuous, steady-state features existing in the speech signal. Multi-scale processing means generating multiple feature representations with complementary characteristics through analysis methods with different time window lengths to enhance the overall speech feature's perception ability of dynamic changes.

[0054] In the process of implementing time resolution analysis, first, perform voice activity detection on the fused Mel band energy to identify the fast-changing speech part and the stable speech part in the signal. The fast-changing speech part usually shows obvious amplitude changes and energy mutations, such as the pronunciation feature areas of plosives and fricatives; the stable speech part usually refers to the vowel areas or continuous vocal segments with small energy changes and stable harmonic structures. Voice activity detection can be jointly determined based on various statistical features such as short-time energy, zero-crossing rate, and spectral slope change rate to ensure the accurate classification of the dynamic characteristics of different types of speech.

[0055] After marking the rapidly changing speech part and the stable speech part, time-frequency analysis is performed on these two speech characteristics respectively using time analysis windows of different lengths. For the rapidly changing speech part, a short time window is used for time-frequency analysis, and the length of the time window can be set in the range of 10 ms to 25 ms to ensure high time resolution capture of short-term dynamic changes. For the stable speech part, a long time window is used for time-frequency analysis, and the length of the time window can be set in the range of 50 ms to 100 ms to improve the frequency resolution and enhance the ability to extract the steady-state structure. By differentiating the selection of the time window, more appropriate feature extraction of the fused Mel band energy can be performed at different scales, avoiding information loss or redundancy problems caused by a single window.

[0056] After applying short-window analysis and long-window analysis, transient Mel components and steady-state Mel components are generated respectively. The transient Mel components correspond to the time-frequency local features of the rapidly changing speech part, have higher time resolution, and can accurately locate sudden changes; the steady-state Mel components correspond to the detailed frequency structure of the stable speech part, have better frequency resolution ability, and help to retain important speech attributes such as formants and fundamental frequencies. In order to form a unified feature representation, it is necessary to perform multi-scale time-frequency fusion on the transient Mel components and the steady-state Mel components. During the fusion process, features at different time scales can be combined through linear superposition, weighted average or feature selection mechanism to generate the final multi-scale Mel spectrum amplitude value, which not only retains the rapidly changing information but also takes into account the integrity of the stable structure.

[0057] Multi-scale time-frequency fusion can also introduce a feature weight adaptive mechanism to dynamically allocate the fusion weights of transient and steady-state features according to the degree of dynamic change of the local signal. For example, in the region where the energy of the speech signal changes violently, the fusion weight of the transient component is increased; in the region where the signal is stable, the fusion weight of the steady-state component is increased. Through this adaptive fusion strategy, the expression accuracy and robustness of the final multi-scale Mel spectrum amplitude value in a complex speech environment can be further improved.

[0058] In the field of medical and health, for the application of respiratory sound monitoring, time resolution analysis can be performed based on the fused Mel band energy. The sharp inspiratory sound in the respiratory cycle is marked as the rapidly changing speech part, and a 15-ms short window is used for analysis to extract the sharp change characteristics of inspiration. The steady expiratory sound is marked as the stable speech part, and an 80-ms long window is used for analysis to extract the continuous expiratory energy characteristics, and a complete respiratory sound Mel feature map is generated by fusion.

[0059] In the scenario of medical and health data processing, for the assessment of the speech ability of the elderly, by detecting the rapidly vibrating sounds and stable vowel regions in the speech, short-window and long-window processing are respectively applied to generate transient and steady-state components for evaluating the speaking stability and frequency drift indicators.

[0060] In the financial technology business area, for intelligent customer service voice analysis, the accuracy of emotion change detection and pronunciation clarity assessment is improved by performing different time scale analyses on the rapid reactive utterances and continuous explanatory utterances in the responses of customer service personnel.

[0061] The time window parameters under different implementations can be flexibly adjusted according to the actual application scenarios and system performance requirements. For example, in medical mobile monitoring equipment with complex noise environments, the short window length can be appropriately shortened to enhance the sensitivity to transient noise changes; in the speech recognition server-side processing flow, a longer long window analysis can be used to improve the frequency resolution to optimize the acoustic model training effect.

[0062] This embodiment performs time resolution analysis based on fused Mel-band energy to generate multi-scale Mel-spectrogram amplitude values, which can effectively solve the problem that traditional single-window time-frequency analysis cannot capture fast-changing features of speech and cannot fully express stable structures. Through differentiated time window processing combined with multi-scale time-frequency fusion, key speech features can be extracted more accurately in a dynamic and drastically changing environment, while maintaining the integrity of frequency details in stable speech areas, thereby comprehensively improving the temporal stability and frequency resolution capabilities of speech features, and providing a higher-quality feature foundation for subsequent noise suppression, perceptual weighting, and vocoder processing.

[0063] S40, performing nonlinear transformation on the multi-scale Mel spectrum amplitude value according to the noise intensity parameter to generate a noise suppressed Mel component;

[0064] In this embodiment, the multi-scale Mel spectrum amplitude value is nonlinearly transformed according to the noise intensity parameter to generate a noise suppression Mel component, aiming to perform adaptive characteristic processing on the noise interference existing in different frequency bands in the speech signal and improve the robustness of the speech feature in a complex noise environment. The noise intensity parameter is an indicator to measure the energy level of the background noise of the current speech signal, which comes from the estimation of the energy of the silent segment of the original audio signal and the calculation of the ratio of the speech segment energy to the background noise energy. The silent segment refers to a low-energy segment that does not contain effective speech pronunciation content during the speech process. The silent segment interval can be accurately extracted by short-time energy threshold detection and zero-crossing rate detection. By performing a fast Fourier transform on the silent segment and extracting the amplitude spectrum, the energy distribution of the silent background noise in the frequency domain can be obtained. The energy of the silent segment amplitude spectrum is accumulated or averaged to generate background noise energy as a reference for the noise intensity parameter.

[0065] On this basis, the average energy of the speech segment is further extracted, and the ratio is calculated with the background noise energy to obtain the signal-to-noise ratio parameter. The signal-to-noise ratio parameter is an index that dynamically reflects the local signal-to-noise level of the speech signal and is used to guide the subsequent non-linear transformation intensity control. According to the preset mapping table, different ranges of signal-to-noise ratio parameters are mapped to different combinations of power exponents. The power exponent defines the transformation intensity when performing non-linear transformation on the amplitude values of different frequency bands. A high exponent corresponds to a stronger suppression effect, and a low exponent corresponds to a milder adjustment.

[0066] For the multi-scale Mel spectrum amplitude values, they are divided into high-frequency components and low-frequency components according to the preset frequency threshold. The high-frequency components mainly correspond to high-frequency details in speech such as fricative and plosive features, and the low-frequency components mainly correspond to the fundamental frequency component and formant structure. After division, the first non-linear transformation is performed on the high-frequency components using the first power exponent, and the second non-linear transformation is performed on the low-frequency components using the second power exponent. The non-linear transformation operation refers to performing a power function transformation on each Mel frequency band amplitude value, such as performing a power operation on the amplitude value, to enhance the relative advantage of high-energy components and suppress low-energy noise components.

[0067] The non-linear transformation not only changes the numerical value of the original amplitude value but also effectively reshapes the dynamic range distribution of the frequency characteristics, making the main components of the speech signal more prominent and naturally suppressing the background noise components. To ensure that the order of the transformed frequency characteristics is not disordered, it is necessary to fuse the high-frequency components after the first non-linear transformation and the low-frequency components after the second non-linear transformation in the Mel frequency band order to generate continuous and consistent noise suppression Mel components, providing better input features for subsequent perceptual weighting processing.

[0068] In the field of medical and health, for the task of sleep apnea monitoring, the background breathing noise level can be extracted based on silent segment detection, and the power exponent can be adaptively adjusted in combination with the night environmental noise. A larger first power exponent (such as 2.8) is applied to the high-frequency components of the snoring signal to enhance the separation effect of high-frequency change signals, and at the same time, a smaller second power exponent (such as 1.5) is applied to the low-frequency respiratory flow signal to retain the respiratory rhythm information.

[0069] In the scenario of medical and health data analysis, for the speech signals of patients with neurological speech disorders, the non-linear transformation strategy can be dynamically adjusted according to the real-time monitored background noise changes to improve the recognizability of weak speech pronunciation segments.

[0070] In the field of fintech business, for the environmental noise interference in the process of remote audio acquisition, by performing strong non-linear suppression on the high-frequency background noise (such as the first power exponent of 3.0) and retaining an appropriate dynamic range for the low-frequency speech signal (such as the second power exponent of 1.7), the stability and accuracy of the speech recognition system in an open environment are improved.

[0071] Under different implementation manners, the grading standard of the signal-to-noise ratio parameter can be flexibly configured according to the application scenario. For example, in a strong noise environment, a finer-grained signal-to-noise ratio partition mapping table can be adopted to define more levels of power exponents to achieve more accurate amplitude dynamic compression.

[0072] In this embodiment, by performing a non-linear transformation on the multi-scale Mel spectrum amplitude value according to the noise intensity parameter, the accuracy of extracting speech features under complex background noise conditions can be significantly improved. The non-linear transformation adaptively adjusts the feature dynamic range according to the local signal-to-noise level, strengthens the speech principal components, effectively suppresses the background noise interference, and thus improves the overall performance of subsequent feature perception weighting and speech reconstruction processing. Compared with the traditional unified power processing method, the dynamic non-linear transformation based on the noise intensity parameter can balance speech clarity and robustness, and exhibits better feature stability and adaptability in various signal-to-noise ratio change environments.

[0073] S50, generate a perceptual weighting coefficient according to the auditory perception model;

[0074] In this embodiment, generating a perceptual weighting coefficient according to the auditory perception model aims to introduce weight adjustment that conforms to human auditory characteristics in the frequency-domain feature processing to optimize the performance of the speech signal in subsequent analysis, recognition, or synthesis tasks. The auditory perception model is a mathematical model established based on the variation law of the sensitivity of the human ear to sounds of different frequencies, and usually includes standard equal-loudness contour data and masking effect modeling. The standard equal-loudness contour data comes from the ISO 226 international standard, which describes the equalization relationship of the subjective perceived loudness of each frequency component by the human ear at different sound pressure levels. The masking effect modeling describes the phenomenon that a strong signal will mask the perception of a weak signal when multiple frequency components exist simultaneously, and is usually quantitatively expressed through the critical band theory and the masking threshold calculation formula.

[0075] In the process of generating the perceptual weighting coefficient, first load the preset auditory perception model to ensure that the model contains a frequency sensitivity weight table and masking threshold parameters. Initially generate an initial frequency sensitivity weight distribution according to the standard equal-loudness contour data to form the basic weighting coefficients corresponding to each Mel frequency band. This initial weighting distribution reflects the perception priority of the human ear for the sound energy of each frequency band without specific environmental noise interference. In order to further adapt to individual differences or the requirements of specific application environments, obtain the hearing threshold test data of the user. The hearing threshold test data can be detected by professional equipment or software, and records the ability of the user to perceive the minimum sound pressure level at different frequencies.

[0076] Adjust the initial frequency sensitivity weight distribution based on the hearing threshold test data. For the frequency range where the user's hearing is sensitive, the weight coefficient can be appropriately increased; for the frequency range with greater hearing loss, the weight coefficient can be decreased or kept fixed to avoid introducing noise enhancement. Through this adjustment process, a personalized auditory weighting curve is generated, which better conforms to the subjective auditory perception characteristics of the target user. Subsequently, analyze the Mel-band energy distribution of the noise suppression Mel components, extract the local energy levels of each frequency band, and calculate the auditory masking threshold of each Mel band in combination with the masking effect parameters provided by the auditory perception model. The masking threshold reflects the minimum energy that a specific frequency component needs to reach to be perceived in the presence of background noise.

[0077] Finally, perform weighted fusion of the personalized auditory weighting curve and the auditory masking threshold. The fusion process can be achieved through linear superposition, multiplicative superposition, or a weighting function based on a perceptual optimization strategy to generate the final perceptual weighting coefficient. The perceptual weighting coefficient is applied to the noise suppression Mel components during the frequency-domain energy adjustment stage to dynamically adjust the energy of each frequency band, making the retained important speech information more conform to the human ear's perception characteristics, suppressing the energy of noise or unnecessary frequency bands further, and making the overall features more in line with the auditory perception optimization goal.

[0078] In the field of medical and health, for the speech assistance system for otological hearing rehabilitation patients, a hearing standard model for a specific age group can be loaded, and personalized perceptual weighting coefficients can be generated by adjusting based on the patient's actual hearing test data to enhance the energy of the audible frequency band of the speech prompt system. In the scenario of medical and health data processing, for the remote pathological speech diagnosis platform, the energy distribution of the noise suppression Mel components can be analyzed in real time based on the characteristics of the patient's call background noise, the masking threshold can be calculated dynamically, and the perceptual weighting coefficients that vary over different time periods can be generated adaptively to enhance the pathological feature signals.

[0079] In the field of fintech business, for the audio identity verification system in the remote account opening video, a standardized auditory perception model can be loaded, the noise spectrum energy can be detected in combination with the characteristics of the microphones of different devices, and the perceptual weighting coefficient can be adjusted to enhance the vocal characteristics in the key frequency bands while suppressing interference noises such as keyboard tapping sounds and environmental echoes, improving the accuracy of identity verification recognition.

[0080] In different implementation manners, the auditory masking threshold can be achieved by a method of dynamically estimating based on frequency grouping. For example, a higher masking gain is applied in the low-frequency band, and more stringent masking requirements are set in the high-frequency band to adapt to different types of background noise.

[0081] In this embodiment, by generating a perceptual weighting coefficient according to the auditory perception model, it is possible to fully combine the sensitivity of the human ear to signals of different frequencies and the characteristics of the masking effect, and achieve dynamic adaptive adjustment of the energy in the speech frequency band. By fusing the personalized auditory weighting curve and the auditory masking threshold, it is possible to effectively enhance the energy in the key information frequency band while naturally suppressing the noise in the invalid frequency band, significantly improving the auditory quality of the speech features and the accuracy of the speech recognition system.

[0082] S60. Perform frequency-domain energy adjustment on the noise-suppressed Mel components based on the perceptual weighting coefficient to generate a Mel spectrogram representation.

[0083] In this embodiment, after generating the noise-suppressed Mel components, the next processing step involves performing frequency-domain energy adjustment on them based on the perceptual weighting coefficient to optimize the Mel spectrogram representation. The goal of the frequency-domain energy adjustment is to further improve the quality of the signal by adjusting the energy in the frequency band after noise suppression, so that it can better adapt to subsequent speech recognition or synthesis tasks while meeting the requirements of the auditory perception model.

[0084] First, by loading and applying the pre-calculated perceptual weighting coefficient, the energy of each Mel band is weighted and adjusted. The role of the perceptual weighting coefficient is to dynamically adjust the energy of each frequency band according to the sensitivity of the human ear to different frequencies and the masking effect. This process can not only enhance the speech features in the target frequency band but also suppress the noise frequency bands irrelevant to the speech.

[0085] During the implementation process, first analyze the generated noise-suppressed Mel components to obtain the energy distribution of each Mel band. Then, weight the energy values of these frequency bands one by one with the perceptual weighting coefficient. The selection of the weighting coefficient is based on the aforementioned personalized auditory weighting curve and the auditory masking threshold. In the low-frequency and high-frequency bands, the perceptual weighting coefficient may be different. Generally, the low-frequency band will receive less weighting, while the high-frequency band may be given a higher weight.

[0086] The adjusted frequency-domain energy will be used to generate a Mel spectrogram representation. As one of the spectral features of the speech signal, the Mel spectrogram representation intuitively shows the energy distribution of the signal at different frequencies. This process is crucial for the speech recognition system because a good Mel spectrogram can effectively highlight the key information in the speech signal while suppressing the irrelevant noise. Especially in a noisy environment, the role of the perceptual weighting coefficient is particularly important, which can enable the speech recognition system to extract effective speech features from a complex environment and improve the recognition accuracy.

[0087] In the field of healthcare, speech recognition technology is often used for the speech analysis of patients' condition feedback. After performing frequency-domain energy adjustment on the noise-suppressed Mel components through the perceptual weighting coefficient, the characteristics of different patients' speech signals can be optimized, especially in the presence of background noise. This adjustment helps improve the quality of remote communication between doctors and patients. For patients with relatively weak hearing, it can reduce the impact of the auditory masking effect on the signal and ensure the accuracy of information transmission in medical records.

[0088] In the financial field, speech recognition is used in remote identity authentication. Through the dynamic adjustment of the perceptual weighting coefficient, the speech recognition accuracy of financial users in a noisy environment is effectively improved. For example, during the process of opening a remote bank account, background noise (such as traffic, mall noise, etc.) often affects the accuracy of speech verification. By performing frequency-domain energy adjustment on the noise-suppressed Mel components, the speech characteristics of users become more accurate in the financial transaction environment, reducing the possibility of identity verification failure and ensuring the security of financial transactions.

[0089] In this embodiment, by performing frequency-domain energy adjustment on the noise-suppressed Mel components based on the perceptual weighting coefficient, the performance of speech signals in a complex noise environment can be effectively improved. This adjustment process optimizes the energy distribution of the Mel frequency bands, strengthens the key information of the speech, and suppresses the background noise, thereby improving the accuracy of speech recognition.

[0090] The present invention relates to the technical field of speech processing and can be applied to business scenarios such as fintech and healthcare. It discloses a speech feature processing method, device, equipment, and medium, including: obtaining an original audio signal, performing adaptive frequency resolution adjustment on the original audio signal to generate fused Mel band energy, performing time resolution analysis based on the fused Mel band energy to generate multi-scale Mel spectral amplitude values, performing non-linear transformation on the multi-scale Mel spectral amplitude values according to the noise intensity parameter to generate noise-suppressed Mel components, generating a perceptual weighting coefficient according to the auditory perception model, and performing frequency-domain energy adjustment on the noise-suppressed Mel components based on the perceptual weighting coefficient to generate a Mel spectrogram representation. Based on adaptive frequency resolution, dynamic time resolution adjustment, and auditory perception modeling, the present invention applies noise adaptive non-linear transformation and perceptual weighting processing to the multi-scale Mel spectral amplitude values, which can effectively reduce the impact of noise interference on speech features, enhance the ability to retain key information of speech signals, improve the accuracy of speech recognition and the quality of speech synthesis at the same time, and achieve the collaborative optimization of noise robustness and auditory perception adaptability.

[0091] In one embodiment, the above step S20 includes:

[0092] S201, dividing the original audio signal into a low-frequency segment signal and a high-frequency segment signal;

[0093] S202. Perform high-resolution Mel filtering on the low-frequency band signal to generate the first Mel band energy distribution.

[0094] S203. Perform low-resolution Mel filtering on the high-frequency band signal to generate the second Mel band energy distribution.

[0095] S204. Merge the band energies of the first Mel band energy distribution and the second Mel band energy distribution to generate the fused Mel band energy.

[0096] In this embodiment, the original audio signal is first divided into a low-frequency band signal and a high-frequency band signal. The purpose of this step is to process the low-frequency and high-frequency parts of the audio signal with different frequency resolutions respectively. The signals in the low-frequency band and the high-frequency band have different spectral characteristics. Therefore, performing differential processing on them helps to more accurately capture the characteristics of the speech signal in different frequency bands.

[0097] The division of low frequency and high frequency can be carried out according to the frequency range of the signal and the processing requirements. One method is to divide based on a fixed frequency threshold. Set a fixed frequency range. For example, the part from 0 Hz to 1000 Hz or below 2000 Hz is regarded as low frequency, and the part above this frequency is classified as high frequency.

[0098] To better match the auditory perception characteristics of the human ear, the Mel scale is used to divide low frequency and high frequency. In the Mel scale, the division of low frequency and high frequency is achieved according to the non-linear perception characteristics of the human ear for frequency. The Mel scale is a way of frequency perception that mimics human hearing. The low-frequency part has a higher resolution, while the high-frequency part has a lower resolution. To better reflect the perception characteristics of the human ear, the original frequency value needs to be first converted to Mel frequency through the Mel scale. Specifically, the Mel frequency is calculated by the formula:

[0099]

[0100] Convert the frequency f (unit: Hz) to the unit on the Mel scale. This formula shows that the change in low frequency has a greater impact on the Mel frequency, while the change in high frequency has a relatively smaller impact on the Mel frequency.

[0101] In the Mel scale, the frequency bands in the low-frequency part are more dense, while those in the high-frequency part are more sparse. This means that the Mel scale will divide more and narrower frequency bands in the low-frequency range to finely capture the details of the speech signal, and in the high-frequency range, wider frequency bands are used to reduce the impact of high-frequency components on auditory perception. When implementing this process, a Mel filter bank can be used. The Mel filter bank filters the spectrum of the signal according to Mel frequencies, dividing the spectrum into multiple sub-bands, with each sub-band corresponding to a Mel frequency band. The design of the Mel filter bank is such that the low-frequency part contains more filters to adapt to the higher sensitivity of the human ear to low frequencies, while the number of filters in the high-frequency part is less. Finally, through the processing of the Mel filter bank, the spectrum of the signal is divided into multiple Mel frequency bands, which can effectively simulate the auditory characteristics of the human ear, especially the processing differences for low and high frequencies, thus better extracting the key features of the speech signal. For example, suppose there is an audio signal with a frequency range from 0 Hz to 8000 Hz. Then, under the Mel scale division:

[0102] Low-frequency band: On the Mel scale, the frequency range from 0 Hz to approximately 1000 Hz will be divided into more and narrower frequency bands. Each band represents the details of the signal, such as the vowel part in human speech. These bands have a high resolution for capturing details.

[0103] High-frequency band: Between 1000 Hz and 8000 Hz, the Mel scale will divide this range into fewer and wider frequency bands. These bands represent the high-frequency components of the signal, mainly related to consonants and high-frequency noise in speech. In the Mel scale, since the impact of high frequencies on the human ear is smaller, these bands will have a lower resolution.

[0104] Suppose 50 Mel filters are used to divide this 8000 Hz frequency band. The low-frequency part (0 Hz - 1000 Hz) may occupy 30 filters, while the remaining 20 filters will cover the high-frequency part (1000 Hz - 8000 Hz). In the low-frequency part, the bandwidth of each filter will be relatively narrow, perhaps dozens of Hertz (e.g., a frequency bandwidth from 20 Hz to 50 Hz), in order to capture the details of the speech signal. In the high-frequency part, the bandwidth of each filter will be relatively wide, perhaps 100 Hz or more, in order to reduce the excessive sensitivity to the high-frequency part.

[0105] Next, perform high-resolution Mel filtering on the low-frequency band signal. The low-frequency band signal usually contains relatively stable speech components. Therefore, using high-resolution Mel filters for processing in this part helps to more finely extract the speech features of the low-frequency band. The main purpose of this step is to improve the time-domain and frequency-domain resolution of the low-frequency band speech, making the low-frequency components in the speech signal clearer.

[0106] For high-frequency signals, low-resolution Mel filtering is adopted. The speech information contained in high-frequency signals is usually rather subtle, and the human ear has a lower sensitivity to high frequencies. Therefore, it can be processed with a lower frequency resolution. Using low-resolution Mel filters can effectively reduce the computational load in the high-frequency part and will not lose the feature information crucial for speech recognition. The purpose of this step is to simplify the spectral representation in the high-frequency band while maintaining its contribution to the speech signal.

[0107] After filtering the above two frequency bands respectively, the Mel band energy distributions of low frequency and high frequency need to be merged. Through the frequency band energy merging step, the high-resolution Mel band energy distribution in the low-frequency band is combined with the low-resolution Mel band energy distribution in the high-frequency band to generate a complete fused Mel band energy. This step combines the information of the two, retains the fine information in the low-frequency band, and compresses the redundant information in the high-frequency band, thereby optimizing the overall representation of the Mel band energy.

[0108] In the field of healthcare, audio signal processing technology can be used in intelligent health monitoring devices, such as intelligent hearing aids. Through adaptive frequency resolution adjustment, the processing of speech signals in different frequency ranges by the device can be optimized. For example, the device can provide a higher frequency resolution in the low-frequency band (such as the fundamental tone in a patient's speech) to improve speech clarity; at the same time, low-resolution filtering is used in the high-frequency band (such as environmental noise or high-frequency details) to reduce the unnecessary computational burden and retain speech features. This kind of processing can help the elderly or patients with hearing impairments to better hear the doctor's instructions and improve the accuracy of the speech recognition system.

[0109] In the financial field, it can be applied to the customer speech recognition system. By adaptively adjusting according to the frequency resolution, the system can better capture the key information in the speech while reducing the influence of background noise. For example, there may be low-frequency background noise (such as traffic noise) in the customer's phone speech. Through high-resolution filtering, the customer's voice can be captured more clearly, while the redundant noise in the high-frequency part is simplified by low-resolution filtering. This can reduce the speech recognition error rate and improve the accuracy of the customer identity verification system.

[0110] In this embodiment, by performing adaptive frequency resolution adjustment on the original audio signal, suitable filters can be used for processing in different frequency bands. High-resolution filtering is adopted in the low-frequency band to enhance the details of speech features, while low-resolution filtering is used in the high-frequency band to simplify the processing, which not only optimizes the computational efficiency but also preserves the key features of the speech signal. Through the frequency band energy merging step, the finally generated fused mel-frequency band energy represents the full-spectrum features of the audio signal, which can not only retain important low-frequency information but also suppress redundant high-frequency noise. This processing method significantly improves the performance of the speech recognition system in a noisy environment, especially in the fields of medical health and finance, and can effectively improve speech clarity and recognition accuracy.

[0111] In one embodiment, step S30 above includes:

[0112] S301, detecting the rapidly changing speech part and the stable speech part in the original audio signal;

[0113] S302, performing short-window time-frequency analysis on the fused mel-frequency band energy marked as the rapidly changing speech part to generate transient mel components;

[0114] S303, performing long-window time-frequency analysis on the fused mel-frequency band energy marked as the stable speech part to generate steady-state mel components;

[0115] S304, performing multi-scale time-frequency fusion on the transient mel components and the steady-state mel components to generate multi-scale mel spectral amplitude values.

[0116] In this embodiment, time resolution analysis is performed based on the fused mel-frequency band energy to generate multi-scale mel spectral amplitude values. This process analyzes the original audio signal, detects the rapidly changing speech part and the stable speech part, and processes them respectively using different window sizes (short window and long window) to extract the transient and steady-state features in the signal. Finally, through multi-scale time-frequency fusion, these features are combined to generate the final mel spectral amplitude values.

[0117] The original audio signal usually contains multiple frequency components, some of which change rapidly (such as consonants in speech), while others are relatively stable (such as vowels). The rapidly changing parts usually contain more high-frequency information, while the stable parts contain more low-frequency components. To distinguish these two parts, the time-domain characteristics of the audio signal, such as short-time energy, zero-crossing rate, etc., can be analyzed to determine the rapidly changing part and the stable part of the audio signal. The rapidly changing part is usually marked as the transient part, and the stable part is marked as the steady-state part. Usually, the rapidly changing part of speech occupies the consonant part of the audio signal, and the stable speech part is often the vowel part. In implementation, the short-time energy and zero-crossing rate of the audio signal are calculated. The short-time energy calculation is performed by dividing the signal into multiple frames and calculating the energy of each frame, while the zero-crossing rate calculation measures the number of times the signal crosses the zero axis in each frame. The part with large energy changes in the signal usually corresponds to the rapidly changing speech part, while the low-energy part corresponds to the stable speech part.

[0118] Once the rapidly changing speech part in the signal is marked, short-window time-frequency analysis is then used to process this part. Time-frequency analysis can help reveal the distribution characteristics of the signal in the time and frequency domains. Short-window time-frequency analysis can capture the instantaneous changes of the signal and is usually used to extract components that change rapidly in time, such as consonants in speech. A short window function (such as a 20-ms short window) is selected for Fourier transform or other time-frequency transforms (such as short-time Fourier transform, STFT). The frequency band width corresponding to each window function is relatively narrow, so as to accurately capture the changes in the high-frequency part. During the analysis process, the fused mel-band energy is used as the input, and the transient mel components are obtained through short-window time-frequency analysis.

[0119] For the region in the signal marked as the stable speech part, the long-window time-frequency analysis method is adopted. This is because the speech components in the stable part usually change slowly (such as vowels), and a larger time window is required to accurately capture their change characteristics. Different from short-window time-frequency analysis, long-window time-frequency analysis can help extract the long-term stability in the low-frequency region. To analyze the stable speech part, a longer window function (such as an 80-ms long window) can be selected. The longer window function can cover a longer time period of the signal and is suitable for capturing the components with lower frequencies and slower changes in the signal. By performing long-window time-frequency analysis on the stable speech part, the steady-state mel components can be obtained.

[0120] Finally, the transient Mel components and the steady-state Mel components are fused to generate multi-scale Mel spectral amplitude values. The purpose of multi-scale fusion is to integrate the features extracted at different frequency and time resolutions to form a richer and more comprehensive feature representation, which is suitable for subsequent speech processing tasks. By performing weighted fusion on the transient and steady-state Mel components, different weights can be assigned according to the importance of different components. The result of weighted fusion will produce a Mel spectral amplitude value containing multi-scale information. Such multi-scale Mel spectral amplitude values can better reflect different features of the signal, especially the comprehensive characteristics of the transient and steady-state parts in speech.

[0121] Through the above steps in this embodiment, the transient and steady-state features in the signal can be effectively separated and captured. Especially in a noisy environment, clearer speech features can be extracted, improving the performance of speech recognition and synthesis. This multi-scale time-frequency fusion method not only improves the time resolution of the speech signal but also enhances the frequency resolution of the speech signal, improving the accuracy and robustness of the speech processing system.

[0122] In one embodiment, the above step S40 includes:

[0123] S401, detecting the silent segments in the original audio signal;

[0124] S402, performing a fast Fourier transform on the silent segments to generate a silent segment amplitude spectrum;

[0125] S403, extracting the energy of the silent segment amplitude spectrum as the background noise energy;

[0126] S404, analyzing the ratio of the average energy of the speech segments in the original audio signal to the background noise energy to generate a signal-to-noise ratio parameter;

[0127] S405, selecting the corresponding first power exponent and second power exponent from a preset mapping table according to the signal-to-noise ratio parameter;

[0128] S406, dividing the multi-scale Mel spectral amplitude values into high-frequency components and low-frequency components according to a preset frequency threshold;

[0129] S407, performing a first non-linear transformation on each Mel band amplitude value of the high-frequency components based on the first power exponent;

[0130] S408, performing a second non-linear transformation on each Mel band amplitude value of the low-frequency components based on the second power exponent;

[0131] S409, merging the high-frequency components after the first non-linear transformation and the low-frequency components after the second non-linear transformation in the order of Mel bands to generate a noise-suppressed Mel component.

[0132] In this embodiment, a non - linear transformation is performed on the multi - scale Mel spectral amplitude values based on the noise intensity parameter to generate noise - suppressed Mel components. This process includes multiple steps, involving operations such as the detection of silent segments, the analysis of background noise, the calculation of the signal - to - noise ratio (SNR), the division of frequency bands, and the non - linear transformation based on the noise intensity.

[0133] Silent segments refer to the parts without speech activity, which are usually background noise or environmental noise. Detecting silent segments is the first step in noise modeling. By calculating features such as the short - time energy and zero - crossing rate of the audio signal, silent segments and speech - containing segments can be distinguished. The detection of silent segments helps the system separate noise and speech during processing, laying the foundation for subsequent noise - suppression operations. It can be judged by setting a threshold to determine whether the short - time energy is lower than a certain set value, or by calculating the zero - crossing rate (i.e., the frequency at which the signal switches from positive to negative). For the part with energy lower than a certain energy threshold and a low zero - crossing rate, it can be determined as a silent segment. Usually, a fixed time window (such as 10 ms) is set for sliding analysis to detect the energy and zero - crossing rate of each window.

[0134] The fast Fourier transform (FFT) is a method for converting a time - domain signal into a frequency - domain signal. Although silent segments do not contain speech components, they still contain background noise information. By performing the FFT transformation on silent segments, the frequency components of the background noise can be obtained, which is crucial for subsequent noise - suppression processing. Perform the fast Fourier transform on the audio signal of the silent segment to calculate the spectrum of this segment. The spectrum shows the amplitude and phase information of different frequency components. The generated amplitude spectrum of the silent segment can help further extract the noise energy for modeling as the background noise energy.

[0135] Extracting energy information from the amplitude spectrum of the silent segment is a key step in noise modeling. The energy of the background noise is estimated based on the intensity of the frequency components of the silent segment, which can provide a basis for the subsequent calculation of the signal - to - noise ratio (SNR). Obtain the background noise energy by calculating the total energy of the amplitude spectrum of the silent segment. This is usually obtained by squaring and summing the amplitude values of each frequency band in the spectrum. The estimation of the background noise energy provides a reference for the noise - suppression algorithm in subsequent steps.

[0136] The Signal-to-Noise Ratio (SNR) is an important parameter that measures the relationship between the strength of a voice signal and background noise. The calculation of SNR helps to determine the clarity and distinguishability of different parts of the voice signal. By comparing the average energy of the voice segment and the background noise energy, SNR can provide a basis for subsequent non-linear transformations, guiding which parts need to be suppressed and which parts need to be retained. First, calculate the average energy of the voice segment in the original voice signal. This can be obtained by averaging the short-time energy of each audio frame. Then, compare the calculated voice segment energy with the background noise energy to obtain the Signal-to-Noise Ratio (SNR). Generally, parts with higher SNR values indicate stronger voice signals and weaker background noise, while parts with lower SNR values reflect stronger noise.

[0137] According to the Signal-to-Noise Ratio (SNR) parameter, select the corresponding power exponent. The selection of the power exponent directly affects the intensity and method of noise suppression. The preset mapping table provides the correspondence between SNR and the power exponent. Generally, higher SNR values correspond to smaller power exponents, meaning less noise suppression, while lower SNR values correspond to larger power exponents, meaning stronger noise suppression. By looking up the preset mapping table, select the corresponding power exponent according to the calculated SNR value. This mapping table is preset based on experience and experimental data and can be adjusted according to different environments or noise types. The mapping table usually sets different power exponents for different SNR ranges. For example, when SNR is greater than 20 dB, use a power exponent of 2; when SNR is less than 10 dB, use a power exponent of 3.

[0138] According to the frequency characteristics of the Mel spectrum, divide the multi-scale Mel spectrum amplitude values into high-frequency components and low-frequency components according to a preset frequency threshold. The low-frequency components usually contain steady-state components such as vowels in speech, while the high-frequency components contain transient components such as consonants. The selection of the frequency threshold usually depends on the characteristics of the signal and the nature of the noise. According to the preset frequency threshold (e.g., 1000 Hz), divide the Mel band energy into low-frequency and high-frequency parts. The low-frequency components contain the parts with frequencies lower than the threshold, while the high-frequency components contain the parts with frequencies higher than the threshold. This division method helps to perform specific suppression or enhancement on noise in different frequency bands.

[0139] The non-linear transformation of the high-frequency part usually uses a larger power exponent to perform stronger suppression on high-frequency noise. Through this transformation, the interference of high-frequency noise on the voice signal can be effectively reduced. Perform non-linear transformation on each Mel band amplitude value of the high-frequency component. The specific transformation method is to raise the amplitude value to the power of the power exponent. For example, when the power exponent is 3, the amplitude value will be cubed, which can enhance the noise suppression effect under lower SNR conditions.

[0140] For the non - linear transformation of the low - frequency part, a smaller power exponent is used, which can ensure the minimum impact on the naturalness and clarity of speech. Through this transformation, low - frequency noise will be appropriately suppressed while retaining the important components in the speech. Similar to the high - frequency part, a non - linear transformation is performed on the magnitude value of each Mel - frequency band of the low - frequency component, and the power exponent is usually set to 2. This can moderately suppress the noise in the low - frequency part without having too much impact on the speech signal.

[0141] Finally, the high - frequency and low - frequency components after non - linear transformation are merged in the order of Mel - frequency bands to generate the final noise - suppressed Mel components, which is to ensure the frequency consistency of the signal and use the noise - suppressed Mel - frequency band energy as the output. The transformed high - frequency component and low - frequency component are spliced or weighted and superimposed in the order of Mel - frequency bands to form the final noise - suppressed Mel components. The merged Mel components will be used as the input for subsequent processing or speech synthesis.

[0142] In this embodiment, by dynamically adjusting the time and frequency resolution, the noise in different frequency bands is appropriately suppressed, which not only avoids the loss of high - frequency information caused by over - suppression but also can effectively reduce the interference of low - frequency noise, thereby improving the robustness and performance of the speech processing system.

[0143] In one embodiment, the above step S50 includes:

[0144] S501, load a preset auditory perception model, and the auditory perception model contains standard equal - loudness curve data;

[0145] S502, generate an initial frequency sensitivity weight distribution according to the standard equal - loudness curve data;

[0146] S503, obtain the hearing threshold test data of the user, and adjust the initial frequency sensitivity weight distribution based on the hearing threshold test data to generate a personalized auditory weighting curve;

[0147] S504, analyze the Mel - frequency band energy distribution of the noise - suppressed Mel components, and determine the auditory masking threshold of each Mel - frequency band through the auditory perception model based on the Mel - frequency band energy distribution;

[0148] S505, perform weighted fusion of the personalized auditory weighting curve and the auditory masking threshold to generate a perceptual weighting coefficient.

[0149] In this embodiment, based on the auditory perception model, considering the human ear's perception characteristics of sound at different frequencies and time resolutions, the noise - suppressed Mel components are optimized.

[0150] The auditory perception model is a mathematical modeling of the physical perception characteristics of the human auditory system. This model is usually based on standard equal-loudness contour data (such as ISO 226 equal-loudness contours), which describes the sensitivity of the human ear to different frequencies. Equal-loudness contours are used to determine the perceived intensity of sound by the human ear at different frequencies, especially the perceptual differences in the low and high frequencies. When loading this model, mainly the data of the equal-loudness contours are stored and used as the basis for subsequent calculations. Predefined standard equal-loudness contour data (such as the frequency response data in the ISO 226 standard) can be used during loading, and these data are stored as the basic parameters of the model in the system. This data usually includes the variation of the auditory sensitivity of the human ear in the range from 20 Hz to 20 kHz.

[0151] Based on the equal-loudness contour data, an initial frequency sensitivity weight distribution is calculated and generated. The frequency sensitivity weight distribution represents the sensitivity of the human ear to sound at different frequencies. The weight differences between the low-frequency band and the high-frequency band are relatively large, while the weights in the middle-frequency band are usually more balanced. This weight distribution will be used to weight-adjust the mel-frequency band energy of the audio signal. The sensitivity weights of each frequency band are calculated using the equal-loudness contour data. By mapping the frequency-perceived sensitivity relationship of the equal-loudness contour to the mel-frequency bands, the corresponding weights can be calculated for each mel-frequency band. This weight distribution is usually a relative value, representing the perceived intensity of the human ear for each frequency band.

[0152] The user's hearing threshold test data includes the individual's hearing sensitivity at different frequencies. These data are used to adjust the initial frequency sensitivity weight distribution to better match the user's personalized auditory characteristics. Through the adjusted weight distribution, the audio signal can be optimized according to the hearing characteristics of each user. The user's hearing threshold test data are usually obtained through professional hearing test equipment, and the test results include the user's hearing thresholds at each frequency. Based on these data, the initial frequency sensitivity weights are adjusted. The specific method can be to calculate the difference between the user's hearing threshold and the standard equal-loudness contour, and adjust the weight distribution of the mel-frequency bands based on this difference. This step helps to optimize the perceived effect of the audio signal in a personalized way.

[0153] The noise-suppressed Mel components are generated in the previous step and have a certain noise suppression effect. In this step, the Mel-band energy distribution of these components is analyzed, and the masking threshold of each band is calculated using an auditory perception model. The masking effect refers to the phenomenon that a strong signal can mask a weaker signal, and the masking threshold is used to describe whether a signal can be masked by another signal at a specific frequency. By analyzing the Mel-band energy of the noise-suppressed Mel components, the energy value of each band is determined. Then, according to the masking effect parameters in the auditory perception model, the auditory masking threshold of each band is calculated. The calculation of the masking threshold usually depends on the relative intensity of the Mel-band energy and the sensitivity of the human ear to that frequency band. In this step, stronger band energy raises the masking threshold, while weaker band energy makes the masking threshold lower.

[0154] Finally, the user's personalized auditory weighting curve is weighted and fused with the calculated auditory masking threshold to generate a perceptual weighting coefficient. The perceptual weighting coefficient is a parameter used to further adjust the audio signal, which can optimize the energy distribution of the signal in the frequency domain, enhance the perceptual quality of the speech signal, and suppress unnecessary noise at the same time. The personalized auditory weighting curve is corresponded and weighted with the masking threshold of each Mel band one by one. Specifically, the personalized curve adjusts the perceptual weight of each band according to the user's hearing characteristics, while the masking threshold adjusts the perceivable threshold of the band according to the energy of the Mel band and the masking effect. The weighted fusion process combines the effects of the two to obtain the final perceptual weighting coefficient. This coefficient will be used for subsequent frequency-domain energy adjustment of the noise-suppressed Mel components.

[0155] Through the above steps, this embodiment can improve the perceptual quality of the speech signal on a personalized basis, reduce the interference of noise, thereby enhancing the clarity and recognition rate of the speech signal. It is applicable to scenarios with strong environmental noise, can dynamically adjust the perceptual effect of the audio signal according to different frequencies and hearing characteristics, and significantly improve the robustness and performance of the speech processing system.

[0156] In one embodiment, after the above step S60, it further includes:

[0157] S701, analyzing the current energy dynamic range of each Mel band in the Mel spectrogram representation;

[0158] S702, obtaining a preset target energy dynamic range;

[0159] S703, determining a compression ratio according to the ratio of the target energy dynamic range to the current energy dynamic range;

[0160] S704, determining a compression threshold based on the maximum and minimum values of the current energy dynamic range;

[0161] S705, for the Mel frequency bands in the Mel spectrogram representation whose energy exceeds the compression threshold, perform logarithmic compression based on the compression ratio;

[0162] S706, perform linear enhancement on the Mel frequency bands in the Mel spectrogram representation whose energy is lower than the compression threshold;

[0163] S707, output the Mel spectrogram representation after energy range adaptation processing.

[0164] In this embodiment, the energy of the Mel spectrogram representation is adjusted to conform to the target dynamic range, and the perceptual quality of the audio signal is optimized by performing compression and enhancement operations on the energy.

[0165] First, it is necessary to analyze the energy of each Mel frequency band in the Mel spectrogram representation. The Mel spectrogram is a spectrogram that compresses the frequency components of an audio signal through the Mel scale and reflects the energy distribution of different frequency bands. The current energy dynamic range represents the range of energy values in the Mel frequency band, that is, the difference between the maximum energy and the minimum energy in the frequency band. By processing the Mel spectrogram, the energy values of each Mel frequency band are extracted, and the maximum and minimum values of the energy are calculated to obtain the current energy dynamic range. This range reflects the energy fluctuation situation of the current Mel frequency band and provides a reference for subsequent adjustment.

[0166] The target energy dynamic range is a preset target range according to the requirements of the audio signal processing system. Usually, in an audio processing system, an ideal energy range is set according to the system requirements to make the intensity and clarity of the audio signal meet the requirements. The target energy dynamic range may be set based on the input requirements of the vocoder, the standardized audio output requirements, or the specific requirements of the application scenario. The target dynamic range can be set through the parameter configuration of the system, usually a fixed range, or dynamically adjusted based on the processing requirements of the audio signal. This value can be adjusted according to different application scenarios (such as speech recognition, speech synthesis, etc.) to ensure the best audio performance.

[0167] The compression ratio is determined by comparing the current energy dynamic range and the target energy dynamic range. This ratio is used to adjust the energy of the current Mel frequency band so that it can adapt to the preset target dynamic range. If the current dynamic range is large and the target dynamic range is small, the compression ratio will be high, and vice versa. The calculation method of the compression ratio is to calculate the ratio of the target energy dynamic range to the current energy dynamic range. This ratio determines the intensity of compression or enhancement of the Mel frequency band energy. The calculation formula of the compression ratio is: compression ratio = target dynamic range / current dynamic range.

[0168] The compression threshold is determined based on the maximum and minimum energy values of the current Mel band. It is used to distinguish between bands with higher energy and those with lower energy, and to decide which bands need to be compressed and which can be enhanced. The compression threshold is set by the difference between the maximum and minimum values of the current energy dynamic range. Usually, the threshold is set to a specific point in the current energy range (such as the maximum value, the minimum value, or a median value in between). This value is used to distinguish which Mel bands have higher energy and need to be compressed, and which bands have lower energy and need to be enhanced.

[0169] For Mel bands with energy exceeding the compression threshold, a logarithmic compression operation is performed. Logarithmic compression is a common signal processing technique used to reduce the dynamic range of stronger signals, thereby avoiding distortion and making the energy distribution of the signal more balanced. For Mel bands with energy exceeding the compression threshold, a logarithmic compression function is applied to adjust the energy. The compression process is achieved through the following formula:

[0170] E new = log(E current + 1) × R

[0171] where E new is the compressed energy value, E current is the current energy value, and R is the calculated compression ratio. It can effectively reduce the energy of stronger bands while avoiding excessive weakening of the signal.

[0172] For Mel bands with energy below the compression threshold, a linear enhancement operation is performed. This enhancement operation is used to increase the energy of weaker signals, making low-energy bands more prominent and improving the perceptibility of the speech signal. For Mel bands with energy below the compression threshold, a linear gain function is used to enhance their energy. The enhancement process is achieved through the following formula:

[0173] E new = E current × Gain

[0174] where E new is the enhanced energy value, E current is the current energy value, and Gain is a preset gain value (such as 1.5 times). This gain value can be adjusted according to the system's requirements and is usually a fixed value or dynamically adjusted according to the actual needs of the audio signal.

[0175] Through the compression and enhancement processes of the foregoing steps, a Mel spectrogram representation adapted to the target energy dynamic range is generated. This step ensures that the energy distribution of the audio signal conforms to the system requirements and can perform better in subsequent processing stages. The output Mel spectrogram representation undergoes compression and enhancement processes to ensure the equalization of the energy in the frequency bands. Finally, the obtained Mel spectrogram can be directly input into subsequent audio processing modules (such as a vocoder model, a speech recognition module, etc.) for further speech generation or recognition.

[0176] In this embodiment, by dynamically adjusting the energy in the frequency bands, both signal distortion is avoided and the enhancement of low-energy frequency bands is ensured, thereby improving the clarity and recognizability of speech. Especially in a multi-noise environment, the accuracy and robustness of the speech recognition system can be significantly improved.

[0177] In one embodiment, after the above step S60, the following steps are further included:

[0178] S801, performing frame sequence segmentation on the Mel spectrogram representation to generate a Mel spectrogram frame sequence;

[0179] S802, performing normalization processing on the Mel spectrogram frame sequence to generate a normalized Mel spectrogram frame;

[0180] S803, inputting the normalized Mel spectrogram frame into a pre-trained vocoder model;

[0181] S804, extracting the time-frequency correlation features of the normalized Mel spectrogram frame through the convolutional neural network module of the vocoder model;

[0182] S805, converting the time-frequency correlation features into a time-domain speech waveform through the waveform generation module of the vocoder model;

[0183] S806, performing overlap-and-add processing on the time-domain speech waveform to generate a continuous target speech signal.

[0184] In this embodiment, the Mel spectrogram representation will be sliced into several small frames for subsequent feature extraction and processing. Frame sequence segmentation is a common technique in speech signal processing and usually divides the signal based on a fixed-length time window. Each frame contains the frequency-domain information of a certain duration and can help the system capture the local features of the signal. The generation of the Mel spectrogram frame sequence is completed by setting a fixed frame length and overlap rate. Usually, the frame length is between 20 ms and 50 ms, and the overlap rate is 50%. For example, if the frame length is set to 25 ms and the overlap rate is 50%, then each frame of data contains half of the data of the previous frame. This method can ensure a high similarity between each frame, thereby ensuring the information correlation between consecutive frames.

[0185] Normalization is to ensure that different Mel bands have similar amplitude ranges to avoid adverse effects on subsequent processing caused by overly large or small values in some bands. The goal of normalization is to adjust the values of each Mel band to a unified range so that the system can better process this data. Mean-variance normalization is usually adopted for the normalization of Mel spectrogram frames. The Mel band values of each frame are subtracted by the mean of that band in the training set and then divided by its standard deviation.

[0186] The normalized Mel spectrogram frames are input into a pre-trained vocoder model. The role of the vocoder model is to convert the Mel spectrogram into a speech waveform. A vocoder is usually a deep learning model that can generate speech signals based on the input Mel spectrogram frames, mimicking the human pronunciation mechanism. Vocoders usually adopt neural network architectures, especially models such as WaveNet and Vocoder. Taking the standardized Mel spectrogram frames as input, the vocoder model will convert them into time-domain speech waveforms. During this process, the vocoder extracts time-frequency correlation features through a Convolutional Neural Network (CNN) module and generates speech waveforms through a waveform generation module.

[0187] The Convolutional Neural Network (CNN) module is an important part of a deep learning model for extracting local features. In this step, the CNN module is used to extract time-frequency domain correlation features from the input Mel spectrogram frames, helping the model understand the variation relationships of each band in the time domain and frequency domain, so as to better reconstruct the speech waveform. The CNN module processes the Mel spectrogram frames through a series of convolutional operations to extract the local time-frequency features. These features reflect the variations of the signal at different times and frequencies and are the basis for generating speech waveforms. Usually, the CNN module will include multiple convolutional layers, pooling layers, and activation functions to extract higher-level features layer by layer.

[0188] The main task of the waveform generation module is to convert the time-frequency correlation features extracted by the CNN into a playable time-domain speech signal. This process is completed through transposed convolution (or called deconvolution neural network), which maps the features back to the time domain to restore the original speech signal. The transposed convolution operation will utilize the learned feature weights to convert the time-frequency domain features into a time-domain waveform. Through these operations, the model can restore the frequency information contained in the Mel spectrogram frames into an audible audio signal. In the vocoder model, generative models such as WaveNet are widely used in this process, which generate speech waveforms through autoregression.

[0189] Overlap-and-add processing is used to reassemble the segmented frame sequence into a continuous speech signal. The purpose of this step is to eliminate the seam noise that may occur during the frame segmentation and reconstruction processes, ensuring the smoothness of the speech signal. In overlap-and-add processing, the overlapping parts between frames are weighted and averaged, and usually a window function (such as Hamming window or Hanning window) is used to smooth the edges of each frame to avoid sudden changes at the splicing points. Specifically, the overlapping part of each frame is added to the adjacent frame, and then a continuous and smooth time-domain waveform is output to generate the final target speech signal.

[0190] Example illustration: In the field of healthcare, there is usually a large amount of background noise in the hospital environment, such as the conversations of other patients, the noise of medical equipment, etc. These noises often interfere with the speech signal, reducing the accuracy and reliability of the speech recognition system. To solve this problem, for the noise interference problem in the medical environment, first, after obtaining the original audio signal, the system performs adaptive frequency resolution adjustment on the signal. By dividing the signal into low-frequency and high-frequency bands and adopting different Mel filtering processing methods (high-resolution Mel filtering for low frequencies and low-resolution Mel filtering for high frequencies), the system can generate fused Mel band energies. This process helps to better retain the key features of the speech signal and compress the redundant information in the high-frequency band, thereby reducing the impact of noise. This processing method is particularly suitable for extracting clear and accurate speech features in the medical environment, avoiding the influence of unnecessary information in the speech signal on subsequent processing. Next, the system performs time resolution analysis based on the fused Mel band energies to generate multi-scale Mel spectrum amplitude values. The system detects the rapidly changing speech part and the stable speech part in the original audio signal, and respectively uses short-window and long-window time-frequency analysis methods to process different speech parts, thereby generating transient Mel components and steady-state Mel components. By performing multi-scale time-frequency fusion on these two parts of the signal, the multi-scale Mel spectrum amplitude values generated by the system can effectively extract the key time-frequency features of the speech signal and remove the interference of background noise. For healthcare applications, this processing method can ensure that the speech recognition system can accurately recognize the speech commands of patients in the presence of high background noise, providing more efficient and accurate services. When performing noise suppression, the system performs a non-linear transformation on the multi-scale Mel spectrum amplitude values according to the noise intensity parameter to generate noise suppression Mel components. First, the system detects the silent segments in the original audio signal, generates the amplitude spectrum of the silent segments through fast Fourier transform, and extracts the energy of the amplitude spectrum of the silent segments as the background noise energy. Then, by analyzing the ratio of the average energy of the speech segments in the original audio signal to the background noise energy, a signal-to-noise ratio parameter is generated, and an appropriate power exponent is selected from a preset mapping table according to this parameter. Based on this information, the system performs non-linear transformations on the high-frequency and low-frequency components respectively, using the first power exponent for the high-frequency components and the second power exponent for the low-frequency components, thereby performing noise suppression. Through this noise suppression process, the noise components in the speech signal are effectively removed, improving the quality and clarity of the speech signal, and providing a cleaner signal source for subsequent speech recognition processing. Then, the system generates perceptual weighting coefficients according to the auditory perception model. By loading the preset auditory perception model, the system first generates an initial frequency sensitivity weight distribution from the standard equal-loudness curve data and adjusts this weight distribution according to the user's hearing threshold test data to generate a personalized auditory weighting curve. Through this personalized adjustment, the system can better adapt to the auditory characteristics of the user and improve the recognition accuracy of the speech signal.Then, the system analyzes the Mel-band energy distribution of the noise-suppressed Mel components, calculates the auditory masking threshold of each Mel band according to the auditory perception model, and finally performs weighted fusion of the personalized auditory weighting curve and the masking threshold to generate the perceptual weighting coefficient. The purpose of this process is to optimize the frequency-domain energy in the speech signal to make it more conform to the perceptual characteristics of human hearing and further improve the accuracy of the speech recognition system. Finally, based on the generated perceptual weighting coefficient, the system adjusts the frequency-domain energy of the noise-suppressed Mel components to generate a Mel spectrogram representation. At this time, the system first analyzes the current energy dynamic range of each Mel band in the Mel spectrogram representation and obtains the preset target energy dynamic range. By calculating the ratio of the target energy dynamic range to the current energy dynamic range, the system determines the compression ratio and determines the compression threshold based on the maximum and minimum values of the current energy dynamic range. For the bands in the Mel spectrogram with energy exceeding the compression threshold, the system performs logarithmic compression based on the compression ratio; for the bands with energy below the compression threshold, the system performs linear enhancement. Through this energy range adaptation process, the system can optimize the dynamic range of the speech signal and ensure that the speech signal can still maintain good recognizability in different noise environments.

[0191] In this embodiment, through frequency-domain energy adjustment, vocoder generation, and overlap-and-add processing, the time-domain performance of the speech signal can be optimized, enabling the speech recognition system to still maintain a high accuracy in complex environments, enhancing the robustness of the system. Especially in the fields of medical health and finance, more reliable speech recognition services can be provided.

[0192] In one embodiment, a speech feature processing device is provided, which corresponds one-to-one to the speech feature processing method in the above embodiment. Refer to Figure 3 , Figure 3 which is a schematic diagram of the functional modules of a preferred embodiment of the speech feature processing device of the present invention. An audio acquisition module 10, a frequency resolution adjustment module 20, a time resolution analysis module 30, a noise suppression module 40, a perceptual weighting generation module 50, and a frequency-domain energy adjustment module 60. The detailed description of each functional module is as follows:

[0193] The audio acquisition module 10 is used to acquire the original audio signal;

[0194] The frequency resolution adjustment module 20 is used to perform adaptive frequency resolution adjustment on the original audio signal to generate fused Mel-band energy;

[0195] The time resolution analysis module 30 is used to perform time resolution analysis based on the fused Mel-band energy to generate multi-scale Mel spectrogram amplitude values;

[0196] A noise suppression module 40, configured to perform a non - linear transformation on the multi - scale Mel spectrum amplitude values according to a noise intensity parameter to generate noise - suppressed Mel components;

[0197] A perceptual weighting generation module 50, configured to generate perceptual weighting coefficients according to an auditory perception model;

[0198] A frequency - domain energy adjustment module 60, configured to perform frequency - domain energy adjustment on the noise - suppressed Mel components based on the perceptual weighting coefficients to generate a Mel spectrogram representation.

[0199] In one embodiment, the frequency resolution adjustment module 20 is specifically configured to:

[0200] Divide the original audio signal into a low - frequency band signal and a high - frequency band signal;

[0201] Perform high - resolution Mel filtering processing on the low - frequency band signal to generate a first Mel band energy distribution;

[0202] Perform low - resolution Mel filtering processing on the high - frequency band signal to generate a second Mel band energy distribution;

[0203] Perform band - energy merging on the first Mel band energy distribution and the second Mel band energy distribution to generate a fused Mel band energy.

[0204] In one embodiment, the time resolution analysis module 30 is specifically configured to:

[0205] Detect fast - changing speech parts and stable speech parts in the original audio signal;

[0206] Perform short - window time - frequency analysis on the fused Mel band energy marked as fast - changing speech parts to generate transient Mel components;

[0207] Perform long - window time - frequency analysis on the fused Mel band energy marked as stable speech parts to generate steady - state Mel components;

[0208] Perform multi - scale time - frequency fusion on the transient Mel components and the steady - state Mel components to generate multi - scale Mel spectrum amplitude values.

[0209] In one embodiment, the noise suppression module 40 is specifically configured to:

[0210] Detect silent segments in the original audio signal;

[0211] Perform a fast Fourier transform on the silent segments to generate a silent - segment amplitude spectrum;

[0212] Extract the energy of the silent - segment amplitude spectrum as background noise energy;

[0213] Analyze the ratio of the average energy of the speech segments in the original audio signal to the background noise energy to generate a signal-to-noise ratio parameter;

[0214] Select corresponding first and second power exponents from a preset mapping table according to the signal-to-noise ratio parameter;

[0215] Divide the multi-scale Mel spectrum magnitude values into high-frequency components and low-frequency components according to a preset frequency threshold;

[0216] Perform a first non-linear transformation on the Mel band magnitude value of each high-frequency component based on the first power exponent;

[0217] Perform a second non-linear transformation on the Mel band magnitude value of each low-frequency component based on the second power exponent;

[0218] Merge the high-frequency components after the first non-linear transformation and the low-frequency components after the second non-linear transformation in the Mel band order to generate noise-suppressed Mel components.

[0219] In one embodiment, the perceptual weighting generation module 50 is specifically configured to:

[0220] Load a preset auditory perception model, and the auditory perception model includes standard equal-loudness curve data;

[0221] Generate an initial frequency sensitivity weight distribution according to the standard equal-loudness curve data;

[0222] Obtain the hearing threshold test data of the user, and adjust the initial frequency sensitivity weight distribution based on the hearing threshold test data to generate a personalized auditory weighting curve;

[0223] Analyze the Mel band energy distribution of the noise-suppressed Mel components, and determine the auditory masking threshold of each Mel band based on the Mel band energy distribution through the auditory perception model;

[0224] Perform weighted fusion on the personalized auditory weighting curve and the auditory masking threshold to generate a perceptual weighting coefficient.

[0225] In one embodiment, the frequency domain energy adjustment module 60 is specifically configured to:

[0226] Analyze the current energy dynamic range of each Mel band in the Mel spectrogram representation;

[0227] Obtain a preset target energy dynamic range;

[0228] Determine the compression ratio according to the ratio of the target energy dynamic range to the current energy dynamic range;

[0229] Determine a compression threshold based on the maximum and minimum values of the current energy dynamic range;

[0230] For the Mel frequency bands in the Mel spectrogram representation whose energy exceeds the compression threshold, perform logarithmic compression based on the compression ratio;

[0231] Perform linear enhancement on the Mel frequency bands in the Mel spectrogram representation whose energy is lower than the compression threshold;

[0232] Output the Mel spectrogram representation after energy range adaptation processing.

[0233] In one embodiment, the frequency-domain energy adjustment module 60 is specifically configured to:

[0234] Perform frame sequence segmentation on the Mel spectrogram representation to generate a Mel spectrogram frame sequence;

[0235] Perform normalization processing on the Mel spectrogram frame sequence to generate a normalized Mel spectrogram frame;

[0236] Input the normalized Mel spectrogram frame into a pre-trained vocoder model;

[0237] Extract the time-frequency correlation features of the normalized Mel spectrogram frame through the convolutional neural network module of the vocoder model;

[0238] Convert the time-frequency correlation features into a time-domain speech waveform through the waveform generation module of the vocoder model;

[0239] Perform overlap-and-add processing on the time-domain speech waveform to generate a continuous target speech signal.

[0240] In one embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 4 shown. The computer device includes a processor, a memory, a network interface, and a database connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes non-volatile and / or volatile storage media, and internal memory. The non-volatile storage media stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage media. The network interface of the computer device is used to communicate with an external client through a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the server side of a speech feature processing method.

[0241] In one embodiment, a computer device is provided. The computer device can be a client, and its internal structure diagram can be as Figure 5As shown. The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected by a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server through a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the user side of a voice feature processing method

[0242] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are realized:

[0243] Obtain an original audio signal;

[0244] Perform adaptive frequency resolution adjustment on the original audio signal to generate fused mel-band energy;

[0245] Perform time resolution analysis based on the fused mel-band energy to generate multi-scale mel-spectrum amplitude values;

[0246] Perform a non-linear transformation on the multi-scale mel-spectrum amplitude values according to a noise intensity parameter to generate noise-suppressed mel components;

[0247] Generate a perceptual weighting coefficient according to an auditory perception model;

[0248] Perform frequency-domain energy adjustment on the noise-suppressed mel components based on the perceptual weighting coefficient to generate a mel-spectrum representation.

[0249] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are realized:

[0250] Obtain an original audio signal;

[0251] Perform adaptive frequency resolution adjustment on the original audio signal to generate fused mel-band energy;

[0252] Perform time resolution analysis based on the fused mel-band energy to generate multi-scale mel-spectrum amplitude values;

[0253] Perform a non-linear transformation on the multi-scale mel-spectrum amplitude values according to a noise intensity parameter to generate noise-suppressed mel components;

[0254] Generate a perceptual weighting coefficient according to an auditory perception model;

[0255] Perform frequency-domain energy adjustment on the noise-suppressed Mel components based on the perception weighting coefficients to generate a Mel spectrogram representation.

[0256] It should be noted that for the functions or steps that can be achieved by the above computer-readable storage medium or computer device, reference can be made to the relevant descriptions on the server side and the client side in the foregoing method embodiments. To avoid repetition, they will not be described in detail here.

[0257] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above method embodiments. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0258] Those skilled in the art can clearly understand that for the convenience and simplicity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules as needed, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above.

[0259] It should be noted that if there are software tools or components of other companies in the embodiments of this application, they are only used for example introduction and do not represent actual use. The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present invention.

Claims

1. A method for processing voice features, characterized in that, Including the following steps: Obtain the original audio signal; Perform adaptive frequency resolution adjustment on the original audio signal to generate fused mel-band energy; Perform time resolution analysis based on the fused mel-band energy to generate multi-scale mel-spectrum amplitude values; Perform non-linear transformation on the multi-scale mel-spectrum amplitude values according to the noise intensity parameter to generate noise-suppressed mel components; Generate a perceptual weighting coefficient according to the auditory perception model; Perform frequency-domain energy adjustment on the noise-suppressed mel components based on the perceptual weighting coefficient to generate a mel-spectrum representation.

2. The voice feature processing method according to claim 1, wherein Performing adaptive frequency resolution adjustment on the original audio signal to generate fused mel-band energy includes: Divide the original audio signal into a low-frequency segment signal and a high-frequency segment signal; Perform high-resolution mel-filtering processing on the low-frequency segment signal to generate a first mel-band energy distribution; Perform low-resolution mel-filtering processing on the high-frequency segment signal to generate a second mel-band energy distribution; Perform band energy merging on the first mel-band energy distribution and the second mel-band energy distribution to generate fused mel-band energy.

3. The voice feature processing method according to claim 1, wherein Performing time resolution analysis based on the fused mel-band energy to generate multi-scale mel-spectrum amplitude values includes: Detect the fast-changing speech part and the stable speech part in the original audio signal; Perform short-window time-frequency analysis on the fused mel-band energy marked as the fast-changing speech part to generate transient mel components; Perform long-window time-frequency analysis on the fused mel-band energy marked as the stable speech part to generate steady-state mel components; Perform multi-scale time-frequency fusion on the transient mel components and the steady-state mel components to generate multi-scale mel-spectrum amplitude values.

4. The voice feature processing method according to claim 1, characterized in that, Performing non-linear transformation on the multi-scale mel-spectrum amplitude values according to the noise intensity parameter to generate noise-suppressed mel components includes: Detect the silent segment in the original audio signal; Perform fast Fourier transform on the silent segment to generate a silent segment amplitude spectrum; Extract the energy of the silent segment amplitude spectrum as the background noise energy; Analyze the ratio of the average energy of the speech segment in the original audio signal to the background noise energy to generate a signal-to-noise ratio parameter; Select corresponding first power exponent and second power exponent from a preset mapping table according to the signal-to-noise ratio parameter; Divide the multi-scale mel-spectrum amplitude values into high-frequency components and low-frequency components according to a preset frequency threshold; Perform a first non-linear transformation on each mel-band amplitude value of the high-frequency components based on the first power exponent; Perform a second non-linear transformation on each mel-band amplitude value of the low-frequency components based on the second power exponent; Merge the high-frequency components after the first non-linear transformation and the low-frequency components after the second non-linear transformation in the order of mel-bands to generate noise-suppressed mel components.

5. The voice feature processing method according to claim 1, wherein Generating a perceptual weighting coefficient according to the auditory perception model includes: Load a preset auditory perception model, and the auditory perception model contains standard equal-loudness curve data; Generate an initial frequency sensitivity weight distribution according to the standard equal-loudness curve data; Obtain the hearing threshold test data of the user, and adjust the initial frequency sensitivity weight distribution based on the hearing threshold test data to generate a personalized auditory weighting curve; Analyze the Mel-frequency band energy distribution of the noise-suppressed Mel components, and determine the auditory masking threshold of each Mel-frequency band based on the Mel-frequency band energy distribution through the auditory perception model; Perform weighted fusion of the personalized auditory weighting curve and the auditory masking threshold to generate a perceptual weighting coefficient.

6. The voice feature processing method according to claim 1, wherein After performing frequency-domain energy adjustment on the noise-suppressed Mel components based on the perceptual weighting coefficient to generate a Mel spectrogram representation, it further includes: Analyze the current energy dynamic range of each Mel-frequency band in the Mel spectrogram representation; Obtain a preset target energy dynamic range; Determine the compression ratio according to the ratio of the target energy dynamic range to the current energy dynamic range; Determine the compression threshold based on the maximum and minimum values of the current energy dynamic range; Perform logarithmic compression on the Mel-frequency bands in the Mel spectrogram representation whose energy exceeds the compression threshold based on the compression ratio; Perform linear enhancement on the Mel-frequency bands in the Mel spectrogram representation whose energy is lower than the compression threshold; Output the Mel spectrogram representation after energy range adaptation processing.

7. The voice feature processing method according to claim 1, wherein After performing frequency-domain energy adjustment on the noise-suppressed Mel components based on the perceptual weighting coefficient to generate a Mel spectrogram representation, it further includes: Perform frame sequence segmentation on the Mel spectrogram representation to generate a Mel spectrogram frame sequence; Perform normalization processing on the Mel spectrogram frame sequence to generate normalized Mel spectrogram frames; Input the normalized Mel spectrogram frames into a pre-trained vocoder model; Extract the time-frequency correlation features of the normalized Mel spectrogram frames through the convolutional neural network module of the vocoder model; Convert the time-frequency correlation features into a time-domain speech waveform through the waveform generation module of the vocoder model; Perform overlap-and-add processing on the time-domain speech waveform to generate a continuous target speech signal.

8. A voice feature processing device, characterized in that The speech feature processing device includes: An audio acquisition module for acquiring an original audio signal; A frequency resolution adjustment module for performing adaptive frequency resolution adjustment on the original audio signal to generate fused Mel-frequency band energy; A time resolution analysis module for performing time resolution analysis based on the fused Mel-frequency band energy to generate multi-scale Mel spectrogram amplitude values; A noise suppression module for performing non-linear transformation on the multi-scale Mel spectrogram amplitude values according to the noise intensity parameter to generate noise-suppressed Mel components; A perceptual weighting generation module for generating a perceptual weighting coefficient according to the auditory perception model; A frequency-domain energy adjustment module for performing frequency-domain energy adjustment on the noise-suppressed Mel components based on the perceptual weighting coefficient to generate a Mel spectrogram representation.

9. A computer device, characterized in that, The computer device includes a memory, a processor, and a speech feature processing program stored in the memory and executable on the processor. When the speech feature processing program is executed by the processor, it implements the steps of the speech feature processing method according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, A voice feature processing program is stored on the storage medium. When the voice feature processing program is executed by a processor, the steps of the voice feature processing method according to any one of claims 1-7 are implemented.

Citation Information

Cited By

  • Heart sound signal processing method and device and wearable equipment

    CN120837117A

  • Audio coding and decoding method based on dynamic sampling

    CN120913573A

  • An audio coding method based on dynamic sampling

    CN120913573B

  • Digital hearing aid automatic gain control method and system

    CN120980427A

  • Intelligent data analysis method and system based on large model

    CN120994811A