A post-processing method, device and equipment for improving speech perception and a medium

By adjusting the amplitude spectrum of the speech frame using frequency-related parameters in the contrast stretching method, reducing the energy of insensitive frequencies and limiting the degree of stretching, the problem of speech signal amplitude overflow is solved, achieving effective enhancement and stability of the speech signal, and improving the quality of speech communication and recognition.

CN119446163BActive Publication Date: 2025-11-21BEIJING UNISOUND INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411569485.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-05
Publication Date
2025-11-21
Estimated Expiration
2044-11-05

AI Technical Summary

Technical Problem

Existing contrast stretching methods, while enhancing speech signals, are prone to amplitude spectrum overflow, affecting the stability of speech signals, especially in real-time speech communication and recognition applications.

Method used

By acquiring the phase spectrum and amplitude spectrum of the speech frame, and using pre-configured frequency-dependent contrast stretching parameters, the energy of frequencies insensitive to the human ear is reduced while the energy of sensitive frequencies is maintained. The stretching degree is limited to avoid numerical overflow. The amplitude spectrum is processed using the target contrast stretching parameters, and the stretched spectrum is synthesized.

Benefits of technology

It effectively avoids numerical overflow problems while maintaining the clarity and naturalness of the speech signal, improving the listening experience of the speech signal, especially the speech quality within the frequency range sensitive to the human ear.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119446163B_ABST
    Figure CN119446163B_ABST
Patent Text Reader

Abstract

The application discloses a post-processing method and device for improving voice perception, equipment and medium. The pre-configured contrast stretching maximum value can limit the degree of contrast stretching, preventing signal distortion caused by excessive stretching. When determining the target contrast stretching parameter of any frequency, the pre-set contrast stretching preset parameter corresponding to the frequency is divided by the pre-determined contrast stretching maximum value, which can effectively reduce the energy of the non-sensitive frequency of the human ear and maintain the energy of the sensitive frequency of the human ear. Subsequently, the amplitude spectrum of the current voice frame is stretched by applying the target contrast stretching parameter corresponding to each frequency to obtain the stretched amplitude spectrum. The stretched amplitude spectrum maintains the energy of the sensitive frequency of the human ear while reducing the energy of the non-sensitive frequency, thereby avoiding excessive stretching of the amplitude of the sensitive frequency of the human ear and the problem of numerical overflow.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of digital signal processing and deep learning technology, and in particular to a post-processing method, apparatus, device and medium for improving the auditory quality of speech. Background Technology

[0002] In the field of speech signal processing, post-processing is crucial for improving speech quality and intelligibility. Among these techniques, contrast stretching based on human auditory characteristics is an effective speech enhancement method. This method deeply simulates the complex perceptual mechanism of the human ear for sound signals. By nonlinearly stretching the amplitude spectrum of the speech signal, particularly focusing on stretching more frequency components that are more sensitive to human hearing, while maintaining relatively less stretching for less sensitive frequency components, it effectively enhances key features in the speech, resulting in clearer and more natural processed speech.

[0003] However, existing contrast stretching methods have a significant problem: the stretched amplitude spectrum values ​​are unbounded and typically too large. This means that during processing, the stretched amplitude spectrum values ​​can easily exceed the maximum value of the PCM (Pulse Code Modulation) code, leading to overflow risks. Numerical overflow not only causes distortion of the speech signal but can also damage the stability of the entire communication system, especially in applications requiring rapid response and high stability, such as real-time voice communication and real-time speech recognition.

[0004] Therefore, how to maintain the effectiveness of contrast stretching in enhancing speech signals while quickly and effectively avoiding numerical overflow has become a key technical challenge that urgently needs to be solved in the field of speech signal processing. Summary of the Invention

[0005] This application provides a post-processing method, apparatus, device, and medium for improving the auditory experience of speech, which solves the problem that existing contrast stretching methods based on human hearing cannot effectively enhance the speech signal while avoiding numerical overflow.

[0006] In a first aspect, this application provides a post-processing method for improving the auditory quality of speech, the method comprising:

[0007] For any acquired speech frame, obtain the phase spectrum and amplitude spectrum of the speech frame;

[0008] The target contrast stretching parameters corresponding to each frequency in the speech frame are obtained to reduce the energy of frequencies that are not sensitive to the human ear and maintain the energy of frequencies that are sensitive to the human ear. The target contrast stretching parameter for any frequency is determined by dividing the preset contrast stretching parameter corresponding to that frequency by a pre-configured maximum contrast stretching value. The preset contrast stretching parameters corresponding to different frequencies are positively correlated with the human ear's auditory sensitivity corresponding to those different frequencies, and the preset contrast stretching parameters corresponding to different frequencies are all greater than 1.

[0009] The amplitude spectrum is stretched by applying the target contrast stretching parameters to obtain the stretched amplitude spectrum.

[0010] The phase spectrum and the stretched amplitude spectrum are combined to obtain the stretched spectrum;

[0011] Based on the stretched spectrum, the processed speech signal is obtained.

[0012] Secondly, this application also provides a post-processing apparatus for improving the auditory quality of speech, the apparatus comprising:

[0013] The acquisition unit is used to acquire the phase spectrum and amplitude spectrum of any acquired speech frame.

[0014] The parameter determination unit is used to obtain the target contrast stretching parameters corresponding to each frequency in the speech frame, so as to reduce the energy of frequencies that are not sensitive to the human ear and maintain the energy of frequencies that are sensitive to the human ear through the target contrast stretching parameters corresponding to each frequency; wherein, the target contrast stretching parameter of any frequency is determined by dividing the preset contrast stretching parameter corresponding to that frequency by a pre-configured maximum contrast stretching value; the preset contrast stretching parameters corresponding to different frequencies are positively correlated with the human ear hearing sensitivity corresponding to the different frequencies, and the preset contrast stretching parameters corresponding to different frequencies are all greater than 1.

[0015] An amplitude stretching unit is used to stretch the amplitude spectrum by applying the target contrast stretching parameters to obtain a stretched amplitude spectrum.

[0016] A synthesis unit is used to synthesize the phase spectrum and the stretched amplitude spectrum to obtain the stretched spectrum;

[0017] The output unit is used to obtain the processed speech signal based on the stretched spectrum.

[0018] Thirdly, this application provides a computer device including a processor, which executes a computer program stored in a memory to implement the steps of the post-processing method for improving speech perception as described above.

[0019] Fourthly, this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the post-processing method for improving speech perception as described above.

[0020] The beneficial effects of this application are as follows:

[0021] 1. By determining the maximum contrast stretching value from the preset contrast stretching parameters corresponding to different frequencies, the degree of contrast stretching can be limited, preventing signal distortion caused by excessive stretching.

[0022] 2. When determining the target contrast stretching parameter for any frequency, the preset contrast stretching parameter corresponding to that frequency is divided by a predetermined maximum contrast stretching value. This effectively reduces the energy of frequencies insensitive to the human ear while preserving the energy of frequencies sensitive to the human ear. Subsequently, by applying the target contrast stretching parameters corresponding to each frequency, the amplitude spectrum of the current speech frame is stretched to obtain a stretched amplitude spectrum. This stretched amplitude spectrum retains the energy of frequencies sensitive to the human ear while reducing the energy of insensitive frequencies, thus avoiding excessive stretching of the amplitude of frequencies sensitive to the human ear and the problem of numerical overflow. Attached Figure Description

[0023] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0024] Figure 1 A schematic diagram illustrating a post-processing method for improving the auditory quality of speech, provided in an embodiment of this application;

[0025] Figure 2 A schematic flowchart illustrating a specific post-processing method for improving the auditory quality of speech provided in this application embodiment;

[0026] Figure 3 A schematic diagram of a post-processing device for improving the listening experience of speech provided in an embodiment of this application;

[0027] Figure 4This is a schematic diagram of the structure of a computer device provided in an optional embodiment of this application. Detailed Implementation

[0028] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0029] In order to maintain the effectiveness of the contrast stretching method in enhancing the speech signal while quickly and effectively avoiding numerical overflow, this application provides a post-processing method, apparatus, device, and medium for improving the auditory experience of speech.

[0030] Example 1:

[0031] This application provides a post-processing method to improve the listening experience of speech. Figure 1 This application provides a schematic diagram of a post-processing method for improving the auditory quality of speech, which includes:

[0032] S101: For any acquired speech frame, obtain the phase spectrum and amplitude spectrum of the speech frame.

[0033] The post-processing method for improving the auditory quality provided in this application is applied to computer equipment, which can be a smart terminal, such as a computer or mobile phone, or a server, such as a business server or application server.

[0034] In this application, a computer device can acquire continuous speech signals in the working environment in real time, or acquire continuous speech signals acquired or stored in real time by other devices communicating with it. Then, it acquires the speech frames contained in the continuous speech signal. Here, a speech frame is the basic unit in speech signal processing, typically containing speech signals over a period of time. For example, a preset framing algorithm, such as framing and windowing, can be used to segment individual speech frames from the continuous speech signal. Specifically, using a window length equal to the duration (or "frame length") of each speech frame and a frame shift as the step size, the window is moved across the continuous speech signal according to the chronological order of the speech frames to obtain a sequence of speech frames for the continuous speech signal. This sequence of speech frames includes multiple speech frames arranged in chronological order. The frame shift refers to the time difference between the starting positions of two adjacent speech frames.

[0035] For any acquired speech frame, the frame is processed to obtain its amplitude spectrum and phase spectrum. The phase spectrum describes the phase information of each frequency component in the speech frame, while the amplitude spectrum describes the amplitude (or energy) information of each frequency component. For example, the speech frame can first be converted from the time domain to the frequency domain to obtain its spectrum. For instance, a Fourier transform (such as a Short-Time Fourier Transform, STFT) can be used to convert the speech frame from the time domain to the frequency domain, obtaining its complex spectrum. This spectrum is then further decomposed to obtain the phase spectrum and amplitude spectrum of the speech frame. For example, the norm of the spectrum can be determined as the amplitude spectrum, and the phase angle of the spectrum can be determined as the phase spectrum.

[0036] For example, the phase spectrum and amplitude spectrum of a speech frame can be determined by the following formula:

[0037] A(t,f)=||X(t,f)||1

[0038] θ(t,f)=angle(X(t,f))

[0039] Where A(t,f) represents the original amplitude value of the t-th speech frame at frequency f, ||||1 represents the norm, X(t,f) represents the spectrum of the t-th speech frame at frequency f, θ(t,f) represents the phase angle of the t-th speech frame at frequency f, and angle(·) is the phase operator.

[0040] S102: Obtain the target contrast stretching parameters corresponding to each frequency in the speech frame, so as to reduce the energy of frequencies that are not sensitive to the human ear and maintain the energy of frequencies that are sensitive to the human ear through the target contrast stretching parameters corresponding to each frequency; wherein, the target contrast stretching parameter of any frequency is determined by dividing the preset contrast stretching parameter corresponding to that frequency by a pre-configured maximum contrast stretching value; the preset contrast stretching parameters corresponding to different frequencies are positively correlated with the human ear hearing sensitivity corresponding to the different frequencies, and the preset contrast stretching parameters corresponding to different frequencies are all greater than 1.

[0041] In this application, preset contrast stretching parameters corresponding to different frequencies are pre-configured. These preset parameters aim to adjust the contrast of the speech signal at different frequencies, where contrast refers to the difference in intensity or dynamic range of the speech signal at different frequencies. When setting these preset parameters, the perceptual characteristics of the human ear at different frequencies, i.e., human auditory sensitivity, can be fully considered. This auditory sensitivity can be determined based on auditory parameters (such as equal loudness curves, masking effects, etc.). Specifically, since the human ear is more sensitive to mid-frequency sounds and relatively less sensitive to low- and high-frequency sounds, a larger preset contrast stretching parameter can be set for the mid-frequency range to better highlight the speech features within these frequency ranges; while a relatively smaller preset contrast stretching parameter is set for the low- and high-frequency ranges. In other words, the setting of these preset contrast stretching parameters is positively correlated with human auditory sensitivity.

[0042] It should be noted that the preset contrast stretching parameters for different frequencies are all greater than 1 to ensure that the voice signal is enhanced during the stretching process.

[0043] For example, based on a 16kHz sampling rate, the preset parameters for contrast stretching of the 512-point Fourier transform (denoted as PCS) are:

[0044] PCS[0:3]=1

[0045] PCS[3:6]=1.070175439

[0046] PCS[6:9] = 1.182456140

[0047] PCS[9:12]=1.287719298

[0048] PCS[12:138] = 1.4

[0049] PCS[138:166]=1.322807018

[0050] PCS[166:200] = 1.238596491

[0051] PCS[200:241]=1.161403509

[0052] PCS[241:256]=1.077192982

[0053] Meanwhile, since the pre-configured contrast stretching preset parameters for different frequencies are all greater than 1, aiming to increase the energy of the sensitive area of ​​the human ear, directly stretching the amplitude spectrum based on these preset parameters would result in an unbounded amplitude spectrum value, which could easily exceed the maximum value of PCM (Pulse Code Modulation) encoding, thus causing an overflow risk. Therefore, in this application, a maximum value (denoted as the maximum contrast stretching value) can be determined from the pre-configured contrast stretching preset parameters for different frequencies to limit the degree of contrast stretching and prevent signal distortion caused by excessive stretching.

[0054] After obtaining the amplitude spectrum of the speech frame based on the above embodiments, the target contrast stretching parameters corresponding to each frequency in the speech frame can be obtained. This allows for the reduction of energy at frequencies insensitive to the human ear while maintaining energy at frequencies sensitive to the human ear. When determining the target contrast stretching parameter for any frequency, it can be determined by dividing a pre-set contrast stretching preset parameter corresponding to that frequency by a pre-determined maximum contrast stretching value.

[0055] For example, the target contrast stretching parameter for any frequency can be determined by the following formula:

[0056] PCS fix (f) = PCS(f) / max(PCS)

[0057] Among them, PCS fix (f) represents the target contrast stretching parameter corresponding to frequency f, PCS(f) represents the preset contrast stretching parameter corresponding to frequency f, and max(PCS) represents the maximum contrast stretching value.

[0058] It should be noted that there are two ways to determine the target contrast stretching parameters:

[0059] 1. Pre-determined: In this method, before processing the speech frame, the target contrast stretching parameters corresponding to different frequencies can be pre-calculated based on the maximum contrast stretching value and the preset contrast stretching parameters corresponding to different frequencies, and stored for later use.

[0060] 2. Real-time determination: Another approach is to calculate the target contrast stretching parameters for each frequency in the current speech frame in real time, based on the preset contrast stretching parameters and the pre-configured maximum contrast stretching value for each frequency in the current speech frame.

[0061] When processing a certain speech frame, based on its frequency components, the preset contrast stretching parameters corresponding to each frequency in the speech frame are retrieved from the preset contrast stretching parameters corresponding to different frequencies.

[0062] S103: The amplitude spectrum is stretched by applying the target contrast stretching parameters to obtain the stretched amplitude spectrum.

[0063] After obtaining the target contrast stretching parameters corresponding to each frequency in the speech frame based on the above embodiments, the amplitude spectrum of the speech frame can be stretched using the target contrast stretching parameters corresponding to each frequency component. For example, this can be achieved by multiplying each frequency component of the amplitude spectrum by its corresponding target contrast stretching parameter.

[0064] In one possible implementation, for each frequency in the speech frame, the amplitude corresponding to that frequency can be obtained from the amplitude spectrum of the speech frame. The sum of this amplitude and 1 is determined to avoid zero or negative values ​​when taking the logarithm. Then, the natural logarithm of this sum is calculated to transform the amplitude value into a range more suitable for stretching. The obtained natural logarithmic value is multiplied by the target contrast stretching parameter corresponding to that frequency. The product is added to 1 to maintain the non-zero nature of the amplitude. Then, the natural exponent of the sum of the product and 1 is calculated to obtain the natural exponent value. This natural exponent value is determined as the stretched amplitude value.

[0065] For example, by applying the target contrast stretching parameters, the amplitude spectrum is stretched to obtain the stretched amplitude spectrum, which is expressed by the following formula:

[0066]

[0067] Among them, A P (t,f) represents the amplitude value of the t-th speech frame after stretching at frequency f, PCS fix (f) represents the target contrast stretching parameter corresponding to frequency f, A(t,f) represents the original amplitude value of the t-th speech frame at frequency f, and e (·) Ln is the natural exponent operator. (·) It is the natural logarithm operator.

[0068] By repeating the above steps, a stretched amplitude value can be calculated for each frequency in the speech frame. Combining these values ​​yields the stretched amplitude spectrum. This stretched amplitude spectrum preserves the energy of frequencies sensitive to the human ear while reducing the energy of insensitive frequencies, avoiding excessive stretching of the amplitude of frequencies sensitive to the human ear and thus preventing numerical overflow.

[0069] S104: Combine the phase spectrum with the stretched amplitude spectrum to obtain the stretched spectrum.

[0070] After obtaining the stretched amplitude spectrum based on the above embodiments, the stretched amplitude spectrum can be synthesized with the original phase spectrum to obtain the stretched spectrum.

[0071] For example, for each frequency in the speech frame, the stretched amplitude corresponding to that frequency is used as the modulus of a complex number, and the phase corresponding to that frequency in the original phase spectrum is used as the argument of a complex number. Then, using the conversion formula between polar and rectangular coordinates (i.e., Euler's formula), this modulus and argument are converted into a complex number. Combining the complex numbers corresponding to all frequencies yields the stretched spectrum. This spectrum contains both the stretched amplitude information and retains the original phase information.

[0072] Another example is to obtain the phase angle corresponding to each frequency in the speech frame from the phase spectrum of the speech frame. Then, determine the product of this phase angle and the imaginary unit, where the imaginary unit is equal to the square root of one. Calculate the natural exponent of this product to obtain the natural exponent value. Next, multiply this natural exponent value by the stretched amplitude value corresponding to the frequency, and determine the stretched spectrum as the resulting product.

[0073] For example,

[0074] X p (t,f)=A P (t,f)·e jθ(t,f)

[0075] Among them, X p (t,f) represents the stretched spectrum of the t-th speech frame at frequency f, A P (t,f) represents the amplitude value of the t-th speech frame after stretching at frequency f, θ(t,f) represents the phase angle of the t-th speech frame at frequency f, and j represents the imaginary unit. e (·) It is the natural exponent operator.

[0076] S105: Based on the stretched spectrum, the processed speech signal is obtained.

[0077] Finally, after obtaining the stretched spectrum based on the above embodiments, the stretched spectrum can be converted back from the frequency domain to the time domain using inverse Fourier transform (such as inverse short-time Fourier transform ISTFT) to obtain the processed speech signal. This processed speech signal, while preserving the original speech content, has higher clarity and naturalness, especially within the frequency range sensitive to human hearing.

[0078] In one possible implementation, the spectrum after contrast stretching has been adjusted as required in the frequency domain. To reflect these adjustments in the time domain signal, the stretched spectrum can be transformed using the inverse Fourier transform (IFT) to convert the frequency domain signal back to the time domain, thus obtaining the processed speech frame time domain signal.

[0079] Since speech signals are typically continuous and long, they are usually segmented into multiple shorter speech frames for processing. However, this segmentation can lead to discontinuities between frames, affecting the quality of the processed speech. To address this issue, adjacent frames can overlap during speech signal segmentation. The length of the overlap typically depends on the window function used (such as the Hanning window, Hamming window, etc.) and the specific signal processing task. After obtaining the processed time-domain signals corresponding to each speech frame, the processing results of adjacent frames can be added in the overlapping region with certain weights, for example, using the window function value as a weight, to smooth the transition and reduce discontinuities between adjacent frames. The processed speech frames (including the overlapping portion) are then concatenated to form a complete processed speech signal. By using the overlap-addition method, discontinuities between frames can be effectively reduced, improving the quality of the processed speech.

[0080] The beneficial effects of this application are as follows:

[0081] 1. By determining the maximum contrast stretching value from the preset contrast stretching parameters corresponding to different frequencies, the degree of contrast stretching can be limited, preventing signal distortion caused by excessive stretching.

[0082] 2. When determining the target contrast stretching parameter for any frequency, the preset contrast stretching parameter corresponding to that frequency is divided by a predetermined maximum contrast stretching value. This effectively reduces the energy of frequencies insensitive to the human ear while preserving the energy of frequencies sensitive to the human ear. Subsequently, by applying the target contrast stretching parameters corresponding to each frequency, the amplitude spectrum of the current speech frame is stretched to obtain a stretched amplitude spectrum. This stretched amplitude spectrum retains the energy of frequencies sensitive to the human ear while reducing the energy of insensitive frequencies, thus avoiding excessive stretching of the amplitude of frequencies sensitive to the human ear and the problem of numerical overflow.

[0083] Example 2:

[0084] The following describes a post-processing method for improving speech perception provided in this application through specific embodiments. Figure 2A flowchart illustrating a specific post-processing method for improving speech perception provided in this application embodiment is shown. The process includes:

[0085] S201: Acquire continuous speech signal.

[0086] S202: Perform framing and windowing processing on the continuous speech signal to obtain each speech frame in the continuous speech signal.

[0087] For each speech frame in this continuous speech signal, the following steps S203 to S208 are executed:

[0088] S203: Perform a Fast Fourier Transform on the speech frame to obtain its spectrum.

[0089] S204: Based on this spectrum, determine the amplitude spectrum and phase spectrum of the speech frame.

[0090] S205: Obtain the target contrast stretching parameters corresponding to each frequency in the speech frame.

[0091] The target contrast stretching parameter for any frequency is determined by dividing the preset contrast stretching parameter for that frequency by a pre-configured maximum contrast stretching value. The preset contrast stretching parameters for different frequencies are positively correlated with the human auditory sensitivity for those frequencies, and the preset contrast stretching parameters for each frequency are all greater than 1.

[0092] S206: The amplitude spectrum is stretched by applying the contrast stretching parameters of each target to obtain the stretched amplitude spectrum.

[0093] In one possible implementation, for each frequency in the speech frame, the amplitude corresponding to that frequency is obtained from the amplitude spectrum; the natural logarithm of the sum of the amplitude and 1 is determined to obtain the natural logarithm value; the product of the target contrast stretching parameter and the natural logarithm value is determined; and the stretched amplitude value is determined based on the natural exponent of the sum of the product and 1.

[0094] S207: Combine the phase spectrum with the stretched amplitude spectrum to obtain the stretched spectrum.

[0095] In one possible implementation, for each frequency in the speech frame, the phase angle corresponding to that frequency is obtained from the phase spectrum; the natural exponent of the product of the phase angle and the imaginary unit is determined to obtain the natural exponent value; wherein the imaginary unit is equal to the square root of one; and the stretched spectrum is determined based on the product of the natural exponent value and the stretched amplitude value corresponding to that frequency.

[0096] S208: Perform an inverse Fourier transform on the stretched spectrum to obtain the processed time-domain signal of the speech frame.

[0097] S209: The time-domain signals corresponding to each speech frame are processed by the overlapping addition method to obtain the processed speech signal.

[0098] Example 3:

[0099] Based on the same inventive concept, this application also provides a post-processing device for improving the auditory experience of speech. Figure 3 This application provides a schematic diagram of a post-processing device for improving speech perception, comprising:

[0100] The acquisition unit 31 is used to acquire the phase spectrum and amplitude spectrum of any acquired speech frame;

[0101] The parameter determination unit 32 is used to obtain the target contrast stretching parameters corresponding to each frequency in the speech frame, so as to reduce the energy of frequencies that are not sensitive to the human ear and maintain the energy of frequencies that are sensitive to the human ear through the target contrast stretching parameters corresponding to each frequency; wherein, the target contrast stretching parameter of any frequency is determined by dividing the preset contrast stretching parameter corresponding to that frequency by a pre-configured maximum contrast stretching value; the preset contrast stretching parameters corresponding to different frequencies are positively correlated with the human ear hearing sensitivity corresponding to the different frequencies, and the preset contrast stretching parameters corresponding to different frequencies are all greater than 1.

[0102] Amplitude stretching unit 33 is used to stretch the amplitude spectrum by applying the target contrast stretching parameters to obtain the stretched amplitude spectrum.

[0103] Synthesis unit 34 is used to synthesize the phase spectrum and the stretched amplitude spectrum to obtain the stretched spectrum;

[0104] Output unit 35 is used to obtain the processed speech signal based on the stretched spectrum.

[0105] In some possible implementations, the amplitude stretching unit 33 is specifically used for:

[0106] For each frequency, the amplitude corresponding to that frequency is obtained from the amplitude spectrum; the natural logarithm of the sum of the amplitude and 1 is determined to obtain the natural logarithm value; the product of the target contrast stretching parameter and the natural logarithm value is determined; and the stretched amplitude value is determined based on the natural exponent of the sum of the product and 1.

[0107] In some possible implementations, the synthesis unit 34 is specifically used for:

[0108] For each frequency, the phase angle corresponding to that frequency is obtained from the phase spectrum; the natural exponent of the product of the phase angle and the imaginary unit is determined to obtain the natural exponent value; wherein the imaginary unit is equal to the square root of one; the stretched spectrum is determined based on the product of the natural exponent value and the stretched amplitude value corresponding to that frequency.

[0109] In some possible implementations, the acquisition unit 31 is specifically used for:

[0110] Obtain the spectrum of the speech frame;

[0111] The norm of the spectrum is determined as the amplitude spectrum, and the phase angle of the spectrum is determined as the phase spectrum.

[0112] In some possible implementations, the acquisition unit 31 is specifically used for:

[0113] The acquired continuous speech signal is segmented and windowed to obtain speech frames in the continuous speech signal.

[0114] In some possible implementations, the output unit 35 is specifically used for:

[0115] Perform an inverse Fourier transform on the stretched spectrum to obtain the processed time-domain signal of the speech frame;

[0116] The time-domain signals corresponding to each of the aforementioned speech frames are processed using an overlap-addition method to obtain the processed speech signals.

[0117] In this embodiment, a post-processing device for improving the listening experience of speech is presented in the form of a functional module. Here, a module refers to an application-specific integrated circuit (ASIC), a processor and memory that execute one or more software or fixed programs, and / or other devices that can provide the above-mentioned functions.

[0118] Further functional descriptions of the above modules and units are the same as those in the corresponding embodiments described above, and will not be repeated here.

[0119] Example 4:

[0120] Please see Figure 4 , Figure 4 This is a schematic diagram of the structure of a computer device provided in an optional embodiment of this application, such as... Figure 4As shown, the computer device includes one or more processors 10, memory 20, and interfaces for connecting the components, including high-speed interfaces and low-speed interfaces. The components communicate with each other via different buses and can be mounted on a common motherboard or otherwise installed as needed. The processors can process instructions executed within the computer device, including instructions stored in or on memory to display graphical information of a GUI on external input / output devices (such as display devices coupled to the interfaces). In some alternative implementations, multiple processors and / or multiple buses can be used with multiple memories and multiple memory modules, if desired. Similarly, multiple computer devices can be connected, each providing some of the necessary operations (e.g., as a server array, a group of blade servers, or a multiprocessor system). Figure 4 Take a processor 10 as an example.

[0121] Processor 10 may be a central processing unit, a network processor, or a combination thereof. Processor 10 may further include a hardware chip. The hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The programmable logic device may be a complex programmable logic device (CAMP), a field-programmable gate array (FPGA), a general-purpose array logic (GDA), or any combination thereof.

[0122] The memory 20 stores instructions executable by at least one processor 10 to cause at least one processor 10 to perform the method shown in the above embodiments.

[0123] The memory 20 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the computer device as shown by a landing page for an app. Furthermore, the memory 20 may include high-speed random access memory and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some alternative embodiments, the memory 20 may optionally include memory remotely located relative to the processor 10, which can be connected to the computer device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0124] The memory 20 may include volatile memory, such as random access memory; the memory may also include non-volatile memory, such as flash memory, hard disk or solid-state drive; the memory 20 may also include a combination of the above types of memory.

[0125] The computer device also includes an input device 30 and an output device 40. The processor 10, memory 20, input device 30, and output device 40 can be connected via a bus or other means. Figure 4 Taking the example of a connection between China and Israel via a bus.

[0126] Input device 30 can receive input numerical or character information, and generate key signal inputs related to user settings and function control of the computer device, such as a touchscreen, keypad, mouse, trackpad, touchpad, joystick, one or more mouse buttons, trackball, joystick, etc. Output device 40 may include display devices, auxiliary lighting devices (e.g., LEDs), and haptic feedback devices (e.g., vibration motors). The aforementioned display devices include, but are not limited to, liquid crystal displays, light-emitting diodes, displays, and plasma displays. In some alternative embodiments, the display device may be a touchscreen.

[0127] Example 5:

[0128] Based on the above embodiments, this application also provides a computer-readable storage medium storing a computer program executable by a processor. When the program runs on the processor, it causes the processor to perform the following steps:

[0129] For any acquired speech frame, obtain the phase spectrum and amplitude spectrum of the speech frame;

[0130] The target contrast stretching parameters corresponding to each frequency in the speech frame are obtained to reduce the energy of frequencies that are not sensitive to the human ear and maintain the energy of frequencies that are sensitive to the human ear. The target contrast stretching parameter for any frequency is determined by dividing the preset contrast stretching parameter corresponding to that frequency by a pre-configured maximum contrast stretching value. The preset contrast stretching parameters corresponding to different frequencies are positively correlated with the human ear's auditory sensitivity corresponding to those different frequencies, and the preset contrast stretching parameters corresponding to different frequencies are all greater than 1.

[0131] The amplitude spectrum is stretched by applying the target contrast stretching parameters to obtain the stretched amplitude spectrum.

[0132] The phase spectrum and the stretched amplitude spectrum are combined to obtain the stretched spectrum;

[0133] Based on the stretched spectrum, the processed speech signal is obtained.

[0134] Since the principle of the computer-readable storage medium in solving the problem is similar to that of a post-processing method for improving the auditory quality of speech, the implementation of the computer-readable storage medium can be found in the embodiments of the method, and repeated details will not be described again.

[0135] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.

Claims

1. A post-processing method for improving the auditory quality of speech, characterized in that, The method includes: For any acquired speech frame, obtain the phase spectrum and amplitude spectrum of the speech frame; The target contrast stretching parameters corresponding to each frequency in the speech frame are obtained to reduce the energy of frequencies that are not sensitive to the human ear and maintain the energy of frequencies that are sensitive to the human ear. The target contrast stretching parameter for any frequency is determined by dividing the preset contrast stretching parameter corresponding to that frequency by a pre-configured maximum contrast stretching value. The preset contrast stretching parameters corresponding to different frequencies are positively correlated with the human ear's auditory sensitivity corresponding to those different frequencies, and the preset contrast stretching parameters corresponding to different frequencies are all greater than 1. The amplitude spectrum is stretched by applying the target contrast stretching parameters to obtain the stretched amplitude spectrum. The phase spectrum and the stretched amplitude spectrum are combined to obtain the stretched spectrum; Based on the stretched spectrum, the processed speech signal is obtained; The step of stretching the amplitude spectrum by applying each of the target contrast stretching parameters to obtain the stretched amplitude spectrum includes: For each frequency, the amplitude corresponding to that frequency is obtained from the amplitude spectrum; the natural logarithm of the sum of the amplitude and 1 is determined to obtain the natural logarithm value; the product of the target contrast stretching parameter and the natural logarithm value is determined; and the stretched amplitude value is determined based on the natural exponent of the sum of the product and 1.

2. The method as described in claim 1, characterized in that, The step of synthesizing the phase spectrum and the stretched amplitude spectrum to obtain the stretched spectrum includes: For each frequency, the phase angle corresponding to that frequency is obtained from the phase spectrum; the natural exponent of the product of the phase angle and the imaginary unit is determined to obtain the natural exponent value; wherein the imaginary unit is equal to the square root of one; the stretched spectrum is determined based on the product of the natural exponent value and the stretched amplitude value corresponding to that frequency.

3. The method as described in claim 1, characterized in that, For any acquired speech frame, obtaining the phase spectrum and amplitude spectrum of the speech frame includes: Obtain the spectrum of the speech frame; The norm of the spectrum is determined as the amplitude spectrum, and the phase angle of the spectrum is determined as the phase spectrum.

4. The method as described in claim 1, characterized in that, Obtaining the audio frame includes: The acquired continuous speech signal is segmented and windowed to obtain speech frames in the continuous speech signal.

5. The method as described in claim 4, characterized in that, The process of obtaining the processed speech signal based on the stretched spectrum includes: Perform an inverse Fourier transform on the stretched spectrum to obtain the processed time-domain signal of the speech frame; The time-domain signals corresponding to each of the aforementioned speech frames are processed using an overlap-addition method to obtain the processed speech signals.

6. A post-processing device for improving the auditory quality of speech, characterized in that, The device includes: The acquisition unit is used to acquire the phase spectrum and amplitude spectrum of any acquired speech frame. The parameter determination unit is used to obtain the target contrast stretching parameters corresponding to each frequency in the speech frame, so as to reduce the energy of frequencies that are not sensitive to the human ear and maintain the energy of frequencies that are sensitive to the human ear through the target contrast stretching parameters corresponding to each frequency; wherein, the target contrast stretching parameter of any frequency is determined by dividing the preset contrast stretching parameter corresponding to that frequency by a pre-configured maximum contrast stretching value; the preset contrast stretching parameters corresponding to different frequencies are positively correlated with the human ear hearing sensitivity corresponding to the different frequencies, and the preset contrast stretching parameters corresponding to different frequencies are all greater than 1. An amplitude stretching unit is used to stretch the amplitude spectrum by applying the target contrast stretching parameters to obtain a stretched amplitude spectrum. A synthesis unit is used to synthesize the phase spectrum and the stretched amplitude spectrum to obtain the stretched spectrum; The output unit is used to obtain the processed speech signal based on the stretched spectrum; Specifically, the amplitude stretching unit is used for: For each frequency, the amplitude corresponding to that frequency is obtained from the amplitude spectrum; the natural logarithm of the sum of the amplitude and 1 is determined to obtain the natural logarithm value; the product of the target contrast stretching parameter and the natural logarithm value is determined; and the stretched amplitude value is determined based on the natural exponent of the sum of the product and 1.

7. A computer device, characterized in that, The computer device includes a processor that executes a computer program stored in a memory to implement the steps of a post-processing method for improving speech perception as described in any one of claims 1-5.

8. A computer-readable storage medium, characterized in that, It stores a computer program that, when executed by a processor, implements the steps of a post-processing method for improving speech perception as described in any one of claims 1-5.

Citation Information

Patent Citations

  • Audio signal processing method and device and readable storage medium

    CN111405419A

  • Audio signal harmony processing method and device, electronic device and storage medium

    CN112086085A