Sound signal processing method and processing device
By extracting the Mel cepspectral coefficients and using the Mel generative adversarial network to generate synthetic sound signals, the problem that hearing impaired people find it difficult to receive high-frequency sounds, and the effective transfer of high-frequency air sounds and the retention of sound characteristics is achieved.
Patent Information
- Application Number
- CN202410109371.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-26
- Publication Date
- 2025-07-29
AI Technical Summary
It is difficult for hearing-impaired people to receive high-frequency sound signals clearly, but the prior art amplification of high-frequency signals can lead to unpleasant sounds and cannot effectively transfer high-frequency air tones to low frequencies for easy understanding.
By extracting the Mel cepspectral coefficients, high-frequency power is mapped to low frequencies using a bandpass filter, and a synthetic sound signal is generated, and a frequency shifted sound signal is generated using the Mel generative adversarial network.
The generated synthetic sound signals can retain sound characteristics, allowing hearing-impaired people to understand the meaning without producing unpleasant sounds, achieving effective transfer of high-frequency air sounds.
Smart Images

Figure CN120388556A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a processing technology for sound signals, and in particular, also relates to a processing method and a processing device for sound signals. Background Art
[0002] Most hearing-impaired people cannot clearly receive some high-frequency sound signals, but can clearly hear low-frequency sound signals. If only using an equalizer to amplify high-frequency speech signals to improve the above problems, many unpleasant sounds (such as whistling, amplified noise, excessive ear pressure, etc.) may occur.
[0003] Research indicates that if the high-frequency aspirated sounds in speech (such as "s"tudy) are transferred to specific low-frequency signals (such as "z"tudy), listeners can distinguish their semantics based on experience and learning. The high-frequency aspirated sounds in general speech are about 6 kilohertz (kHz). According to the different degrees of ear damage of hearing-impaired people, we can downshift the high-frequency aspirated sounds by one-half or one-quarter to generate sound signals corresponding to 3 kHz or 1.5 kHz. Summary of the Invention
[0004] The present invention is directed to a processing method and a processing device for sound signals, which can provide a frequency-shifting effect for sound signals.
[0005] According to an embodiment of the present invention, the processing method for sound signals includes (but is not limited to) the following steps: extracting a plurality of Mel-Frequency Cepstrum Coefficients (MFCC) from a sound signal to be processed, including: obtaining the power of the sound signal to be processed corresponding to a plurality of Mel frequencies through a plurality of band-pass filters, where each band-pass filter corresponds to a Mel frequency, and the Mel frequencies corresponding to these band-pass filters are different; mapping a first frequency among these Mel frequencies to a second frequency among these Mel frequencies, and replacing the power corresponding to the second frequency with the power corresponding to the first frequency, where the second frequency is lower than the first frequency; and generating these Mel-Frequency Cepstrum Coefficients using the power corresponding to these Mel frequencies. Generating a synthesized sound signal using the Mel-Frequency Cepstrum Coefficients of the sound signal to be processed, where the sound signal to be processed and the synthesized sound signal are used for playing through a speaker.
[0006] According to an embodiment of the present invention, a processing device includes a memory and a processor. The memory is used to store program code. The processor is coupled to the memory. The processor is used to load the program code to execute: extracting a plurality of mel cepstral coefficients from a sound signal to be processed, and generating a synthesized sound signal using the mel cepstral coefficients of the sound signal to be processed, where the sound signal to be processed and the synthesized sound signal are for playing through a speaker. The processor is configured to: obtain the power of the sound signal to be processed corresponding to a plurality of mel frequencies through a plurality of band-pass filters, where each band-pass filter corresponds to a mel frequency, and the mel frequencies corresponding to these band-pass filters are different; map a first frequency among these mel frequencies to a second frequency among these mel frequencies, and replace the power corresponding to the second frequency with the power corresponding to the first frequency, where the second frequency is lower than the first frequency; and generate these mel cepstral coefficients using the power corresponding to these mel frequencies.
[0007] Based on the above, the method and device for processing a sound signal according to an embodiment of the present invention can replace the power of a lower mel frequency band with the power of a higher mel frequency, and generate shifted mel cepstral coefficients and a corresponding synthesized sound signal accordingly. This synthesized sound signal can be used for hearing-impaired persons to listen to and can retain the sound characteristics. Description of the Drawings
[0008] The drawings are included to provide a further understanding of the present invention, and the drawings are incorporated into and constitute a part of this specification. The drawings illustrate embodiments of the present invention and, together with the description, are used to explain the principles of the present invention.
[0009] Figure 1 is a block diagram of components of a device for processing a sound signal according to an embodiment of the present invention;
[0010] Figure 2 is a flowchart of a method for processing a sound signal according to an embodiment of the present invention;
[0011] Figure 3 is a flowchart of a method for generating mel cepstral coefficients according to an embodiment of the present invention;
[0012] Figure 4 is a schematic diagram of mel frequency mapping according to an embodiment of the present invention;
[0013] Figure 5 is a schematic diagram of machine learning training according to an embodiment of the present invention;
[0014] Figure 6 is a flowchart of a method for processing a sound signal according to an embodiment of the present invention.
[0015] Description of the Reference Numerals in the Drawings
[0016] 100: Processing device;
[0017] 110: Loudspeaker;
[0018] 120: Memory;
[0019] 130: Processor;
[0020] S210~S220, S310~S320, S610~S640: Steps;
[0021] S B1 : Sound signal to be processed;
[0022] Mel cepstral coefficients;
[0023] |X[k]| 2 , Y[m], Y t [m]~Y t [M]: Power;
[0024] BPF: Band-pass filter;
[0025] Δ FS : Displacement;
[0026] DS: Discriminator;
[0027] GS: Generator;
[0028] z: Basic sound signal;
[0029] s, y: Mel cepstral coefficients;
[0030] G(s,z), G(y,z): Estimated sound signal;
[0031] S in : Initial sound signal;
[0032] S B2 : Bypass sound signal;
[0033] Synthesized sound signal;
[0034] S out : Output sound signal. Detailed implementation
[0035] Reference will now be made in detail to the exemplary embodiments of the present invention, examples of which are illustrated in the accompanying drawings. Whenever possible, the same component symbols are used in the drawings and the description to represent the same or similar parts.
[0036] Figure 1 is a block diagram of the components of a sound signal processing device 10 according to an embodiment of the present invention. Please refer to Figure 1, the processing device 10 includes (but is not limited to) a speaker 110 (optional), a memory 120, and a processor 130. The processing device 100 can be a smart phone, a tablet computer, a smart assistant device, a wearable device, a vehicle-mounted system, a laptop computer, a hearing aid, a conference call device, or other electronic devices.
[0037] The speaker 110 can be a transducer or an electronic component that converts an electronic sound signal into sound. One or more speakers 110 can also form an audio group. In one embodiment, the speaker 110 is used to play a sound signal.
[0038] The memory 120 can be any type of fixed or removable random access memory (RAM), read only memory (ROM), flash memory, traditional hard disk drive (HDD), solid state drive (SSD), or similar components. In one embodiment, the memory 120 is used to store program codes, software modules, configuration settings, data (such as sound signals, coefficients, algorithm parameters, etc.), or files, and embodiments thereof will be described in detail later.
[0039] The processor 130 is coupled to the speaker 110 and the memory 120. The processor 130 can be a central processing unit (CPU), a graphics processing unit (GPU), or other programmable general-purpose or special-purpose microprocessors, digital signal processors (DSPs), programmable controllers, field programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), neural network accelerators, or other similar components, or a combination of the above components. In one embodiment, the processor 130 is used to execute all or part of the operations of the processing device 100, and can load and execute each program code, software module, file, and data stored in the memory 120. In some embodiments, the functions of the processor 130 can be implemented by software or a chip.
[0040] In the following, the method described in the embodiments of the present invention will be described in conjunction with each component and module in the processing device 100. Each process of this method can be adjusted according to the implementation situation and is not limited thereto.
[0041] Figure 2 is a flowchart of a method for processing a sound signal according to an embodiment of the present invention. Please refer to Figure 2 , the processor 130 extracts a plurality of Mel-Frequency Cepstrum Coefficients (MFCCs) from the sound signal to be processed (step S210). Specifically, the sound signal to be processed may be a voice signal obtained by recording, receiving, or synthesizing. For example, a voice signal generated by recording human voices through a microphone; a voice signal received by a communication transceiver circuit from a conference / call server; or a voice signal edited by audio software. In one embodiment, the sound signal to be processed is a voice signal obtained in application scenarios such as calls, video conferences, conversations, movie viewing, music playing, etc. In other application scenarios, the sound signal to be processed is not limited to voice signals.
[0042] In one embodiment, the sound signal to be processed may belong to or only include signals in a specific frequency band. For example, the frequency band corresponding to aspirated sounds in speech is 4 kilohertz (kHz) to 8 kHz. However, the frequency band corresponding to the sound signal to be processed is still adjusted according to the actual application scenario.
[0043] In one embodiment, the sound signal to be processed is a signal generated by noise suppression, gain amplification, filtering, sound source extraction, or other sound signal processing.
[0044] In one embodiment, the sound signal to be processed is used to be played through the speaker 100. That is, the speaker 110 can play the sound signal to be processed.
[0045] On the other hand, the Mel-Frequency Cepstrum Coefficients are the coefficients that make up the Mel-Frequency Cepstrum. This Mel-Frequency Cepstrum is formed by the cepstrum of an audio segment and divides its frequency band according to the Mel Scale (each frequency band corresponds to a Mel frequency). An approximate mathematical conversion can be performed between the Mel Scale and the linear frequency scale Hertz (Hz). For example, the formula for converting x Hertz to y Mel (frequency) is:
[0046] y = 2595 log 10 (1 + x / 700)…(1).
[0047] One of the multiple characteristics of the Mel-Frequency Cepstrum Coefficients lies in the sound characteristics that are close to the auditory characteristics of the human ear. For example, the Mel Scale can represent the human ear's perception of equally spaced changes in pitch. Therefore, in some application scenarios, the Mel-Frequency Cepstrum Coefficients can be applied to speech recognition functions, but are not limited thereto.
[0048] Figure 3 is a flowchart of a method for generating Mel-Frequency Cepstrum Coefficients according to an embodiment of the present invention. Please refer to Figure 3, the processor 130 preprocesses the sound signal S to be processed B1 (step S310).
[0049] Specifically, pre-emphasis (step S311) is that the processor 130 passes the sound signal S to be processed B1 through a high-pass filter (S B1 (n) - a·S B1 (n), where a is a coefficient between 0.9 and 1) to eliminate the effects caused by the vocal cords and lips during the sound production process and compensate for the high-frequency part of the speech signal suppressed by the pronunciation system.
[0050] Frame Blocking (step S312) is to define a set of i sampling points as a frame, and the processor 130 can divide the sound signal S to be processed B1 into one frame for every i sampling points, thereby obtaining multiple frames (or called frames).
[0051] Windowing (step S313) is that the processor 130 multiplies each frame of the sound signal S to be processed B1 by a Hamming window to increase the continuity of the left and right ends of the frame.
[0052] Next, the processor 130 extracts Mel-frequency cepstral coefficients from the sound signal S to be processed that has been preprocessed (step S310) B1 (step S320).
[0053] Specifically, the Fast Fourier Transform (FFT) (step S321) is that the processor 130 converts each frame of the sound signal S to be processed B1 from the time domain to the frequency domain to obtain the energy distribution of each frame on the spectrum. For example, the power or energy on the spectrum. In another embodiment, a discrete Fourier transform or other time-domain to frequency-domain conversion can be used.
[0054] Filtering process (step S322) is that the processor 130 filters the spectral energy obtained in step S321 through multiple band-pass filters (of triangular or cosine windows) respectively to obtain the energy spectrum on the Mel scale. Each band-pass filter corresponds to a Mel Frequency in units of the Mel scale, and the Mel frequencies corresponding to these band-pass filters are different. Therefore, the energy spectrum on the Mel scale represents the power / energy corresponding to multiple Mel frequencies.
[0055] For example, Figure 4 is a schematic diagram of the Mel-frequency mapping according to an embodiment of the present invention. Please refer to Figure 4 , |X[k]|2 is the power on the spectrum obtained through step S321. k is the identification number (e.g., bin number) of the frequency unit (bin) used for fast / discrete Fourier transform, and k = 1 to N (N is a positive integer). Y[m] is the energy spectrum on the Mel scale. m is the identification number (e.g., bank number) of the group unit of the band-pass filter BPF (e.g., Mel filter), and m = 1 to M (M is a positive integer). Each group (bank) corresponds to a band-pass filter BPF and also corresponds to a Mel frequency. The Mel frequencies corresponding to these band-pass filters BPF are different.
[0056] The formula for energy spectrum conversion is:
[0057]
[0058] where W[k] is the weight of the band-pass filter BPF. Y t [m] is the power on the energy spectrum on the Mel scale and corresponding to the group number m of the Mel frequency. The power Y[m] is one of Y t [1] to Y t [M].
[0059] The logarithmic operation (step S323) is that the processor 130 takes the logarithm of the power at each Mel frequency of the energy spectrum on the Mel scale to obtain the log energy or the logarithm of the power.
[0060] The discrete cosine operation (step S324) is that the processor 130 brings the log energy corresponding to multiple Mel frequencies into the discrete cosine transform (DCT) to obtain the Mel-scale cepstral coefficients of order L, and L is, for example, 12 (but not limited thereto).
[0061] The dynamic feature extraction (step S325) is that the processor 130 superimposes the frame energy obtained from the preprocessing in step S310 and the cepstral coefficients obtained from the discrete cosine operation to obtain the Mel cepstral coefficients
[0062] Please refer to Figure 2 , the processor 130 obtains the power corresponding to multiple Mel frequencies of the sound signal to be processed through multiple band-pass filters (step S211). Specifically, as described in the aforementioned Figure 3 step S322 and Figure 4 description, the power corresponding to each Mel frequency is converted from the power of the sound signal to be processed on the spectrum into the power corresponding to the Mel frequency in Mel scale units through its corresponding band-pass filter (such as the Figure 4 band-pass filter BPF).
[0063] The processor 130 maps the first frequency among multiple Mel frequencies to the second frequency among these Mel frequencies, and replaces the power corresponding to the first frequency with the power corresponding to the second frequency (step S212). Specifically, the second frequency is lower than the first frequency. That is to say, the first frequency is mapped to the lower second frequency. The power of the second frequency will be directly changed to the power of the first frequency. Similarly, other frequencies among these Mel frequencies (for example, the third frequency different from the first frequency) can also be mapped to a still lower frequency (for example, the fourth frequency different from the second frequency), and accordingly, the power of the higher frequency is directly replaced with the power of the lower frequency. The above-mentioned "frequency mapping" and "power replacement" will be collectively referred to as "frequency shift" in the context.
[0064] In one embodiment, the processor 130 can shift the first frequency to the second frequency according to the displacement amount. The unit of this displacement amount corresponds to the Mel scale. Figure 4 For example, the group number (m) is from 1 to M. The higher the group number, the higher the frequency; the lower the group number, the lower the frequency. The group number corresponding to the first frequency is greater than the group number corresponding to the second frequency. The formula for frequency shift is:
[0065]
[0066] , where Y FS [j] is the power of the Mel frequency with group number j after frequency shift, Y t [j] is the power corresponding to the Mel frequency with group number j (as shown in formula (2)), Δ FS is the displacement amount, f is the frequency (unit: Hertz) of the Mel frequency with group number j after conversion as shown in formula (1), F1 and F2 are respectively the lower limit and the upper limit of the target frequency band. The target frequency band refers to the frequency band where frequency shift is desired. For example, 2 kHz (corresponding to F1) to 4 kHz (corresponding to F2) corresponding to the high-frequency aspirated sound in speech, but not limited thereto. That is to say, frequency shift is performed on the power corresponding to the group number of the target frequency band, the power corresponding to the frequency below the target frequency band remains unchanged (i.e., no frequency shift), and the power corresponding to the frequency above the target frequency band remains unchanged (i.e., no frequency shift) or is zeroed (i.e., filtered).
[0067] For example, the displacement amount Δ FS is 3. Assume that the first frequency is within the target frequency band, and the group number corresponding to the first frequency is M, then the group number corresponding to the second frequency is M - 3. Therefore, Y FS [M - 3] = Y t [(M - 3) + 3] = Y t [M].
[0068] Similarly, the processor 130 can shift a third frequency among multiple Mel frequencies to a fourth frequency among these Mel frequencies according to the (same) displacement amount, and the fourth frequency is lower than the third frequency. For example, assuming that the group number corresponding to the third frequency is M-1, then the group number corresponding to the fourth frequency is M-4. Therefore, Y FS [M-4]=Y t [(M-4)+3]=Y t [M-1].
[0069] In another embodiment, the displacement amounts of multiple Mel frequencies may be different. For example, the displacement amount of the first frequency is 3, and the displacement amount of the third frequency is 4, but not limited thereto.
[0070] Please refer to Figure 2 , the processor 130 generates multiple Mel cepstral coefficients using the powers corresponding to multiple Mel frequencies (step S213). Specifically, as Figure 3 the description of steps S323 to S325 in
[0071] performs logarithmic operation, discrete cosine operation, and dynamic feature extraction on the powers corresponding to the frequency-shifted Mel frequencies, and then Mel cepstral coefficients can be generated.
[0072] Next, the processor 130 generates a synthesized sound signal using multiple Mel cepstral coefficients of the sound signal to be processed (step S220). Specifically, the Mel cepstral coefficients belong to a kind of sound feature. The processor 130 can synthesize a sound signal according to the sound feature.
[0073] Figure 5 is a schematic diagram of machine learning training according to an embodiment of the present invention. Please refer to Figure 5, the MelGAN is a non-autoregressive convolutional neural network architecture that uses a Generative Adversarial Network (GAN) to invert Mel spectrograms back to waveforms, with a faster speed and smaller size than the previously used WaveNet while still maintaining computational efficiency. The MelGAN includes a Discriminator DS and a Generator GS. The Discriminator DS is a deep neural network composed of multiple convolutional layers and residual blocks. The convolutional layers of the Discriminator DS are used to extract sound features from the audio waveform, while the residual blocks are used to improve the ability of the Discriminator DS to distinguish between real and synthetic audio waveforms. The Generator GS is also a deep neural network composed of multiple convolutional layers and residual blocks. The convolutional layers of the Generator GS are used to extract features from the Mel spectrogram, while the residual blocks are used to improve the quality of the synthesized audio waveform.
[0074] By inputting a base sound signal z and Mel cepstral coefficients s into the Generator GS, an estimated sound signal G(s,z) (e.g., a frequency-shifted sound signal) can be output / generated. The base sound signal z can be a white noise signal, a brown noise signal, a pink noise signal, or other sound signals. The Mel cepstral coefficients s are the Mel cepstral coefficients of the downsampled sound signal x. The downsampled sound signal x is obtained by downsampling the training sound signal. For example, in Patent TW I557729, the frequency is reduced to one-fourth or one-half of the sampled speech signal (e.g., the training sound signal). Alternatively, the frequency of the training sound signal can be reduced by other downsampling algorithms. The training sound signal is a speech signal. For example, a speech signal generated by recording, receiving, or software editing.
[0075] On the other hand, during the training of the MelGAN, by inputting the downsampled sound signal x and the estimated sound signal G(s,z) into the Discriminator DS, the Discriminator DS can use the downsampled sound signal x to determine the authenticity of the estimated sound signal G(s,z) generated by the Generator GS. That is, to determine whether the estimated sound signal G(s,z) is the downsampled sound signal x. The Generator GS and the Discriminator DS continuously compete with each other, and based on this, the parameters of the Generator GS are updated. The trained Generator GS can be used to input Mel cepstral coefficients y and output a frequency-shifted estimated sound signal G(y,z). For example, by inputting multiple Mel cepstral coefficients of the sound signal to be processed into the trained Generator GS, a synthesized sound signal (i.e., a frequency-shifted sound signal) can be output.
[0076] In other embodiments, a synthesized sound signal can be generated by other neural networks that have been trained to know the relationship between Mel cepstral coefficients and the synthesized sound signal.
[0077] In one embodiment, the synthesized sound signal is used to be played through the speaker 110. That is, the speaker 110 can play the synthesized sound signal. Alternatively, the processing device 10 can transmit the synthesized sound signal to other devices and play the synthesized sound signal through other devices.
[0078] Figure 6 is a flowchart of a method for processing a sound signal according to an embodiment of the present invention. Please refer to Figure 6 , the processor 130 performs band-pass filtering on the initial sound signal S in to output the sound signal S to be processed B1 and the bypass sound signal S B2 . The initial sound signal S in can be a voice signal generated by recording, receiving or editing. Step S610 and step S615 correspond to different frequency bands respectively. The sound signal S to be processed B1 corresponds to the first frequency band. For example, it corresponds to 4 kHz to 8 kHz of the aspirated sound in speech, but is not limited thereto. The bypass sound signal S B2 corresponds to the second frequency band. For example, it does not correspond to below 4 kHz of the aspirated sound, but is not limited thereto. The first frequency band is different from the second frequency band.
[0079] The processor 130 performs frequency shift processing on the sound signal S to be processed B1 and generates the frequency-shifted sound signal S to be processed accordingly B1 corresponding mel-frequency cepstral coefficients (step S620). For a detailed description of step S620, please refer to the description of the aforementioned step S210, which will not be repeated here.
[0080] The processor 130 generates a synthesized sound signal according to the mel-frequency cepstral coefficients (step S630). For a detailed description of step S630, please refer to the description of the aforementioned step S220, which will not be repeated here.
[0081] The processor 130 combines the synthesized sound signal and the bypass sound signal S B2 into an output sound signal S out (step S640). For example, the two sound signals are superimposed in the frequency domain or the time domain. The output sound signal S out is used to be played through the speaker 110. That is, the speaker 110 can play the output sound signal S out . Alternatively, the processing device 10 can transmit the output sound signal S out to other devices and play the output sound signal S out .
[0082] In an application scenario, a to-be-processed sound signal played in a video conference, call, or conversation can be converted into an output sound signal S out , enabling hearing-impaired persons to distinguish the complete semantics without being unable to understand the semantics due to the aspiration in the speech sound.
[0083] In summary, in the sound signal processing method and processing device according to the embodiments of the present invention, during the process of extracting Mel cepstral coefficients, the power of the Mel spectrum is frequency-shifted, and a synthesized sound signal is generated based on the frequency-shifted Mel cepstral coefficients. Other frequency-shifting algorithms may elongate the sound signal during the process, but finally only the length before elongation can be retained, and some features will be ignored. However, the embodiments of the present invention can retain the complete sound features and also achieve the effect of frequency shifting.
[0084] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features. And these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for processing a voice signal, characterized in that, Including: Extracting a plurality of mel cepstral coefficients from a sound signal to be processed, including: Obtaining the power of the sound signal to be processed corresponding to a plurality of mel frequencies through a plurality of band-pass filters, where each of the band-pass filters corresponds to one of the mel frequencies, and the mel frequencies corresponding to the band-pass filters are different; Mapping a first frequency among the mel frequencies to a second frequency among the mel frequencies, and replacing the power corresponding to the second frequency with the power corresponding to the first frequency, where the second frequency is lower than the first frequency; and Generating the mel cepstral coefficients using the power corresponding to the mel frequencies; and Generating a synthesized sound signal using the mel cepstral coefficients of the sound signal to be processed, where the sound signal to be processed and the synthesized sound signal are played by a speaker.
2. The method for processing a sound signal according to claim 1, wherein the step of mapping the first frequency among the mel frequencies to the second frequency among the mel frequencies includes: Displacing the first frequency to the second frequency according to a displacement amount, where the unit of the displacement amount corresponds to a mel scale; And Displacing a third frequency among the mel frequencies to a fourth frequency among the mel frequencies according to the displacement amount, where the fourth frequency is lower than the third frequency.
3. The method for processing a sound signal according to claim 1, wherein before the step of obtaining the mel cepstral coefficients of the sound signal to be processed, further included is: Performing band-pass filtering on an initial sound signal to output the sound signal to be processed and a bypass sound signal, where the sound signal to be processed corresponds to a first frequency band, the bypass sound signal corresponds to a second frequency band, and the first frequency band is different from the second frequency band.
4. The method for processing a sound signal according to claim 3, wherein after the step of generating the synthesized sound signal using the mel cepstral coefficients of the sound signal to be processed, further included is: Combining the synthesized sound signal and the bypass sound signal into an output sound signal, where the output sound signal is used to be played by the speaker.
5. The method for processing a sound signal according to claim 1, wherein the step of generating the synthesized sound signal using the mel cepstral coefficients of the sound signal to be processed includes: Inputting the mel cepstral coefficients of the sound signal to be processed into a mel generative adversarial network to output the synthesized sound signal, where During the training of the mel generative adversarial network, a discriminator uses a downsampled sound signal to determine the authenticity of an estimated sound signal generated by a generator, where the downsampled sound signal is obtained by downsampling a training sound signal.
6. A processing device for a voice signal, characterized in that, Including: A memory for storing program code; And A processor coupled to the memory and configured to load the program code to execute: Extracting a plurality of mel cepstral coefficients from a sound signal to be processed, where the processor is configured to: Obtain the power of the sound signal to be processed corresponding to multiple Mel frequencies through multiple band-pass filters, where each of the band-pass filters corresponds to one of the Mel frequencies, and the Mel frequencies corresponding to the band-pass filters are different; Map a first frequency among the Mel frequencies to a second frequency among the Mel frequencies, and replace the power corresponding to the second frequency with the power corresponding to the first frequency, where the second frequency is lower than the first frequency; and Generate the Mel cepstrum coefficients using the power corresponding to the Mel frequencies; and Generate a synthesized sound signal using the Mel cepstrum coefficients of the sound signal to be processed, where the sound signal to be processed and the synthesized sound signal are for playing through a speaker.
7. The sound signal processing device according to claim 6, wherein the step of mapping the first frequency among the Mel frequencies to the second frequency among the Mel frequencies includes: Displace the first frequency to the second frequency according to a displacement amount, where the unit of the displacement amount corresponds to the Mel scale; And Displace a third frequency among the Mel frequencies to a fourth frequency among the Mel frequencies according to the displacement amount, where the fourth frequency is lower than the third frequency.
8. The sound signal processing device according to claim 6, wherein before the step of obtaining the Mel cepstrum coefficients of the sound signal to be processed, further includes: Perform band-pass filtering on an initial sound signal to output the sound signal to be processed and a bypass sound signal, where the sound signal to be processed corresponds to a first frequency band, the bypass sound signal corresponds to a second frequency band, and the first frequency band is different from the second frequency band.
9. The sound signal processing device according to claim 8, wherein after the step of generating the synthesized sound signal using the Mel cepstrum coefficients of the sound signal to be processed, further includes: Combine the synthesized sound signal and the bypass sound signal into an output sound signal, where the output sound signal is for playing through the speaker.
10. The sound signal processing device according to claim 6, wherein the step of generating the synthesized sound signal using the Mel cepstrum coefficients of the sound signal to be processed includes: Input the Mel cepstrum coefficients of the sound signal to be processed into a Mel generative adversarial network to output the synthesized sound signal, where During the training of the Mel generative adversarial network, a discriminator uses a downsampled sound signal to determine the authenticity of the estimated sound signal generated by a generator, where the downsampled sound signal is obtained by downsampling a training sound signal.