Speech coding, speech decoding methods, devices, computer equipment, and storage media
By performing frequency band division and feature compression of voice signals, the problem of sampling rate limitation in traditional voice encoding and decoding technologies is solved, and the effect of flexible adjustment and improvement of sound quality is achieved.
Patent Information
- Application Number
- CN202110693160.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-06-22
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2041-06-22
AI Technical Summary
In traditional voice encoding and decoding technology, the acquisition and playback of voice signals are limited by the sampling rate supported by the voice encoder and decoder, resulting in high limitations in the sampling rate and the inability to freely adjust and improve the sound quality.
By obtaining the initial band characteristic information of the voice signal, band division and feature compression are performed, the sampling rate is reduced to the range supported by the voice encoder, and band expansion is performed during decoding, so as to achieve flexible adjustment of the sampling rate.
It realizes that without changing the existing system, improve the sampling rate and sound quality of voice signals, enrich the playback of voice information, and reduce operational costs.
Smart Images

Figure CN115512711B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technologies, and particularly to a method and apparatus for voice encoding and decoding, a computer device, and a storage medium. Background Art
[0002] With the development of computer technologies, voice encoding and decoding technologies have emerged. Voice encoding and decoding technologies can be applied to voice storage and voice transmission.
[0003] In traditional technologies, a voice acquisition device needs to be used in conjunction with a voice encoder, and the sampling rate of the voice acquisition device needs to be within the range supported by the voice encoder. In this way, the voice signal collected by the voice acquisition device can be encoded by the voice encoder for storage or transmission. In addition, the playback of the voice signal also depends on the voice decoder. The voice encoder can only decode and play the voice signal whose sampling rate is within the range supported by itself. Therefore, it can only play the voice signal whose sampling rate is within the range supported by the voice encoder.
[0004] However, in traditional methods, the acquisition of voice signals is restricted by the sampling rate supported by the existing voice encoder, and the playback of voice signals is also restricted by the sampling rate supported by the existing voice decoder, with relatively large limitations. Summary of the Invention
[0005] Based on this, in view of the above technical problems, it is necessary to provide a method and apparatus for voice encoding and decoding, a computer device, and a storage medium, in which the acquisition and playback of voice signals are not restricted by the sampling rate supported by the voice encoder.
[0006] A voice encoding method, the method comprising:
[0007] Obtaining initial band feature information corresponding to a voice signal to be processed;
[0008] Obtaining target feature information corresponding to a first frequency band based on the initial feature information corresponding to the first frequency band in the initial band feature information;
[0009] Performing feature compression on the initial feature information corresponding to a second frequency band in the initial band feature information to obtain target feature information corresponding to a compressed frequency band, where the frequency of the first frequency band is less than the frequency of the second frequency band, and the frequency range of the second frequency band is greater than the frequency range of the compressed frequency band;
[0010] Obtaining intermediate band feature information based on the target feature information corresponding to the first frequency band and the target feature information corresponding to the compressed frequency band, and obtaining a compressed voice signal corresponding to the voice signal to be processed based on the intermediate band feature information;
[0011] The compressed speech signal is encoded by a speech encoding module to obtain encoded speech data corresponding to the speech signal to be processed. The target sampling rate corresponding to the compressed speech signal is less than or equal to the supported sampling rate corresponding to the speech encoding module, and the target sampling rate is less than the sampling rate corresponding to the speech signal to be processed.
[0012] A speech encoding device, the device comprising:
[0013] A frequency band feature information acquisition module, configured to acquire initial frequency band feature information corresponding to a speech signal to be processed;
[0014] A first target feature information determination module, configured to obtain target feature information corresponding to a first frequency band based on initial feature information corresponding to the first frequency band in the initial frequency band feature information;
[0015] A second target feature information determination module, configured to perform feature compression on initial feature information corresponding to a second frequency band in the initial frequency band feature information to obtain target feature information corresponding to a compressed frequency band. The frequency of the first frequency band is less than the frequency of the second frequency band, and the frequency range of the second frequency band is greater than the frequency range of the compressed frequency band;
[0016] A compressed speech signal generation module, configured to obtain intermediate frequency band feature information based on the target feature information corresponding to the first frequency band and the target feature information corresponding to the compressed frequency band, and obtain a compressed speech signal corresponding to the speech signal to be processed based on the intermediate frequency band feature information;
[0017] A speech signal encoding module, configured to encode the compressed speech signal through a speech encoding module to obtain encoded speech data corresponding to the speech signal to be processed. The target sampling rate corresponding to the compressed speech signal is less than or equal to the supported sampling rate corresponding to the speech encoding module, and the target sampling rate is less than the sampling rate corresponding to the speech signal to be processed.
[0018] A computer device, comprising a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0019] Obtain initial frequency band feature information corresponding to a speech signal to be processed;
[0020] Obtain target feature information corresponding to a first frequency band based on initial feature information corresponding to the first frequency band in the initial frequency band feature information;
[0021] Perform feature compression on the initial feature information corresponding to the second frequency band in the initial frequency band feature information to obtain target feature information corresponding to the compressed frequency band. The frequency of the first frequency band is less than the frequency of the second frequency band, and the frequency range of the second frequency band is greater than the frequency range of the compressed frequency band;
[0022] Obtain intermediate frequency band feature information based on the target feature information corresponding to the first frequency band and the target feature information corresponding to the compressed frequency band, and obtain a compressed speech signal corresponding to the speech signal to be processed based on the intermediate frequency band feature information;
[0023] Perform encoding processing on the compressed speech signal through a speech encoding module to obtain encoded speech data corresponding to the speech signal to be processed. The target sampling rate of the compressed speech signal is less than or equal to the supported sampling rate corresponding to the speech encoding module, and the target sampling rate is less than the sampling rate corresponding to the speech signal to be processed.
[0024] A computer-readable storage medium stores a computer program thereon. When the computer program is executed by a processor, the following steps are implemented:
[0025] Obtain initial frequency band feature information corresponding to a speech signal to be processed;
[0026] Obtain target feature information corresponding to the first frequency band based on the initial feature information corresponding to the first frequency band in the initial frequency band feature information;
[0027] Perform feature compression on the initial feature information corresponding to the second frequency band in the initial frequency band feature information to obtain target feature information corresponding to the compressed frequency band. The frequency of the first frequency band is less than the frequency of the second frequency band, and the frequency range of the second frequency band is greater than the frequency range of the compressed frequency band;
[0028] Obtain intermediate frequency band feature information based on the target feature information corresponding to the first frequency band and the target feature information corresponding to the compressed frequency band, and obtain a compressed speech signal corresponding to the speech signal to be processed based on the intermediate frequency band feature information;
[0029] Perform encoding processing on the compressed speech signal through a speech encoding module to obtain encoded speech data corresponding to the speech signal to be processed. The target sampling rate of the compressed speech signal is less than or equal to the supported sampling rate corresponding to the speech encoding module, and the target sampling rate is less than the sampling rate corresponding to the speech signal to be processed.
[0030] The above voice encoding method, device, computer device, and storage medium obtain the initial frequency band feature information corresponding to the voice signal to be processed, obtain the target feature information corresponding to the first frequency band based on the initial feature information corresponding to the first frequency band in the initial frequency band feature information, perform feature compression on the initial feature information corresponding to the second frequency band in the initial frequency band feature information to obtain the target feature information corresponding to the compressed frequency band. The frequency of the first frequency band is less than the frequency of the second frequency band, and the frequency range of the second frequency band is greater than the frequency range of the compressed frequency band. The intermediate frequency band feature information is obtained based on the target feature information corresponding to the first frequency band and the target feature information corresponding to the compressed frequency band, and the compressed voice signal corresponding to the voice signal to be processed is obtained based on the intermediate frequency band feature information. The compressed voice signal is encoded by a voice encoding module to obtain the encoded voice data corresponding to the voice signal to be processed. The target sampling rate of the compressed voice signal is less than or equal to the supported sampling rate of the voice encoding module, and the target sampling rate is less than the sampling rate of the voice signal to be processed. In this way, before voice encoding, the voice signal to be processed with any sampling rate can be compressed through the frequency band feature information, and the sampling rate of the voice signal to be processed can be reduced to the sampling rate supported by the voice encoder to obtain a compressed voice signal with a low sampling rate. Since the sampling rate of the compressed voice signal is less than or equal to the sampling rate supported by the voice encoder, the compressed voice signal can be successfully encoded by the voice encoder.
[0031] A voice decoding method, the method comprising:
[0032] Obtain encoded voice data, where the encoded voice data is obtained by performing voice compression processing on a voice signal to be processed;
[0033] Decode the encoded voice data through a voice decoding module to obtain a decoded voice signal, where the target sampling rate of the decoded voice signal is less than or equal to the supported sampling rate of the voice decoding module;
[0034] Generate the target frequency band feature information corresponding to the decoded voice signal, and obtain the extended feature information corresponding to the first frequency band based on the target feature information corresponding to the first frequency band in the target frequency band feature information;
[0035] Perform feature extension on the target feature information corresponding to the compressed frequency band in the target frequency band feature information to obtain the extended feature information corresponding to the second frequency band; the frequency of the first frequency band is less than the frequency of the compressed frequency band, and the frequency range of the compressed frequency band is less than the frequency range of the second frequency band;
[0036] Obtain extended frequency band feature information based on the extended feature information corresponding to the first frequency band and the extended feature information corresponding to the second frequency band, and obtain a target voice signal corresponding to the voice signal to be processed based on the extended frequency band feature information, where the sampling rate of the target voice signal is greater than the target sampling rate;
[0037] Play the target voice signal.
[0038] A voice decoding device, the device includes:
[0039] A voice data acquisition module, configured to acquire encoded voice data, where the encoded voice data is obtained by performing voice compression processing on a voice signal to be processed;
[0040] A voice signal decoding module, configured to perform decoding processing on the encoded voice data through a voice decoding module to obtain a decoded voice signal, where the target sampling rate of the decoded voice signal is less than or equal to the supported sampling rate corresponding to the voice decoding module;
[0041] A first extended feature information determination module, configured to generate target frequency band feature information corresponding to the decoded voice signal, and obtain extended feature information corresponding to the first frequency band based on the target feature information corresponding to the first frequency band in the target frequency band feature information;
[0042] A second extended feature information determination module, configured to perform feature extension on the target feature information corresponding to the compressed frequency band in the target frequency band feature information to obtain extended feature information corresponding to the second frequency band; the frequency of the first frequency band is less than the frequency of the compressed frequency band, and the frequency range of the compressed frequency band is less than the frequency range of the second frequency band;
[0043] A target voice signal determination module, configured to obtain extended frequency band feature information based on the extended feature information corresponding to the first frequency band and the extended feature information corresponding to the second frequency band, and obtain a target voice signal corresponding to the voice signal to be processed based on the extended frequency band feature information, where the sampling rate of the target voice signal is greater than the target sampling rate;
[0044] A voice signal playback module, configured to play the target voice signal.
[0045] A computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0046] Obtain encoded voice data, where the encoded voice data is obtained by performing voice compression processing on a voice signal to be processed;
[0047] The encoded speech data is decoded by a speech decoding module to obtain a decoded speech signal, and the target sampling rate corresponding to the decoded speech signal is less than or equal to the supported sampling rate corresponding to the speech decoding module;
[0048] Generate target band feature information corresponding to the decoded speech signal, and obtain extended feature information corresponding to the first frequency band based on the target feature information corresponding to the first frequency band in the target band feature information;
[0049] Perform feature expansion on the target feature information corresponding to the compressed frequency band in the target band feature information to obtain extended feature information corresponding to the second frequency band; the frequency of the first frequency band is less than the frequency of the compressed frequency band, and the frequency range of the compressed frequency band is less than the frequency range of the second frequency band;
[0050] Obtain extended band feature information based on the extended feature information corresponding to the first frequency band and the extended feature information corresponding to the second frequency band, and obtain a target speech signal corresponding to the speech signal to be processed based on the extended band feature information, and the sampling rate of the target speech signal is greater than the target sampling rate;
[0051] Play the target speech signal.
[0052] A computer-readable storage medium stores a computer program thereon, and when the computer program is executed by a processor, the following steps are implemented:
[0053] Obtain encoded speech data, where the encoded speech data is obtained by performing speech compression processing on a speech signal to be processed;
[0054] The encoded speech data is decoded by a speech decoding module to obtain a decoded speech signal, and the target sampling rate corresponding to the decoded speech signal is less than or equal to the supported sampling rate corresponding to the speech decoding module;
[0055] Generate target band feature information corresponding to the decoded speech signal, and obtain extended feature information corresponding to the first frequency band based on the target feature information corresponding to the first frequency band in the target band feature information;
[0056] Perform feature expansion on the target feature information corresponding to the compressed frequency band in the target band feature information to obtain extended feature information corresponding to the second frequency band; the frequency of the first frequency band is less than the frequency of the compressed frequency band, and the frequency range of the compressed frequency band is less than the frequency range of the second frequency band;
[0057] Obtain extended frequency band feature information based on the extended feature information corresponding to the first frequency band and the extended feature information corresponding to the second frequency band, and obtain a target voice signal corresponding to the voice signal to be processed based on the extended frequency band feature information, where the sampling rate of the target voice signal is greater than the target sampling rate;
[0058] Play the target voice signal.
[0059] The above voice decoding method, device, computer device and storage medium obtain encoded voice data by acquiring the encoded voice data, which is obtained by performing voice compression processing on the voice signal to be processed. The encoded voice data is decoded by a voice decoding module to obtain a decoded voice signal. The target sampling rate corresponding to the decoded voice signal is less than or equal to the supported sampling rate corresponding to the voice decoding module, and generate target frequency band feature information corresponding to the decoded voice signal. Based on the target feature information corresponding to the first frequency band in the target frequency band feature information, obtain the extended feature information corresponding to the first frequency band, and perform feature extension on the target feature information corresponding to the compressed frequency band in the target frequency band feature information to obtain the extended feature information corresponding to the second frequency band; the frequency of the first frequency band is less than the frequency of the compressed frequency band, and the frequency range of the compressed frequency band is less than the frequency range of the second frequency band. Obtain extended frequency band feature information based on the extended feature information corresponding to the first frequency band and the extended feature information corresponding to the second frequency band, and obtain a target voice signal corresponding to the voice signal to be processed based on the extended frequency band feature information, where the sampling rate of the target voice signal is greater than the target sampling rate, and play the target voice signal. In this way, after obtaining the encoded voice data obtained by voice compression processing, the encoded voice data can be decoded to obtain a decoded voice signal. Through the extension of the frequency band feature information, the sampling rate of the decoded voice signal can be increased to obtain a target voice signal and play it. The playback of the voice signal is not restricted by the sampling rate supported by the voice decoder. When playing the voice, a high sampling rate voice signal with richer information can also be played. Description of the Drawings
[0060] Figure 1 It is an application environment diagram of a voice encoding and voice decoding method in an embodiment;
[0061] Figure 2 It is a flowchart of a voice encoding method in an embodiment;
[0062] Figure 3 It is a flowchart of obtaining target feature information by performing feature compression on initial feature information in an embodiment;
[0063] Figure 4 It is a schematic diagram of the mapping relationship between an initial sub-frequency band and a target sub-frequency band in an embodiment;
[0064] Figure 5Schematic flowchart of a voice decoding method in an embodiment;
[0065] Figure 6A Schematic flowchart of a voice encoding and decoding method in an embodiment;
[0066] Figure 6B Schematic diagram of frequency domain signals before and after compression in an embodiment;
[0067] Figure 6C Schematic diagram of voice signals before and after compression in an embodiment;
[0068] Figure 6D Schematic diagram of frequency domain signals before and after expansion in an embodiment;
[0069] Figure 6E Schematic diagram of voice signals to be processed and target voice signals in an embodiment;
[0070] Figure 7A Structural block diagram of a voice encoding device in an embodiment;
[0071] Figure 7B Structural block diagram of a voice encoding device in another embodiment;
[0072] Figure 8 Structural block diagram of a voice decoding device in an embodiment;
[0073] Figure 9 Internal structure diagram of a computer device in an embodiment. Detailed implementation manners
[0074] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0075] The voice encoding and voice decoding methods provided by the present application can be applied to an application environment as shown in Figure 1 In the figure. Among them, the voice sending end 102 communicates with the voice receiving end 104 through a network. The voice sending end 102 and the voice receiving end 104 can be terminals, and the terminals can be but are not limited to various personal computers, laptop computers, smart phones, tablet computers, and portable wearable devices.
[0076] Specifically, the voice sending end acquires the initial frequency band feature information corresponding to the voice signal to be processed. The voice sending end can obtain the target feature information corresponding to the first frequency band based on the initial feature information corresponding to the first frequency band in the initial frequency band feature information, and perform feature compression on the initial feature information corresponding to the second frequency band in the initial frequency band feature information to obtain the target feature information corresponding to the compressed frequency band. Among them, the frequency of the first frequency band is less than the frequency of the second frequency band, and the frequency range of the second frequency band is greater than the frequency range of the compressed frequency band. The voice sending end obtains the intermediate frequency band feature information based on the target feature information corresponding to the first frequency band and the target feature information corresponding to the compressed frequency band, obtains the compressed voice signal corresponding to the voice signal to be processed based on the intermediate frequency band feature information, and performs encoding processing on the compressed voice signal through the voice encoding module to obtain the encoded voice data corresponding to the voice signal to be processed. Among them, the target sampling rate of the compressed voice signal is less than or equal to the supported sampling rate corresponding to the voice encoding module, and the target sampling rate is less than the sampling rate corresponding to the voice signal to be processed. The voice sending end can send the encoded voice data to the voice receiving end so that the voice receiving end performs voice restoration processing on the encoded voice data to obtain the target voice signal corresponding to the voice signal to be processed and play the target voice signal. The voice sending end can also store the encoded voice data locally. When playback is needed, the voice sending end performs voice restoration processing on the encoded voice data to obtain the target voice signal corresponding to the voice signal to be processed and play the target voice signal.
[0077] The voice receiving end acquires the encoded voice data and performs decoding processing on the encoded voice data through the voice decoding module to obtain the decoded voice signal. Among them, the encoded voice data can be sent by the voice sending end or obtained by the voice receiving end through local voice compression processing on the voice signal to be processed. The voice receiving end generates the target frequency band feature information corresponding to the decoded voice signal, obtains the extended feature information corresponding to the first frequency band based on the target feature information corresponding to the first frequency band in the target frequency band feature information corresponding to the decoded voice signal, and performs feature extension on the target feature information corresponding to the compressed frequency band in the target frequency band feature information to obtain the extended feature information corresponding to the second frequency band. Among them, the frequency of the first frequency band is less than the frequency of the compressed frequency band, and the frequency range of the compressed frequency band is less than the frequency range of the second frequency band. The voice receiving end obtains the extended frequency band feature information based on the extended feature information corresponding to the first frequency band and the extended feature information corresponding to the second frequency band, obtains the target voice signal corresponding to the voice signal to be processed based on the extended frequency band feature information, and the sampling rate of the target voice signal is greater than the target sampling rate of the decoded voice signal. Finally, the voice receiving end plays the target voice signal.
[0078] It can be understood that during the transmission of encoded voice data, the encoded voice data may pass through a server. The server can be implemented by an independent server, a server cluster composed of multiple servers, or a cloud server. The voice receiving end and the voice sending end can be mutually converted, that is, the voice receiving end can also be used as the voice sending end, and the voice sending end can also be used as the voice receiving end.
[0079] In one embodiment, as Figure 2 shown, a voice encoding method is provided. Taking the application of this method to a terminal as an example for illustration, the terminal can be Figure 1 the voice sending end in
[0080] Step S202, obtain the initial frequency band feature information corresponding to the voice signal to be processed.
[0081] Among them, the voice signal to be processed refers to the voice signal collected by the voice acquisition device on the terminal, which is the voice signal to be played. The voice signal to be processed can be the voice signal collected in real time by the voice acquisition device. The terminal can perform frequency band compression and encoding processing on the latest collected voice signal in real time to obtain encoded voice data. The voice signal to be processed can also be the voice signal collected by the voice acquisition device in the past. The voice sending end can obtain the voice signal collected at a historical time from the database as the voice signal to be processed, perform frequency band compression and encoding processing on the voice signal to be processed, and obtain encoded voice data. The terminal can store the encoded voice data and, when needed for playback, decode and play the encoded voice data. If the terminal is the voice sending end, the terminal can also send the encoded voice signal to the voice receiving end, and the voice receiving end decodes and plays the encoded voice data. Send the processed voice signal to the voice receiving end. The voice signal to be processed is a time-domain signal, which can reflect the change of the voice signal over time.
[0082] Frequency band compression can reduce the sampling rate of the voice signal while keeping the voice content intelligible. Frequency band compression refers to compressing a voice signal with a large frequency band into a voice signal with a small frequency band, where the voice signal with the small frequency band and the voice signal with the large frequency band have the same low-frequency information.
[0083] The initial frequency band feature information refers to the feature information of the voice signal to be processed in the frequency domain. The feature information of the voice signal in the frequency domain includes the amplitudes and phases of multiple frequency points within a frequency bandwidth (i.e., frequency band). One frequency point represents a specific frequency. According to the Shannon theorem, the sampling rate and the frequency band of the voice signal are in a two-fold relationship. For example, if the sampling rate of the voice signal is 48 kHz, then the frequency band of this voice signal is 24 kHz, specifically 0 - 24 kHz; if the sampling rate of the voice signal is 16 kHz, then the frequency band of this voice signal is 8 kHz, specifically 0 - 8 kHz.
[0084] Specifically, the terminal can use the voice signal collected by the local voice acquisition device as the voice signal to be processed, and extract the frequency-domain features of the voice signal to be processed locally as the initial frequency-band feature information corresponding to the voice signal to be processed. Among them, the terminal can use a time-domain to frequency-domain conversion algorithm to convert the time-domain signal into a frequency-domain signal, so as to extract the frequency-domain features of the voice signal to be processed. For example, a custom time-domain to frequency-domain conversion algorithm, Laplace transform algorithm, Z-transform algorithm, Fourier transform algorithm, etc.
[0085] Step S204: Obtain the target feature information corresponding to the first frequency band based on the initial feature information corresponding to the first frequency band in the initial frequency-band feature information.
[0086] Among them, a frequency band is a frequency interval composed of some frequencies in a frequency band. A frequency band can be composed of at least one frequency band. The initial frequency band corresponding to the voice signal to be processed includes a first frequency band and a second frequency band, and the frequency of the first frequency band is less than the frequency of the second frequency band. The terminal can divide the initial frequency-band feature information into the initial feature information corresponding to the first frequency band and the initial feature information corresponding to the second frequency band. That is, the initial frequency-band feature information can be divided into the initial feature information corresponding to the low-frequency band and the initial feature information corresponding to the high-frequency band. The initial feature information corresponding to the low-frequency band mainly determines the content information of the voice. For example, the specific semantic content "what time do you get off work", and the initial feature information corresponding to the high-frequency band mainly determines the texture of the voice. For example, a hoarse and deep voice.
[0087] The initial feature information refers to the feature information corresponding to each frequency before frequency-band compression, and the target feature information refers to the feature information corresponding to each frequency after frequency-band compression.
[0088] Specifically, if the sampling rate of the speech signal to be processed is higher than the sampling rate supported by the speech encoder, it is impossible to directly encode the speech signal to be processed through the speech encoder. Therefore, it is necessary to perform band compression on the speech signal to be processed to reduce the sampling rate of the speech signal to be processed. When performing band compression, in addition to reducing the sampling rate of the speech signal to be processed, it is also necessary to ensure that the semantic content remains unchanged and is naturally intelligible. Since the semantic content of speech depends on the low-frequency information in the speech signal, the terminal can divide the initial band feature information into the initial feature information corresponding to the first frequency band and the initial feature information corresponding to the second frequency band. The initial feature information corresponding to the first frequency band is the low-frequency information in the speech signal to be processed, and the initial feature information corresponding to the second frequency band is the high-frequency information in the speech signal to be processed. To ensure the intelligibility and readability of speech, when performing band compression, the terminal can keep the low-frequency information unchanged and compress the high-frequency information. Therefore, the terminal can obtain the target feature information corresponding to the first frequency band based on the initial feature information corresponding to the first frequency band in the initial band feature information, and use the initial feature information corresponding to the first frequency band in the initial band feature information as the target feature information corresponding to the first frequency band in the intermediate band feature information. That is, before and after band compression, the low-frequency information remains unchanged and is consistent.
[0089] In one embodiment, the terminal can divide the initial band into a first frequency band and a second frequency band based on a preset frequency. The preset frequency can be set based on expert knowledge. For example, the preset frequency is set to 6 kHz. If the sampling rate of the speech signal is 48 kHz, then the initial band corresponding to the speech signal is 0 - 24 kHz, the first frequency band is 0 - 6 kHz, and the second frequency band is 6 - 24 kHz.
[0090] Step S206: Perform feature compression on the initial feature information corresponding to the second frequency band in the initial band feature information to obtain the target feature information corresponding to the compressed frequency band. The frequency of the first frequency band is less than the frequency of the second frequency band, and the frequency range of the second frequency band is greater than the frequency range of the compressed frequency band.
[0091] Among them, feature compression is to compress the feature information corresponding to a large frequency band into the feature information corresponding to a small frequency band, refining and concentrating the feature information. The second frequency band represents the large frequency band, and the compression frequency band represents the small frequency band, that is, the frequency range of the second frequency band is greater than that of the compression frequency band, that is, the length of the second frequency band is greater than that of the compression frequency band. It can be understood that considering the seamless connection between the first frequency band and the compression frequency band, the minimum frequency in the second frequency band can be the same as the minimum frequency in the compression frequency band. At this time, the maximum frequency in the second frequency band is obviously greater than the maximum frequency in the compression frequency band. For example, if the first frequency band is 0-6 kHz and the second frequency band is 6-24 kHz, then the compression frequency band can be 6-8 kHz, 6-16 kHz, etc. Feature compression can also be considered as compressing the feature information corresponding to the high frequency band into the feature information corresponding to the low frequency band.
[0092] Specifically, when performing frequency band compression, the terminal mainly compresses the high frequency information in the voice signal. The terminal can perform feature compression on the initial feature information corresponding to the second frequency band in the initial frequency band feature information to obtain the target feature information corresponding to the compression frequency band.
[0093] In one embodiment, the initial frequency band feature information includes the amplitudes and phases corresponding to multiple initial voice frequency points. When performing feature compression, the terminal can compress both the amplitudes and phases of the initial voice frequency points corresponding to the second frequency band in the initial frequency band feature information to obtain the amplitudes and phases of the target voice frequency points corresponding to the compression frequency band, and obtain the target feature information corresponding to the compression frequency band based on the amplitudes and phases of the target voice frequency points. Compressing the amplitude or phase can be calculating the average value of the amplitude or phase of the initial voice frequency points corresponding to the second frequency band as the amplitude or phase of the target voice frequency points corresponding to the compression frequency band, or calculating the weighted average value of the amplitude or phase of the initial voice frequency points corresponding to the second frequency band as the amplitude or phase of the target voice frequency points corresponding to the compression frequency band, or other compression methods. In addition to overall compression, the compression of the amplitude or phase can also be further segmented compression.
[0094] Furthermore, in order to reduce the difference between the target feature information and the initial feature information, the terminal can only compress the amplitude of the initial voice frequency points corresponding to the second frequency band in the initial frequency band feature information to obtain the amplitude of the target voice frequency points corresponding to the compression frequency band, search for the initial voice frequency points with the same frequency as the target voice frequency points corresponding to the compression frequency band in the initial voice frequency points corresponding to the second frequency band as the intermediate voice frequency points, use the phase corresponding to the intermediate voice frequency points as the phase of the target voice frequency points, and obtain the target feature information corresponding to the compression frequency band based on the amplitudes and phases of the target voice frequency points. For example, if the second frequency band is 6-24 kHz and the compression frequency band is 6-8 kHz, then the phase of the initial voice frequency points corresponding to 6-8 kHz in the second frequency band can be used as the phase of each target voice frequency point corresponding to 6-8 kHz in the compression frequency band.
[0095] Step S208: Obtain intermediate frequency band feature information based on the target feature information corresponding to the first frequency band and the target feature information corresponding to the compressed frequency band, and obtain a compressed speech signal corresponding to the speech signal to be processed based on the intermediate frequency band feature information.
[0096] Among them, the intermediate frequency band feature information refers to the feature information obtained after frequency band compression of the initial frequency band feature information. The compressed speech signal refers to the speech signal obtained after frequency band compression of the speech signal to be processed. Frequency band compression can reduce the sampling rate of the speech signal while keeping the speech content intelligible. It can be understood that the sampling rate of the speech signal to be processed is greater than the sampling rate corresponding to the compressed speech signal.
[0097] Specifically, the terminal can obtain intermediate frequency band feature information based on the target feature information corresponding to the first frequency band and the target feature information corresponding to the compressed frequency band. The intermediate frequency band feature information is a frequency domain signal. After obtaining the intermediate frequency band feature information, the terminal can convert the frequency domain signal into a time domain signal to obtain a compressed speech signal. Among them, the terminal can use a frequency domain-time domain conversion algorithm to convert the frequency domain signal into a time domain signal, for example, a custom frequency domain-time domain conversion algorithm, an inverse Laplace transform algorithm, an inverse Z transform algorithm, a Fourier inverse transform algorithm, etc.
[0098] For example, the sampling rate of the speech signal to be processed is 48 kHz, and the initial frequency band is 0 - 24 kHz. The terminal can obtain the initial feature information corresponding to 0 - 6 kHz from the initial frequency band feature information and directly use the initial feature information corresponding to 0 - 6 kHz as the target feature information corresponding to 0 - 6 kHz. The terminal can obtain the initial feature information corresponding to 6 - 24 kHz from the initial frequency band feature information and compress the initial feature information corresponding to 6 - 24 kHz into the target feature information corresponding to 6 - 8 kHz. The terminal can generate a compressed speech signal based on the target feature information corresponding to 0 - 8 kHz, and the target sampling rate corresponding to the compressed speech signal is 16 kHz.
[0099] It can be understood that if the sampling rate of the speech signal to be processed is higher than the sampling rate supported by the speech encoder, then the terminal's frequency band compression of the speech signal to be processed can be to compress the high-sampling rate speech signal to be processed into the sampling rate supported by the speech encoder, so that the speech encoder can successfully perform encoding processing on the speech signal to be processed. Of course, if the sampling rate of the speech signal to be processed is also equal to or less than the sampling rate supported by the speech encoder, then the terminal's frequency band compression of the speech signal to be processed can be to compress the normally sampled speech signal to be processed into a speech signal with a lower sampling rate, thereby reducing the computational amount during the encoding process of the speech encoder, reducing the data transmission amount, and finally quickly transmitting the speech signal through the network to the speech receiving end.
[0100] In one embodiment, the frequency band corresponding to the intermediate frequency band feature information and the frequency band corresponding to the initial frequency band feature information may be the same or different. When the frequency band corresponding to the intermediate frequency band feature information is the same as the frequency band corresponding to the initial frequency band feature information, in the intermediate frequency band feature information, there are specific feature information in the first frequency band and the compressed frequency band, and the feature information corresponding to each frequency greater than the compressed frequency band is zero. For example, the initial frequency band feature information includes the amplitudes and phases of multiple frequency points in the range of 0 - 24 kHz, the intermediate frequency band feature information includes the amplitudes and phases of multiple frequency points in the range of 0 - 24 kHz, the first frequency band is 0 - 6 kHz, the second frequency band is 8 - 24 kHz, and the compressed frequency band is 6 - 8 kHz. In the initial frequency band feature information, there are corresponding amplitudes and phases for each frequency point in the range of 0 - 24 kHz. In the intermediate frequency band feature information, there are corresponding amplitudes and phases for each frequency point in the range of 0 - 8 kHz, and the amplitudes and phases corresponding to each frequency point in the range of 8 - 24 kHz are all zero. If the frequency band corresponding to the intermediate frequency band feature information is the same as the frequency band corresponding to the initial frequency band feature information, the terminal needs to first convert the intermediate frequency band feature information into a time-domain signal, and then perform downsampling processing on the time-domain signal to obtain the compressed speech signal.
[0101] When the frequency band corresponding to the intermediate frequency band feature information is different from the frequency band corresponding to the initial frequency band feature information, the frequency band corresponding to the intermediate frequency band feature information consists of the first frequency band and the compressed frequency band, and the frequency band corresponding to the initial frequency band feature information consists of the first frequency band and the second frequency band. For example, the initial frequency band feature information includes the amplitudes and phases of multiple frequency points in the range of 0 - 24 kHz, the intermediate frequency band feature information includes the amplitudes and phases of multiple frequency points in the range of 0 - 8 kHz, the first frequency band is 0 - 6 kHz, the second frequency band is 8 - 24 kHz, and the compressed frequency band is 6 - 8 kHz. In the initial frequency band feature information, there are corresponding amplitudes and phases for each frequency point in the range of 0 - 24 kHz. In the intermediate frequency band feature information, there are corresponding amplitudes and phases for each frequency point in the range of 0 - 8 kHz. If the frequency band corresponding to the intermediate frequency band feature information is different from the frequency band corresponding to the initial frequency band feature information, the terminal can directly convert the intermediate frequency band feature information into a time-domain signal to obtain the compressed speech signal.
[0102] Step S210, perform encoding processing on the compressed speech signal through the speech encoding module to obtain the encoded speech data corresponding to the speech signal to be processed. The target sampling rate corresponding to the compressed speech signal is less than or equal to the supported sampling rate corresponding to the speech encoding module, and the target sampling rate is less than the sampling rate corresponding to the speech signal to be processed.
[0103] Among them, the voice encoding module is a module for encoding and processing voice signals. The voice encoding module can be hardware or software. The supported sampling rate corresponding to the voice encoding module refers to the maximum sampling rate supported by the voice encoding module, that is, the sampling rate upper limit. It can be understood that if the supported sampling rate corresponding to the voice encoding module is 16 kHz, then the voice encoding module can encode and process voice signals with a sampling rate less than or equal to 16 kHz.
[0104] Specifically, by performing frequency band compression on the voice signal to be processed, the terminal can compress the voice signal to be processed into a compressed voice signal, so that the sampling rate of the compressed voice signal meets the sampling rate requirement of the voice encoding module. The voice encoding module supports processing voice signals with a sampling rate less than or equal to the sampling rate upper limit. The terminal can encode and process the compressed voice signal through the voice encoding module to obtain the encoded voice data corresponding to the voice signal to be processed. The encoded voice data is bitstream data. If the encoded voice data is only stored locally and does not need to be transmitted over the network, then the terminal can perform voice encoding on the compressed voice signal through the voice encoding module to obtain the encoded voice data. If the encoded voice data needs to be further transmitted to the voice receiving end, then the terminal can perform voice encoding on the compressed voice signal through the voice encoding module to obtain the first voice data, and perform channel encoding on the first voice data to obtain the encoded voice data.
[0105] For example, in a voice chat scenario, friends can have a voice chat on the instant messaging application of the terminal. The user can send a voice message to a friend on the conversation interface in the instant messaging application. When friend A sends a voice message to friend B, the terminal corresponding to friend A is the voice sending end, and the terminal corresponding to friend B is the voice receiving end. The voice sending end can obtain the trigger operation of friend A on the voice collection control on the conversation interface to collect the voice signal, and collect the voice signal of friend A through the microphone to obtain the voice signal to be processed. When a high-quality microphone is used to collect the voice message, the initial sampling rate corresponding to the voice signal to be processed can be 48 kHz, the voice signal to be processed has good sound quality and has an ultra-wide frequency band, specifically 0 - 24 kHz. The voice sending end performs Fourier transform processing on the voice signal to be processed to obtain the initial frequency band feature information corresponding to the voice signal to be processed. The initial frequency band feature information includes the frequency domain information within the range of 0 - 24 kHz. After the frequency domain information of 0 - 24 kHz is subjected to non-linear frequency band compression by the voice sending end, the frequency domain information of 0 - 24 kHz is concentrated on 0 - 8 kHz. Specifically, the initial feature information corresponding to 0 - 6 kHz in the initial frequency band feature information can be kept unchanged, and the initial feature information corresponding to 6 - 24 kHz is compressed to 6 - 8 kHz. The voice sending end generates a compressed voice signal based on the frequency domain information of 0 - 8 kHz obtained after non-linear frequency band compression. The target sampling rate corresponding to the compressed voice signal is 16 kHz. Then, the voice sending end can encode the compressed voice signal through a conventional voice encoder that supports 16 kHz to obtain encoded voice data, and send the encoded voice data to the voice receiving end. The sampling rate of the encoded voice data is consistent with the target sampling rate. After receiving the encoded voice data, the voice receiving end can obtain the target voice signal through decoding processing and non-linear frequency band expansion processing. The sampling rate of the target voice signal is consistent with the initial sampling rate. The voice receiving end can obtain the trigger operation of friend B on the voice message on the conversation interface to play the voice signal, and play the target voice signal with a high sampling rate through the speaker.
[0106] In a recording scenario, when the terminal obtains a recording operation triggered by the user, the terminal can collect the user's voice signal through the microphone to obtain a voice signal to be processed. The terminal performs Fourier transform processing on the voice signal to be processed to obtain initial frequency band feature information corresponding to the voice signal to be processed. The initial frequency band feature information includes frequency domain information in the range of 0 - 24 kHz. After the frequency domain information of 0 - 24 kHz is subjected to non-linear frequency band compression, the frequency domain information of 0 - 24 kHz is concentrated on 0 - 8 kHz. Specifically, the initial feature information corresponding to 0 - 6 kHz in the initial frequency band feature information can be kept unchanged, and the initial feature information corresponding to 6 - 24 kHz is compressed to 6 - 8 kHz. The terminal generates a compressed voice signal based on the frequency domain information of 0 - 8 kHz obtained after non-linear frequency band compression. The target sampling rate of the compressed voice signal is 16 kHz. Then, the terminal can perform encoding processing on the compressed voice signal through a conventional voice encoder that supports 16 kHz to obtain encoded voice data, and store the encoded voice data. When the terminal obtains a recording playback operation triggered by the user, the terminal can perform voice restoration processing on the encoded voice data to obtain a target voice signal, and play the target voice signal.
[0107] In one embodiment, the encoded voice data can carry compression identification information, and the compression identification information is used to identify the frequency band mapping information between the second frequency band and the compressed frequency band. Then, when performing voice restoration processing, the terminal can perform voice restoration processing on the encoded voice data based on the compression identification information to obtain a target voice signal.
[0108] In one embodiment, the maximum frequency in the compressed frequency band can be determined based on the supported sampling rate corresponding to the voice encoding module on the terminal. For example, the supported sampling rate corresponding to the voice encoding module is 16 kHz. When the sampling rate of the voice signal is 16 kHz, the corresponding frequency band is 0 - 8 kHz. Then, the maximum frequency in the compressed frequency band can be 8 kHz. Of course, the maximum frequency in the compressed frequency band can also be less than 8 kHz. Even if the maximum frequency in the compressed frequency band is less than 8 kHz, the voice encoding module with a supported sampling rate of 16 kHz can encode the corresponding compressed voice signal. The maximum frequency in the compressed frequency band can also be a default frequency, and the default frequency can be determined based on the supported sampling rates corresponding to various existing voice encoding modules. For example, among the supported sampling rates corresponding to various known voice encoding modules, the minimum value is 16 kHz, then the default frequency can be set to 8 kHz.
[0109] In the above voice encoding method, by obtaining the initial frequency band feature information corresponding to the voice signal to be processed, the target feature information corresponding to the first frequency band is obtained based on the initial feature information corresponding to the first frequency band in the initial frequency band feature information, and the initial feature information corresponding to the second frequency band in the initial frequency band feature information is compressed to obtain the target feature information corresponding to the compressed frequency band. The frequency of the first frequency band is less than the frequency of the second frequency band, and the frequency range of the second frequency band is greater than the frequency range of the compressed frequency band. The intermediate frequency band feature information is obtained based on the target feature information corresponding to the first frequency band and the target feature information corresponding to the compressed frequency band. The compressed voice signal corresponding to the voice signal to be processed is obtained based on the intermediate frequency band feature information. The compressed voice signal is encoded by a voice encoding module to obtain the encoded voice data corresponding to the voice signal to be processed. The target sampling rate corresponding to the compressed voice signal is less than or equal to the supported sampling rate corresponding to the voice encoding module. In this way, before voice encoding, the voice signal to be processed with any sampling rate can be compressed through the frequency band feature information, and the sampling rate of the voice signal to be processed can be reduced to the sampling rate supported by the voice encoder. The target sampling rate is less than the sampling rate corresponding to the voice signal to be processed, and a compressed voice signal with a low sampling rate is obtained. Since the sampling rate of the compressed voice signal is less than or equal to the sampling rate supported by the voice encoder, the compressed voice signal can be successfully encoded by the voice encoder, and finally the encoded voice data obtained through the encoding process can be transmitted to the voice receiving end.
[0110] In one embodiment, obtaining the initial frequency band feature information corresponding to the voice signal to be processed includes:
[0111] Obtaining the voice signal to be processed collected by a voice collection device; performing Fourier transform processing on the voice signal to be processed to obtain the initial frequency band feature information, where the initial frequency band feature information includes the initial amplitude and initial phase corresponding to multiple initial voice frequency points.
[0112] Among them, the voice collection device refers to a device used to collect voices, for example, a microphone. The Fourier transform processing refers to performing a Fourier transform on the voice signal to be processed to convert the time-domain signal into a frequency-domain signal, and the frequency-domain signal can reflect the feature information of the voice signal to be processed in the frequency domain. The initial frequency band feature information is the frequency-domain signal. The initial voice frequency point refers to the frequency point in the initial frequency band feature information corresponding to the voice signal to be processed.
[0113] Specifically, the terminal can acquire the to-be-processed voice signal collected by the voice acquisition device, perform Fourier transform processing on the to-be-processed voice signal, convert the time-domain signal into a frequency-domain signal, extract the feature information of the to-be-processed voice signal in the frequency domain, and obtain the initial frequency-band feature information. The initial frequency-band feature information is composed of the initial amplitudes and initial phases respectively corresponding to multiple initial voice frequency points. Among them, the phase of the frequency point determines the smoothness of the voice, the amplitude of the low-frequency frequency point determines the specific semantic content of the voice, and the amplitude of the high-frequency frequency point determines the texture of the voice. The frequency range composed of all the initial voice frequency points is the initial frequency band corresponding to the to-be-processed voice signal.
[0114] In one embodiment, the to-be-processed voice signal can obtain N initial voice frequency points through fast Fourier transform. Usually, N is an integer power of 2, and the N initial voice frequency points are evenly distributed. For example, if N is 1024 and the initial frequency band corresponding to the to-be-processed voice signal is 24 kHz, then the resolution of the initial voice frequency point is 24k / 1024 = 23.4375, that is, there is an initial voice frequency point every 23.4375 kHz. It can be understood that in order to ensure a high resolution, voice signals with different sampling rates can obtain different numbers of voice frequency points through fast Fourier transform. The higher the sampling rate of the voice signal, the more initial voice frequency points are obtained through fast Fourier transform.
[0115] In this embodiment, by performing Fourier transform processing on the to-be-processed voice signal, the initial frequency-band feature information corresponding to the to-be-processed voice signal can be quickly obtained.
[0116] In one embodiment, as Figure 3 shown, feature compression is performed on the initial feature information corresponding to the second frequency band in the initial frequency-band feature information to obtain the target feature information corresponding to the compressed frequency band, including:
[0117] Step S302: Divide the second frequency band into at least two initial sub-bands arranged in order.
[0118] Step S304: Divide the compressed frequency band into at least two target sub-bands arranged in order.
[0119] Among them, frequency band division refers to splitting a frequency band into multiple sub-frequency bands. The terminal can perform frequency band division on the second frequency band or the compressed frequency band in a linear or non-linear manner. Taking the second frequency band as an example, the terminal can perform linear frequency band division on the second frequency band, that is, evenly split the second frequency band. For example, if the second frequency band is 6 - 24 kHz, the second frequency band can be evenly divided into three equal-sized initial sub-frequency bands, namely 6 - 12 kHz, 12 - 18 kHz, and 18 - 24 kHz. The terminal can also perform non-linear frequency band division on the second frequency band, that is, not evenly split the second frequency band. For example, if the second frequency band is 6 - 24 kHz, the second frequency band can be non-linearly divided into five initial sub-frequency bands, namely 6 - 8 kHz, 8 - 10 kHz, 10 - 12 kHz, 12 - 18 kHz, and 18 - 24 kHz.
[0120] Specifically, the terminal can perform frequency band division on the second frequency band to obtain at least two sequentially arranged initial sub-frequency bands, and perform frequency band division on the compressed frequency band to obtain at least two sequentially arranged target sub-frequency bands. The number of initial sub-frequency bands and the number of target sub-frequency bands can be the same or different. When the number of initial sub-frequency bands is the same as the number of target sub-frequency bands, the initial sub-frequency bands and the target sub-frequency bands correspond one by one. When the number of initial sub-frequency bands is different from the number of target sub-frequency bands, multiple initial sub-frequency bands can correspond to one target sub-frequency band, or one initial sub-frequency band can correspond to multiple target sub-frequency bands.
[0121] Step S306: Determine the target sub-frequency band corresponding to each initial sub-segment based on the sub-frequency band sorting of the initial sub-frequency band and the target sub-frequency band.
[0122] Specifically, the terminal can determine the target sub-frequency band corresponding to each initial sub-segment based on the sub-frequency band sorting of the initial sub-frequency band and the target sub-frequency band. When the number of initial sub-frequency bands is the same as the number of target sub-frequency bands, the terminal can establish an association relationship between the initial sub-frequency band and the target sub-frequency band with the same sorting. Refer to Figure 4, the sorted initial sub-bands are 6 - 8 kHz, 8 - 10 kHz, 10 - 12 kHz, 12 - 18 kHz, 18 - 24 kHz, and the sorted target sub-bands are 6 - 6.4 kHz, 6.4 - 6.8 kHz, 6.8 - 7.2 kHz, 7.2 - 7.6 kHz, 7.6 - 8 kHz. Then, 6 - 8 kHz corresponds to 6 - 6.4 kHz, 8 - 10 kHz corresponds to 6.4 - 6.8 kHz, 10 - 12 kHz corresponds to 6.8 - 7.2 kHz, 12 - 18 kHz corresponds to 7.2 - 7.6 kHz, and 18 - 24 kHz corresponds to 7.6 - 8 kHz. When the number of initial sub-bands is different from the number of target sub-bands, the terminal can establish a one-to-one correspondence relationship between the initial sub-bands and target sub-bands with higher rankings, establish a one-to-one correspondence relationship between the initial sub-bands and target sub-bands with lower rankings, and establish a one-to-many or many-to-one correspondence relationship between the initial sub-bands and target sub-bands with middle rankings. For example, when the number of initial sub-bands with middle rankings is greater than the number of target sub-bands, a many-to-one correspondence relationship is established.
[0123] Step S308: Use the initial feature information of the current initial sub-band corresponding to the current target sub-band as the first intermediate feature information. From the initial frequency band feature information, obtain the initial feature information of the sub-band corresponding to the current target sub-band with the same frequency band information as the second intermediate feature information. Based on the first intermediate feature information and the second intermediate feature information, obtain the target feature information corresponding to the current target sub-band.
[0124] Specifically, the feature information corresponding to a frequency band includes the amplitudes and phases corresponding to at least one frequency point. When performing feature compression, the terminal can only compress the amplitudes, while the phases remain the original phases. The current target sub-band refers to the target sub-band for generating the current target feature information. When generating the target feature information corresponding to the current target sub-band, the terminal can use the initial feature information of the current initial sub-band corresponding to the current target sub-band as the first intermediate feature information, and the first intermediate feature information is used to determine the amplitudes of the frequency points in the target feature information corresponding to the current target sub-band. The terminal can obtain the initial feature information of the sub-band corresponding to the current target sub-band with the same frequency band information as the second intermediate feature information from the initial frequency band feature information, and the second intermediate feature information is used to determine the phases of the frequency points in the target feature information corresponding to the current target sub-band. Therefore, the terminal can obtain the target feature information corresponding to the current target sub-band based on the first intermediate feature information and the second intermediate feature information.
[0125] For example, the initial frequency band feature information includes the initial feature information corresponding to 0 - 24 kHz. The current target sub - frequency band is 6 - 6.4 kHz, and the corresponding initial sub - frequency band of the current target sub - frequency band is 6 - 8 kHz. The terminal can obtain the target feature information corresponding to 6 - 6.4 kHz based on the initial feature information corresponding to 6 - 8 kHz and the initial feature information corresponding to 6 - 6.4 kHz in the initial frequency band feature information.
[0126] Step S310, obtain the target feature information corresponding to the compressed frequency band based on the target feature information corresponding to each target sub - frequency band.
[0127] Specifically, after obtaining the target feature information corresponding to each target sub - frequency band, the terminal can obtain the target feature information corresponding to the compressed frequency band based on the target feature information corresponding to each target sub - frequency band, and the target feature information corresponding to the compressed frequency band is composed of the target feature information corresponding to each target sub - frequency band.
[0128] In this embodiment, by further subdividing the second frequency band and the compressed frequency band for feature compression, the reliability of feature compression can be improved, and the difference between the initial feature information corresponding to the second frequency band and the target feature information corresponding to the compressed frequency band can be reduced. In this way, a target speech signal with a relatively high similarity to the speech signal to be processed can be restored during subsequent frequency band expansion.
[0129] In one embodiment, both the first intermediate feature information and the second intermediate feature information include the initial amplitudes and initial phases corresponding to multiple initial speech frequency points. Obtaining the target feature information corresponding to the current target sub - frequency band based on the first intermediate feature information and the second intermediate feature information includes:
[0130] Based on the statistical values of the initial amplitudes corresponding to each initial speech frequency point in the first intermediate feature information, obtain the target amplitudes of each target speech frequency point corresponding to the current target sub - frequency band; based on the initial phases corresponding to each initial speech frequency point in the second intermediate feature information, obtain the target phases of each target speech frequency point corresponding to the current target sub - frequency band; based on the target amplitudes and target phases of each target speech frequency point corresponding to the current target sub - frequency band, obtain the target feature information corresponding to the current target sub - frequency band.
[0131] Specifically, for the amplitude of the frequency points, the terminal can statistically analyze the initial amplitudes corresponding to the respective initial speech frequency points in the first intermediate feature information, and use the calculated statistical value as the target amplitude of the respective target speech frequency points corresponding to the current target sub-band. For the phase of the frequency points, the terminal can obtain the target phase of the respective target speech frequency points corresponding to the current target sub-band based on the initial phases corresponding to the respective initial speech frequency points in the second intermediate feature information. The terminal can obtain the initial phase of the initial speech frequency point with the same frequency as the target speech frequency point from the second intermediate feature information as the target phase of the target speech frequency point, that is, the target phase corresponding to the target speech frequency point follows the original phase. Among them, the statistical value can be an arithmetic mean, a weighted mean, etc.
[0132] For example, the terminal can calculate the arithmetic mean of the initial amplitudes corresponding to the respective initial speech frequency points in the first intermediate feature information, and use the calculated arithmetic mean as the target amplitude of the respective target speech frequency points corresponding to the current target sub-band.
[0133] The terminal can also calculate the weighted mean of the initial amplitudes corresponding to the respective initial speech frequency points in the first intermediate feature information, and use the calculated weighted mean as the target amplitude of the respective target speech frequency points corresponding to the current target sub-band. For example, generally speaking, the importance of the center frequency point is relatively high. The terminal can assign a higher weight to the initial amplitude of the center frequency point of a frequency band, assign a lower weight to the initial amplitudes of other frequency points in that frequency band, and then perform weighted averaging on the initial amplitudes of each frequency band to obtain the weighted mean.
[0134] The terminal can further divide the initial sub-band corresponding to the current target sub-band and the current target sub-band to obtain at least two sorted first frequency bands corresponding to the initial sub-band and at least two sorted second frequency bands corresponding to the current target sub-band. The terminal can establish an association relationship between the first frequency band and the second frequency band according to the sorting of the first frequency band and the second frequency band, and use the statistical value of the initial amplitude corresponding to each initial voice frequency point in the current first frequency band as the target amplitude of each target voice frequency point in the second frequency band corresponding to the current first frequency band. For example, the current target sub-band is 6 - 6.4 khz, and the initial sub-band corresponding to the current target sub-band is 6 - 8 khz. The initial sub-band and the current target sub-band are equally divided to obtain two first frequency bands (6 - 7 khz and 7 - 8 khz) and two second frequency bands (6 - 6.2 khz and 6.2 khz - 6.4 khz). 6 - 7 khz corresponds to 6 - 6.2 khz, and 7 - 8 khz corresponds to 6.2 khz - 6.4 khz. Calculate the arithmetic mean of the initial amplitudes corresponding to each initial voice frequency point in 6 - 7 khz as the target amplitude of each target voice frequency point in 6 - 6.2 khz. Calculate the arithmetic mean of the initial amplitudes corresponding to each initial voice frequency point in 7 - 8 khz as the target amplitude of each target voice frequency point in 6.2 khz - 6.4 khz.
[0135] In one embodiment, if the frequency band corresponding to the initial frequency band feature information is equal to the frequency band corresponding to the intermediate frequency band feature information, then the number of initial voice frequency points corresponding to the initial frequency band feature information is equal to the number of target voice frequency points corresponding to the intermediate frequency band feature information. For example, the frequency bands corresponding to the initial frequency band feature information and the intermediate frequency band feature information are both 24 khz. In the initial frequency band feature information and the intermediate frequency band feature information, the amplitudes and phases of the voice frequency points corresponding to 0 - 6 khz are the same. In the intermediate frequency band feature information, the target amplitude of the target voice frequency points corresponding to 6 - 8 khz is calculated based on the initial amplitudes of the initial voice frequency points corresponding to 6 - 24 khz in the initial frequency band feature information, and the target phase of the target voice frequency points corresponding to 6 - 8 khz follows the initial phase of the initial voice frequency points corresponding to 6 - 8 khz in the initial frequency band feature information. In the intermediate frequency band feature information, the target amplitudes and target phases of the target voice frequency points corresponding to 8 - 24 khz are zero.
[0136] If the frequency band corresponding to the initial frequency band feature information is greater than the frequency band corresponding to the intermediate frequency band feature information, then the number of initial voice frequency points corresponding to the initial frequency band feature information is greater than the number of target voice frequency points corresponding to the intermediate frequency band feature information. Further, the ratio of the number of initial voice frequency points to the number of target voice frequency points can be the same as the ratio of the bandwidths of the initial frequency band feature information and the target frequency band feature information, so as to convert the amplitudes and phases between the frequency points. For example, if the frequency band corresponding to the initial frequency band feature information is 24 kHz and the frequency band corresponding to the intermediate frequency band feature information is 12 kHz, then the number of initial voice frequency points corresponding to the initial frequency band feature information can be 1024, and the number of target voice frequency points corresponding to the intermediate frequency band feature information can be 512. In the initial frequency band feature information and the intermediate frequency band feature information, the amplitudes and phases of the voice frequency points corresponding to 0 - 6 kHz are the same. In the intermediate frequency band feature information, the target amplitude of the target voice frequency points corresponding to 6 - 12 kHz is calculated based on the initial amplitude of the initial voice frequency points corresponding to 6 - 24 kHz in the initial frequency band feature information, and the target phase of the target voice frequency points corresponding to 6 - 12 kHz follows the initial phase of the initial voice frequency points corresponding to 6 - 12 kHz in the initial frequency band feature information.
[0137] In this embodiment, in the target feature information corresponding to the compression frequency band, the amplitude of the target voice frequency point is the statistical value of the amplitude of the corresponding initial voice frequency point, and the phase of the target voice frequency point follows the original phase, which can further reduce the difference between the initial feature information corresponding to the second frequency band and the target feature information corresponding to the compression frequency band.
[0138] In one embodiment, obtaining intermediate frequency band feature information based on the target feature information corresponding to the first frequency band and the target feature information corresponding to the compression frequency band, and obtaining a compressed voice signal corresponding to the voice signal to be processed based on the intermediate frequency band feature information, includes:
[0139] Determining a third frequency band based on the frequency difference between the compression frequency band and the second frequency band, and setting the target feature information corresponding to the third frequency band as invalid information; obtaining intermediate frequency band feature information based on the target feature information corresponding to the first frequency band, the target feature information corresponding to the compression frequency band, and the target feature information corresponding to the third frequency band; performing an inverse Fourier transform process on the intermediate frequency band feature information to obtain an intermediate voice signal, where the sampling rate corresponding to the intermediate voice signal is the same as the sampling rate corresponding to the voice signal to be processed; performing downsampling processing on the intermediate voice signal based on the supported sampling rate to obtain a compressed voice signal.
[0140] Among them, the third frequency band is composed of the frequencies between the maximum frequency of the compression frequency band and the maximum frequency of the second frequency band. The inverse Fourier transform process performs an inverse Fourier transform on the intermediate frequency band feature information to convert the frequency domain signal into a time domain signal. Both the intermediate voice signal and the compressed voice signal are time domain signals.
[0141] Downsampling processing refers to filtering and sampling the speech signal in the time domain. For example, if the sampling rate of the signal is 48 kHz, it means that 48k points are collected in one second. If the sampling rate of the signal is 16 kHz, it means that 16k points are collected in one second.
[0142] Specifically, in order to improve the conversion speed between the frequency-domain signal and the time-domain signal, when performing frequency band compression, the terminal can keep the number of speech frequency points unchanged and change the amplitudes and phases of some speech frequency points to obtain intermediate frequency band feature information. Furthermore, the terminal can quickly perform inverse Fourier transform processing on the intermediate frequency band feature information to obtain an intermediate speech signal, and the sampling rate of the intermediate speech signal is the same as that of the speech signal to be processed. Then, the terminal performs downsampling processing on the intermediate speech signal to reduce the sampling rate of the intermediate speech signal to the supported sampling rate of the speech encoder or below, obtaining a compressed speech signal. Among them, in the intermediate frequency band feature information, the target feature information corresponding to the first frequency band follows the initial feature information corresponding to the first frequency band in the initial frequency band feature information, the target feature information corresponding to the compressed frequency band is obtained based on the initial feature information corresponding to the second frequency band in the initial frequency band feature information, and the target feature information corresponding to the third frequency band is set to invalid information, that is, the target feature information corresponding to the third frequency band is cleared.
[0143] In this embodiment, when processing the frequency-domain signal, keeping the frequency band unchanged, converting the frequency-domain signal into a time-domain signal, and then reducing the sampling rate of the signal through downsampling processing can reduce the complexity of frequency-domain signal processing.
[0144] In one embodiment, the speech coding module performs coding processing on the compressed speech signal to obtain the coded speech data corresponding to the speech signal to be processed, including:
[0145] The speech coding module performs speech coding on the compressed speech signal to obtain first speech data; performs channel coding on the first speech data to obtain coded speech data.
[0146] Among them, voice coding is used to compress the data rate of the voice signal and remove the redundancy in the signal. Voice coding is to encode the analog voice signal, convert the analog signal into a digital signal, thereby reducing the transmission code rate and enabling digital transmission. Voice coding can also be called source coding. It should be noted that voice coding does not change the sampling rate of the voice signal. The coded bitstream data can be completely restored to the voice signal before coding through decoding processing. Band compression, on the other hand, will change the sampling rate of the voice signal. The voice signal after band compression cannot be exactly restored to the voice signal before band compression through band expansion. However, the semantic content conveyed by the voice signals before and after band compression is the same, which does not affect the listener's understanding. The terminal can use voice coding methods such as waveform coding, parametric coding (source coding), and hybrid coding to perform voice coding on the compressed voice signal.
[0147] Channel coding is used to improve the stability of data transmission. Due to interference and fading in mobile communication and network transmission, errors may occur during the transmission of the voice signal. Therefore, error correction and detection techniques, that is, error correction and detection coding techniques, need to be applied to the digital signal to enhance the ability of the data to resist various interferences during transmission in the channel and improve the reliability of voice transmission. The error correction and detection coding performed on the digital signal to be transmitted in the channel is channel coding. The terminal can use channel coding methods such as convolutional coding and Turbo coding to perform channel coding on the first voice data.
[0148] Specifically, when performing coding processing, the terminal can perform voice coding on the compressed voice signal through a voice coding module to obtain the first voice data, and then perform channel coding on the first voice data to obtain the coded voice data. It can be understood that the voice coding module may only integrate voice coding algorithms. Then, the terminal can perform voice coding on the compressed voice signal through the voice coding module and perform channel coding on the first voice data through other modules or software programs. The voice coding module can also integrate both voice coding algorithms and channel coding algorithms. The terminal performs voice coding on the compressed voice signal through the voice coding module to obtain the first voice data, and performs channel coding on the first voice data through the voice coding module to obtain the coded voice data.
[0149] In this embodiment, performing voice coding and channel coding on the compressed voice signal can reduce the data volume of the voice signal and ensure the stability of voice signal transmission.
[0150] In one embodiment, the method further includes:
[0151] Sending the coded voice data to a voice receiving end so that the voice receiving end performs voice restoration processing on the coded voice data to obtain the target voice signal corresponding to the voice signal to be processed, and plays the target voice signal.
[0152] Among them, the voice receiving end refers to a device for receiving voice signals and playing voice signals. Voice restoration processing is used to restore the encoded voice data into playable voice signals. For example, restoring the decoded voice signal with a low sampling rate into a voice signal with a high sampling rate, and decoding the bitstream data with a small amount of data into a voice signal with a large amount of data.
[0153] Specifically, if the terminal is used as the voice sending end, the voice sending end can send the encoded voice data to the voice receiving end. After receiving the encoded voice data, the voice receiving end can perform voice restoration processing on the encoded voice data to obtain the target voice signal corresponding to the voice signal to be processed, and then play the target voice signal.
[0154] When performing voice restoration processing, the voice receiving end can only perform decoding processing on the encoded voice data to obtain a compressed voice signal, use the compressed voice signal as the target voice signal, and play the compressed voice signal. At this time, although the sampling rate of the compressed voice signal is lower than that of the originally collected voice signal to be processed, the semantic content reflected by the compressed voice signal and the voice signal to be processed is the same, and the compressed voice signal can also be understood by the listener.
[0155] Of course, in order to further improve the playback clarity and intelligibility of the voice signal, when performing voice restoration processing, the voice receiving end can perform decoding processing on the encoded voice data to obtain a compressed voice signal, restore the compressed voice signal with a low sampling rate into a voice signal with a high sampling rate, and use the restored voice signal as the target voice signal. At this time, the target voice signal refers to the voice signal obtained by performing bandwidth expansion on the compressed voice signal corresponding to the voice signal to be processed, and the sampling rate of the target voice signal is the same as that of the voice signal to be processed. It can be understood that during bandwidth compression, there is a certain loss of information. Therefore, the target voice signal restored by bandwidth expansion is not exactly the same as the original voice signal to be processed, but the semantic content reflected by the target voice signal and the voice signal to be processed is the same. Moreover, compared with the compressed voice signal, the target voice signal has a wider bandwidth, contains more information, has better sound quality, and is clear and intelligible.
[0156] In this embodiment, the encoded voice data can be applied to voice communication and voice transmission. Compressing the voice signal with a high sampling rate into a voice signal with a low sampling rate and then transmitting it can reduce the voice transmission cost.
[0157] In one embodiment, sending the encoded voice data to the voice receiving end so that the voice receiving end performs voice restoration processing on the encoded voice data to obtain the target voice signal corresponding to the voice signal to be processed and play the target voice signal includes:
[0158] Obtain the compression identification information corresponding to the speech signal to be processed based on the second frequency band and the compressed frequency band; send the encoded speech data and the compression identification information to the speech receiving end, so that the speech receiving end decodes the encoded speech data to obtain a compressed speech signal, performs band expansion on the compressed speech signal based on the compression identification information to obtain a target speech signal, and plays the target speech signal.
[0159] Among them, the compression identification information is used to identify the frequency band mapping information between the second frequency band and the compressed frequency band. The frequency band mapping information includes the sizes of the second frequency band and the compressed frequency band, and the mapping relationship (corresponding relationship, association relationship) between the sub-frequency bands of the second frequency band and the compressed frequency band. Band expansion can increase the sampling rate of the speech signal while keeping the speech content intelligible. Band expansion refers to expanding a speech signal with a small frequency band into a speech signal with a large frequency band, where the speech signal with a small frequency band and the speech signal with a large frequency band have the same low-frequency information.
[0160] Specifically, after receiving the encoded speech data, the speech receiving end can default that the encoded speech data has undergone band compression, automatically decode the encoded speech data to obtain a compressed speech signal, and perform band expansion on the compressed speech signal to obtain a target speech signal. However, considering the compatibility with traditional speech processing methods and the diversity of frequency band mapping information during feature compression, when sending the encoded speech data to the speech receiving end, the speech sending end can synchronously send the compression identification information to the speech receiving end, so that the speech receiving end can quickly identify whether the encoded speech data has undergone band compression and the frequency band mapping information during band compression, so as to decide whether to directly decode and play the encoded speech data or play it after decoding and corresponding band expansion. It can be understood that in order to save the computing resources of the speech sending end, for speech signals whose sampling rate is originally less than or equal to that of the speech encoder, the speech sending end can choose to directly encode and process them using traditional speech processing methods and then send them to the speech receiving end.
[0161] If the speech sending end performs band compression on the speech signal to be processed, the speech sending end can generate the compression identification information corresponding to the speech signal to be processed based on the second frequency band and the compressed frequency band, and send the encoded speech data and the compression identification information to the speech receiving end, so that the speech receiving end performs band expansion on the compressed speech signal based on the frequency band mapping information corresponding to the compression identification information to obtain a target speech signal. The compressed speech signal is obtained by the speech receiving end decoding the encoded speech data.
[0162] In addition, if default frequency band mapping information is agreed upon between the voice sending end and the voice receiving end, when generating compression identification information corresponding to the voice signal to be processed based on the second frequency band and the compressed frequency band, the voice sending end can directly obtain a pre-agreed special identifier as the compression identification information. The special identifier is used to indicate that the compressed voice signal is obtained by performing frequency band compression based on the default frequency band mapping information. After receiving the encoded voice data and the compression identification information, the voice receiving end can perform decoding processing on the encoded voice data to obtain the compressed voice signal, and perform frequency band expansion on the compressed voice signal based on the default frequency band mapping information to obtain the target voice signal. If there are multiple frequency band mapping information stored between the voice sending end and the voice receiving end, the voice sending end and the voice receiving end can agree on preset identifiers corresponding to various frequency band mapping information respectively. Different frequency band mapping information can be that the sizes of the second frequency band and the compressed frequency band are different, the sub-band division methods are different, etc. When generating compression identification information corresponding to the voice signal to be processed based on the second frequency band and the compressed frequency band, the voice sending end can obtain the corresponding preset identifier as the compression identification information based on the frequency band mapping information used during feature compression for the second frequency band and the compressed frequency band. After receiving the encoded voice data and the compression identification information, the voice receiving end can perform frequency band expansion on the decoded compressed voice signal based on the frequency band mapping information corresponding to the compression identification information to obtain the target voice signal. Of course, the compression identification information can also directly include specific frequency band mapping information.
[0163] It can be understood that the specific process of performing frequency band expansion on the compressed voice signal can refer to the methods described in various relevant embodiments in the subsequent voice decoding method, such as the methods described in steps S506 to S510.
[0164] In one embodiment, dedicated frequency band mapping information can be designed for different application programs. For example, for application programs with high audio quality requirements (such as singing application programs), a larger number of sub-bands can be designed to be used during feature compression, so as to retain the overall frequency domain characteristics of the original voice signal and the overall change trend of the frequency point amplitudes to the greatest extent. For application programs with low audio quality requirements (such as instant messaging application programs), a smaller number of sub-bands can be designed to be used during feature compression, so as to speed up the compression speed while ensuring semantic intelligibility. Therefore, the compression identification information can also be an application program identifier. After receiving the encoded voice data and the compression identification information, the voice receiving end can perform corresponding frequency band expansion on the decoded compressed voice signal based on the frequency band mapping information corresponding to the application program identifier to obtain the target voice signal.
[0165] In this embodiment, sending the encoded voice data and the compression identification information to the voice receiving end can enable the voice receiving end to more accurately perform frequency band expansion on the decoded compressed voice signal to obtain a target voice signal with high restoration.
[0166] In one embodiment, as Figure 5 shown, a voice decoding method is provided. Taking the terminal in Figure 1 as an example for illustration, the terminal can be Figure 1 the voice sending end in
[0167] Step S502, obtain encoded voice data, which is obtained by performing voice compression processing on the voice signal to be processed.
[0168] Among them, voice compression processing is used to compress the voice signal to be processed into a code stream data that can be transmitted. For example, compress a voice signal with a high sampling rate into a voice signal with a low sampling rate, and then encode the low sampling rate voice signal into code stream data, or encode a voice signal with a large amount of data into code stream data with a small amount of data.
[0169] Specifically, the terminal obtains encoded voice data. Among them, the encoded voice data can be obtained by the terminal performing encoding processing on the voice signal to be processed, or can be received by the terminal from the voice sending end. If the terminal is a voice receiving end, the encoded voice data can be obtained by the voice sending end performing encoding processing on the voice signal to be processed, or can be obtained by the voice sending end performing band compression on the voice signal to be processed to obtain a compressed voice signal, and then performing encoding processing on the compressed voice signal.
[0170] Step S504, perform decoding processing on the encoded voice data through a voice decoding module to obtain a decoded voice signal, and the target sampling rate of the decoded voice signal is less than or equal to the supported sampling rate corresponding to the voice decoding module.
[0171] Among them, the voice decoding module is a module for performing decoding processing on voice signals. The voice decoding module can be hardware or software. The voice encoding module and the voice decoding module can be integrated on one module. The supported sampling rate corresponding to the voice decoding module refers to the maximum sampling rate supported by the voice decoding module, that is, the sampling rate upper limit. It can be understood that if the supported sampling rate corresponding to the voice decoding module is 16 kHz, then the voice decoding module can perform decoding processing on voice signals with a sampling rate less than or equal to 16 kHz.
[0172] Specifically, after the terminal obtains the encoded voice data, it can perform decoding processing on the encoded voice data through the voice decoding module to obtain a decoded voice signal, and restore the voice signal before encoding. The voice decoding module supports processing voice signals with a sampling rate less than or equal to the sampling rate upper limit. The decoded voice signal is a time-domain signal.
[0173] In one embodiment, performing decoding processing on the encoded voice data through a voice decoding module to obtain a decoded voice signal includes:
[0174] Perform channel decoding on the encoded speech data to obtain second speech data; perform speech decoding on the second speech data through a speech decoding module to obtain a decoded speech signal.
[0175] Specifically, channel decoding can be considered as the inverse process of channel encoding. Speech decoding can be considered as the inverse process of speech encoding. When the terminal performs decoding processing on the encoded speech data, it first performs channel decoding on the encoded speech data to obtain second speech data, and then performs speech decoding on the second speech data through a speech decoding module to obtain a decoded speech signal. It can be understood that the speech decoding module may only integrate a speech decoding algorithm, so the terminal can perform channel decoding on the encoded speech data through other modules or software programs, and then perform speech decoding on the second speech data through the speech decoding module. The speech decoding module may also integrate both a speech decoding algorithm and a channel decoding algorithm, so the terminal can perform channel decoding on the encoded speech data through the speech decoding module to obtain second speech data, and perform speech decoding on the second speech data through the speech decoding module to obtain a decoded speech signal.
[0176] It can be understood that if the encoded speech data is generated locally at the terminal, the terminal's decoding process on the encoded speech data can also be to perform speech decoding on the encoded speech data to obtain a decoded speech signal.
[0177] Step S506: Generate target band feature information corresponding to the decoded speech signal, and obtain extended feature information corresponding to the first frequency band based on the target feature information corresponding to the first frequency band in the target band feature information.
[0178] Among them, the target band corresponding to the decoded speech signal includes a first frequency band and a compressed frequency band, and the frequency of the first frequency band is less than the frequency of the compressed frequency band. The terminal can divide the target band feature information into target feature information corresponding to the first frequency band and target feature information corresponding to the compressed frequency band. That is, the target band feature information can be divided into target feature information corresponding to the low frequency band and target feature information corresponding to the high frequency band. The target feature information refers to the feature information corresponding to each frequency before band expansion, and the extended feature information refers to the feature information corresponding to each frequency after band expansion.
[0179] Specifically, the terminal can extract the frequency-domain features of the decoded speech signal, convert the time-domain signal into a frequency-domain signal, and obtain the target frequency-band feature information corresponding to the decoded speech signal. It can be understood that if the sampling rate of the speech signal to be processed is higher than the supported sampling rate corresponding to the speech coding module, then the terminal or the speech sending end has performed frequency-band compression on the speech signal to be processed to reduce the sampling rate of the speech signal to be processed. At this time, the terminal needs to perform frequency-band expansion on the decoded speech signal to restore the speech signal to be processed with a high sampling rate. At this time, the decoded speech signal is a compressed speech signal. If the speech signal to be processed has not undergone frequency-band compression, the terminal can also perform frequency-band expansion on the decoded speech signal to increase the sampling rate of the decoded speech signal and enrich the frequency-domain information.
[0180] When performing frequency-band expansion, in order to ensure that the semantic content remains unchanged, natural, and intelligible, the terminal can keep the low-frequency information unchanged and expand the high-frequency information. Therefore, the terminal can obtain the extended feature information corresponding to the first frequency band based on the target feature information corresponding to the first frequency band in the target frequency-band feature information, and use the initial feature information corresponding to the first frequency band in the target frequency-band feature information as the extended feature information corresponding to the first frequency band in the extended frequency-band feature information. That is, before and after frequency-band expansion, the low-frequency information remains unchanged and is consistent. Similarly, the terminal can divide the target frequency band into a first frequency band and a compressed frequency band based on a preset frequency.
[0181] Step S508: Perform feature expansion on the target feature information corresponding to the compressed frequency band in the target frequency-band feature information to obtain the extended feature information corresponding to the second frequency band; the frequency of the first frequency band is less than the frequency of the compressed frequency band, and the frequency range of the compressed frequency band is less than the frequency range of the second frequency band.
[0182] Among them, feature expansion is to expand the feature information corresponding to a small frequency band into the feature information corresponding to a large frequency band to enrich the feature information. The compressed frequency band represents a small frequency band, and the second frequency band represents a large frequency band, that is, the frequency range of the compressed frequency band is less than the frequency range of the second frequency band, that is, the length of the compressed frequency band is less than the length of the second frequency band.
[0183] Specifically, when performing frequency-band expansion, the terminal mainly expands the high-frequency information in the speech signal. The terminal can perform feature expansion on the target feature information corresponding to the compressed frequency band in the target frequency-band feature information to obtain the extended feature information corresponding to the second frequency band.
[0184] In one embodiment, the target frequency band feature information includes the amplitudes and phases corresponding to multiple target voice frequency points. When performing feature expansion, the terminal can copy the amplitudes of the target voice frequency points corresponding to the compressed frequency band in the target frequency band feature information to obtain the amplitudes of the initial voice frequency points corresponding to the second frequency band, and copy or randomly assign the phases of the target voice frequency points corresponding to the compressed frequency band in the target frequency band feature information to obtain the phases of the initial voice frequency points corresponding to the second frequency band, so as to obtain the expanded feature information corresponding to the second frequency band. In addition to overall copying, the amplitude can also be further copied in segments.
[0185] Step S510: Obtain the expanded frequency band feature information based on the expanded feature information corresponding to the first frequency band and the expanded feature information corresponding to the second frequency band, and obtain the target voice signal corresponding to the voice signal to be processed based on the expanded frequency band feature information. The sampling rate of the target voice signal is greater than the target sampling rate.
[0186] Among them, the expanded frequency band feature information refers to the feature information obtained after expanding the target frequency band feature information. The target voice signal refers to the voice signal obtained after frequency band expansion of the decoded voice signal. Frequency band expansion can increase the sampling rate of the voice signal while keeping the speech content intelligible. It can be understood that the sampling rate of the target voice signal is greater than the sampling rate corresponding to the decoded voice signal.
[0187] Specifically, the terminal obtains the expanded frequency band feature information based on the expanded feature information corresponding to the first frequency band and the expanded feature information corresponding to the second frequency band. The expanded frequency band feature information is a frequency domain signal. After obtaining the expanded frequency band feature information, the terminal can convert the frequency domain signal into a time domain signal to obtain the target voice signal. For example, the terminal performs an inverse Fourier transform process on the expanded frequency band feature information to obtain the target voice signal.
[0188] For example, the sampling rate of the decoded voice signal is 16 kHz, and the target frequency band is 0 - 8 kHz. The terminal can obtain the target feature information corresponding to 0 - 6 kHz from the target frequency band feature information, and directly use the target feature information corresponding to 0 - 6 kHz as the expanded feature information corresponding to 0 - 6 kHz. The terminal can obtain the target feature information corresponding to 6 - 8 kHz from the target frequency band feature information, and expand the target feature information corresponding to 6 - 8 kHz into the expanded feature information corresponding to 6 - 24 kHz. The terminal can generate a target voice signal based on the expanded feature information corresponding to 0 - 24 kHz, and the sampling rate of the target voice signal is 48 kHz.
[0189] Step S512: Play the target voice signal.
[0190] Specifically, after obtaining the target voice signal, the terminal can play the target voice signal through a speaker.
[0191] In the above voice decoding method, by obtaining encoded voice data, which is obtained by performing voice compression processing on the voice signal to be processed, and performing decoding processing on the encoded voice data through a voice decoding module to obtain a decoded voice signal, the target sampling rate corresponding to the decoded voice signal is less than or equal to the supported sampling rate corresponding to the voice decoding module. Generate target band feature information corresponding to the decoded voice signal, obtain extended feature information corresponding to the first frequency band based on the target feature information corresponding to the first frequency band in the target band feature information, perform feature extension on the target feature information corresponding to the compressed frequency band in the target band feature information to obtain extended feature information corresponding to the second frequency band; the frequency of the first frequency band is less than the frequency of the compressed frequency band, and the frequency range of the compressed frequency band is less than the frequency range of the second frequency band. Based on the extended feature information corresponding to the first frequency band and the extended feature information corresponding to the second frequency band, obtain extended band feature information, and based on the extended band feature information, obtain a target voice signal corresponding to the voice signal to be processed. The sampling rate of the target voice signal is greater than the target sampling rate, and play the target voice signal. In this way, after the terminal obtains the encoded voice data obtained by voice compression processing, it can perform decoding processing on the encoded voice data to obtain a decoded voice signal. Through the extension of the band feature information, the sampling rate of the decoded voice signal can be increased to obtain a target voice signal and play it. The playback of the voice signal is not restricted by the sampling rate supported by the voice decoder. During voice playback, a high-sampling-rate voice signal with richer information can also be played.
[0192] In one embodiment, performing feature extension on the target feature information corresponding to the compressed frequency band in the target band feature information to obtain extended feature information corresponding to the second frequency band includes:
[0193] Obtain band mapping information, where the band mapping information is used to determine the mapping relationship between at least two target sub-bands corresponding to the compressed frequency band and at least two initial sub-bands corresponding to the second frequency band; perform feature extension on the target feature information corresponding to the compressed frequency band in the target band feature information based on the band mapping information to obtain extended feature information corresponding to the second frequency band.
[0194] Among them, the band mapping information is used to determine the mapping relationship between at least two target sub-bands corresponding to the compressed frequency band and at least two initial sub-bands corresponding to the second frequency band. When performing feature compression, the terminal or the voice sending end performs feature compression on the initial feature information corresponding to the second frequency band in the initial band feature information based on this mapping relationship to obtain the target feature information corresponding to the compressed frequency band. Then, when performing feature extension, the terminal performs feature extension on the target feature information corresponding to the compressed frequency band in the target band feature information based on this mapping relationship to maximize the restoration of the initial feature information corresponding to the second frequency band and obtain the extended feature information corresponding to the second frequency band.
[0195] Specifically, the terminal can obtain frequency band mapping information, and perform feature expansion on the target feature information corresponding to the compressed frequency band in the target frequency band feature information based on the frequency band mapping information to obtain the expanded feature information corresponding to the second frequency band. The voice receiving end and the voice sending end can pre-agree on the default frequency band mapping information. The voice sending end performs feature compression based on the default frequency band mapping information, and the voice receiving end performs feature expansion based on the default frequency band mapping information. The voice receiving end and the voice sending end can also pre-agree on multiple candidate frequency band mapping information. The voice sending end selects one of the frequency band mapping information for feature compression and generates compression identification information to send to the voice receiving end, so that the voice receiving end can determine the corresponding frequency band mapping information based on the compression identification information, and then perform feature expansion based on the frequency band mapping information.
[0196] In this embodiment, performing feature expansion on the target feature information corresponding to the compressed frequency band in the target frequency band feature information based on the frequency band mapping information to obtain the expanded feature information corresponding to the second frequency band can obtain relatively accurate expanded feature information, which helps to obtain a target voice signal with a higher restoration degree.
[0197] In one embodiment, the encoded voice data carries compression identification information, and obtaining the frequency band mapping information includes:
[0198] Obtaining the frequency band mapping information based on the compression identification information.
[0199] Specifically, when the terminal performs frequency band compression, it can generate compression identification information based on the frequency band mapping information used during feature compression, and associate the encoded voice data corresponding to the compressed voice signal with the corresponding compression identification information. Thus, when performing frequency band expansion later, the terminal can obtain the corresponding frequency band mapping information based on the compression identification information carried in the encoded voice data, and perform frequency band expansion on the decoded voice signal obtained by decoding based on the frequency band mapping information. For example, when the voice sending end performs frequency band compression, it can generate compression identification information based on the frequency band mapping information used during feature compression. Subsequently, the voice sending end sends the encoded voice data and the compression identification information to the voice receiving end together. The voice receiving end can then obtain the frequency band mapping information based on the compression identification information to perform frequency band expansion on the decoded voice signal obtained by decoding.
[0200] In one embodiment, performing feature expansion on the target feature information corresponding to the compressed frequency band in the target frequency band feature information based on the frequency band mapping information to obtain the expanded feature information corresponding to the second frequency band includes:
[0201] Use the target feature information of the current target sub-band corresponding to the current initial sub-band as the third intermediate feature information. From the target band feature information, obtain the target feature information of the sub-band whose band information is consistent with that of the current initial sub-band as the fourth intermediate feature information. Based on the third intermediate feature information and the fourth intermediate feature information, obtain the extended feature information corresponding to the current initial sub-band; based on the extended feature information corresponding to each initial sub-band, obtain the extended feature information corresponding to the second band.
[0202] Specifically, the terminal can determine the mapping relationship between at least two target sub-bands corresponding to the compressed band and at least two initial sub-bands corresponding to the second band based on the band mapping information. Thus, by performing feature expansion based on the target feature information corresponding to each target sub-band, the extended feature information of the initial sub-band corresponding to each target sub-band can be obtained, and finally the extended feature information corresponding to the second band can be obtained. The current initial sub-band refers to the initial sub-band for which the extended feature information is currently generated. When generating the extended feature information corresponding to the current initial sub-band, the terminal can use the target feature information of the current target sub-band corresponding to the current initial sub-band as the third intermediate feature information, and the third intermediate feature information is used to determine the amplitude of the frequency point in the extended feature information corresponding to the current initial sub-band. The terminal can obtain the target feature information of the sub-band whose band information is consistent with that of the current initial sub-band from the target band feature information as the fourth intermediate feature information, and the fourth intermediate feature information is used to determine the phase of the frequency point in the extended feature information corresponding to the current initial sub-band. Therefore, the terminal can obtain the extended feature information corresponding to the current initial sub-band based on the third intermediate feature information and the fourth intermediate feature information. After obtaining the extended feature information corresponding to each initial sub-band, the terminal can obtain the extended feature information corresponding to the second band based on the extended feature information corresponding to each initial sub-band, and the extended feature information corresponding to the second band is composed of the extended feature information corresponding to each initial sub-band.
[0203] For example, the target band feature information includes the target feature information corresponding to 0 - 8 kHz. The current initial sub-band is 6 - 8 kHz, and the target sub-band corresponding to the current initial sub-band is 6 - 6.4 kHz. The terminal can obtain the extended feature information corresponding to 6 - 8 kHz based on the target feature information corresponding to 6 - 6.4 kHz and the target feature information corresponding to 6 - 8 kHz in the target band feature information.
[0204] In this embodiment, by further subdividing the compressed band and the second band for feature expansion, the reliability of feature expansion can be improved, and the difference between the extended feature information corresponding to the second band and the initial feature information corresponding to the second band can be reduced. In this way, finally, a target voice signal with a relatively high similarity to the voice signal to be processed can be restored.
[0205] In one embodiment, both the third intermediate feature information and the fourth intermediate feature information include the target amplitudes and target phases corresponding to multiple target voice frequency points. Obtaining the extended feature information corresponding to the current initial sub-band based on the third intermediate feature information and the fourth intermediate feature information includes:
[0206] Based on the target amplitudes corresponding to the respective target voice frequency points in the third intermediate feature information, obtain the reference amplitudes of the respective initial voice frequency points corresponding to the current initial sub-band; when the fourth intermediate feature information is empty, add a random perturbation value to the phases of the respective initial voice frequency points corresponding to the current initial sub-band to obtain the reference phases of the respective initial voice frequency points corresponding to the current initial sub-band; when the fourth intermediate feature information is not empty, obtain the reference phases of the respective initial voice frequency points corresponding to the current initial sub-band based on the target phases corresponding to the respective target voice frequency points in the fourth intermediate feature information; and obtain the extended feature information corresponding to the current initial sub-band based on the reference amplitudes and reference phases of the respective initial voice frequency points corresponding to the current initial sub-band.
[0207] Specifically, for the amplitude of the frequency point, the terminal may use the target amplitudes corresponding to the respective target voice frequency points in the third intermediate feature information as the reference amplitudes of the respective initial voice frequency points corresponding to the current initial sub-band. For the phase of the frequency point, if the fourth intermediate feature information is empty, the terminal adds a random perturbation value to the target phases of the respective target voice frequency points corresponding to the current target sub-band to obtain the reference phases of the respective initial voice frequency points corresponding to the current initial sub-band. It can be understood that if the fourth intermediate feature information is empty, it means that the current initial sub-band does not exist in the target frequency band feature information, and this part has no energy and no phase. However, when converting the frequency domain signal to the time domain signal, the frequency point needs to have an amplitude and a phase. The amplitude can be obtained by copying, and the phase can be obtained by adding a random perturbation value. Moreover, the human ear is not sensitive to the high-frequency phase, and randomly assigning a value to the high-frequency part of the phase has little impact. If the fourth intermediate feature information is not empty, the terminal may obtain the target phase of the target voice frequency point whose frequency is the same as that of the initial voice frequency point from the fourth intermediate feature information as the reference phase of the initial voice frequency point. That is, the reference phase corresponding to the initial voice frequency point may follow the original phase. Among them, the random perturbation value is a random phase value. It can be understood that the value of the reference phase needs to be within the value range of the phase.
[0208] For example, the target band feature information includes the target feature information corresponding to 0 - 8 kHz, and the extended band feature information includes the extended feature information corresponding to 0 - 24 kHz. If the current initial sub - band is 6 - 8 kHz and the target sub - band corresponding to the current initial sub - band is 6 - 6.4 kHz, the terminal can use the target amplitudes of the target voice frequency points corresponding to 6 - 6.4 kHz as the reference amplitudes of the respective initial voice frequency points corresponding to 6 - 8 kHz, and use the target phases of the target voice frequency points corresponding to 6 - 6.4 kHz as the reference phases of the respective initial voice frequency points corresponding to 6 - 8 kHz. If the current initial sub - band is 8 - 10 kHz and the target sub - band corresponding to the current initial sub - band is 6.4 - 6.8 kHz, the terminal can use the target amplitudes of the target voice frequency points corresponding to 6.4 - 6.8 kHz as the reference amplitudes of the respective initial voice frequency points corresponding to 8 - 10 kHz, and use the target phases of the target voice frequency points corresponding to 6.4 - 6.8 kHz plus a random perturbation value as the reference phases of the respective initial voice frequency points corresponding to 8 - 10 kHz.
[0209] It can be understood that the number of initial voice frequency points in the extended band feature information can be equal to the number of initial voice frequency points in the initial band feature information. The number of initial voice frequency points corresponding to the second band in the extended band feature information is greater than the number of target voice frequency points corresponding to the compressed band in the target band feature information, and the ratio of the number of initial voice frequency points to the number of target voice frequency points is the frequency band ratio of the extended band feature information to the target band feature information.
[0210] In this embodiment, in the extended feature information corresponding to the second band, the amplitude of the initial voice frequency point is the amplitude of the corresponding target voice frequency point, and the phase of the initial voice frequency point follows the original phase or is a random value, which can reduce the difference between the extended feature information corresponding to the second band and the initial feature information corresponding to the second band.
[0211] This application also provides an application scenario that applies the above - mentioned voice encoding and voice decoding methods. Specifically, the application of the voice encoding and voice decoding methods in this application scenario is as follows:
[0212] The encoding and decoding of voice signals play an important role in modern communication systems. The encoding and decoding of voice signals can effectively reduce the bandwidth of voice signal transmission, and play a decisive role in saving the storage and transmission costs of voice information and ensuring the integrity of voice information during the transmission process of the communication network.
[0213] The clarity of speech is directly related to the speech spectrum bandwidth. Traditional landline phones use narrowband speech, with a sampling rate of 8 kHz, resulting in poor sound quality, blurred voices, and low intelligibility. In contrast, current VoIP (Voice over Internet Protocol) phones typically use broadband speech, with a sampling rate of 16 kHz, offering better sound quality and clear, intelligible voices. Even better sound quality can be achieved with ultra-wideband or full-band speech, which can have a sampling rate of up to 48 kHz, providing higher sound fidelity. The speech encoders used at different sampling rates are either different or different modes of the same encoder, and the corresponding speech coding bitstream sizes also vary. Traditional speech encoders only support processing speech signals at specific sampling rates. For example, the AMR-NB (Adaptive MultiRate-Narrow Band Speech Codec) encoder only supports input signals of 8 kHz and below, while the AMR-WB (Adaptive Multi-Rate-Wideband Speech Codec) encoder only supports input signals of 16 kHz and below.
[0214] In addition, generally, the higher the sampling rate, the greater the bandwidth consumed by the speech coding bitstream. To achieve a better speech experience, the speech frequency band needs to be increased. For example, the sampling rate can be increased from 8 kHz to 16 kHz or even 48 kHz. However, existing solutions require modifying and replacing the speech codecs in the existing client and backend transmission systems. At the same time, the increase in speech transmission bandwidth will inevitably lead to an increase in operating costs. It can be understood that the end-to-end speech sampling rate in existing solutions is restricted by the settings of the speech encoder and cannot break through the speech frequency band to obtain a better sound quality experience. To improve the sound quality experience, the parameters of the speech codec must be modified or other speech codecs that support higher sampling rates must be replaced. This will inevitably lead to system upgrades, increased operating costs, as well as a large amount of development work and development cycle.
[0215] However, by using the speech coding and speech decoding methods of the present application, without changing the speech codec and signal transmission system of the existing call system, the speech sampling rate of the existing call system can be upgraded, achieving a call experience beyond the existing speech frequency band, effectively improving speech clarity and intelligibility, and the operating costs are basically not affected.
[0216] Reference Figure 6A, The voice sending end collects high-quality voice signals, performs non-linear frequency band compression processing on the voice signals, and compresses the original high-sampling-rate voice signals into low-sampling-rate voice signals supported by the voice encoder of the call system through non-linear frequency band compression processing. The voice sending end then performs voice encoding and channel encoding on the compressed voice signals and finally transmits them to the voice receiving end through the network.
[0217] 1. Non-linear frequency band compression processing
[0218] In view of the characteristic that the human ear is sensitive to low-frequency signals and insensitive to high-frequency signals, the voice sending end can compress the signals in the high-frequency part. For example, after a full-band 48 kHz signal (i.e., the sampling rate is 48 kHz and the frequency band range is within 24 kHz) undergoes non-linear frequency band compression, all frequency band information is concentrated into a 16 kHz signal range (i.e., the sampling rate is 16 kHz and the frequency band range is within 8 kHz), and the high-frequency signals above the 16 kHz sampling range are suppressed to zero, and then downsampled to a 16 kHz signal. The low-sampling-rate signal obtained through non-linear frequency band compression processing can be encoded using a conventional 16 kHz voice encoder to obtain bitstream data.
[0219] Taking the full-band 48 kHz signal as an example, the essence of non-linear frequency band compression is not to modify the signals below 6 kHz in the speech spectrum (i.e., the frequency spectrum), but only to compress the speech spectrum signals from 6 kHz to 24 kHz. If the full-band 48 kHz signal is compressed to a 16 kHz signal, when performing frequency band compression, the frequency band mapping information can be as Figure 6B shown. Before compression, the frequency band of the voice signal is 0 - 24 kHz, the first frequency band is 0 - 6 kHz, and the second frequency band is 6 - 24 kHz. The second frequency band can be further divided into 6 - 8 kHz, 8 - 10 kHz, 10 - 12 kHz, 12 - 18 kHz, and 18 - 24 kHz, a total of 5 sub-frequency bands. After compression, the frequency band of the voice signal can still be 0 - 24 kHz, the first frequency band is 0 - 6 kHz, the compressed frequency band is 6 - 8 kHz, and the third frequency band is 8 - 24 kHz. The compressed frequency band can be further divided into 6 - 6.4 kHz, 6.4 - 6.8 kHz, 6.8 - 7.2 kHz, 7.2 - 7.6 kHz, and 7.6 - 8 kHz, a total of 5 sub-frequency bands. 6 - 8 kHz corresponds to 6 - 6.4 kHz, 8 - 10 kHz corresponds to 6.4 - 6.8 kHz, 10 - 12 kHz corresponds to 6.8 - 7.2 kHz, 12 - 18 kHz corresponds to 7.2 - 7.6 kHz, and 18 - 24 kHz corresponds to 7.6 - 8 kHz.
[0220] First, perform a fast Fourier transform on the high-sampling-rate voice signal to obtain the amplitude and phase of each frequency point. The information in the first frequency band remains unchanged. Figure 6BThe statistical value of the amplitudes of the intermediate frequency points in each sub - frequency band on the left is used as the amplitude of the corresponding intermediate frequency points in the sub - frequency band on the right, and the phase of the intermediate frequency points in the sub - frequency band on the right can follow the original phase value. For example, the amplitudes of each frequency point in the range of 6 kHz - 8 kHz on the left are added and averaged, and this average value is used as the amplitude of each frequency point in the range of 6 kHz - 6.4 kHz on the right, while the phase values of each frequency point in the range of 6 kHz - 6.4 kHz on the right are the original phase values. The information of the third frequency band is cleared. The frequency - domain signal of 0 - 24 kHz on the right undergoes inverse Fourier transform and down - sampling processing to obtain the compressed speech signal. Reference Figure 6C , (a) is the speech signal before compression, and (b) is the speech signal after compression. Figure 6C In the upper part is the time - domain signal, and in the lower part is the frequency - domain signal.
[0221] It can be understood that although the clarity of the low - sampling - rate speech signal after non - linear frequency - band compression is not as good as that of the original high - sampling - rate speech signal, the sound signal is naturally intelligible without perceptible noise and discomfort. Therefore, even if the speech receiving end is an in - network device without modification, it does not affect the call experience. Thus, the method of this application has good compatibility.
[0222] Reference Figure 6A , after the speech receiving end receives the bitstream data, it performs channel decoding and speech decoding on the bitstream data, and then through non - linear frequency - band expansion processing, restores the low - sampling - rate speech signal to a high - sampling - rate speech signal, and finally plays the high - sampling - rate speech signal.
[0223] 2. Non - linear frequency - band expansion processing
[0224] Reference Figure 6D , contrary to non - linear frequency - band compression processing, non - linear frequency - band expansion processing is to re - expand the compressed 6 kHz - 8 kHz signal to a 6 kHz - 24 kHz speech spectrum signal. That is, after Fourier transform, the amplitude of the intermediate frequency point in the sub - frequency band before expansion will be used as the amplitude of the corresponding intermediate frequency point in the sub - frequency band after expansion, and the phase either follows the original phase or adds a random perturbation value to the phase value of the intermediate frequency point in the sub - frequency band before expansion. The expanded spectrum signal can obtain a high - sampling - rate speech signal after inverse Fourier transform. Although it is not a perfect restoration, it is relatively close to the original high - sampling speech signal in terms of listening perception, and there is a significant improvement in subjective experience. Reference Figure 6E , (a) is the spectrum of the original high - sampling - rate speech signal (i.e., the spectrum information corresponding to the speech signal to be processed), and (b) is the spectrum of the high - sampling speech signal after expansion (i.e., the spectrum information corresponding to the target speech signal).
[0225] In this embodiment, it is possible to achieve the effect of improving the sound quality with a small modification based on the existing call system, without affecting the call cost. Through the voice encoding and decoding methods of this application, the original voice codec can achieve the super-wideband encoding and decoding effect, realizing a call experience that transcends the existing voice frequency band, effectively improving the voice clarity and intelligibility.
[0226] It can be understood that the voice encoding and decoding methods of this application can be applied not only to voice calls but also to the storage of voice content, such as the voice in a video, as well as scenarios involving voice encoding and decoding applications such as voice messages.
[0227] It should be understood that although Figure 2 、 Figure 3 、 Figure 5 the steps in the flowcharts of Figure 2 、 Figure 3 、 Figure 5 are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover,
[0228] In one embodiment, as Figure 7A shown, a voice encoding device is provided. This device can be a software module, a hardware module, or a combination of both to become part of a computer device. Specifically, the device includes: a frequency band feature information acquisition module 702, a first target feature information determination module 704, a second target feature information determination module 706, a compressed voice signal generation module 708, and a voice signal encoding module 710, where:
[0229] The frequency band feature information acquisition module 702 is used to acquire the initial frequency band feature information corresponding to the voice signal to be processed.
[0230] The first target feature information determination module 704 is used to obtain the target feature information corresponding to the first frequency band based on the initial feature information corresponding to the first frequency band in the initial frequency band feature information.
[0231] The second target feature information determination module 706 is configured to perform feature compression on the initial feature information corresponding to the second frequency band in the initial frequency band feature information to obtain the target feature information corresponding to the compressed frequency band. The frequency of the first frequency band is less than the frequency of the second frequency band, and the frequency range of the second frequency band is greater than the frequency range of the compressed frequency band.
[0232] The compressed speech signal generation module 708 is configured to obtain the intermediate frequency band feature information based on the target feature information corresponding to the first frequency band and the target feature information corresponding to the compressed frequency band, and obtain the compressed speech signal corresponding to the speech signal to be processed based on the intermediate frequency band feature information.
[0233] The speech signal encoding module 710 is configured to perform encoding processing on the compressed speech signal through the speech encoding module to obtain the encoded speech data corresponding to the speech signal to be processed. The target sampling rate corresponding to the compressed speech signal is less than or equal to the supported sampling rate corresponding to the speech encoding module, and the target sampling rate is less than the sampling rate corresponding to the speech signal to be processed.
[0234] In one embodiment, the frequency band feature information acquisition module is further configured to acquire the speech signal to be processed collected by the speech acquisition device, perform Fourier transform processing on the speech signal to be processed, and obtain the initial frequency band feature information. The initial frequency band feature information includes the initial amplitude and initial phase corresponding to multiple initial speech frequency points.
[0235] In one embodiment, the second target feature information determination module includes:
[0236] The frequency band division unit is configured to perform frequency band division on the second frequency band to obtain at least two initial sub-frequency bands arranged in sequence; perform frequency band division on the compressed frequency band to obtain at least two target sub-frequency bands arranged in sequence.
[0237] The frequency band association unit is configured to determine the target sub-frequency band corresponding to each initial sub-segment based on the sub-frequency band sorting of the initial sub-frequency band and the target sub-frequency band;
[0238] The information conversion unit is configured to use the initial feature information of the current initial sub-frequency band corresponding to the current target sub-frequency band as the first intermediate feature information, obtain the initial feature information of the sub-frequency band whose frequency band information is consistent with that of the current target sub-frequency band from the initial frequency band feature information as the second intermediate feature information, and obtain the target feature information corresponding to the current target sub-frequency band based on the first intermediate feature information and the second intermediate feature information;
[0239] The information determination unit is configured to obtain the target feature information corresponding to the compressed frequency band based on the target feature information corresponding to each target sub-frequency band.
[0240] In one embodiment, both the first intermediate feature information and the second intermediate feature information include the initial amplitudes and initial phases corresponding to a plurality of initial voice frequency points. The information conversion unit is further configured to obtain the target amplitudes of the respective target voice frequency points corresponding to the current target sub-band based on the statistical values of the initial amplitudes corresponding to the respective initial voice frequency points in the first intermediate feature information, obtain the target phases of the respective target voice frequency points corresponding to the current target sub-band based on the initial phases corresponding to the respective initial voice frequency points in the second intermediate feature information, and obtain the target feature information corresponding to the current target sub-band based on the target amplitudes and target phases of the respective target voice frequency points corresponding to the current target sub-band.
[0241] In one embodiment, the compressed voice signal generation module is further configured to determine a third frequency band based on the frequency difference between the compressed frequency band and the second frequency band, set the target feature information corresponding to the third frequency band as invalid information, obtain the intermediate frequency band feature information based on the target feature information corresponding to the first frequency band, the target feature information corresponding to the compressed frequency band, and the target feature information corresponding to the third frequency band, perform an inverse Fourier transform process on the intermediate frequency band feature information to obtain an intermediate voice signal, where the sampling rate corresponding to the intermediate voice signal is the same as the sampling rate corresponding to the voice signal to be processed, and perform downsampling processing on the intermediate voice signal based on the supported sampling rate to obtain a compressed voice signal.
[0242] In one embodiment, the voice signal encoding module is further configured to perform voice encoding on the compressed voice signal through a voice encoding module to obtain first voice data, and perform channel encoding on the first voice data to obtain encoded voice data.
[0243] In one embodiment, as Figure 7B shown, the voice encoding device further includes:
[0244] A voice data sending module 712, configured to send the encoded voice data to a voice receiving end, so that the voice receiving end performs voice restoration processing on the encoded voice data to obtain a target voice signal corresponding to the voice signal to be processed.
[0245] In one embodiment, the voice data sending module is further configured to obtain compression identification information corresponding to the voice signal to be processed based on the second frequency band and the compressed frequency band, send the encoded voice data and the compression identification information to the voice receiving end, so that the voice receiving end performs decoding processing on the encoded voice data to obtain a compressed voice signal, and perform frequency band expansion on the compressed voice signal based on the compression identification information to obtain a target voice signal.
[0246] In one embodiment, as Figure 8As shown, a voice decoding device is provided. This device can be a software module, a hardware module, or a combination of both to form part of a computer device. Specifically, the device includes: a voice data acquisition module 802, a voice signal decoding module 804, a first extended feature information determination module 806, a second extended feature information determination module 808, a target voice signal determination module 810, and a voice signal playback module 812, where:
[0247] The voice data acquisition module 802 is used to acquire encoded voice data, which is obtained by performing voice compression processing on the voice signal to be processed.
[0248] The voice signal decoding module 804 is used to perform decoding processing on the encoded voice data through the voice decoding module to obtain a decoded voice signal, and the target sampling rate corresponding to the decoded voice signal is less than or equal to the supported sampling rate corresponding to the voice decoding module.
[0249] The first extended feature information determination module 806 is used to generate target frequency band feature information corresponding to the decoded voice signal, and obtain extended feature information corresponding to the first frequency band based on the target feature information corresponding to the first frequency band in the target frequency band feature information.
[0250] The second extended feature information determination module 808 is used to perform feature extension on the target feature information corresponding to the compressed frequency band in the target frequency band feature information to obtain extended feature information corresponding to the second frequency band; the frequency of the first frequency band is less than the frequency of the compressed frequency band, and the frequency range of the compressed frequency band is less than the frequency range of the second frequency band.
[0251] The target voice signal determination module 810 is used to obtain extended frequency band feature information based on the extended feature information corresponding to the first frequency band and the extended feature information corresponding to the second frequency band, and obtain the target voice signal corresponding to the voice signal to be processed based on the extended frequency band feature information. The sampling rate of the target voice signal is greater than the target sampling rate.
[0252] The voice signal playback module 812 is used to play the target voice signal.
[0253] In one embodiment, the voice signal decoding module is further used to perform channel decoding on the encoded voice data to obtain second voice data, and perform voice decoding on the second voice data through the voice decoding module to obtain a decoded voice signal.
[0254] In one embodiment, the second extended feature information determination module includes:
[0255] A mapping information acquisition unit, which is used to acquire frequency band mapping information. The frequency band mapping information is used to determine the mapping relationship between at least two target sub - frequency bands corresponding to the compressed frequency band and at least two initial sub - frequency bands corresponding to the second frequency band;
[0256] A feature expansion unit, configured to perform feature expansion on target feature information corresponding to a compressed frequency band in target frequency band feature information based on frequency band mapping information, so as to obtain expansion feature information corresponding to a second frequency band.
[0257] In one embodiment, the encoded speech data carries compression identification information, and the mapping information acquisition unit is further configured to obtain frequency band mapping information based on the compression identification information.
[0258] In one embodiment, the feature expansion unit is further configured to use the target feature information of the current target sub-band corresponding to the current initial sub-band as third intermediate feature information, and obtain, from the target frequency band feature information, target feature information corresponding to a sub-band whose frequency band information is the same as that of the current initial sub-band as fourth intermediate feature information, obtain the expansion feature information corresponding to the current initial sub-band based on the third intermediate feature information and the fourth intermediate feature information, and obtain the expansion feature information corresponding to the second frequency band based on the expansion feature information corresponding to each initial sub-band.
[0259] In one embodiment, both the third intermediate feature information and the fourth intermediate feature information include target amplitudes and target phases corresponding to multiple target speech frequency points. The feature expansion unit is further configured to obtain reference amplitudes of each initial speech frequency point corresponding to the current initial sub-band based on the target amplitudes corresponding to each target speech frequency point in the third intermediate feature information. When the fourth intermediate feature information is empty, add a random perturbation value to the phases of each initial speech frequency point corresponding to the current initial sub-band to obtain reference phases of each initial speech frequency point corresponding to the current initial sub-band. When the fourth intermediate feature information is not empty, obtain reference phases of each initial speech frequency point corresponding to the current initial sub-band based on the target phases corresponding to each target speech frequency point in the fourth intermediate feature information, and obtain the expansion feature information corresponding to the current initial sub-band based on the reference amplitudes and reference phases of each initial speech frequency point corresponding to the current initial sub-band.
[0260] For the specific limitations on the speech encoding and speech decoding devices, reference may be made to the limitations on the speech encoding and speech decoding methods in the foregoing text, which will not be elaborated here. Each module in the foregoing speech encoding and speech decoding devices may be implemented in whole or in part by software, hardware, and their combination. The foregoing modules may be embedded in the processor in the computer device in hardware form or be independent of it, or may be stored in the memory in the computer device in software form, so that the processor can call and execute the operations corresponding to the foregoing modules.
[0261] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as Figure 9As shown in the figure. The computer device includes a processor, a memory, a communication interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner. The wireless manner can be achieved through WIFI, carrier network, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements a voice decoding method, and when the computer program is executed by the processor, it implements a voice encoding method. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covered on the display screen, or a button, trackball, or touchpad set on the shell of the computer device, or an external keyboard, touchpad, or mouse, etc.
[0262] Those skilled in the art can understand that Figure 9 the structure shown in the figure is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0263] In one embodiment, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, the steps in the above method embodiments are implemented.
[0264] In one embodiment, a computer-readable storage medium is provided, storing a computer program. When the computer program is executed by the processor, the steps in the above method embodiments are implemented.
[0265] In one embodiment, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the steps in the above method embodiments.
[0266] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above various methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0267] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0268] The above-described embodiments only represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.
Claims
1. A voice coding method, characterized in that, The method includes: Obtaining initial band feature information corresponding to a speech signal to be processed; Obtaining target feature information corresponding to a first frequency band based on initial feature information corresponding to the first frequency band in the initial band feature information; Performing feature compression on initial feature information corresponding to a second frequency band in the initial band feature information to obtain target feature information corresponding to a compressed frequency band, where the frequency of the first frequency band is less than the frequency of the second frequency band, and the frequency range of the second frequency band is greater than the frequency range of the compressed frequency band; Obtaining intermediate band feature information based on the target feature information corresponding to the first frequency band and the target feature information corresponding to the compressed frequency band, and obtaining a compressed speech signal corresponding to the speech signal to be processed based on the intermediate band feature information; Performing encoding processing on the compressed speech signal through a speech encoding module to obtain encoded speech data corresponding to the speech signal to be processed, where the target sampling rate corresponding to the compressed speech signal is less than or equal to the supported sampling rate corresponding to the speech encoding module, and the target sampling rate is less than the sampling rate corresponding to the speech signal to be processed.
2. The method according to claim 1, characterized in that, The obtaining of the initial band feature information corresponding to the speech signal to be processed includes: Obtaining the speech signal to be processed collected by a speech acquisition device; Performing Fourier transform processing on the speech signal to be processed to obtain the initial band feature information, where the initial band feature information includes initial amplitudes and initial phases corresponding to multiple initial speech frequency points.
3. The method according to claim 1, wherein The performing of feature compression on the initial feature information corresponding to the second frequency band in the initial band feature information to obtain the target feature information corresponding to the compressed frequency band includes: Dividing the second frequency band into at least two initially sub - frequency bands arranged in order; Dividing the compressed frequency band into at least two target sub - frequency bands arranged in order; Determining the target sub - frequency band corresponding to each initial sub - segment based on the sub - frequency band sorting of the initial sub - frequency band and the target sub - frequency band; Taking the initial feature information of the current initial sub - frequency band corresponding to the current target sub - frequency band as the first intermediate feature information, obtaining the initial feature information of the sub - frequency band with the same band information as the current target sub - frequency band from the initial band feature information as the second intermediate feature information, and obtaining the target feature information corresponding to the current target sub - frequency band based on the first intermediate feature information and the second intermediate feature information; Obtaining the target feature information corresponding to the compressed frequency band based on the target feature information corresponding to each target sub - frequency band.
4. The method according to claim 3, wherein Both the first intermediate feature information and the second intermediate feature information include initial amplitudes and initial phases corresponding to multiple initial speech frequency points; The obtaining of the target feature information corresponding to the current target sub - frequency band based on the first intermediate feature information and the second intermediate feature information includes: Obtaining the target amplitudes of each target speech frequency point corresponding to the current target sub - frequency band based on the statistical value of the initial amplitudes corresponding to each initial speech frequency point in the first intermediate feature information; Obtaining the target phases of each target speech frequency point corresponding to the current target sub - frequency band based on the initial phases corresponding to each initial speech frequency point in the second intermediate feature information; Obtain the target feature information corresponding to the current target sub-band based on the target amplitudes and target phases of the respective target speech frequency points corresponding to the current target sub-band.
5. The method according to claim 1, wherein The obtaining of the intermediate frequency band feature information based on the target feature information corresponding to the first frequency band and the target feature information corresponding to the compressed frequency band, and obtaining the compressed speech signal corresponding to the speech signal to be processed based on the intermediate frequency band feature information, includes: Determine a third frequency band based on the frequency difference between the compressed frequency band and the second frequency band, and set the target feature information corresponding to the third frequency band as invalid information; Obtain the intermediate frequency band feature information based on the target feature information corresponding to the first frequency band, the target feature information corresponding to the compressed frequency band, and the target feature information corresponding to the third frequency band; Perform an inverse Fourier transform process on the intermediate frequency band feature information to obtain an intermediate speech signal, where the sampling rate corresponding to the intermediate speech signal is the same as the sampling rate corresponding to the speech signal to be processed; Perform downsampling processing on the intermediate speech signal based on the supported sampling rate to obtain the compressed speech signal.
6. The method according to claim 1, characterized in that, The encoding the compressed speech signal through a speech encoding module to obtain the encoded speech data corresponding to the speech signal to be processed, includes: Perform speech encoding on the compressed speech signal through the speech encoding module to obtain first speech data; Perform channel encoding on the first speech data to obtain the encoded speech data.
7. The method according to any one of claims 1 to 6, characterized in that The method further includes: Send the encoded speech data to a speech receiving end, so that the speech receiving end performs speech restoration processing on the encoded speech data to obtain the target speech signal corresponding to the speech signal to be processed, and play the target speech signal.
8. The method according to claim 7, wherein The sending the encoded speech data to a speech receiving end, so that the speech receiving end performs speech restoration processing on the encoded speech data to obtain the target speech signal corresponding to the speech signal to be processed, and play the target speech signal, includes: Obtain the compression identification information corresponding to the speech signal to be processed based on the second frequency band and the compressed frequency band; Send the encoded speech data and the compression identification information to the speech receiving end, so that the speech receiving end performs decoding processing on the encoded speech data to obtain a compressed speech signal, perform frequency band expansion on the compressed speech signal based on the compression identification information to obtain the target speech signal, and play the target speech signal.
9. A voice decoding method, characterized in that, The method includes: Obtain encoded speech data, where the encoded speech data is obtained by performing speech compression processing on a speech signal to be processed; Perform decoding processing on the encoded speech data through a speech decoding module to obtain a decoded speech signal, where the target sampling rate corresponding to the decoded speech signal is less than or equal to the supported sampling rate corresponding to the speech decoding module; Generate the target frequency band feature information corresponding to the decoded speech signal, and obtain the extended feature information corresponding to the first frequency band based on the target feature information corresponding to the first frequency band in the target frequency band feature information; Perform feature expansion on the target feature information corresponding to the compressed frequency band in the target frequency band feature information to obtain the expanded feature information corresponding to the second frequency band; the frequency of the first frequency band is less than the frequency of the compressed frequency band, and the frequency range of the compressed frequency band is less than the frequency range of the second frequency band; Obtain the expanded frequency band feature information based on the expanded feature information corresponding to the first frequency band and the expanded feature information corresponding to the second frequency band, and obtain the target voice signal corresponding to the voice signal to be processed based on the expanded frequency band feature information, and the sampling rate of the target voice signal is greater than the target sampling rate; Play the target voice signal.
10. The method according to claim 9, wherein The decoding the encoded voice data by the voice decoding module to obtain a decoded voice signal includes: Perform channel decoding on the encoded voice data to obtain second voice data; Perform voice decoding on the second voice data by the voice decoding module to obtain the decoded voice signal.
11. The method according to claim 9, wherein The performing feature expansion on the target feature information corresponding to the compressed frequency band in the target frequency band feature information to obtain the expanded feature information corresponding to the second frequency band includes: Obtain frequency band mapping information, where the frequency band mapping information is used to determine the mapping relationship between at least two target sub-bands corresponding to the compressed frequency band and at least two initial sub-bands corresponding to the second frequency band; Perform feature expansion on the target feature information corresponding to the compressed frequency band in the target frequency band feature information based on the frequency band mapping information to obtain the expanded feature information corresponding to the second frequency band.
12. The method according to claim 11, wherein The encoded voice data carries compression identification information, and the obtaining the frequency band mapping information includes: Obtain the frequency band mapping information based on the compression identification information.
13. The method according to claim 11, wherein The performing feature expansion on the target feature information corresponding to the compressed frequency band in the target frequency band feature information based on the frequency band mapping information to obtain the expanded feature information corresponding to the second frequency band includes: Use the target feature information of the current target sub-band corresponding to the current initial sub-band as the third intermediate feature information, and obtain the target feature information of the sub-band with the same frequency band information as the current initial sub-band from the target frequency band feature information as the fourth intermediate feature information, and obtain the expanded feature information corresponding to the current initial sub-band based on the third intermediate feature information and the fourth intermediate feature information; Obtain the expanded feature information corresponding to the second frequency band based on the expanded feature information corresponding to each initial sub-band.
14. The method according to claim 13, wherein Both the third intermediate feature information and the fourth intermediate feature information include the target amplitude and target phase corresponding to multiple target voice frequency points; The obtaining the expanded feature information corresponding to the current initial sub-band based on the third intermediate feature information and the fourth intermediate feature information includes: Based on the target amplitudes corresponding to each target voice frequency point in the third intermediate feature information, obtain the reference amplitude of each initial voice frequency point corresponding to the current initial sub-band; When the fourth intermediate feature information is empty, add a random perturbation value to the phase of each initial voice frequency point corresponding to the current initial sub-band to obtain the reference phase of each initial voice frequency point corresponding to the current initial sub-band; When the fourth intermediate feature information is not empty, obtain the reference phases of the respective initial speech frequency points corresponding to the current initial sub-band based on the target phases corresponding to the respective target speech frequency points in the fourth intermediate feature information; Obtain the extended feature information corresponding to the current initial sub-band based on the reference amplitudes and reference phases of the respective initial speech frequency points corresponding to the current initial sub-band.
15. A voice coding device, characterized in that, The device includes: A band feature information acquisition module, configured to acquire initial band feature information corresponding to a speech signal to be processed; A first target feature information determination module, configured to obtain target feature information corresponding to a first frequency band based on the initial feature information corresponding to the first frequency band in the initial band feature information; A second target feature information determination module, configured to perform feature compression on the initial feature information corresponding to a second frequency band in the initial band feature information to obtain target feature information corresponding to a compressed frequency band, where the frequency of the first frequency band is less than the frequency of the second frequency band, and the frequency range of the second frequency band is greater than the frequency range of the compressed frequency band; A compressed speech signal generation module, configured to obtain intermediate band feature information based on the target feature information corresponding to the first frequency band and the target feature information corresponding to the compressed frequency band, and obtain a compressed speech signal corresponding to the speech signal to be processed based on the intermediate band feature information; A speech signal encoding module, configured to perform encoding processing on the compressed speech signal through a speech encoding module to obtain encoded speech data corresponding to the speech signal to be processed, where the target sampling rate of the compressed speech signal is less than or equal to the supported sampling rate corresponding to the speech encoding module, and the target sampling rate is less than the sampling rate corresponding to the speech signal to be processed.
16. A voice decoding device, characterized in that, The device includes: A speech data acquisition module, configured to acquire encoded speech data, where the encoded speech data is obtained by performing speech compression processing on a speech signal to be processed; A speech signal decoding module, configured to perform decoding processing on the encoded speech data through a speech decoding module to obtain a decoded speech signal, where the target sampling rate of the decoded speech signal is less than or equal to the supported sampling rate corresponding to the speech decoding module; A first extended feature information determination module, configured to generate target band feature information corresponding to the decoded speech signal, and obtain extended feature information corresponding to the first frequency band based on the target feature information corresponding to the first frequency band in the target band feature information; A second extended feature information determination module, configured to perform feature extension on the target feature information corresponding to the compressed frequency band in the target band feature information to obtain extended feature information corresponding to the second frequency band; the frequency of the first frequency band is less than the frequency of the compressed frequency band, and the frequency range of the compressed frequency band is less than the frequency range of the second frequency band; A target speech signal determination module, configured to obtain extended band feature information based on the extended feature information corresponding to the first frequency band and the extended feature information corresponding to the second frequency band, and obtain a target speech signal corresponding to the speech signal to be processed based on the extended band feature information, where the sampling rate of the target speech signal is greater than the target sampling rate; A speech signal playback module, configured to play the target speech signal.
17. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 8 or 9 to 14 are implemented.
18. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 or 9 to 14 are implemented.
19. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 or 9 to 14 are implemented.
Citation Information
Patent Citations
An encoding method of audio data SBC algorithm and Bluetooth stereo subsystem
CN101217038A
Speech audio encoding device, speech audio decoding device, speech audio encoding method, and speech audio decoding method
CN104737227A