Voice signal processing method, device, electronic device and computer storage medium

By extracting frame encoding at the sending end and restoring the neural network model at the receiving end, the problem of high bandwidth consumption in voice calls is solved, and high-quality voice transmission at low code rates is achieved, reducing operational costs.

CN115910081BActive Publication Date: 2025-08-05TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110898046.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-08-05
Publication Date
2025-08-05
Estimated Expiration
2041-08-05

AI Technical Summary

Technical Problem

The prior art is difficult to effectively reduce transmission bandwidth in voice calls, especially in scenarios where bandwidth is limited or consumption is high, resulting in increased operating costs.

Method used

The voice signal is extracted and encoded at the sending end, and interpolated and restored using the neural network model. The receiver compensates and repairs the voice signal through the interpolation and neural network model, reducing bandwidth transmission and maintaining high voice quality.

Benefits of technology

Maintaining good voice intelligibility at extremely low code rates significantly reduces operating costs, especially suitable for scenarios with limited call bandwidth such as large-scale voice conferences and live voice broadcasts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115910081B_ABST
    Figure CN115910081B_ABST
Patent Text Reader

Abstract

The embodiments of the present application provide a method, device, electronic device and computer storage medium for voice signal processing, which relate to the fields of artificial intelligence and cloud technology. It includes: receiving a coded stream corresponding to a voice signal to be processed, the coded stream is obtained by encoding and processing the original voice signal of each frame in the discontinuous voice signal by the transmitting end device, and the discontinuous voice signal is obtained by extracting the frame of the voice signal to be processed according to the set frame interval; decoding the coded stream to obtain the original voice signal of each frame, and determining the frequency domain characteristics of the original voice signal of each frame; restoring (interpolation and neural network model) the frequency domain characteristics of the original voice signal of each frame to obtain the frequency domain characteristics of the reconstructed voice signal of each frame; performing frequency-time transformation on the frequency domain characteristics of the reconstructed voice signal of each frame to obtain the target voice signal. In the present application, the transmitting end device encodes the voice signal after frame extraction and sends it to the receiving end, effectively reducing the bandwidth.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the fields of artificial intelligence and cloud technology. Specifically, the present application relates to a speech signal processing method, device, electronic device and computer storage medium. Background Art

[0002] Voice coding and decoding technology plays a crucial role in modern communication systems. For example, in voice communication applications, the transmitting device converts the collected analog voice signal into a digital voice signal through an analog-to-digital conversion circuit. The digital signal is then compressed by a voice encoder and packaged according to the communication network transmission format and protocol. The receiving device decodes the data packets and processes them through a voice decoder to obtain a digital voice signal, which is then played back. Voice coding and decoding technology effectively reduces the bandwidth required for voice signal transmission, playing a crucial role in saving voice signal storage and transmission costs and ensuring the integrity of voice information during transmission over the communication network. Therefore, for scenarios where call bandwidth is limited or bandwidth consumption is high, how to more effectively reduce transmission bandwidth is a pressing technical issue. Summary of the Invention

[0003] The present application provides a speech signal processing method, device, electronic device and computer storage medium, which can effectively reduce bandwidth.

[0004] In one aspect, an embodiment of the present application provides a method for processing a speech signal, the method comprising:

[0005] Receive the coded stream corresponding to the voice signal to be processed. The coded stream is obtained by encoding each frame of the original voice signal in the discontinuous voice signal by the transmitting device. The discontinuous voice signal is obtained by extracting the frame of the voice signal to be processed according to the set frame interval;

[0006] Decoding the encoded code stream to obtain the original speech signal of each frame, and determining the frequency domain characteristics of the original speech signal of each frame;

[0007] Based on the frequency domain characteristics of the original speech signals of each frame, the original speech signals of each frame are interpolated to obtain the frequency domain characteristics of the compensation frame signals between the original speech signals of each frame;

[0008] Inputting the frequency domain features of each frame of original speech signal and each compensated frame signal into the trained neural network model to obtain the spectral gain of each frame of speech signal to be reconstructed, wherein each frame of speech signal to be reconstructed includes each frame of original speech signal and each compensated frame signal;

[0009] Determining the frequency domain characteristics of the reconstructed speech signal of each frame based on the spectral gain of the speech signal to be reconstructed in each frame and the frequency domain characteristics of the speech signal to be reconstructed in each frame;

[0010] The frequency domain features of the reconstructed speech signals of each frame are subjected to frequency-time transformation to obtain the target speech signal.

[0011] On the other hand, an embodiment of the present application provides a method for processing a speech signal, the method comprising:

[0012] Acquire a speech signal to be processed, perform frame extraction on the speech signal to be processed according to a set frame interval, and obtain a discontinuous speech signal;

[0013] Encoding each frame of the original speech signal in the discontinuous speech signal to obtain the encoded code stream corresponding to the speech signal to be processed;

[0014] The coded code stream is sent to the receiving device, so that the receiving device performs the following processing on the coded code stream to obtain the target speech signal:

[0015] Decoding the encoded code stream to obtain the original speech signal of each frame, and determining the frequency domain characteristics of the original speech signal of each frame;

[0016] Based on the frequency domain characteristics of the original speech signals of each frame, the original speech signals of each frame are interpolated to obtain the frequency domain characteristics of the compensation frame signals between the original speech signals of each frame;

[0017] Inputting the frequency domain features of each frame of original speech signal and each compensated frame signal into the trained neural network model to obtain the spectral gain of each frame of speech signal to be reconstructed, wherein each frame of speech signal to be reconstructed includes each frame of original speech signal and each compensated frame signal;

[0018] Determining the frequency domain characteristics of the reconstructed speech signal of each frame based on the spectral gain of the speech signal to be reconstructed in each frame and the frequency domain characteristics of the speech signal to be reconstructed in each frame;

[0019] The target speech signal is determined based on the frequency domain characteristics of the speech signal reconstructed in each frame.

[0020] In another aspect, an embodiment of the present application provides a speech signal processing device, the device comprising:

[0021] A voice signal receiving module for receiving a coded stream corresponding to a voice signal to be processed, which is obtained by encoding each frame of the original voice signal in the discontinuous voice signal by the transmitting device. The discontinuous voice signal is obtained by extracting the frame of the voice signal to be processed according to a set frame interval;

[0022] A decoding module is used to decode the above-mentioned coded code stream to obtain the original speech signal of each frame and determine the frequency domain characteristics of the original speech signal of each frame;

[0023] An interpolation processing module is used to perform interpolation processing on each frame of original speech signal based on the frequency domain characteristics of each frame of original speech signal to obtain the frequency domain characteristics of the compensation frame signal between each frame of original speech signal;

[0024] A spectral gain determination module is used to input the frequency domain features of each frame of the original speech signal and the frequency domain features of each compensated frame signal into a trained neural network model to obtain the spectral gain of each frame of the speech signal to be reconstructed, where each frame of the speech signal to be reconstructed includes each frame of the original speech signal and each compensated frame signal;

[0025] The reconstructed speech signal determination module is used to determine the frequency domain characteristics of the reconstructed speech signal of each frame based on the spectral gain of the speech signal to be reconstructed in each frame and the frequency domain characteristics of the speech signal to be reconstructed in each frame;

[0026] The speech processing module is used to perform frequency-time transformation on the frequency domain features of the reconstructed speech signal of each frame to obtain the target speech signal.

[0027] Optionally, when the above-mentioned interpolation processing module performs interpolation processing on each frame of original voice signals based on the frequency domain characteristics of each frame of original voice signals to obtain the frequency domain characteristics of the compensation frame signals between each frame of original voice signals, it is specifically used to: for each pair of adjacent frame signals in each frame of original voice signals, based on the frequency domain characteristics of each frame of original voice signals in the adjacent frame signals, perform interpolation processing on the adjacent frame signals to obtain the frequency domain characteristics of the compensation frame signals between adjacent frame signals; and determine the frequency domain characteristics of the compensation frame signals between each frame of original voice signals according to the frequency domain characteristics of the compensation frame signals between each adjacent frame signals.

[0028] Optionally, for each pair of adjacent frame signals in each frame signal, the above-mentioned interpolation processing module performs interpolation processing on the adjacent frame signals based on the frequency domain characteristics of the original speech signal of each frame in the adjacent frame signals to obtain the frequency domain characteristics of the compensation frame signals between the adjacent frame signals. It is specifically used to: obtain a first correlation relationship between the frequency domain characteristics of the compensation frame signals between the adjacent frame signals and the frequency domain characteristics of the original speech signals of each frame in the adjacent frame signals; based on the first correlation relationship and the frequency domain characteristics of the original speech signal of each frame in the adjacent frame signals, perform interpolation processing on the adjacent frame signals to obtain the frequency domain characteristics of the compensation frame signals between the adjacent frame signals.

[0029] Optionally, for each pair of adjacent frame signals in each frame of original speech signal, the adjacent frame signals include a first signal and a second signal, and the first signal precedes the second signal; when the interpolation processing module performs interpolation processing on the adjacent frame signals based on the first association relationship and the frequency domain characteristics of each frame of original speech signal in the adjacent frame signals to obtain the frequency domain characteristics of the compensation frame signal between the adjacent frame signals, the interpolation processing module is specifically used to:

[0030] Based on the first association relationship and the frequency domain characteristics of each frame of the original speech signal in the adjacent frame signals, the adjacent frame signals are interpolated to obtain the frequency domain characteristics of the interpolated signals between the adjacent frame signals; the second association relationship between the frequency domain characteristics of the compensation frame signal between the adjacent frame signals, the frequency domain characteristics of the third signal of the adjacent frame signals and the frequency domain characteristics of the first signal is obtained, and the third signal is the previous frame signal of the first signal; based on the second association relationship, the frequency domain characteristics of the first signal and the frequency domain characteristics of the third signal, the adjacent frame signals are extrapolated to obtain the frequency domain characteristics of the extrapolated signals between the adjacent frame signals; the frequency domain characteristics of the interpolated signals of each frame are fused with the frequency domain characteristics of the extrapolated signals of each frame to obtain the frequency domain characteristics of the compensation frame signal between the adjacent frame signals.

[0031] Optionally, when the above-mentioned interpolation processing module fuses the frequency domain characteristics of the interpolation signal of each frame with the frequency domain characteristics of the extrapolation signal of each frame to obtain the frequency domain characteristics of the compensation frame signal between the signals of adjacent frames, it is specifically used to: obtain the first weight corresponding to the interpolation signal of each frame and the second weight corresponding to the extrapolation signal of each frame; for each frame interpolation signal in each frame interpolation signal, perform weighted processing on the frequency domain characteristics of the interpolation signal and the first weight corresponding to the interpolation signal to obtain the frequency domain characteristics of the weighted interpolation signal; for each frame extrapolation signal in each frame extrapolation signal, perform weighted processing on the frequency domain characteristics of the extrapolation signal and the second weight corresponding to the extrapolation signal to obtain the frequency domain characteristics of the weighted extrapolation signal; based on the frequency domain characteristics of the weighted interpolation signal of each frame and the frequency domain characteristics of the weighted extrapolation signal of each frame, determine the frequency domain characteristics of the compensation frame signal between adjacent frame signals.

[0032] Optionally, the above neural network model is trained through the following model training modules:

[0033] The model training module is used to obtain sample data, where the sample data includes multiple sample speech signals; for each sample speech signal, the sample speech signal is framed to obtain framed speech signals of each frame, and each framed speech signal is subjected to frame extraction processing according to a set frame interval to obtain discontinuous framed sample speech signals; the frequency domain characteristics of each frame in the discontinuous framed sample speech signals are determined; the frequency domain characteristics of each frame in the discontinuous framed sample speech signals are interpolated to obtain the frequency domain characteristics of each frame of the sample speech signal to be reconstructed; based on the frequency domain characteristics of each frame of the sample speech signal to be reconstructed and each framed speech signal, the real spectrum gain of each frame of the sample speech signal to be reconstructed is determined;

[0034] Repeat the following training steps until the loss value meets the training end condition to obtain the neural network model: input the frequency domain features of the sample speech signal to be reconstructed in each frame into the initial neural network model to obtain the predicted spectral gain corresponding to the sample speech signal to be reconstructed in each frame; based on each predicted spectral gain and each true spectral gain, determine the loss value corresponding to the initial neural network model. If the loss value meets the training end condition, end the training and obtain the neural network model; if not, adjust the model parameters of the initial neural network model and repeat the training steps.

[0035] Optionally, when determining the frequency domain features of each frame in the discontinuous frame-sampled speech signal, the above-mentioned model training module is specifically used to: perform linear time-frequency transformation on each frame-sampled speech signal in the discontinuous frame-sampled speech signal to obtain the linear frequency domain features of each frame-sampled speech signal; perform feature extraction on the linear frequency domain features of each frame-sampled speech signal to obtain the logarithmic frequency domain features of each frame-sampled speech signal, and use the logarithmic frequency domain features of each frame-sampled speech signal as the frequency domain features of each frame-sampled speech signal.

[0036] Optionally, when determining the frequency domain features of each frame of the original speech signal, the decoding module is specifically configured to: perform time-frequency transformation on each frame of the original speech signal to obtain the frequency domain features and phase features of each frame of the original speech signal;

[0037] When the above-mentioned speech processing module performs frequency-time transformation on the frequency domain features of the reconstructed speech signal of each frame to obtain the target speech signal, it is specifically used to: based on the phase features of the original speech signal of each frame and the frequency domain features of the reconstructed speech signal of each frame, perform frequency-time transformation on the frequency domain features of the reconstructed speech signal of each frame to obtain the time domain features of the reconstructed speech signal of each frame; and use the time domain features of the reconstructed speech signal of each frame as the target speech signal.

[0038] Optionally, if the above-mentioned frequency domain features are Mel-spectrum frequency domain features, when the above-mentioned speech processing module performs frequency-time transformation on the frequency domain features of the reconstructed speech signals of each frame to obtain the target speech signal, it is specifically used to: based on the Mel-spectrum frequency domain features of the reconstructed speech signals of each frame, perform frequency-time transformation on the frequency domain features of the restored speech signals of each frame to obtain the target speech signal.

[0039] In another aspect, an embodiment of the present application provides a speech signal processing device, the device comprising:

[0040] The voice signal acquisition module is used to obtain the voice signal to be processed, and perform frame extraction processing on the voice signal to be processed according to the set frame interval to obtain a discontinuous voice signal;

[0041] The encoding module is used to encode each frame of the original speech signal in the discontinuous speech signal to obtain the encoded code stream corresponding to the speech signal to be processed;

[0042] The sending module is used to send the coded code stream to the receiving device, so that the receiving device performs the following processing on the coded code stream to obtain the target speech signal:

[0043] Decoding the encoded code stream to obtain the original speech signal of each frame, and determining the frequency domain characteristics of the original speech signal of each frame;

[0044] Based on the frequency domain characteristics of the original speech signals of each frame, the original speech signals of each frame are interpolated to obtain the frequency domain characteristics of the compensation frame signals between the original speech signals of each frame;

[0045] Inputting the frequency domain features of each frame of original speech signal and each compensated frame signal into the trained neural network model to obtain the spectral gain of each frame of speech signal to be reconstructed, wherein each frame of speech signal to be reconstructed includes each frame of original speech signal and each compensated frame signal;

[0046] Determining the frequency domain characteristics of the reconstructed speech signal of each frame based on the spectral gain of the speech signal to be reconstructed in each frame and the frequency domain characteristics of the speech signal to be reconstructed in each frame;

[0047] The target speech signal is determined based on the frequency domain characteristics of the speech signal reconstructed in each frame.

[0048] On the other hand, an embodiment of the present application provides an electronic device, including a processor and a memory: the memory is configured to store a computer program, and when the computer program is executed by the processor, the processor executes a speech signal processing method.

[0049] On the other hand, an embodiment of the present application provides a computer-readable storage medium, which is used to store a computer program. When the computer program runs on a computer, the computer can execute the speech signal processing method.

[0050] In another aspect, the present application provides a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the speech signal processing method provided in any of the optional embodiments of the present application.

[0051] The beneficial effects of the technical solution provided by the embodiments of the present application are:

[0052] In an embodiment of the present application, at the transmitting end of the voice signal, the voice signal to be processed is subjected to frame extraction processing, and part of the signal in the voice signal (the original voice signal of each frame) is extracted for encoding, and the encoded code stream obtained by encoding is transmitted to the receiving end. At the receiving end of the voice signal, this part of the signal is first decoded to obtain the original voice signal of each frame, and then restored through interpolation processing and a neural network model based on the frequency domain characteristics of the original voice signal of each frame to obtain the frequency domain characteristics of the reconstructed voice signal of each frame. The purpose of the restoration is to restore the part of the signal to the complete frequency domain voice signal before the frame extraction processing, and finally the target voice signal is obtained based on the frequency domain characteristics of the reconstructed voice signal of each frame. In the scheme of the present application, since part of the signal of the voice signal to be processed is encoded and transmitted at the transmitting end, the bandwidth can be effectively reduced, especially for scenarios with limited call bandwidth and high call bandwidth consumption, and the operating costs can also be reduced. Furthermore, at the receiving end, the original speech signal of each frame is restored through interpolation processing and a neural network model, so that the final target speech signal has a higher speech quality, close to the speech quality of the speech signal to be processed. The solution of the present application can achieve good speech intelligibility even at an extremely low bit rate. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for describing the embodiments of the present application.

[0054] Figure 1 A flowchart of a method for processing a speech signal provided in an embodiment of the present application;

[0055] Figure 2 A flowchart of another method for processing speech signals provided in an embodiment of the present application;

[0056] Figure 3 A schematic diagram of the structure of an initial neural network provided in an embodiment of the present application;

[0057] Figure 4A schematic diagram of a spectrogram of a speech signal to be processed provided in an embodiment of the present application;

[0058] Figure 5 A schematic diagram of a processing flow of a voice signal processing method provided in an embodiment of the present application;

[0059] Figure 6 A flowchart of another method for processing speech signals provided in an embodiment of the present application;

[0060] Figure 7 A schematic diagram of the structure of a speech signal processing device provided in an embodiment of the present application;

[0061] Figure 8 A schematic diagram of the structure of another speech signal processing device provided in an embodiment of the present application;

[0062] Figure 9 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0063] The following describes in detail embodiments of the present application, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application, and are not to be construed as limiting the present application.

[0064] It will be understood by those skilled in the art that, unless expressly stated otherwise, the singular forms "a", "an", "said" and "the" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the specification of the present application refers to the presence of the features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof. It should be understood that when we refer to an element as being "connected" or "coupled" to another element, it may be directly connected or coupled to the other element, or there may be intermediate elements. In addition, "connected" or "coupled" as used herein may include wireless connections or wireless couplings. The term "and / or" used herein includes all or any units and all combinations of one or more associated listed items.

[0065] In related technologies, for some large-scale concurrent voice conferencing, voice live broadcast services, and some business applications with extremely limited bandwidth (for example, 2G baseband mode), since call bandwidth has a significant impact on operating costs, in order to save operating costs, or when bandwidth is very limited, how to more effectively reduce bandwidth is very worthwhile to study. Figure 1The flow chart of the existing speech coding and decoding scheme shown in FIG. 1 is a flow chart of the existing speech coding and decoding scheme shown in FIG. 1 , in which a frame (eg 20ms) of audio signal is collected by a collection device at the sending end ( Figure 1 ), continuously encode the audio signal of this frame (including Figure 1 The speech coding and channel coding shown in the figure, where speech coding belongs to the source coding in communication, mainly source digitization and compression; channel coding is mainly used to detect transmission errors and correct errors, compress data rate, and remove redundancy in the signal) to obtain code stream data. The sending device sends the code stream data to the receiving device through the network, and the receiving end decodes the received code stream data (including Figure 1 Channel decoding and speech decoding as shown in ) and audio signal playback ( Figure 1 In the existing solution, each frame of the signal must be encoded and transmitted, which consumes a lot of call bandwidth.

[0066] To address the aforementioned technical issues, embodiments of the present application propose a voice signal processing method. This method is applicable to any scenario requiring the transmission of voice signals from a transmitter to a receiver, particularly scenarios with limited call bandwidth or scenarios with high call bandwidth consumption (e.g., large-scale concurrent voice conferences and live voice broadcast services). Based on the method of this application, the transmitter encodes and transmits a portion of the voice signal to be processed (e.g., a portion of the signal obtained by extracting frames according to a set frame interval). This effectively reduces bandwidth in scenarios with limited call bandwidth. For example, if the set frame interval is set to N, the bit rate used to transmit the voice signal to be processed using this method is only one-Nth of that used in existing methods. Furthermore, at the receiver, each frame of the original voice signal is restored through interpolation and a neural network model, resulting in a target voice signal with high voice quality, close to that of the voice signal to be processed. This method of this application enables good voice intelligibility to be maintained even at extremely low bit rates.

[0067] The methods provided in the embodiments of the present application can be performed by any electronic device, which can be a server or a terminal device. The server can be a standalone physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server providing cloud computing services. The terminal device can be a smartphone, tablet computer, laptop computer, desktop computer, etc., but is not limited thereto and is not limited to these.

[0068] Optionally, the electronic device may be a cloud device, and the data processing / computation involved in the embodiments of the present application may be implemented based on cloud technology. For example, the step of interpolating the frequency domain features of each frame of the original speech signal to obtain the frequency domain features of the compensation frame signal between the frames of the original speech signal may be implemented using cloud computing.

[0069] Cloud computing is a computing model that distributes computing tasks across a resource pool consisting of a large number of computers, enabling various application systems to access computing power, storage space, and information services as needed. The network that provides these resources is called the "cloud." To users, these resources appear infinitely scalable and can be accessed at any time, used on demand, expanded at any time, and paid for on a per-use basis.

[0070] As a provider of cloud computing infrastructure, a cloud computing resource pool (referred to as a cloud platform, generally referred to as an IaaS (Infrastructure as a Service) platform) is established. Various types of virtual resources are deployed in the resource pool for external customers to choose and use. The cloud computing resource pool mainly includes: computing devices (virtualized machines, including operating systems), storage devices, and network devices.

[0071] Based on logical functional divisions, the PaaS (Platform as a Service) layer can be deployed on top of the IaaS (Infrastructure as a Service) layer, and the SaaS (Software as a Service) layer can be deployed on top of the PaaS layer. SaaS can also be deployed directly on top of IaaS. PaaS is a platform for software execution, such as databases and web containers. SaaS is a variety of business software, such as web portals and text messaging apps. Generally speaking, SaaS and PaaS are layers above IaaS.

[0072] Optionally, the neural network model in the methods provided in the embodiments of this application is trained using artificial intelligence (AI). Artificial intelligence (AI) refers to theories, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce new intelligent machines that can react in a manner similar to human intelligence. AI is the study of the design principles and implementation methods of various intelligent machines, enabling them to have perception, reasoning, and decision-making capabilities. AI technology is an interdisciplinary discipline encompassing a wide range of fields, encompassing both hardware and software technologies. Basic AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, speech processing, natural language processing, as well as machine learning / deep learning, autonomous driving, and smart transportation.

[0073] The following specific embodiments describe in detail the technical solution of the present application and how the technical solution of the present application solves the above-mentioned technical problems. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below in conjunction with the accompanying drawings.

[0074] Figure 2 A flow chart of a method for processing a voice signal provided by an embodiment of the present application is shown. The method can be executed by an electronic device, wherein the electronic device can be a transmitter of a voice signal or a receiver of a voice signal. Figure 2 As shown in , in this example, the execution subject may be a receiving end, and the method may include the following steps:

[0075] Step S110, receiving the coded stream corresponding to the voice signal to be processed, the coded stream is obtained by the transmitting end device encoding each frame of the original voice signal in the discontinuous voice signal, and the discontinuous voice signal is obtained by extracting the frame of the voice signal to be processed according to the set frame interval.

[0076] The voice signal to be processed refers to the voice signal that needs to be transmitted from the transmitting end to the receiving end. The voice signal to be processed can be an analog voice signal collected by a sound collection device, or a digital voice signal obtained by performing analog-to-digital conversion on the analog voice signal. For the coded bitstream corresponding to the voice signal to be processed received by the receiving end device, the transmitting end device performs the following processing on the voice signal to be processed to obtain the coded bitstream: obtain the voice signal to be processed, perform frame extraction processing on the voice signal to be processed according to a set frame interval to obtain a discontinuous voice signal; perform encoding processing on each frame of the original voice signal in the discontinuous voice signal to obtain the coded bitstream corresponding to the voice signal to be processed; and send the coded bitstream to the receiving end device.

[0077] Before extracting frames from the processed voice signal according to the set frame interval (M), the processed voice signal is first divided into frames based on the frame length, resulting in multiple frames of the original voice signal. The frame length is typically 10ms to 30ms, which translates to approximately 33 to 100 frames per second. The value of M is affected by the frame length and can range from 2 to 4. A larger M value results in a lower final voice coding bitrate and transmission bitrate, resulting in lower quality. Therefore, the value of M should be selected based on actual needs. For example, setting a frame interval of 2 means that one frame is extracted from every two frames of the original voice signal.

[0078] Speech signals, especially voiced ones, have strong correlation and similarity between frames. For long-duration speech signals (e.g., 90ms), the spectrum changes relatively slowly over time. This application leverages this characteristic of speech and employs sampling-based encoding and decoding. By extracting only a portion of the signal for encoding and then compensating and repairing the decoded discontinuous speech signal (partial signal), this approach maintains good speech intelligibility even at lower bit rates.

[0079] The discontinuous speech signal obtained by frame extraction refers to the fact that the extracted speech signal is discontinuous within the speech signal to be processed. For example, if the speech signal to be processed has 20 frames after framing, and a frame is extracted every two frames, the discontinuous speech signals after extraction are 7 frames: 1, 4, 7, 10, 13, 16, and 19. Within these 7 frames, every two frames are discontinuous.

[0080] In the present application scheme, after extracting multiple frames of discontinuous voice signals, each frame of the original voice signal in the discontinuous voice signal can be encoded separately. Since the present application scheme uses a non-continuous encoding method, a voice codec with an independent frame encoding method is selected, such as the iLBC (internet Low Bitrate Codec) voice codec. iLBC is an existing independent frame-based voice codec. Based on this voice codec, each frame of the original voice signal can be encoded.

[0081] Optionally, in the optional scheme of the present application, the processed voice signal can be subjected to frame extraction processing using a non-uniform frame extraction method. For example, a frame is extracted after a first set number of frames, and then a frame is extracted after a second set number of frames, and then another frame is extracted after the first set number of frames. The first set number of frames is not equal to the second set number of frames. The specific method of frame extraction is not limited in the scheme of the present application, and all of them are within the scheme protected by the present application.

[0082] Step S120: Decode the encoded code stream to obtain each frame of original speech signal, and determine the frequency domain characteristics of each frame of original speech signal.

[0083] After receiving the coded stream corresponding to the voice signal to be processed, the receiving device first decodes the coded stream. Typically, encoding corresponds to a number of frames of original voice signals, and decoding yields a corresponding number of frames of original voice signals. Time-frequency transformation is then performed on each frame of the original voice signal to obtain the frequency domain characteristics of each frame of the original voice signal. Time-frequency transformation refers to converting the time domain signal to the frequency domain. Optionally, the frequency domain characteristics of each frame of the original voice signal can be obtained by performing Fourier transformation on each frame of the original voice signal. Frequency domain characteristics include at least one of a power spectrum or a frequency spectrum. When performing time-frequency transformation, phase characteristics of each frame of the original voice signal can also be obtained.

[0084] Step S130 , performing interpolation processing on the original speech signals of the frames based on the frequency domain features of the original speech signals of the frames, and obtaining the frequency domain features of the compensation frame signals between the original speech signals of the frames.

[0085] In step S140, the frequency domain features of each frame of the original speech signal and the frequency domain features of each compensated frame signal are input into the trained neural network model to obtain the spectral gain of each frame of the speech signal to be reconstructed, where each frame of the speech signal to be reconstructed includes each frame of the original speech signal and each compensated frame signal.

[0086] Step S150, based on the spectral gain of each frame of the speech signal to be reconstructed and the frequency domain characteristics of each frame of the speech signal to be reconstructed, determining the frequency domain characteristics of each frame of the reconstructed speech signal.

[0087] Restoring (interpolating and using a neural network model) each frame of the original speech signal refers to compensating for and repairing the frequency domain characteristics of each frame of the original speech signal, so that the number of frames of the restored speech signal is close to that of the speech signal before frame extraction, and the frequency domain characteristics of the restored speech signal are close to the frequency domain characteristics of the speech signal before frame extraction. Ideally, both the number of frames and the frequency domain characteristics are the same after restoration as before frame extraction, so that the final target speech signal is closer to the speech signal to be processed. The main purpose of the interpolation process is to make the number of frames of the speech signal after interpolation close to that of the speech signal before frame extraction, and the main purpose of processing through the neural network model is to make the frequency domain characteristics of the processed speech signal close to the frequency domain characteristics of the speech signal before frame extraction.

[0088] The compensation frames between the original speech frames refer to the additional frames inserted between the original speech frames in a pair of adjacent frames. Ideally, the sum of the number of inserted frames and the number of frames in the original speech signals equals the number of frames before the extraction.

[0089] Optionally, an optional implementation method based on the spectral gain of each frame of the speech signal to be reconstructed and the frequency domain characteristics of each frame of the speech signal to be reconstructed is: multiplying each frame of the speech signal to be reconstructed with the corresponding spectral gain to obtain the reconstructed speech signal of each frame.

[0090] The neural network model described above is used to predict the spectral gain of speech signals. The model's input is the frequency domain features of each frame of speech signal, and its output is the spectral gain of each frame of speech signal. The training process of the neural network model is described in detail below.

[0091] In an optional solution of the present application, the frequency domain features of each frame of the original speech signal are interpolated, which can be at least one of interpolation and extrapolation. The specific interpolation process will be further described below and will not be repeated here.

[0092] Step S160: performing frequency-time transformation on the frequency domain features of the reconstructed speech signal of each frame to obtain a target speech signal.

[0093] Among them, after obtaining the frequency domain features of the reconstructed speech signal of each frame, it is also necessary to perform frequency-time transformation on the frequency domain features of the reconstructed speech signal of each frame to obtain the target speech signal.

[0094] In an optional solution of the present application, the frequency-time transformation of the frequency domain features of the reconstructed speech signal of each frame to obtain the target speech signal includes:

[0095] Based on the phase characteristics of the original speech signal of each frame and the frequency domain characteristics of the reconstructed speech signal of each frame, the frequency domain characteristics of the reconstructed speech signal of each frame are transformed into a time domain characteristic of the reconstructed speech signal of each frame;

[0096] The time domain features of the reconstructed speech signal of each frame are used as the target speech signal.

[0097] Frequency-time transformation refers to converting frequency-domain features into time-domain features. Frequency-time transformation can be an inverse Fourier transform. During the frequency-time transformation, the phase features of the reconstructed speech signal of each frame can be the phase features corresponding to the original speech signal of each frame.

[0098] Optionally, the frequency domain feature may be any one of a linear frequency domain feature, a logarithmic frequency domain feature, or a Bark domain feature. If the frequency domain feature is a logarithmic frequency domain feature, the frequency domain feature may be a Mel spectrum frequency domain feature. If the frequency domain feature is a Mel spectrum frequency domain feature, the frequency domain feature of the reconstructed speech signal of each frame is subjected to frequency-time transformation to obtain the target speech signal, including:

[0099] Based on the Mel spectrum frequency domain features of the reconstructed speech signal of each frame, the frequency domain features of the restored speech signal of each frame are transformed into a frequency-time transform to obtain the target speech signal.

[0100] When performing frequency-time transformation on the Mel spectrum frequency domain features, it is not necessary to reconstruct the phase features of the speech signal for each frame.

[0101] In an embodiment of the present application, at the transmitting end of the voice signal, the voice signal to be processed is subjected to frame extraction processing, and part of the signal in the voice signal (the original voice signal of each frame) is extracted for encoding, and the encoded code stream obtained by encoding is transmitted to the receiving end. At the receiving end of the voice signal, this part of the signal is first decoded to obtain the original voice signal of each frame, and then restoration (interpolation and neural network model) processing is performed based on the frequency domain characteristics of the original voice signal of each frame to obtain the frequency domain characteristics of the reconstructed voice signal of each frame. The purpose of the restoration is to restore the part of the signal to the complete frequency domain voice signal before the frame extraction processing, and finally the target voice signal is obtained based on the frequency domain characteristics of the reconstructed voice signal of each frame. In the scheme of the present application, since part of the signal of the voice signal to be processed is encoded and transmitted at the transmitting end, the bandwidth can be effectively reduced, especially for scenarios with limited call bandwidth and high call bandwidth consumption, and the operating costs can also be reduced.

[0102] Furthermore, at the receiving end, the original speech signal of each frame is restored through interpolation processing and a neural network model, so that the final target speech signal has a higher speech quality, close to the speech quality of the speech signal to be processed. The solution of the present application can achieve good speech intelligibility even at an extremely low bit rate.

[0103] The solution of this application can be used to solve the problem of voice transmission at extremely low bit rates. The solution of this application can transmit voice signals at only half or even a fraction of the existing voice coding bit rate. This can significantly reduce operating costs for application scenarios with very limited call bandwidth, large-scale concurrent voice conferencing, and voice live broadcasts.

[0104] In an optional embodiment of the present application, for the interpolation processing method described above (including at least one of the interpolation processing method and the extrapolation processing method), the interpolation processing of each frame of original speech signal based on the frequency domain characteristics of each frame of original speech signal to obtain the frequency domain characteristics of the compensation frame signal between the frames of original speech signal includes:

[0105] For each pair of adjacent frame signals in each frame of the original speech signal, interpolation processing is performed on the adjacent frame signals based on the frequency domain characteristics of each frame of the original speech signal in the adjacent frame signals to obtain the frequency domain characteristics of the compensation frame signal between the adjacent frame signals;

[0106] The frequency domain features of the compensated frame signals between the frames of original speech signals are determined according to the frequency domain features of the compensated frame signals between the adjacent frame signals.

[0107] Among them, the compensation frame signal between adjacent frame signals refers to other frame signals inserted between a pair of adjacent frame signals, and a pair of adjacent frame signals refers to two adjacent frame signals in each frame of original speech signals. For each pair of adjacent frame signals, based on the frequency domain characteristics of each frame of original speech signal in the adjacent frame signals, the adjacent frame signals are interpolated to predict the frequency domain characteristics of the compensation frame signals between the adjacent frame signals.

[0108] It can be understood that the compensation frame signal between each frame of the original speech signal is equal to the sum of the compensation frame signals between adjacent frame signals, and the number of compensation frames inserted between adjacent frames is equal to the set frame interval.

[0109] As an example, the frame interval is set to 2. If the adjacent frame signals are the fourth frame signal and the seventh frame signal, at least two frames of compensation frame signals can be inserted between these two frame signals, and the frequency domain characteristics of each compensation frame signal can be predicted based on the frequency domain characteristics of the fourth frame signal and the frequency domain characteristics of the seventh frame signal.

[0110] If the above-mentioned interpolation processing mode is an extrapolation processing mode, in an optional embodiment of the present application, for each pair of adjacent frame signals in each frame signal, the adjacent frame signals include a first signal and a second signal, and the first signal precedes the second signal; the above-mentioned interpolation processing is performed on the adjacent frame signals based on the frequency domain characteristics of each frame of the original speech signal in the adjacent frame signals to obtain the frequency domain characteristics of the compensated frame signal between the adjacent frame signals, including:

[0111] Acquire a second correlation relationship between a frequency domain feature of a compensation frame signal between adjacent frame signals, a frequency domain feature of a third signal of the adjacent frame signals, and a frequency domain feature of the first signal, where the third signal is a frame signal preceding the first signal;

[0112] Based on the second correlation relationship, the frequency domain characteristics of the first signal and the frequency domain characteristics of the third signal, adjacent frame signals are interpolated to obtain the frequency domain characteristics of the extrapolated signals between adjacent frame signals, and the frequency domain characteristics of the extrapolated signals are used as the frequency domain characteristics of the compensation frame signal.

[0113] Among them, considering the relationship between the signals of each frame, the restoration processing can be carried out by interpolation processing. The interpolation processing refers to the method of interpolating between adjacent frames based on the first signal in the adjacent frames and the previous frame signal (third signal) of the first signal.

[0114] As an example, the frame interval is set to 2. If the third signal is the first frame signal, and the adjacent frame signals are the fourth frame signal (first signal) and the seventh frame signal (second signal), then at least two frames of compensation frame signals can be inserted between these two frame signals, and the frequency domain characteristics of each compensation frame signal can be predicted based on the frequency domain characteristics of the fourth frame signal and the frequency domain characteristics of the first frame signal.

[0115] If the interpolation processing mode is an interpolation processing mode, in an optional embodiment of the present application, for each pair of adjacent frame signals in each frame signal, interpolation processing is performed on the adjacent frame signals based on the frequency domain characteristics of each frame of the original speech signal in the adjacent frame signals to obtain the frequency domain characteristics of the compensated frame signal between the adjacent frame signals, including:

[0116] Acquire a first correlation relationship between the frequency domain features of the compensated frame signal between adjacent frame signals and the frequency domain features of the original speech signal of each frame in the adjacent frame signals;

[0117] Based on the first association relationship and the frequency domain characteristics of each frame of the original speech signal in the adjacent frame signals, interpolation processing is performed on the adjacent frame signals to obtain the frequency domain characteristics of the compensated frame signals between the adjacent frame signals.

[0118] Among them, considering the relationship between the signals of each frame, interpolation processing can also be used for restoration processing. Interpolation processing refers to a method of interpolating between adjacent frames based on the first signal and the second signal in the adjacent frames.

[0119] As an example, the frame interval is set to 2. If the adjacent frame signals are the fourth frame signal (first signal) and the seventh frame signal (second signal), at least two frames of compensation frame signals can be inserted between these two frame signals, and the frequency domain characteristics of each compensation frame signal can be predicted based on the frequency domain characteristics of the first signal and the frequency domain characteristics of the second signal.

[0120] If the interpolation processing mode is an interpolation and extrapolation processing mode, in an optional embodiment of the present application, for each pair of adjacent frame signals in each frame of the original speech signal, interpolation processing is performed on the adjacent frame signals based on the first association relationship and the frequency domain characteristics of each frame of the original speech signal in the adjacent frame signals to obtain the frequency domain characteristics of the compensated frame signal between the adjacent frame signals, including:

[0121] Based on the first association relationship and the frequency domain characteristics of each frame of the original speech signal in the adjacent frame signals, interpolation processing is performed on the adjacent frame signals to obtain the frequency domain characteristics of the interpolated signal between the adjacent frame signals;

[0122] Acquire a second correlation relationship between a frequency domain feature of a compensation frame signal between adjacent frame signals, a frequency domain feature of a third signal of the adjacent frame signals, and a frequency domain feature of the first signal, where the third signal is a frame signal preceding the first signal;

[0123] Based on the second correlation relationship, the frequency domain characteristics of the first signal, and the frequency domain characteristics of the third signal, performing extrapolation processing on adjacent frame signals to obtain frequency domain characteristics of extrapolated signals between adjacent frame signals;

[0124] The frequency domain features of the interpolation signal of each frame are fused with the frequency domain features of the extrapolation signal of each frame to obtain the frequency domain features of the compensation frame signal between adjacent frame signals.

[0125] The present application does not limit whether the interpolation process or the extrapolation process is performed first. Interpolation can be performed first and then extrapolation, or extrapolation can be performed first and then interpolation. Regardless of whether interpolation or extrapolation is performed, the interpolation process results in a compensated frame signal between adjacent frame signals.

[0126] Optionally, fusing the frequency domain features of each frame interpolation signal with the frequency domain features of each frame extrapolation signal generally involves fusing the frequency domain features of each frame interpolation signal with the frequency domain features of a corresponding frame in the frequency domain features of each frame extrapolation signal.

[0127] In an optional solution of the present application, the above-mentioned fusion of the frequency domain features of the interpolated signal of each frame and the frequency domain features of the extrapolated signal of each frame to obtain the frequency domain features of the compensated frame signal between the signals of adjacent frames includes:

[0128] Obtaining a first weight corresponding to each frame interpolation signal and a second weight corresponding to each frame extrapolation signal;

[0129] For each frame interpolation signal, performing weighted processing on the frequency domain feature of the interpolation signal and a first weight corresponding to the interpolation signal to obtain the weighted frequency domain feature of the interpolation signal;

[0130] For each frame of the extrapolated signal, performing weighted processing on the frequency domain characteristics of the extrapolated signal and a second weight corresponding to the extrapolated signal to obtain the frequency domain characteristics of the weighted extrapolated signal;

[0131] Based on the frequency domain characteristics of the weighted interpolation signal of each frame and the frequency domain characteristics of the weighted extrapolation signal of each frame, the frequency domain characteristics of the compensation frame signal between adjacent frame signals are determined.

[0132] In this case, because the correlation between the two frame signals (the first signal and the second signal) in adjacent frame signals and the compensated frame signal is different from the correlation between the first signal and the third signal and the compensated frame signal, and thus the importance of the interpolated signal and the extrapolated signal relative to the compensated frame signal is different, a corresponding first weight is configured for each frame interpolated signal, and a corresponding second weight is configured for each frame extrapolated signal. The first weight represents the importance of the interpolated signal relative to the compensated frame signal, and the second weight represents the importance of the extrapolated signal relative to the compensated frame signal. The first weight and the second weight can be configured based on actual needs. The weight corresponding to each frame interpolated signal in each frame interpolated signal is the first weight, and the weight corresponding to each frame extrapolated signal in each frame extrapolated signal is the second weight. The sum of the first weight and the second weight is 1.

[0133] It can be understood that when determining the frequency domain characteristics of the compensation frame signal between adjacent frame signals based on the frequency domain characteristics of the weighted interpolation signal of each frame and the frequency domain characteristics of the weighted extrapolation signal of each frame, for each compensation frame signal, the frequency domain characteristics corresponding to the compensation frame signal are determined based on the frequency domain characteristics of the weighted interpolation signal and the frequency domain characteristics of the weighted extrapolation signal corresponding to the compensation frame signal.

[0134] In an optional solution of the present application, the above neural network model is trained by the following method:

[0135] Acquire sample data, where the sample data includes a plurality of sample speech signals;

[0136] For each sample speech signal, the sample speech signal is framed to obtain framed speech signals of each frame, and the framed speech signals of each frame are extracted according to a set frame interval to obtain discontinuous framed sample speech signals;

[0137] Determining frequency domain features of each frame in the discontinuous frame sample speech signal;

[0138] Restoring the frequency domain features of each frame in the discontinuous frame-drawing sample speech signal to obtain the frequency domain features of the sample speech signal to be reconstructed in each frame;

[0139] Determining the true spectrum gain of each frame of the sample speech signal to be reconstructed based on the frequency domain characteristics of each frame of the sample speech signal to be reconstructed and the framed speech signal of each frame;

[0140] Repeat the following training steps until the loss value meets the training end condition to obtain the neural network model:

[0141] Inputting the frequency domain features of each frame of the sample speech signal to be reconstructed into the initial neural network model to obtain the predicted spectral gain corresponding to each frame of the sample speech signal to be reconstructed;

[0142] Based on each predicted spectrum gain and each true spectrum gain, the loss value corresponding to the initial neural network model is determined. If the loss value meets the training end condition, the training is ended and the neural network model is obtained; if not, the model parameters of the initial neural network model are adjusted and the training steps are repeated.

[0143] Among them, for each sample speech signal, the sample speech signal can be subjected to frame extraction processing based on the same frame extraction processing method as the speech signal to be processed to obtain a discontinuous frame-extracted sample speech signal. Optionally, for each frame of the sample speech signal to be reconstructed, the true spectral gain of each frame of the sample speech signal to be reconstructed is determined based on the frequency domain characteristics of each frame of the sample speech signal to be reconstructed and the framed speech signal of each frame. Specifically, it may include: dividing the frequency domain characteristics of the framed speech signal corresponding to the sample speech signal to be reconstructed by the frequency domain characteristics of the sample speech signal to be reconstructed by the frame to be reconstructed to obtain the true spectral gain of the sample speech signal to be reconstructed.

[0144] Among them, if the frame interval is set to N, when the value of N is small and the fitting ability of the neural network model is strong enough, the quality of the final audio signal can reach a high level (close to the voice quality of the speech signal to be processed, with good speech intelligibility), and the bit rate used is only one-Nth of the existing solution.

[0145] The above loss value represents the difference between each predicted spectrum gain and each actual spectrum gain.

[0146] Optionally, the above initial neural network model can be a different structure such as LSTM (Long Short-Term Memory), RNN (Recurrent Neural Network), CNN (Convolutional Neural Networks), GRU (Gated Recurrent Units, Recurrent Neural Network), etc.

[0147] As an example, see Figure 3 The network structure diagram shown in Figure 2 shows that in this example, the initial neural network model ( Figure 3The deep learning network shown in is a gru network, which includes a cascaded input DENSE unit, two layers of GRU network units and an output DENSE unit. The input of the network is the power spectrum of the sample speech signal to be reconstructed in each frame ( Figure 3 The power spectrum of the N-frame interpolation signal shown in , in this example, the spectral feature is the power spectrum, N is the number of frames corresponding to the frame processing, and the interpolation signal is the sample speech signal to be reconstructed), the output is the gain value corresponding to the sample speech signal to be reconstructed in each frame ( Figure 3 The frequency point gains shown in ), when restoring, each gain value can be multiplied by the power spectrum of the speech signal to be reconstructed of the corresponding frame, and finally the power spectrum of the reconstructed speech signal of each frame restored by the neural network model (N frame signal power spectrum, ) is obtained.

[0148] In an optional solution of the present application, the above-mentioned determination of the frequency domain features of each frame in the discontinuous frame sample speech signal includes:

[0149] Performing a linear time-frequency transform on each frame of the discontinuous frame-sampled speech signal to obtain a linear frequency domain feature of each frame of the frame-sampled speech signal;

[0150] The linear frequency domain features of each frame sample speech signal are extracted to obtain the logarithmic frequency domain features of each frame sample speech signal, and the logarithmic frequency domain features of each frame sample speech signal are used as the frequency domain features of each frame sample speech signal.

[0151] Among them, during the training process of the neural network model, the frequency domain features of the sampled speech signals of each frame can be logarithmic frequency domain features. Compared with linear frequency domain features, logarithmic frequency domain features are closer to the effect of the real speech signal. Moreover, the number of logarithmic frequency domain features corresponding to the same speech signal is less than that of linear frequency domain features. Therefore, based on the logarithmic frequency domain features, the data processing amount is reduced and the model complexity is reduced during the training process of the model.

[0152] like Figure 6 As shown in , in this example, the execution subject may be a sending end, and the method may include the following steps:

[0153] Step S210: obtaining a speech signal to be processed, performing frame extraction processing on the speech signal to be processed according to a set frame interval to obtain a discontinuous speech signal;

[0154] Step S220, encoding each frame of the original speech signal in the discontinuous speech signal to obtain an encoded code stream corresponding to the speech signal to be processed;

[0155] Step S230: Send the coded stream to the receiving device, so that the receiving device performs the following processing on the coded stream to obtain the target speech signal:

[0156] The encoded code stream is decoded to obtain the original speech signal of each frame, and the frequency domain characteristics of the original speech signal of each frame are determined; based on the frequency domain characteristics of the original speech signal of each frame, the original speech signal of each frame is interpolated to obtain the frequency domain characteristics of the compensation frame signal between the original speech signals of each frame; the frequency domain characteristics of the original speech signal of each frame and the frequency domain characteristics of each compensation frame signal are input into the trained neural network model to obtain the spectrum gain of the speech signal to be reconstructed of each frame, and the speech signal to be reconstructed of each frame includes the original speech signal of each frame and the compensation frame signal; based on the spectrum gain of the speech signal to be reconstructed of each frame and the frequency domain characteristics of the speech signal to be reconstructed of each frame, the frequency domain characteristics of the reconstructed speech signal of each frame are determined; based on the frequency domain characteristics of the reconstructed speech signal of each frame, the target speech signal is determined.

[0157] In an embodiment of the present application, at the transmitting end of the voice signal, the voice signal to be processed is subjected to frame extraction processing, and part of the signal in the voice signal (the original voice signal of each frame) is extracted for encoding, and the encoded code stream obtained by encoding is transmitted to the receiving end. At the receiving end of the voice signal, this part of the signal is first decoded to obtain the original voice signal of each frame, and then the restoration processing is performed based on the frequency domain characteristics of the original voice signal of each frame to obtain the frequency domain characteristics of the reconstructed voice signal of each frame. The purpose of the restoration is to restore the part of the signal to the complete frequency domain voice signal before the frame extraction processing, and finally the target voice signal is obtained based on the frequency domain characteristics of the reconstructed voice signal of each frame. In the scheme of the present application, since part of the signal of the voice signal to be processed is encoded and transmitted at the transmitting end, the bandwidth can be effectively reduced, especially for scenarios with limited call bandwidth and high call bandwidth consumption, and the operating costs can also be reduced.

[0158] Furthermore, at the receiving end, the original speech signal of each frame is restored through interpolation processing and a neural network model, so that the final target speech signal has a higher speech quality, close to the speech quality of the speech signal to be processed. The solution of the present application can achieve good speech intelligibility even at an extremely low bit rate.

[0159] The above scheme and Figure 2 The speech signal processing method shown in is just a scheme with different execution subjects, and the implementation principles are the same. For details, please refer to the scheme described above, which will not be repeated here.

[0160] In order to better illustrate and understand the principles of the method provided by this application, the solution of this application is described below in conjunction with an optional specific embodiment. It should be noted that the specific implementation of each step in this specific embodiment should not be understood as a limitation to the solution of this application. Based on the principles of the solution provided by this application, other implementations that can be thought of by those skilled in the art should also be considered within the scope of protection of this application.

[0161] In this example, a trained neural network model is obtained based on the training method of the above neural network model. Figure 4 and Figure 5 Further explanation of this application plan:

[0162] Step 1: In the voice conference scenario, the sending end device collects the voice signal in the voice conference, converts the collected voice signal into a digital voice signal, and uses the digital voice signal as the voice signal to be processed ( Figure 4 The speech signal shown in Figure 3 is collected continuously).

[0163] Step 2: Perform a time-frequency transform (Fourier transform in this example) on the speech signal to be processed to obtain the frequency domain features and phase features corresponding to the speech signal to be processed, and determine the speech spectrogram of the speech signal to be processed (frame the speech signal to be processed). For details, see Figure 5 (a) shows a schematic diagram of a spectrogram. In this spectrogram, the horizontal axis represents the frame number, and the vertical axis represents the amplitude of different frequency points. For example, the signal to be processed contains 7 frames of signal (for example), and each frame is divided into 10 frequency points in the frequency domain (this is just an illustration; in actual applications, the number of frequency points can be determined based on the number of points in the fast Fourier transform). Figure 5 Each small grid in (a) represents the power spectrum amplitude information corresponding to different frame numbers at different frequency points (in this example, the frequency domain feature is the power spectrum).

[0164] In this example, the speech signal to be processed may be first subjected to time-frequency transformation and then to frame processing, or may be subjected to frame processing first and then to time-frequency transformation. The present application does not limit the order of execution.

[0165] Step 3: Every 2 frames, extract a frame of signal from the speech signal to be processed (7 frames of signal) after frame processing to obtain a discontinuous speech signal ( Figure 4 The interval N frame speech coding shown in Figure 5 (b) Schematic diagram of the discontinuous speech signal shown in the figure, Figure 5 The original speech signals of the discontinuous speech signals shown in (b) are the first frame speech signal, the fourth frame speech signal and the seventh frame speech signal.

[0166] Step 4: Encode the three frames of original speech signals separately through the iLBC speech codec to obtain the encoded bit stream ( Figure 4 ), and sends the encoded code stream to the receiving device through the network.

[0167] The above steps 1 to 4 are performed by the sending device.

[0168] Step 5: After receiving the coded stream, the receiving device decodes the coded stream to obtain the original voice signals of each frame, namely the first frame voice signal, the fourth frame voice signal and the seventh frame voice signal ( Figure 4 channel decoding as shown in ).

[0169] Figure 5 The sampled codec signal (spectrum) shown in (b) refers to encoding and decoding of each frame signal obtained by the sampled spectrum.

[0170] Step 6: Obtain a first correlation between the frequency domain characteristics of the compensated frame signal between adjacent frame signals and the frequency domain characteristics of the original speech signal of each frame in the adjacent frame signals; based on the first correlation and the frequency domain characteristics of each frame of the original speech signal in the adjacent frame signals, perform interpolation processing on the adjacent frame signals to obtain the frequency domain characteristics of the interpolated signal between the adjacent frame signals.

[0171] In this example, the following description is given by taking the adjacent frame signals being the fourth frame speech signal and the seventh frame speech signal as an example.

[0172] The first association relationship can be expressed by the following formula (1):

[0173] (1)

[0174] in, Represents the interpolated estimated value of the power spectrum amplitude of the ith frequency point of the nth compensation frame after the jth frame of the original speech signal in each frame of the original speech signal, Represents the true value of the power spectrum amplitude of the i-th frequency point of the j-th frame original speech signal (real power spectrum value), Represents the true value of the power spectrum amplitude of the i-th frequency point of the original speech signal of the j+1-th frame. is the result of the first interpolation (the power spectrum value of the interpolated signal). Where n is an integer ranging from 1 to N-1, and N is the set frame interval.

[0175] If N is 2, interpolation processing is performed on adjacent frame signals to obtain 2 frames of interpolated signals between adjacent frame signals, that is, two frames of compensation frame signals are inserted between the fourth frame speech signal and the seventh frame speech signal.

[0176] Based on the first association relationship, the power spectrum value of the fourth frame speech signal and the power spectrum value of the seventh frame speech signal, the power spectrum values of the two compensation frame signals (the power spectrum values of the interpolated signals) can be obtained.

[0177] Step 7: Obtain a second correlation relationship between the frequency domain characteristics of the compensation frame signal between adjacent frame signals, the frequency domain characteristics of a third signal of the adjacent frame signals, and the frequency domain characteristics of the first signal, where the third signal is a frame signal previous to the first signal; based on the second correlation relationship, the frequency domain characteristics of the first signal, and the frequency domain characteristics of the third signal, perform extrapolation processing on the adjacent frame signals to obtain frequency domain characteristics of the extrapolated signal between the adjacent frame signals;

[0178] Continuing with the above example, after interpolation processing is performed to obtain two frames of compensation frame signals (interpolation signals), the power spectrum values of the two frames of compensation frame signals (extrapolation signals) inserted between the fourth frame of voice signal and the seventh frame of voice signal can be determined based on the second association relationship, the power spectrum value of the fourth frame of voice signal (the frequency domain characteristics of the first signal), and the power spectrum value of the first frame of voice signal (the frequency domain characteristics of the third signal).

[0179] The second correlation relationship can be expressed by the following formula (2):

[0180] (2)

[0181] in, Represents the interpolated estimated value of the power spectrum amplitude of the ith frequency point of the nth compensation frame after the jth frame of the original speech signal in each frame of the original speech signal, Represents the true value of the power spectrum amplitude of the i-th frequency point of the j-th frame original speech signal (real power spectrum value), Represents the true value of the power spectrum amplitude of the i-th frequency point of the original speech signal of the j-1-th frame. The original speech signal of the j-1-th frame is the frame signal before the j-th frame signal. is the result of the second interpolation (the power spectrum value of the extrapolated signal). Where n is an integer ranging from 1 to N-1, and N is the set frame interval.

[0182] Based on the second association relationship, the power spectrum value of the fourth frame speech signal and the power spectrum value of the first frame speech signal, the power spectrum values of the two frames of compensation frame signals (power spectrum values of the extrapolated signal) can be obtained.

[0183] Step 8: Fusing the frequency domain features of the interpolation signal of each frame with the frequency domain features of the extrapolation signal of each frame to obtain the frequency domain features of the compensation frame signal between adjacent frame signals.

[0184] The specific fusion method can be found in the following formula (3):

[0185] (3)

[0186] in, Represents the interpolated estimated value of the power spectrum amplitude of the ith frequency point of the nth compensation frame after the jth frame of the original speech signal in each frame of the original speech signal, is the first weight, is the second weight, is the power spectrum value of the interpolated signal, is the power spectrum value of the extrapolated signal, represents the power spectrum value of the weighted interpolated signal, Represents the power spectrum value of the weighted extrapolated signal. Indicates the power spectrum value of each compensated frame signal.

[0187] In this example, a=0.7 is optional.

[0188] Based on the above formula (3), the power spectrum value of each compensation frame signal can be obtained ( Figure 4 Speech decoding as shown in ).

[0189] After obtaining the power spectrum value of each compensation frame signal ( Figure 5 After the quadratic weighted interpolation signal shown in (c), there are 7 frames of speech signals. Figure 5 (c) is a schematic diagram of each frame signal, wherein the dotted line portion represents the power spectrum value of each compensated frame signal, and the solid line portion represents each adjacent frame signal.

[0190] In step 9, the frequency domain features of each frame of the original speech signal and the frequency domain features of each compensated frame signal (the power spectrum values of the 7 frames of speech signals) are input into the trained neural network model to obtain the spectral gain of each frame of the speech signal to be reconstructed. Each frame of the speech signal to be reconstructed includes each frame of the original speech signal and each compensated frame signal.

[0191] Step 10: Multiply the spectral gain of each frame of the speech signal to be reconstructed (7 frames of speech signal) by the corresponding frequency domain feature to determine the frequency domain feature of each frame of the reconstructed speech signal. Figure 5 (d) The reconstructed speech signal of each frame shown ( Figure 5 (c) The frequency domain features of the deep learning restored signal) Figure 4 Deep learning N-frame restoration processing as shown in ).

[0192] Step 11: Based on the phase characteristics of the original speech signal of each frame and the frequency domain characteristics of the reconstructed speech signal of each frame, the frequency domain characteristics of the reconstructed speech signal of each frame are transformed into a time domain characteristic of the reconstructed speech signal of each frame; the time domain characteristics of the reconstructed speech signal of each frame are used as the target speech signal, and the target speech signal is played through the receiving device ( Figure 4 (see the speech signal playback shown in ).

[0193] In this example, steps 5 to 11 are performed on the receiving device. In this example, the receiving device may be a terminal device of each user participating in the voice conference, and the sending device may be a terminal device of the user who is speaking in the voice conference.

[0194] The present application provides a speech signal processing device, such as Figure 7 As shown, the speech signal processing device 30 may include: a speech signal receiving module 310, a decoding module 320, an interpolation processing module 330, a spectral gain determination module 340, a reconstructed speech signal determination module 350 and a speech processing module 360, wherein:

[0195] The voice signal receiving module 310 is used to receive the coded stream corresponding to the voice signal to be processed. The coded stream is obtained by encoding the original voice signal of each frame in the discontinuous voice signal at the transmitting end device. The discontinuous voice signal is obtained by extracting the frame of the voice signal to be processed according to the set frame interval;

[0196] The decoding module 320 is used to decode the coded bit stream to obtain the original speech signal of each frame and determine the frequency domain characteristics of the original speech signal of each frame;

[0197] The interpolation processing module 330 is used to perform interpolation processing on each frame of the original speech signal based on the frequency domain characteristics of each frame of the original speech signal to obtain the frequency domain characteristics of the compensation frame signal between the frames of the original speech signal;

[0198] a spectral gain determination module 340 for inputting the frequency domain features of each frame of the original speech signal and the frequency domain features of each compensated frame signal into a trained neural network model to obtain the spectral gain of each frame of the speech signal to be reconstructed, where each frame of the speech signal to be reconstructed includes the original speech signal and the compensated frame signal;

[0199] The reconstructed speech signal determination module 350 is configured to determine the frequency domain characteristics of the reconstructed speech signal of each frame based on the spectral gain of the speech signal to be reconstructed of each frame and the frequency domain characteristics of the speech signal to be reconstructed of each frame;

[0200] The speech processing module 360 is used to perform frequency-time transformation on the frequency domain features of the reconstructed speech signal of each frame to obtain the target speech signal.

[0201] In an embodiment of the present application, at the transmitting end of the voice signal, the voice signal to be processed is subjected to frame extraction processing, and part of the signal in the voice signal (the original voice signal of each frame) is extracted for encoding, and the encoded code stream obtained by encoding is transmitted to the receiving end. At the receiving end of the voice signal, this part of the signal is first decoded to obtain the original voice signal of each frame, and then the restoration processing is performed based on the frequency domain characteristics of the original voice signal of each frame to obtain the frequency domain characteristics of the reconstructed voice signal of each frame. The purpose of the restoration is to restore the part of the signal to the complete frequency domain voice signal before the frame extraction processing, and finally the target voice signal is obtained based on the frequency domain characteristics of the reconstructed voice signal of each frame. In the scheme of the present application, since part of the signal of the voice signal to be processed is encoded and transmitted at the transmitting end, the bandwidth can be effectively reduced, especially for scenarios with limited call bandwidth and high call bandwidth consumption, and the operating costs can also be reduced.

[0202] Furthermore, at the receiving end, the original speech signal of each frame is restored through interpolation processing and a neural network model, so that the final target speech signal has a higher speech quality, close to the speech quality of the speech signal to be processed. The solution of the present application can achieve good speech intelligibility even at an extremely low bit rate.

[0203] Optionally, when the interpolation processing module 330 performs interpolation processing on each frame of original speech signals based on the frequency domain characteristics of each frame of original speech signals to obtain the frequency domain characteristics of the compensation frame signals between each frame of original speech signals, it is specifically used to: for each pair of adjacent frame signals in each frame of original speech signals, based on the frequency domain characteristics of each frame of original speech signals in the adjacent frame signals, perform interpolation processing on the adjacent frame signals to obtain the frequency domain characteristics of the compensation frame signals between the adjacent frame signals; and determine the frequency domain characteristics of the compensation frame signals between each frame of original speech signals according to the frequency domain characteristics of the compensation frame signals between each adjacent frame signals.

[0204] Optionally, for each pair of adjacent frame signals in each frame signal, the above-mentioned interpolation processing module 330 performs interpolation processing on the adjacent frame signals based on the frequency domain characteristics of the original speech signal of each frame in the adjacent frame signals to obtain the frequency domain characteristics of the compensation frame signals between the adjacent frame signals. It is specifically used to: obtain a first correlation relationship between the frequency domain characteristics of the compensation frame signals between the adjacent frame signals and the frequency domain characteristics of the original speech signals of each frame in the adjacent frame signals; based on the first correlation relationship and the frequency domain characteristics of the original speech signal of each frame in the adjacent frame signals, perform interpolation processing on the adjacent frame signals to obtain the frequency domain characteristics of the compensation frame signals between the adjacent frame signals.

[0205] Optionally, for each pair of adjacent frame signals in each frame of original speech signal, the adjacent frame signals include a first signal and a second signal, and the first signal precedes the second signal; when the interpolation processing module 330 performs interpolation processing on the adjacent frame signals based on the first association relationship and the frequency domain characteristics of each frame of original speech signal in the adjacent frame signals to obtain the frequency domain characteristics of the compensation frame signal between the adjacent frame signals, it is specifically configured to:

[0206] Based on the first association relationship and the frequency domain characteristics of each frame of the original speech signal in the adjacent frame signals, the adjacent frame signals are interpolated to obtain the frequency domain characteristics of the interpolated signals between the adjacent frame signals; the second association relationship between the frequency domain characteristics of the compensation frame signal between the adjacent frame signals, the frequency domain characteristics of the third signal of the adjacent frame signals and the frequency domain characteristics of the first signal is obtained, and the third signal is the previous frame signal of the first signal; based on the second association relationship, the frequency domain characteristics of the first signal and the frequency domain characteristics of the third signal, the adjacent frame signals are extrapolated to obtain the frequency domain characteristics of the extrapolated signals between the adjacent frame signals; the frequency domain characteristics of the interpolated signals of each frame are fused with the frequency domain characteristics of the extrapolated signals of each frame to obtain the frequency domain characteristics of the compensation frame signal between the adjacent frame signals.

[0207] Optionally, when the interpolation processing module 330 fuses the frequency domain characteristics of the interpolation signal of each frame with the frequency domain characteristics of the extrapolation signal of each frame to obtain the frequency domain characteristics of the compensation frame signal between the signals of adjacent frames, it is specifically used to: obtain the first weight corresponding to the interpolation signal of each frame and the second weight corresponding to the extrapolation signal of each frame; for each interpolation signal in each frame interpolation signal, perform weighted processing on the frequency domain characteristics of the interpolation signal and the first weight corresponding to the interpolation signal to obtain the frequency domain characteristics of the weighted interpolation signal; for each extrapolation signal in each frame extrapolation signal, perform weighted processing on the frequency domain characteristics of the extrapolation signal and the second weight corresponding to the extrapolation signal to obtain the frequency domain characteristics of the weighted extrapolation signal; based on the frequency domain characteristics of the weighted interpolation signal of each frame and the frequency domain characteristics of the weighted extrapolation signal of each frame, determine the frequency domain characteristics of the compensation frame signal between adjacent frame signals.

[0208] Optionally, the above neural network model is trained through the following model training modules:

[0209] The model training module is used to obtain sample data, where the sample data includes multiple sample speech signals; for each sample speech signal, the sample speech signal is framed to obtain framed speech signals of each frame, and each framed speech signal is subjected to frame extraction processing according to a set frame interval to obtain discontinuous framed sample speech signals; the frequency domain characteristics of each frame in the discontinuous framed sample speech signals are determined; the frequency domain characteristics of each frame in the discontinuous framed sample speech signals are interpolated to obtain the frequency domain characteristics of each frame of the sample speech signal to be reconstructed; based on the frequency domain characteristics of each frame of the sample speech signal to be reconstructed and each framed speech signal, the real spectrum gain of each frame of the sample speech signal to be reconstructed is determined;

[0210] Repeat the following training steps until the loss value meets the training end condition to obtain the neural network model: input the frequency domain features of the sample speech signal to be reconstructed in each frame into the initial neural network model to obtain the predicted spectral gain corresponding to the sample speech signal to be reconstructed in each frame; based on each predicted spectral gain and each true spectral gain, determine the loss value corresponding to the initial neural network model. If the loss value meets the training end condition, end the training and obtain the neural network model; if not, adjust the model parameters of the initial neural network model and repeat the training steps.

[0211] Optionally, when determining the frequency domain features of each frame in the discontinuous frame-sampled speech signal, the above-mentioned model training module is specifically used to: perform linear time-frequency transformation on each frame-sampled speech signal in the discontinuous frame-sampled speech signal to obtain the linear frequency domain features of each frame-sampled speech signal; perform feature extraction on the linear frequency domain features of each frame-sampled speech signal to obtain the logarithmic frequency domain features of each frame-sampled speech signal, and use the logarithmic frequency domain features of each frame-sampled speech signal as the frequency domain features of each frame-sampled speech signal.

[0212] Optionally, when determining the frequency domain features of each frame of the original speech signal, the restoration module 320 is specifically configured to: perform time-frequency transformation on each frame of the original speech signal to obtain the frequency domain features and phase features of each frame of the original speech signal;

[0213] When the above-mentioned speech processing module 360 performs frequency-time transformation on the frequency domain features of the reconstructed speech signal of each frame to obtain the target speech signal, it is specifically used to: based on the phase features of the original speech signal of each frame and the frequency domain features of the reconstructed speech signal of each frame, perform frequency-time transformation on the frequency domain features of the reconstructed speech signal of each frame to obtain the time domain features of the reconstructed speech signal of each frame; and use the time domain features of the reconstructed speech signal of each frame as the target speech signal.

[0214] Optionally, if the frequency domain feature is a Mel spectrum frequency domain feature, the above-mentioned speech processing module 360 performs a frequency-time transform on the frequency domain features of the reconstructed speech signal of each frame to obtain the target speech signal. Specifically, it is used to: based on the Mel spectrum frequency domain features of the reconstructed speech signal of each frame, perform a frequency-time transform on the frequency domain features of the restored speech signal of each frame to obtain the target speech signal.

[0215] The present application provides a speech signal processing device, such as Figure 8 As shown, the speech signal processing device 40 may include: a speech signal acquisition module 410, an encoding module 420 and a sending module 430, wherein:

[0216] The voice signal acquisition module 410 is used to acquire the voice signal to be processed, and perform frame extraction processing on the voice signal to be processed according to a set frame interval to obtain a discontinuous voice signal;

[0217] The encoding module 420 is used to encode each frame of the original speech signal in the discontinuous speech signal to obtain an encoded code stream corresponding to the speech signal to be processed;

[0218] The sending module 430 is configured to send the encoded code stream to a receiving device, so that the receiving device decodes the encoded code stream to obtain each frame of the original speech signal, determines the frequency domain characteristics of each frame of the original speech signal, and performs the following processing on the frequency domain characteristics of each frame of the original signal to obtain the target speech signal:

[0219] The encoded code stream is decoded to obtain the original speech signal of each frame, and the frequency domain characteristics of the original speech signal of each frame are determined; based on the frequency domain characteristics of the original speech signal of each frame, the original speech signal of each frame is interpolated to obtain the frequency domain characteristics of the compensation frame signal between the original speech signals of each frame; the frequency domain characteristics of the original speech signal of each frame and the frequency domain characteristics of each compensation frame signal are input into the trained neural network model to obtain the spectrum gain of the speech signal to be reconstructed of each frame, and the speech signal to be reconstructed of each frame includes the original speech signal of each frame and the compensation frame signal; based on the spectrum gain of the speech signal to be reconstructed of each frame and the frequency domain characteristics of the speech signal to be reconstructed of each frame, the frequency domain characteristics of the reconstructed speech signal of each frame are determined; based on the frequency domain characteristics of the reconstructed speech signal of each frame, the target speech signal is determined.

[0220] In an embodiment of the present application, at the transmitting end of the voice signal, the voice signal to be processed is subjected to frame extraction processing, and part of the signal in the voice signal (the original voice signal of each frame) is extracted for encoding, and the encoded code stream obtained by encoding is transmitted to the receiving end. At the receiving end of the voice signal, this part of the signal is first decoded to obtain the original voice signal of each frame, and then the restoration processing is performed based on the frequency domain characteristics of the original voice signal of each frame to obtain the frequency domain characteristics of the reconstructed voice signal of each frame. The purpose of the restoration is to restore the part of the signal to the complete frequency domain voice signal before the frame extraction processing, and finally the target voice signal is obtained based on the frequency domain characteristics of the reconstructed voice signal of each frame. In the scheme of the present application, since part of the signal of the voice signal to be processed is encoded and transmitted at the transmitting end, the bandwidth can be effectively reduced, especially for scenarios with limited call bandwidth and high call bandwidth consumption, and the operating costs can also be reduced.

[0221] Furthermore, at the receiving end, the original speech signal of each frame is restored through interpolation processing and a neural network model, so that the final target speech signal has a higher speech quality, close to the speech quality of the speech signal to be processed. The solution of the present application can achieve good speech intelligibility even at an extremely low bit rate.

[0222] The speech signal processing device of the embodiment of the present application can execute a speech signal processing method provided in the embodiment of the present application. The implementation principle is similar and will not be repeated here.

[0223] In some embodiments, the speech signal processing device provided by the embodiment of the present invention can be implemented by a combination of software and hardware. As an example, the speech signal processing device provided by the embodiment of the present invention can be a processor in the form of a hardware decoding processor, which is programmed to execute the speech signal processing method provided by the embodiment of the present invention. For example, the processor in the form of a hardware decoding processor can adopt one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components.

[0224] The present application embodiment provides an electronic device, such as Figure 9 As shown, Figure 9The electronic device 2000 shown includes a processor 2001 and a memory 2003. The processor 2001 and the memory 2003 are connected, for example, via a bus 2002. Optionally, the electronic device 2000 may further include a transceiver 2004. It should be noted that in actual applications, the number of transceivers 2004 is not limited to one, and the structure of the electronic device 2000 does not constitute a limitation on the embodiments of the present application.

[0225] The processor 2001 is used in the embodiment of the present application to implement Figure 7 and Figure 8 The functions of each module are shown.

[0226] Processor 2001 may be a CPU, a general-purpose processor, a DSP, an ASIC, an FPGA, or other programmable logic device, a transistor logic device, a hardware component, or any combination thereof. It may implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. Processor 2001 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, and the like.

[0227] The bus 2002 may include a path for transmitting information between the above components. The bus 2002 may be a PCI bus or an EISA bus, etc. The bus 2002 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 9 Only one thick line is used in the diagram, but this does not mean that there is only one bus or one type of bus.

[0228] The memory 2003 may be a ROM or other type of static storage device capable of storing static information and computer programs, a RAM or other type of dynamic storage device capable of storing information and computer programs, or an EEPROM, a CD-ROM or other optical disk storage, an optical disc storage (including a compact disc, a laser disc, an optical disc, a digital versatile disc, a Blu-ray disc, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium capable of carrying or storing a desired computer program in the form of a data structure and capable of being accessed by a computer, but not limited thereto.

[0229] The memory 2003 is used to store the computer program for executing the application program of the present application solution, and the execution is controlled by the processor 2001. The processor 2001 is used to execute the computer program of the application program stored in the memory 2003 to implement Figure 7 and Figure 8 The illustrated embodiment provides operations of the speech signal processing apparatus.

[0230] An embodiment of the present application provides an electronic device, including a processor and a memory: the memory is configured to store a computer program, and when the computer program is executed by the processor, the processor performs any one of the methods in the above embodiments.

[0231] An embodiment of the present application provides a computer-readable storage medium for storing a computer program. When the computer program is run on a computer, the computer can execute any one of the methods in the above embodiments.

[0232] According to one aspect of the present application, a computer program product or computer program is provided, the computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in the various optional implementations described above.

[0233] The nouns and implementation principles involved in a computer-readable storage medium in this application can be specifically referred to a speech signal processing method in an embodiment of this application, and will not be repeated here.

[0234] It should be understood that although the steps in the flowcharts of the accompanying drawings are shown in sequence as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some of the steps in the flowcharts of the accompanying drawings may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed in turn or alternately with other steps or at least a portion of the sub-steps or stages of other steps.

[0235] The above description is only part of the implementation methods of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.

Claims

1. A speech signal processing method, characterized in that: Includes: Receiving a coded stream corresponding to a voice signal to be processed, the coded stream is obtained by encoding each frame of the original voice signal in the discontinuous voice signal by the transmitting end device, and the discontinuous voice signal is obtained by performing frame extraction processing on the voice signal to be processed according to a set frame interval; Decoding the encoded code stream to obtain the original speech signal of each frame, and determining the frequency domain characteristics of the original speech signal of each frame; For each pair of adjacent frame signals in the original speech signals of each frame, interpolation processing is performed on the adjacent frame signals based on the frequency domain characteristics of each frame of the original speech signal in the adjacent frame signals to obtain the frequency domain characteristics of the compensated frame signal between the adjacent frame signals; Determining the frequency domain characteristics of the compensated frame signals between the original speech signals of each frame according to the frequency domain characteristics of the compensated frame signals between the adjacent frame signals; wherein the number of the compensated frame signals between the adjacent frame signals is equal to the set frame interval; Inputting the frequency domain features of the original speech signal of each frame and the frequency domain features of each compensated frame signal into a trained neural network model to obtain the spectral gain of each frame of the speech signal to be reconstructed, wherein each frame of the speech signal to be reconstructed includes the original speech signal of each frame and the compensated frame signal; Determining the frequency domain features of the reconstructed speech signal of each frame based on the spectral gain of the to-be-reconstructed speech signal of each frame and the frequency domain features of the to-be-reconstructed speech signal of each frame; Performing frequency-time transformation on the frequency domain features of the reconstructed speech signal of each frame to obtain the target speech signal.

2. The method according to claim 1, characterized in that For each pair of adjacent frame signals in the signals of each frame, interpolation processing is performed on the adjacent frame signals based on the frequency domain characteristics of the original speech signal of each frame in the adjacent frame signals to obtain the frequency domain characteristics of the compensated frame signal between the adjacent frame signals, including: Acquire a first correlation relationship between the frequency domain features of the compensated frame signal between adjacent frame signals and the frequency domain features of the original speech signal of each frame in the adjacent frame signals; Based on the first association relationship and the frequency domain characteristics of each frame of the original speech signal in the adjacent frame signals, interpolation processing is performed on the adjacent frame signals to obtain the frequency domain characteristics of the compensated frame signals between the adjacent frame signals.

3. The method according to claim 2, characterized in that For each pair of adjacent frame signals in the original speech signal of each frame, the adjacent frame signals include a first signal and a second signal, and the first signal is before the second signal; performing interpolation processing on the adjacent frame signals based on the first association relationship and the frequency domain characteristics of each frame of the original speech signal in the adjacent frame signals to obtain the frequency domain characteristics of the compensated frame signal between the adjacent frame signals, including: Based on the first association relationship and the frequency domain characteristics of each frame of the original speech signal in the adjacent frame signals, interpolation processing is performed on the adjacent frame signals to obtain frequency domain characteristics of the interpolated signal between the adjacent frame signals; Acquire a second correlation relationship between a frequency domain feature of a compensation frame signal between adjacent frame signals, a frequency domain feature of a third signal of the adjacent frame signals, and a frequency domain feature of the first signal, where the third signal is a frame signal preceding the first signal; performing extrapolation processing on the adjacent frame signals based on the second association relationship, the frequency domain characteristics of the first signal, and the frequency domain characteristics of the third signal to obtain frequency domain characteristics of the extrapolated signal between the adjacent frame signals; The frequency domain features of the interpolation signal of each frame are fused with the frequency domain features of the extrapolation signal of each frame to obtain the frequency domain features of the compensation frame signal between the adjacent frame signals.

4. The method according to claim 3, characterized in that The fusing the frequency domain features of the interpolated signal of each frame with the frequency domain features of the extrapolated signal of each frame to obtain the frequency domain features of the compensated frame signal between the signals of adjacent frames includes: Obtaining a first weight corresponding to the interpolation signal of each frame and a second weight corresponding to the extrapolation signal of each frame; For each interpolation signal of each frame, performing weighted processing on the frequency domain feature of the interpolation signal and a first weight corresponding to the interpolation signal to obtain the frequency domain feature of the weighted interpolation signal; For each frame of the extrapolated signal, performing weighted processing on the frequency domain feature of the extrapolated signal and the second weight corresponding to the extrapolated signal to obtain the frequency domain feature of the weighted extrapolated signal; Based on the frequency domain characteristics of the weighted interpolation signal of each frame and the frequency domain characteristics of the weighted extrapolation signal of each frame, the frequency domain characteristics of the compensation frame signal between the adjacent frame signals are determined.

5. The method according to any one of claims 1 to 4, characterized in that The neural network model is trained in the following way: Acquiring sample data, wherein the sample data includes a plurality of sample speech signals; For each sample speech signal, the sample speech signal is framed to obtain framed speech signals of each frame, and the framed speech signals of each frame are de-framed according to a set frame interval to obtain discontinuous de-framed sample speech signals; Determining frequency domain features of each frame in the discontinuous frame sample speech signal; Performing interpolation processing on the frequency domain features of each frame in the discontinuous frame-drawing sample speech signal to obtain the frequency domain features of the sample speech signal to be reconstructed in each frame; Determining a true spectrum gain of the sample speech signal to be reconstructed in each frame based on the frequency domain characteristics of the sample speech signal to be reconstructed in each frame and the framed speech signal in each frame; Repeat the following training steps until the loss value meets the training end condition to obtain the neural network model: Inputting the frequency domain features of the sample speech signal to be reconstructed in each frame into the initial neural network model to obtain the predicted spectral gain corresponding to the sample speech signal to be reconstructed in each frame; Based on each of the predicted spectral gains and each of the real spectral gains, determine the loss value corresponding to the initial neural network model. If the loss value meets the training end condition, end the training and obtain the neural network model; if not, adjust the model parameters of the initial neural network model and repeat the training steps.

6. The method according to claim 5, characterized in that The determining of the frequency domain features of each frame in the discontinuous frame sample speech signal includes: Performing a linear time-frequency transform on each frame of the discontinuous frame-sampled speech signal to obtain a linear frequency domain feature of each frame of the frame-sampled speech signal; Feature extraction is performed on the linear frequency domain features of the frame-sampled speech signals of each frame to obtain the logarithmic frequency domain features of the frame-sampled speech signals of each frame, and the logarithmic frequency domain features of the frame-sampled speech signals of each frame are used as the frequency domain features of the frame-sampled speech signals of each frame.

7. The method according to any one of claims 1 to 4, characterized in that The determining of the frequency domain features of the original speech signal of each frame includes: Performing time-frequency transformation on the original speech signal of each frame to obtain frequency domain features and phase features of the original speech signal of each frame; The performing frequency-time transformation on the frequency domain features of the reconstructed speech signal of each frame to obtain the target speech signal includes: Based on the phase characteristics of the original speech signal of each frame and the frequency domain characteristics of the reconstructed speech signal of each frame, frequency-time transformation is performed on the frequency domain characteristics of the reconstructed speech signal of each frame to obtain the time domain characteristics of the reconstructed speech signal of each frame; The time domain features of the speech signal reconstructed in each frame are used as the target speech signal.

8. The method according to any one of claims 1 to 4, characterized in that If the frequency domain feature is a Mel-spectrogram frequency domain feature, performing a frequency-time transform on the frequency domain feature of the reconstructed speech signal of each frame to obtain a target speech signal includes: Based on the Mel spectrum frequency domain features of the reconstructed speech signal of each frame, a frequency-time transform is performed on the frequency domain features of the reconstructed speech signal of each frame to obtain a target speech signal.

9. A speech signal processing method, characterized in that: include: Acquire a speech signal to be processed, and perform frame extraction processing on the speech signal to be processed according to a set frame interval to obtain a discontinuous speech signal; Encoding each frame of the original speech signal in the discontinuous speech signal to obtain a coded code stream corresponding to the speech signal to be processed; The coded code stream is sent to a receiving device, so that the receiving device performs the following processing on the coded code stream to obtain a target speech signal: Decoding the encoded code stream to obtain the original speech signal of each frame, and determining the frequency domain characteristics of the original speech signal of each frame; For each pair of adjacent frame signals in the original speech signals of each frame, interpolation processing is performed on the adjacent frame signals based on the frequency domain characteristics of each frame of the original speech signal in the adjacent frame signals to obtain the frequency domain characteristics of the compensated frame signal between the adjacent frame signals; Determining the frequency domain characteristics of the compensated frame signals between the original speech signals of each frame according to the frequency domain characteristics of the compensated frame signals between the adjacent frame signals; wherein the number of the compensated frame signals between the adjacent frame signals is equal to the set frame interval; Inputting the frequency domain features of the original speech signal of each frame and the frequency domain features of each compensated frame signal into a trained neural network model to obtain the spectral gain of each frame of the speech signal to be reconstructed, wherein each frame of the speech signal to be reconstructed includes the original speech signal of each frame and the compensated frame signal; Determining the frequency domain features of the reconstructed speech signal of each frame based on the spectral gain of the to-be-reconstructed speech signal of each frame and the frequency domain features of the to-be-reconstructed speech signal of each frame; A target speech signal is determined based on the frequency domain features of the reconstructed speech signal in each frame.

10. A speech signal processing device, characterized in that: include: A voice signal receiving module for receiving a coded stream corresponding to a voice signal to be processed, wherein the coded stream is obtained by encoding each frame of the original voice signal in the discontinuous voice signal by the transmitting end device, and the discontinuous voice signal is obtained by extracting the frame of the voice signal to be processed according to a set frame interval; A decoding module, configured to decode the encoded code stream to obtain the original speech signal of each frame, and determine the frequency domain characteristics of the original speech signal of each frame; an interpolation processing module, configured to perform interpolation processing on each pair of adjacent frame signals in the original speech signal of each frame based on the frequency domain characteristics of each frame of the adjacent frame signals to obtain the frequency domain characteristics of the compensated frame signal between the adjacent frame signals; Determining the frequency domain characteristics of the compensated frame signals between the original speech signals of each frame according to the frequency domain characteristics of the compensated frame signals between the adjacent frame signals; wherein the number of the compensated frame signals between the adjacent frame signals is equal to the set frame interval; a spectral gain determination module, configured to input the frequency domain features of the original speech signal of each frame and the frequency domain features of each compensated frame signal into a trained neural network model to obtain the spectral gain of each frame of the speech signal to be reconstructed, wherein each frame of the speech signal to be reconstructed includes the original speech signal of each frame and the compensated frame signal; A reconstructed speech signal determination module is used to determine the frequency domain characteristics of the reconstructed speech signal of each frame based on the spectral gain of the speech signal to be reconstructed in each frame and the frequency domain characteristics of the speech signal to be reconstructed in each frame; The speech processing module is used to perform frequency-time transformation on the frequency domain features of the reconstructed speech signal of each frame to obtain the target speech signal.

11. The device according to claim 10, characterized in that For each pair of adjacent frame signals in the signal of each frame, the interpolation processing module performs interpolation processing on the adjacent frame signals based on the frequency domain characteristics of each frame of the original speech signal in the adjacent frame signals to obtain the frequency domain characteristics of the compensated frame signal between the adjacent frame signals, specifically for: Acquire a first correlation relationship between the frequency domain features of the compensated frame signal between adjacent frame signals and the frequency domain features of the original speech signal of each frame in the adjacent frame signals; Based on the first association relationship and the frequency domain characteristics of each frame of the original speech signal in the adjacent frame signals, interpolation processing is performed on the adjacent frame signals to obtain the frequency domain characteristics of the compensated frame signals between the adjacent frame signals.

12. The device according to claim 11, characterized in that For each pair of adjacent frame signals in the original speech signal of each frame, the adjacent frame signals include a first signal and a second signal, and the first signal is before the second signal; when the interpolation processing module is used to perform interpolation processing on the adjacent frame signals based on the first association relationship and the frequency domain characteristics of each frame of the original speech signal in the adjacent frame signals to obtain the frequency domain characteristics of the compensation frame signal between the adjacent frame signals, it is specifically used to: Based on the first association relationship and the frequency domain characteristics of each frame of the original speech signal in the adjacent frame signals, interpolation processing is performed on the adjacent frame signals to obtain frequency domain characteristics of the interpolated signal between the adjacent frame signals; Acquire a second correlation relationship between a frequency domain feature of a compensation frame signal between adjacent frame signals, a frequency domain feature of a third signal of the adjacent frame signals, and a frequency domain feature of the first signal, where the third signal is a frame signal preceding the first signal; performing extrapolation processing on the adjacent frame signals based on the second association relationship, the frequency domain characteristics of the first signal, and the frequency domain characteristics of the third signal to obtain frequency domain characteristics of the extrapolated signal between the adjacent frame signals; The frequency domain features of the interpolation signal of each frame are fused with the frequency domain features of the extrapolation signal of each frame to obtain the frequency domain features of the compensation frame signal between the adjacent frame signals.

13. The device according to claim 12, characterized in that When the interpolation processing module is used to fuse the frequency domain features of the interpolation signal of each frame with the frequency domain features of the extrapolation signal of each frame to obtain the frequency domain features of the compensation frame signal between the signals of adjacent frames, it is specifically used to: Obtaining a first weight corresponding to the interpolation signal of each frame and a second weight corresponding to the extrapolation signal of each frame; For each interpolation signal of each frame, performing weighted processing on the frequency domain feature of the interpolation signal and a first weight corresponding to the interpolation signal to obtain the frequency domain feature of the weighted interpolation signal; For each frame of the extrapolated signal, performing weighted processing on the frequency domain feature of the extrapolated signal and the second weight corresponding to the extrapolated signal to obtain the frequency domain feature of the weighted extrapolated signal; Based on the frequency domain characteristics of the weighted interpolation signal of each frame and the frequency domain characteristics of the weighted extrapolation signal of each frame, the frequency domain characteristics of the compensation frame signal between the adjacent frame signals are determined.

14. The device according to any one of claims 10 to 13, characterized in that The device also includes a model training module, which is used to: Acquiring sample data, wherein the sample data includes a plurality of sample speech signals; For each sample speech signal, the sample speech signal is framed to obtain framed speech signals of each frame, and the framed speech signals of each frame are de-framed according to a set frame interval to obtain discontinuous de-framed sample speech signals; Determining frequency domain features of each frame in the discontinuous frame sample speech signal; Performing interpolation processing on the frequency domain features of each frame in the discontinuous frame-drawing sample speech signal to obtain the frequency domain features of the sample speech signal to be reconstructed in each frame; Determining a true spectrum gain of the sample speech signal to be reconstructed in each frame based on the frequency domain characteristics of the sample speech signal to be reconstructed in each frame and the framed speech signal in each frame; Repeat the following training steps until the loss value meets the training end condition to obtain the neural network model: Inputting the frequency domain features of the sample speech signal to be reconstructed in each frame into the initial neural network model to obtain the predicted spectral gain corresponding to the sample speech signal to be reconstructed in each frame; Based on each of the predicted spectral gains and each of the real spectral gains, determine the loss value corresponding to the initial neural network model. If the loss value meets the training end condition, end the training and obtain the neural network model; if not, adjust the model parameters of the initial neural network model and repeat the training steps.

15. The device according to claim 14, characterized in that When the model training module is used to determine the frequency domain features of each frame in the discontinuous frame sample speech signal, it is specifically used to: Performing a linear time-frequency transform on each frame of the discontinuous frame-sampled speech signal to obtain a linear frequency domain feature of each frame of the frame-sampled speech signal; Feature extraction is performed on the linear frequency domain features of the frame-sampled speech signals of each frame to obtain the logarithmic frequency domain features of the frame-sampled speech signals of each frame, and the logarithmic frequency domain features of the frame-sampled speech signals of each frame are used as the frequency domain features of the frame-sampled speech signals of each frame.

16. The device according to any one of claims 10 to 13, characterized in that When the decoding module is used to determine the frequency domain features of the original speech signal of each frame, it is specifically used to: Performing time-frequency transformation on the original speech signal of each frame to obtain frequency domain features and phase features of the original speech signal of each frame; When the speech processing module performs frequency-time transformation on the frequency domain features of the reconstructed speech signal of each frame to obtain the target speech signal, it is specifically used to: Based on the phase characteristics of the original speech signal of each frame and the frequency domain characteristics of the reconstructed speech signal of each frame, frequency-time transformation is performed on the frequency domain characteristics of the reconstructed speech signal of each frame to obtain the time domain characteristics of the reconstructed speech signal of each frame; The time domain features of the speech signal reconstructed in each frame are used as the target speech signal.

17. The device according to any one of claims 10 to 13, characterized in that If the frequency domain feature is a Mel-spectrum frequency domain feature, the speech processing module is used to perform frequency-time transformation on the frequency domain feature of the reconstructed speech signal of each frame to obtain the target speech signal, specifically for: Based on the Mel spectrum frequency domain features of the reconstructed speech signal of each frame, a frequency-time transform is performed on the frequency domain features of the reconstructed speech signal of each frame to obtain a target speech signal.

18. A speech signal processing device, characterized in that: include: The voice signal acquisition module is used to acquire the voice signal to be processed, and perform frame extraction processing on the voice signal to be processed according to the set frame interval to obtain a discontinuous voice signal; An encoding module, configured to encode each frame of the original speech signal in the discontinuous speech signal to obtain an encoded code stream corresponding to the speech signal to be processed; The sending module is used to send the encoded code stream to a receiving device, so that the receiving device performs the following processing on the encoded code stream to obtain a target speech signal: Decoding the encoded code stream to obtain the original speech signal of each frame, and determining the frequency domain characteristics of the original speech signal of each frame; For each pair of adjacent frame signals in the original speech signals of each frame, interpolation processing is performed on the adjacent frame signals based on the frequency domain characteristics of each frame of the original speech signal in the adjacent frame signals to obtain the frequency domain characteristics of the compensated frame signal between the adjacent frame signals; Determining the frequency domain characteristics of the compensated frame signals between the original speech signals of each frame according to the frequency domain characteristics of the compensated frame signals between the adjacent frame signals; wherein the number of the compensated frame signals between the adjacent frame signals is equal to the set frame interval; Inputting the frequency domain features of the original speech signal of each frame and the frequency domain features of each compensated frame signal into a trained neural network model to obtain the spectral gain of each frame of the speech signal to be reconstructed, wherein each frame of the speech signal to be reconstructed includes the original speech signal of each frame and the compensated frame signal; Determining the frequency domain features of the reconstructed speech signal of each frame based on the spectral gain of the to-be-reconstructed speech signal of each frame and the frequency domain features of the to-be-reconstructed speech signal of each frame; A target speech signal is determined based on the frequency domain features of the speech signals reconstructed from the frames.

19. An electronic device, characterized in that: Including processor and memory: The memory is configured to store a computer program which, when executed by the processor, causes the processor to perform the method of any one of claims 1 to 9.

20. A computer-readable storage medium, characterized in that The computer-readable storage medium is used to store a computer program. When the computer program is run on a computer, the computer can execute the method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Device and method for expanding speech bandwidth based on audio watermarking

    CN102543086A

  • Low-bit-rate voice coder and decoder

    CN103854655A