Audio processing method and device, electronic equipment, storage medium and program product
Patent Information
- Application Number
- CN202210566670.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-23
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2042-05-23
AI Technical Summary
[0002]相关技术中,语音编解码主要是基于语音模型的语言编解码方法,例如码激励线性预测语音编码(CELP)方法可以应用于不同的场景,包括单人语音、多人语音、人声与其它信号并存等场景,但语音编码器码率较高,需要占用较多的网络带宽或存储空间
根据音频帧的多个基音周期和基音周期对应的相关特征确定出目标基音特征,并对倒谱特征和目标基音特征进行压缩编码得到码流数据,从而实现不需要将大量的编码参数向对端的解码器进行传输,减少了编码数据在传输过程中的码率,进而减小占有的网络宽带或存储空间。
Smart Images

Figure CN117153173B_ABST
Abstract
Description
Technical Field
[0001] This application relates to audio and video processing technology, and more particularly to an audio processing method, apparatus, electronic device, computer-readable storage medium, and computer program product. Background Technology
[0002] In related technologies, speech coding and decoding are mainly language coding and decoding methods based on speech models. For example, the code-excited linear predictive speech coding (CELP) method can be applied to different scenarios, including single-person speech, multi-person speech, and scenarios where human voice and other signals coexist. However, speech encoders have high bit rates and require a lot of network bandwidth or storage space. Summary of the Invention
[0003] This application provides an audio processing method, apparatus, electronic device, computer-readable storage medium, and computer program product that reduces the bit rate of the bitstream data during transmission, thereby reducing the network bandwidth or storage space occupied.
[0004] The technical solution of this application embodiment is implemented as follows: This application provides an audio processing method, including: The audio data is parsed and processed to obtain multiple sampling features and multiple pitch periods for each audio frame, wherein the audio data includes multiple audio frames; Based on multiple pitch periods of each audio frame, autocorrelation processing is performed on multiple sampling features of each audio frame to obtain the correlation features corresponding to the pitch periods; Based on multiple pitch periods of each audio frame and the relevant features corresponding to the pitch periods, the target pitch features corresponding to each audio frame are determined; The audio frames are subjected to cepstral processing to obtain the cepstral features of each audio frame; The cepstral features of the audio frame and the target pitch features are compressed and encoded to obtain the bitstream data.
[0005] This application provides an audio processing method, including: Obtain the bitstream data; The bitstream data is decompressed to obtain the cepstral features and target pitch features of each audio frame included in the bitstream data; Based on the target pitch features of the audio frame, the cepstral features of the audio frame are decoded to obtain the decoded audio data.
[0006] This application provides an audio processing apparatus, including: The data parsing module is used to parse and process the audio data to obtain multiple sampling features and multiple pitch periods for each audio frame, wherein the audio data includes multiple audio frames; The first feature calculation module is used to perform autocorrelation processing on multiple sampling features of each audio frame based on multiple pitch periods of each audio frame to obtain the correlation features corresponding to the pitch periods; The second feature calculation module is used to determine the target pitch feature corresponding to each audio frame based on multiple pitch periods of each audio frame and the relevant features corresponding to the pitch periods; The cepstral processing module is used to perform cepstral processing on the audio frames to obtain the cepstral features of each audio frame; The compression encoding processing module is used to perform compression encoding processing on the cepstral features of the audio frame and the target pitch features to obtain bitstream data.
[0007] This application provides an audio processing apparatus, including: The acquisition and processing module is used to acquire the bitstream data; The decompression processing module is used to decompress the bitstream data to obtain the cepstral features and target pitch features of each audio frame included in the bitstream data; The decoding processing module is used to decode the cepstral features of the audio frame based on the target pitch features of the audio frame to obtain the decoded audio data.
[0008] In the above technical solution, the data parsing module is also used to perform frame segmentation processing on the audio data to obtain multiple audio frames included in the audio data; For any one of the plurality of audio frames, the following processing is performed: The audio frame is sampled based on the sampling frequency to obtain multiple sampling features corresponding to the audio frame; Based on the minimum fundamental frequency corresponding to the audio frame and the sampling frequency, multiple fundamental frequency periods of the audio frame are determined.
[0009] In the above technical solution, the data parsing module is also used to calculate the reciprocal of the minimum fundamental frequency corresponding to the audio frame to obtain the maximum fundamental frequency value corresponding to the audio frame; The maximum pitch value, the sampling frequency, and the pitch period coefficient are multiplied to obtain the maximum period value of the pitch period. The plurality of fundamental periods are obtained from the range of values from the minimum period value to the maximum period value of the fundamental period.
[0010] In the above technical solution, the first feature calculation module is further used to sort multiple sampling features of the audio frame in ascending order based on the sampling points of the audio frame, and take the first LM sampling features in the ascending order sorting result as candidate sampling features, wherein L represents the number of sampling features of the audio frame, and M represents the maximum period value of the pitch period; Perform the following processing on any of the candidate sampling features: Obtain the next sampling feature that is spaced from the candidate sampling feature by the pitch period; The product of the candidate sampling feature and the next sampling feature is used as the relevant feature of the candidate sampling feature; The sum of the relevant features of the LM candidate sampling features is taken as the relevant feature corresponding to the pitch period.
[0011] In the above technical solution, the second feature calculation module is further used to perform a first filtering process on the plurality of pitch periods and the related features corresponding to the pitch periods to obtain a first filtering result; Based on the maximum value of the pitch period in the first filtering result, the first filtering result is subjected to a second filtering process to obtain a second filtering result; Based on the relevant features in the second filtering result, the second filtering result is sorted in descending order to obtain the descending sort result; Based on the descending sorting result, the target pitch feature corresponding to each audio frame is determined.
[0012] In the above technical solution, the second feature calculation module is further used to delete the relevant features that do not reach the first threshold to obtain the deleted relevant features, and to delete the pitch period corresponding to the relevant features that do not reach the first threshold to obtain the deleted pitch period; The deleted pitch period and the deleted related features are used as the first filtering result.
[0013] In the above technical solution, the second feature calculation module is also used to round down the ratio of the maximum value to the set value to obtain the filter value; The pitch period that is equal to the filter value in the first filtering result is deleted to obtain the deleted pitch period, and the related features corresponding to the pitch period that is equal to the filter value are deleted to obtain the deleted related features. The deleted pitch period and the deleted related features are used as the second filtering result.
[0014] In the above technical solution, the second feature calculation module is further used to, when the number of pitch periods in the descending sorting result reaches a number threshold, take the first N pitch periods in the descending sorting result and the related features corresponding to the first N pitch periods as the target pitch features corresponding to the audio frame, where N represents the number threshold. When the number of fundamental frequencies in the descending sorting result does not reach the number threshold, the set value is supplemented with the fundamental frequencies in the descending sorting result between the Kth and Nth frequencies, and the related features corresponding to the fundamental frequencies in the descending sorting result between the Kth and Nth frequencies, where K represents the number of fundamental frequencies. The first N fundamental frequencies in the supplemented descending sorting result and the related features corresponding to the first N fundamental frequencies are used as the target fundamental frequency features corresponding to the audio frame.
[0015] In the above technical solution, the cepstral processing module is also used to obtain the power spectrum of the audio frame; The power of multiple frequency points in the same frequency domain in the power spectrum is weighted and summed to obtain the frequency band power corresponding to the audio frame. Logarithmic processing is performed on the bandwidth power to obtain the logarithmic processing result; The logarithmic processing result is subjected to discrete cosine transformation to obtain the cepstral features of each audio frame.
[0016] In the above technical solution, the compression encoding processing module is further used to aggregate the cepstral features of the audio frame and the target pitch features corresponding to the audio frame to obtain an aggregation result; The aggregation result is quantized and encoded to obtain bitstream data, and the bitstream data is used as codestream data.
[0017] In the above technical solution, the electronic device is used to acquire the bitstream data; The bitstream data is decompressed to obtain the cepstral features and target pitch features of each audio frame included in the bitstream data; Based on the target pitch features of the audio frame, the cepstral features of the audio frame are decoded to obtain the decoded audio data.
[0018] This application provides an electronic device for audio processing, the electronic device comprising: Memory, used to store executable instructions; The processor, when executing executable instructions stored in the memory, implements the audio processing method provided in the embodiments of this application.
[0019] This application provides a computer-readable storage medium storing executable instructions for inducing a processor to execute and implement the audio processing method provided in this application.
[0020] This application provides a computer program product, including a computer program or instructions, characterized in that the computer program or instructions, when executed by a processor, implement the audio processing method provided in this application.
[0021] The embodiments of this application have the following beneficial effects: The target pitch feature is determined by multiple pitch periods of the audio frame and the corresponding features of the pitch period. The cepstral features and the target pitch feature are compressed and encoded to obtain the bitstream data. This eliminates the need to transmit a large number of encoding parameters to the decoder at the other end, reduces the bit rate of the encoded data during transmission, and thus reduces the network bandwidth or storage space occupied. Attached Figure Description
[0022] Figure 1 This is a schematic diagram of the encoding process of the speech CELP model provided in the embodiments of this application; Figure 2 This is a schematic diagram of an application scenario for the speech encoding and decoding processing system 10 provided in this application embodiment; Figure 3 This is a schematic diagram of the structure of an electronic device 500 for audio processing provided in an embodiment of this application; Figure 4 This is a flowchart illustrating an audio processing method provided in an embodiment of this application; Figure 5 This is a schematic diagram of a process for obtaining the sampling features and pitch period of an audio frame according to an embodiment of this application; Figure 6 This is a flowchart illustrating the process of determining multiple cycles of an audio frame according to an embodiment of this application; Figure 7 This is a flowchart illustrating the process of determining the relevant features corresponding to the pitch period provided in an embodiment of this application; Figure 8 This is a flowchart illustrating the process of determining the target pitch feature corresponding to each audio frame, as provided in an embodiment of this application. Figure 9 This is a schematic flowchart illustrating the process of obtaining a first filtering result through the first filtering process provided in an embodiment of this application. Figure 10 This is a schematic flowchart illustrating the process of obtaining a second filtering result through a second filtering process provided in an embodiment of this application. Figure 11 This is a flowchart illustrating the process of determining the target pitch feature based on the descending sorting result, provided in an embodiment of this application. Figure 12 This is a schematic diagram of a process for obtaining cepstral features by performing cepstral analysis on an audio frame, provided in an embodiment of this application. Figure 13 This is another schematic flowchart of the audio processing method provided in the embodiments of this application; Figure 14 This is a schematic diagram of a process for obtaining decoded audio data provided in an embodiment of this application; Figure 15 This is a schematic diagram of an encoding / decoding principle provided in an embodiment of this application; Figure 16 This is a schematic flowchart of an audio processing method provided in an embodiment of this application; Figure 17 This is a schematic diagram of a process for obtaining multiple candidate pitches, multiple cross-correlation values, and Barker cepstral coefficient features provided in an embodiment of this application. Detailed Implementation
[0023] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0024] In the following description, the terms "first" and "second" are used merely to distinguish similar objects and do not represent a specific ordering of objects. It is understood that "first" and "second" may be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.
[0025] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0026] Before providing a further detailed description of the embodiments of this application, the nouns and terms involved in the embodiments of this application will be explained, and the nouns and terms involved in the embodiments of this application shall be interpreted as follows.
[0027] 1) Vocoder: Vocoder is an abbreviation of Voice Encoder, also known as a speech signal analysis and synthesis system. Its function is to convert acoustic features into sound.
[0028] 2) Spectrograms: These refer to the frequency domain representation of a signal in the time domain. They can be obtained by performing a Fourier transform on the signal. The result is two graphs with amplitude and phase on the vertical axis and frequency on the horizontal axis, respectively. In speech synthesis technology, phase information is often omitted, and only the amplitude information corresponding to different frequencies is retained. 3) Long Short-Term Memory Network (LSTM) Term Memory (LSTM) is a recurrent neural network that incorporates a cell into its algorithm to determine the usefulness of information. Each cell contains an input gate, a forget gate, and an output gate. After information enters the LSTM, its usefulness is determined according to rules. Only information that meets the algorithm's criteria is retained; otherwise, it is forgotten through the forget gate. This network is suitable for processing and predicting important events in time series with relatively long intervals and delays.
[0029] 4) Gate Recurrent Unit (GRU): This is a type of recurrent neural network, similar to Long Short-Term Memory (LSTM), proposed to address issues related to long-term memory and gradients during backpropagation. Compared to LSTM, GRU has one less gate and fewer parameters, achieving comparable performance to LSTM in most cases while significantly reducing computational time.
[0030] 5) Pitch: Speech signals can generally be divided into two categories. One category is voiced sounds with a short periodicity. When a person produces a voiced sound, airflow passes through the glottis, causing the vocal cords to vibrate in a relaxation-oscillation manner, generating a quasi-periodic pulse of airflow. This airflow excites the vocal tract to produce a voiced sound, also known as spoken speech. It carries most of the energy in speech, and its period is called the pitch. The other category is unvoiced sounds with random noise characteristics, produced by the air forced into the oral cavity when the glottis is closed.
[0031] 6) Linear Predictive Coding (LPC): Primarily used in audio signal processing and speech processing, this tool represents the spectrum of digital speech signals in compressed form based on information from a linear prediction model. It is one of the most effective speech analysis techniques.
[0032] 7) Linear Predictive Coding Network (LPCNet): This is a network that cleverly combines digital signal processing and neural networks. It is used as a vocoder in speech synthesis to synthesize high-quality speech in real time.
[0033] In related technologies, many speech coding and decoding techniques originate from the Code-Excited Linear Predictive Speech Coding (CELP) model or its improved versions. See also Figure 1 , Figure 1 This is a schematic diagram of the encoding process of the speech CELP model provided in the embodiments of this application.
[0034] The CELP model encoding process involves preprocessing the original speech signal through high-pass filtering, followed by linear predictive analysis and parameter quantization. The difference between the original signal and the LPC predictive filtering result is then determined as the residual signal. After a series of processing steps, the residual signal yields the gain parameter 'a' and the fixed codebook gain parameter 'c'. Finally, the encoding parameters are transmitted to the receiving end, where the end decoder parses all the encoding parameters and processes them to obtain the final speech signal.
[0035] During the implementation of the embodiments of this application, the applicant discovered the following problems with the related technology: In related technologies, speech encoding and decoding can be applied to different scenarios, including single-person speech, multi-person speech, and scenarios where human voice and other signals coexist. However, speech encoders have high bit rates and require a lot of network bandwidth or storage space.
[0036] To address the aforementioned issues, embodiments of this application provide an audio processing method, apparatus, device, and computer-readable storage medium that can reduce the bitrate of the data stream during transmission, thereby occupying less network bandwidth or storage space. The following describes exemplary applications of the audio processing device provided in this application. The device provided in this application can be implemented as various types of user terminals such as laptops, tablets, desktop computers, set-top boxes, and mobile devices (e.g., mobile phones, portable music players, personal digital assistants, dedicated messaging devices, portable gaming devices), or as a server.
[0037] See Figure 2 , Figure 2 This is a schematic diagram of the application scenario of the voice encoding and decoding processing system 10 provided in the embodiments of this application. Terminal 200-1 and terminal 200-2 are connected to server 100 through network 300. Network 300 can be a wide area network or a local area network, or a combination of the two.
[0038] In some embodiments, server 100 may be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. Terminals 200-1 and 200-2 may be smartphones, tablets, laptops, desktop computers, smart speakers, smartwatches, in-vehicle terminals, etc., but are not limited to these. Terminals and servers can be directly or indirectly connected via wired or wireless communication, which is not limited in this embodiment of the invention.
[0039] Terminals 200-1 and 200-2 have a client with a voice application installed and running, which is used to provide audio data playback to users and can also receive related audio data. Terminal 200-1 sends an access request carrying audio data to server 100.
[0040] Server 100 is used to parse and process audio data to obtain multiple sampling features and multiple pitch periods for each audio frame. The audio data includes multiple audio frames. Based on the multiple pitch periods of each audio frame, autocorrelation processing is performed on the multiple sampling features of each audio frame to obtain the correlation features corresponding to the pitch periods. Based on the multiple pitch periods and the correlation features corresponding to the pitch periods of each audio frame, the target pitch feature corresponding to each audio frame is determined. Cepstral processing is performed on the audio frames to obtain the cepstral features of each audio frame. The cepstral features and target pitch features of the audio frames are compressed and encoded to obtain the bitstream data.
[0041] Server 100 transmits the bitstream data to terminal 200-2, which decompresses the bitstream data to obtain the cepstral features and target pitch features of each audio frame. Based on the target pitch features of the audio frames, the cepstral features are decoded to obtain the decoded audio data. Terminal 200-2 can then play back the audio smoothly and naturally on the client side based on the decoded audio data. Because server 200 can compress and encode the bitstream data using only the cepstral features and target pitch features corresponding to the audio frames throughout the entire processing of the speech encoding and decoding system 10, it avoids transmitting a large number of encoding parameters to the decoder at the other end. This reduces the latency of the server's backend speech transmission, allowing terminal 200-2 to immediately obtain the compressed bitstream data. This enables users of terminal 200-2 to hear the speech content expressed by terminal 200-1 more efficiently and in a shorter time.
[0042] In some embodiments, the server 100 can implement the audio processing method provided in this application embodiment by running a computer program. For example, the computer program can be a native program or software module in an operating system; it can be a native application (APP), that is, a program that needs to be installed in the operating system to run, such as a voice call APP (i.e., the terminal 200-1 mentioned above); it can also be a mini-program, that is, a program that only needs to be downloaded into a browser environment to run; or it can be a voice call mini-program that can be embedded in any APP. In short, the above-mentioned computer program can be any form of application, module or plugin.
[0043] Taking computer programs as an example, in actual implementation, the terminal has applications that support virtual scenarios installed and running. These application scenarios can be any of the following: high-concurrency voice conferencing, live voice broadcasting services, or some bandwidth-limited (2G baseband mode) business applications.
[0044] In some embodiments, terminal 200-1 can be a large-scale conference voice publishing device, while terminal 200-2 can be a smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, or vehicle terminal. When terminal 200-1 acquires audio data from the conference room and sends it to server 100, server 100 parses and processes the audio data to obtain multiple sampling features and multiple pitch periods for each audio frame. Based on the multiple pitch periods of each audio frame, autocorrelation processing is performed on the multiple sampling features of each audio frame to obtain the correlation features corresponding to the pitch periods. Based on the multiple pitch periods of each audio frame and the correlation features corresponding to the pitch periods, the target pitch feature corresponding to each audio frame is determined. Cepstral processing is performed on the audio frames to obtain the cepstral features of each audio frame. The cepstral features and target pitch features of the audio frames are then compressed and encoded to obtain the bitstream data. Server 100 sends the bitstream data to terminal 200-2, where client 200-2 further decompresses the bitstream data to obtain the cepstral features and target pitch features of each audio frame. Based on the target pitch features of the audio frames, the cepstral features of the audio frames are decoded to obtain the decoded audio data. Thus, client 200-2 can encode only the cepstral features and target pitch features corresponding to the audio frames without transmitting a large number of encoding parameters to the decoder at the other end. This reduces the latency of the server's backend voice transmission service, allowing users of client 200-2 to hear the audio data of the venue transmitted by terminal 200-1 more efficiently and in a shorter time.
[0045] In some embodiments, the embodiments of this application can also be implemented with the aid of cloud technology, which refers to a hosting technology that unifies a series of resources such as hardware, software, and networks within a wide area network or local area network to realize the computation, storage, processing, and sharing of data.
[0046] Cloud technology is a general term encompassing network technology, information technology, integration technology, management platform technology, and application technology based on the cloud computing business model. It can form resource pools, allowing for on-demand use with flexibility and convenience. Cloud computing technology will become a crucial support. The backend services of cloud computing systems require substantial computing and storage resources.
[0047] Figure 3 This is a schematic diagram of the structure of an electronic device 500 for audio processing provided in an embodiment of this application. The example given is of the electronic device 500 being a server. Figure 3 The illustrated electronic device 500 for flow limiting includes at least one processor 510, a memory 550, and at least one network interface 520. The various components in the electronic device 500 are coupled together via a bus system 530. It is understood that the bus system 530 is used to implement communication between these components. In addition to a data bus, the bus system 530 also includes a power bus, a control bus, and a status signal bus. However, for clarity, in... Figure 3 The general labeled all buses as Bus System 530.
[0048] The processor 510 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc.
[0049] Memory 550 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), and the volatile memory may be random access memory (RAM). The memory 550 described in this application embodiment is intended to include any suitable type of memory. Memory 550 may optionally include one or more storage devices physically located away from processor 510.
[0050] In some embodiments, memory 550 is capable of storing data to support various operations, examples of which include programs, modules, and data structures or subsets or supersets thereof, as illustrated below.
[0051] Operating system 551 includes system programs for handling various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, driver layer, etc., for implementing various basic business functions and handling hardware-based tasks; The network communication module 553 is used to reach other computing devices via one or more (wired or wireless) network interfaces 520, exemplary network interfaces 520 including: Bluetooth, WiFi, and Universal Serial Bus (USB), etc. In some embodiments, the audio processing apparatus provided in this application can be implemented in software. Figure 3 A speech encoding processing device 555 stored in memory 550 is shown. It can be software in the form of programs and plug-ins, including the following software modules: a data parsing module 5551, a first feature calculation module 5552, a second feature calculation module 5553, a cepstral processing module 5554, and a compression encoding processing module 5555, or an acquisition processing module 5556, a decompression processing module 5557, and a decoding processing module 5558. The data parsing module 5551, the first feature calculation module 5552, the second feature calculation module 5553, the cepstral processing module 5554, and the compression encoding processing module 5555 are used to implement audio encoding functions, and the acquisition processing module 5556, the decompression processing module 5557, and the decoding processing module 5558 are used to implement audio decoding functions. These modules are logically related and can therefore be arbitrarily combined or further split according to the functions implemented.
[0052] In other embodiments, the audio processing apparatus provided in this application can be implemented in hardware. As an example, the audio processing apparatus provided in this application can be a processor in the form of a hardware decoding processor, which is programmed to execute the audio processing method provided in this application. For example, the processor in the form of a hardware decoding processor can be one or more application-specific integrated circuits (ASICs), digital signal processing (DSPs), programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components.
[0053] As mentioned above, the data processing method provided in this application can be implemented by various types of electronic devices, such as servers. See also Figure 4 , Figure 4 This is a flowchart illustrating an audio processing method provided in an embodiment of this application, combined with... Figure 4 The steps shown are explained.
[0054] In step 201, the audio data is parsed to obtain multiple sampling features and multiple pitch periods for each audio frame.
[0055] Here, audio is an important medium in multimedia, representing a form of sound signal. As a carrier of information, audio can be categorized into three types: speech, music, and other sounds. Different types will have different inherent characteristics, including physical sample characteristics and acoustic characteristics. Physical sample characteristics include sampling frequency, time scale, samples, format, encoding, etc., while acoustic characteristics include pitch, melody, rhythm, LPC coefficients, and the structured representation of audio. Sampling characteristics are a parameter used to characterize the fluctuations of sound; the larger the parameter value, the stronger the sound produced. The sampling characteristics of each audio frame record the sound amplitude, and the fundamental period is used to characterize the fundamental period characteristics, that is, it can characterize the trajectory characteristics of the vocal cord vibration frequency in the audio data.
[0056] It should be noted that audio data is obtained through the process of digitizing sound. Specifically, this involves performing analog-to-digital conversion (ADC) on continuous analog audio signals from electronic devices at a certain frequency to obtain audio data. Playing digitized sound involves converting the audio data into analog audio signals via digital-to-analog conversion (DAC).
[0057] In some embodiments, see Figure 5 , Figure 5 This is a schematic diagram of a process for obtaining the sampling features and pitch period of an audio frame according to an embodiment of this application. Figure 4 In step 201 shown, the audio data is parsed to obtain multiple sampling features and multiple pitch periods for each audio frame, which can be obtained through... Figure 5 Steps 2011A-2013A are implemented.
[0058] In step 2011A, the audio data is processed by frame segmentation to obtain multiple audio frames included in the audio data.
[0059] For example, the input audio data is processed into frames with a length of 40ms. For instance, audio data with a total length of 1s is processed into 25 audio frames with a length of 40ms each.
[0060] In step 2012A, the audio frame is sampled based on the sampling frequency to obtain multiple sampling features corresponding to the audio frame.
[0061] Here, in digitizing sound, the sampling frequency represents the number of samples per unit time. The higher the sampling frequency, the smaller the interval between sampling points, and the more realistic the digitized sound.
[0062] Continuing from the previous example, after dividing a 40ms frame into 25 audio frames of 40ms each, the sampling frequency is used to sample one of the 25 audio frames. For example, when the sampling frequency is 16,000 times per second, 640 sampling points are obtained after sampling the 40ms audio frame. The sound amplitude corresponding to each sampling point represents a sampling feature of the sampling point.
[0063] In step 2013A, multiple pitch periods of the audio frame are determined based on the minimum fundamental frequency and sampling frequency corresponding to the audio frame.
[0064] Here, the fundamental frequency is closely related to the structure of an individual's vocal cords, and therefore is used to identify the source of sound. Generally speaking, male speakers have lower fundamental frequencies, while female speakers and children have relatively higher fundamental frequencies.
[0065] Continuing from the previous example, the minimum fundamental frequency of the selected audio frame is determined, and multiple fundamental frequency periods of the current audio frame are determined based on the sampling frequency.
[0066] In some embodiments, see Figure 6 , Figure 6 This is a flowchart illustrating the determination of multiple periods of an audio frame according to an embodiment of this application. Figure 5 The step 2013A shown can also be achieved through Figure 6 Steps 2011B-2013B are implemented.
[0067] In step 2011B, the reciprocal of the minimum fundamental frequency corresponding to the audio frame is taken to obtain the maximum fundamental frequency value corresponding to the audio frame.
[0068] For example, when the fundamental frequency of the selected audio frame is identified and the minimum fundamental frequency of the audio frame is determined to be f, the maximum fundamental frequency of the audio frame is T=1 / f.
[0069] It should be noted that the determination of the fundamental tone value represents the fundamental tone detection, and the purpose of the fundamental tone detection is to find the trajectory curve that is completely consistent with or as close as possible to the vocal cord vibration frequency. The maximum fundamental tone value is represented by the maximum value of the vocal cord vibration period.
[0070] In step 2012B, the maximum pitch value, the sampling frequency, and the pitch period coefficient are multiplied to obtain the maximum period value of the pitch period.
[0071] Continuing the previous example, here the pitch period coefficient includes data and is directly proportional to the period value, representing the pitch period characteristics. The period value and candidate period values are also directly proportional, representing the pitch period characteristics. Multiplying the maximum pitch value T = 1 / f, the sampling frequency c, and the pitch period coefficient a yields the maximum pitch period value M, for example, 1 / f. c a=M.
[0072] In step 2013B, multiple pitch periods are obtained from the range of the minimum pitch period value to the maximum pitch period value.
[0073] Continuing from the previous example, once the maximum period value of the obtained pitch period is determined to be M, the range of pitch period values is [0, M]. Multiple positive integers (e.g., 1, 2, 3) are selected from 0 to M as multiple pitch periods corresponding to the audio frame.
[0074] In step 202, based on multiple pitch periods of each audio frame, autocorrelation processing is performed on multiple sampling features of each audio frame to obtain the correlation features corresponding to the pitch periods.
[0075] Here, each audio frame has multiple pitch periods corresponding to it, and each audio frame also has multiple sampling features corresponding to it. Then, based on the multiple pitch periods determined by each audio frame, autocorrelation is calculated by the multiple sampling features of each audio frame to obtain the correlation features corresponding to the multiple pitch periods.
[0076] In some embodiments, see Figure 7 , Figure 7 This is a flowchart illustrating the process of determining the relevant features corresponding to the fundamental period, as provided in an embodiment of this application. Figure 4 Step 202 shown can also be achieved through Figure 7 Steps 2021A-2024A are implemented.
[0077] In step 2021A, based on the sampling points of the audio frame, multiple sampling features of the audio frame are sorted in ascending order, and the first LM sampling features in the ascending sort result are used as candidate sampling features.
[0078] For example, L represents the number of sampling features of the audio frame, and M represents the maximum period value of the pitch period. The multiple sampling features of the audio frame are sorted according to the sampling order of the sampling points, and the sampling features corresponding to the first LM sampling points in the sorting result are taken as candidate sampling features.
[0079] In step 2022A, the next sampling feature is obtained at a pitch period interval from the candidate sampling feature.
[0080] Continuing from the previous example, after taking the sampling features corresponding to the first LM sampling points in the sorting result as candidate sampling features, we push forward e pitch period sampling points from the sampling point corresponding to the last candidate sampling feature and obtain the sampling feature corresponding to that sampling point as the next sampling feature.
[0081] In step 2023A, the product of the candidate sampling feature and the next sampling feature is used as the relevant feature of the candidate sampling feature.
[0082] Continuing from the previous example, the product of the first LM candidate sampling features and the corresponding LM next sampling features in the ascending sort result is taken as the relevant feature of the candidate sampling feature.
[0083] It should be noted that since the pitch period between the candidate sampling feature and the next sampling feature is a fixed value, the pitch period is between the sampling feature corresponding to the first candidate sampling point and the sampling feature of the next sampling point corresponding to the first candidate sampling point, and the pitch period is also between the sampling feature corresponding to the second candidate sampling point and the sampling feature of the next sampling point corresponding to the second candidate sampling point. Similarly, the pitch period is between the sampling feature corresponding to the Nth candidate sampling point and the sampling feature of the next sampling point corresponding to the Nth candidate sampling point.
[0084] In step 2024A, the sum of the relevant features of the LM candidate sampling features is used as the relevant feature corresponding to the pitch period.
[0085] Continuing from the previous example, the relevant features of the LM candidate sampling features include the product of the first LM candidate sampling features in the ascending sorted result and the corresponding LM next sampling features. The LM relevant features are added together, and the result of the addition is used as the relevant feature corresponding to the pitch period.
[0086] In step 203, the target pitch feature corresponding to each audio frame is determined based on multiple pitch periods of each audio frame and the related features corresponding to the pitch periods.
[0087] Here, the target pitch feature corresponding to each audio frame is characterized by multiple pitch periods and multiple target pitch features in each audio frame.
[0088] In some embodiments, see Figure 8 , Figure 8 This is a flowchart illustrating the process of determining the target pitch feature corresponding to each audio frame, as provided in an embodiment of this application. Figure 4 Step 203 shown can also be achieved through Figure 8Steps 2031A-2034A are implemented.
[0089] In step 2031A, a first filtering process is performed on multiple pitch periods and the related features corresponding to the pitch periods to obtain a first filtering result.
[0090] For example, a first matching rule is set for the relevant features corresponding to the pitch period, and the relevant features that do not match the first matching rule and the pitch period corresponding to the relevant features are deleted to obtain the first filtering result.
[0091] In step 2032A, based on the maximum value of the pitch period in the first filtering result, the first filtering result is subjected to a second filtering process to obtain a second filtering result.
[0092] Continuing from the previous example, based on the maximum value of the pitch period in the first filtering result, a second matching rule is set for multiple pitch periods. Pitch periods that do not match the second matching rule, as well as related features corresponding to the pitch periods, are deleted to obtain the second filtering result. The second matching rule is related to the maximum value of the pitch period in the first filtering result.
[0093] In step 2033A, based on the relevant features in the second filtering result, the second filtering result is sorted in descending order to obtain the descending sort result.
[0094] Continuing from the previous example, the second filtering results are sorted in descending order of the relevant features. Then, the second filtering results are ranked first by the largest relevant feature and the pitch period corresponding to the largest relevant feature, and last by the smallest relevant feature and the pitch period corresponding to the smallest relevant feature.
[0095] In step 2034A, the target pitch feature corresponding to each audio frame is determined based on the descending sorting result.
[0096] Continuing from the previous example, based on the relevant features of different values in the descending sort results and the pitch period corresponding to the relevant features, the target pitch feature corresponding to each audio frame is determined.
[0097] In some embodiments, see Figure 9 , Figure 9 This is a schematic flowchart illustrating the process of obtaining a first filtering result through the first filtering process provided in an embodiment of this application. Figure 8 The step 2031A shown can be achieved through Figure 9 Steps 2031B-2032B are implemented.
[0098] In step 2031B, the relevant features that do not reach the first threshold are deleted to obtain the deleted relevant features, and the pitch period corresponding to the relevant features that do not reach the first threshold is deleted to obtain the deleted pitch period.
[0099] For example, if the first threshold is set to X, then a search is performed on multiple pitch periods and the related features corresponding to the pitch periods. When the related features are less than the first threshold X, the related features and the pitch periods corresponding to the related features are deleted, thereby retaining the related features and the pitch periods corresponding to the related features that are greater than or equal to the first threshold X.
[0100] In step 2032B, the deleted pitch period and the deleted related features are used as the first filtering result.
[0101] Continuing from the previous example, the relevant features that are greater than or equal to the first threshold X, and the pitch period corresponding to the relevant features, are used as the first filtering result.
[0102] In some embodiments, see Figure 10 , Figure 10 This is a schematic flowchart illustrating the process of obtaining a second filtering result through a second filtering process provided in an embodiment of this application. Figure 8 Step 2032A shown can be achieved through Figure 10 Steps 2031C-2033C are implemented.
[0103] In step 2031C, the ratio of the maximum value to the set value is rounded down to obtain the filtered value.
[0104] For example, the maximum value is divided by several different set values to obtain multiple corresponding filter values. The set values represent the least common divisor of the minimum fundamental frequency. For instance, when the minimum fundamental frequency is 60 Hz, the set values can be 2, 3, 4, or 5. When the maximum value of the fundamental frequency period is M, the maximum value is rounded down using the floor function: y1 = floor(M / 2), y2 = floor(M / 3), y3 = floor(M / 4), y4 = floor(M / 5), resulting in different filter values for y1, y2, y3, and y4.
[0105] In step 2032C, the pitch period that is equal to the filter value in the first filtering result is deleted to obtain the deleted pitch period, and the related features corresponding to the pitch period that is equal to the filter value are deleted to obtain the deleted related features.
[0106] Continuing from the previous example, when the pitch period in the first filtering result is equal to any one of the filtering values y1, y2, y3, and y4, the pitch period in the first filtering result that is equal to the filtering value is deleted, and the related features corresponding to the pitch period that is equal to the filtering value are also deleted.
[0107] In step 2033C, the deleted pitch period and the deleted related features are used as the second filtering result.
[0108] Continuing with the previous example, after deleting the pitch period that is equal to the filter value and the related features corresponding to the pitch period that is equal to the filter value in the first filtering result, the remaining pitch period and the related features corresponding to the remaining pitch period are used as the second filtering result.
[0109] In some embodiments, see Figure 11 , Figure 11 This is a flowchart illustrating a process for determining the target pitch feature based on descending order sorting results, provided in an embodiment of this application. Figure 8 Step 2034A shown can be achieved through Figure 11 Steps 2031D-2032D are implemented.
[0110] In step 2031D, when the number of pitch periods in the descending sorting result reaches the number threshold, the first N pitch periods in the descending sorting result and the related features corresponding to the first N pitch periods are used as the target pitch features of the audio frame, where N represents the number threshold.
[0111] For example, when the quantity threshold N is 3, the first 3 pitch periods and their corresponding features in the descending sort result are (x1, y1), (x2, y2), and (x3, y3), respectively. Then, the combined result (x1, y1, x2, y2, x3, y3) is used as the target pitch feature corresponding to the audio frame.
[0112] In step 2032D, when the number of pitch periods in the descending sorting result does not reach the number threshold, the set value is supplemented with the pitch periods in the descending sorting result between the Kth and Nth pitch periods, and the related features corresponding to the pitch periods in the descending sorting result between the Kth and Nth pitch periods, where K represents the number of pitch periods. The first N pitch periods in the supplemented descending sorting result and the related features corresponding to the first N pitch periods are used as the target pitch features corresponding to the audio frame.
[0113] For example, when the quantity threshold N is 4 and the number of pitch periods K is 2, the first 4 pitch periods in the descending sorting result and the corresponding features of the first 4 pitch periods are (x1, y1), (x2, y2), (x3, y3), (x4, y4), respectively. The set values are supplemented to the pitch periods in the descending sorting result between the 2nd and 4th pitch periods and the corresponding features of the pitch periods in the descending sorting result. The set values include 0, that is, (x1, y1), (x2, y2), (x3, y3), (x4, y4) are supplemented to (x1, y1), (x2, y2), (0, 0), (0, 0). The first 4 pitch periods in the descending sorting result and the corresponding features of the first 4 pitch periods are combined and the combined result (x1, y1, x2, y2, 0, 0, 0, 0) is used as the target pitch feature of the audio frame.
[0114] In step 204, cepstral processing is performed on the audio frames to obtain the cepstral features of each audio frame.
[0115] Here, cepstral processing can linearly separate two or more separate signals after convolution. The audio frame is the audio frame obtained after the audio data is segmented according to the time interval in step 2011A. Cepstral processing is performed on the segmented audio frames to obtain the cepstral features corresponding to each audio frame.
[0116] In some embodiments, see Figure 12 , Figure 12 This is a schematic diagram of a process for obtaining cepstral features from an audio frame, provided in an embodiment of this application. Figure 4 The illustrated step 204 can be achieved through Figure 12 Steps 2041A-2044A are implemented.
[0117] In step 2041A, the power spectrum of the audio frame is obtained.
[0118] Here, power spectrum is short for power spectral density function. In the case of a finite signal, it represents the change in signal power within a unit frequency band. Power varies with frequency, that is, the distribution of signal power in the frequency domain, which is expressed as the power spectrum. It is specifically used to analyze the energy of a finite signal with available power, containing some amplitude information of the spectrum, while phase information is discarded. The specific process for obtaining the power spectrum of an audio frame is as follows: the audio frame is subjected to a Fast Fourier Transform (FFT) to obtain its energy distribution in the spectrum. Then, the FFT is performed on each frame signal after windowing to obtain the spectrum of each audio frame. Finally, the power spectrum of the audio frame is obtained by taking the square of the modulus of the audio frame's spectrum.
[0119] It should be noted that windowing refers to multiplying each frame by a Hamming window to increase the continuity between the left and right ends of the audio frame. Assume the signal after framing is S(n), n=0,1,…,N-1, where N is the number of frames. Different values of the Hamming window multiplier will produce different Hamming windows.
[0120] In step 2042A, the power of multiple frequency points in the same frequency domain in the power spectrum is weighted and summed to obtain the frequency band power corresponding to the audio frame.
[0121] Continuing from the previous example, the power spectrum of the audio frame is obtained by taking the square of the modulus of the spectrum. The power corresponding to multiple frequency points in the same frequency domain is determined from the power spectrum, and the multiple powers are weighted and summed to obtain the frequency band power of the audio frame in the frequency band including multiple frequency points.
[0122] In step 2043A, the bandwidth power is logarithmically processed to obtain the logarithmic processing result.
[0123] In step 2044A, the logarithmic processing result is subjected to discrete cosine transformation to obtain the cepstral features of each audio frame.
[0124] Here, the cepstral features of each audio frame contain a preset number of feature dimensions. The feature dimensions represent the number of vectors in the features, and the vectors in the features are used to describe various types of feature information, such as pitch, formants, spectrum, and articulation domain functions.
[0125] For example, the cepstral features of an audio frame can be a Mel-scale spectrogram, a linear logarithmic amplitude spectrogram, or a Buck-scale spectrogram.
[0126] In step 205, the cepstral features and target pitch features of the audio frame are compressed and encoded to obtain the bitstream data.
[0127] In some embodiments, the cepstral features of the audio frame and the target pitch features corresponding to the audio frame are aggregated to obtain an aggregation result. The aggregation result is then quantized and encoded to obtain bitstream data, which is used as the bitstream data.
[0128] For example, the cepstral feature matrix of the audio frame is added to the target pitch feature matrix corresponding to the audio frame, or a weighted addition is performed to obtain an aggregation result. Then, the aggregation result is binary encoded to obtain bitstream data, and the binary encoded bitstream data is used as the code stream data.
[0129] As mentioned above, the data processing method provided in this application can be implemented by various types of electronic devices, such as a client. See also Figure 13 , Figure 13 This is another flowchart illustrating the audio processing method provided in the embodiments of this application, combined with... Figure 13 The steps shown are explained.
[0130] In step 301, the bitstream data is acquired.
[0131] Here, the bitstream data can be obtained in the following ways, see [link / reference]. Figure 2 The server 100 receives the bitstream data obtained by compressing and encoding the cepstral features and target pitch features of the audio frame.
[0132] In step 302, the bitstream data is decompressed to obtain the cepstral features and target pitch features of each audio frame included in the bitstream data.
[0133] For example, after decoding the bitstream data, a quantization value sequence is obtained, and the cepstral features and target pitch features of the audio frame are determined based on the quantization value sequence.
[0134] In step 303, the cepstral features of the audio frame are decoded based on the target pitch features of the audio frame to obtain the decoded audio data.
[0135] In some embodiments, see Figure 14 , Figure 14 This is a schematic diagram of a process for obtaining decoded audio data provided in an embodiment of this application. Figure 13 Step 303 shown can be achieved through Figure 14 Steps 3031-3034 are implemented.
[0136] In step 3031, multiple conditional features for each audio frame are determined based on the target pitch features and cepstral features.
[0137] In some embodiments, the target pitch features and cepstral features of each audio frame are subjected to multi-layer convolution processing to obtain multiple conditional features for each audio frame.
[0138] For example, the target pitch features of each audio frame include multiple pitch periods, related features corresponding to the pitch periods, and cepstral features. Multi-layer convolution is performed to extract multiple sets of high-level speech features for each audio frame as multiple conditional features corresponding to the audio frame. For example, the convolutional layer can be two convolutional layers of size 3 (conv3x1). For an acoustic feature frame containing 18-dimensional BFCC features plus 2-dimensional tone features, the 20-dimensional features in each frame are first passed through two convolutional layers to generate 5 receptive fields based on the acoustic feature frames of the two frames before and the two frames after the frame. The receptive fields of the 5 frames are added to the residual and connected to multiple target pitch features. Then, multiple 129-dimensional conditional vectors are output through two fully connected layers as multiple conditional features for each audio frame, which are used for forward residual prediction.
[0139] In step 3032, autocorrelation processing is performed on the cepstral features to obtain the coefficient features of each audio frame.
[0140] In some embodiments, the cepstral features are subjected to inverse discrete cosine transform to obtain the processing result; the processing result is subjected to exponential processing to obtain the band power; the band power is subjected to interpolation processing to obtain the power spectrum, and the power spectrum is subjected to inverse Fourier transform processing to obtain the autocorrelation coefficient; historical sampling features and historical sample correction values are obtained; the historical sampling features are filtered based on the autocorrelation coefficient to obtain the current sample filtering result; the current sample filtering result, historical sample correction values, and historical sampling features are used as the coefficient features of each audio frame.
[0141] For example, a discrete cosine transform is performed on the cepstral features of each audio frame, followed by processing with the exponential function to obtain the band power. Linear interpolation is then performed on the band power to obtain the power spectrum, followed by an inverse Fourier transform to obtain the autocorrelation coefficient. Finally, the autocorrelation coefficient is calculated using the Levinson-Durbin algorithm. This process involves obtaining historical sampling features S_(t-16)…S_(t-1) and historical sample correction values e_(t...). 1) After that, based on the autocorrelation coefficient, a linear filtering module is used to filter the historical sampling features to obtain the current sample filtering result, and the current sample filtering result, the historical sample correction value and the historical sampling features are used as the coefficient features of each audio frame.
[0142] In step 3033, based on multiple conditional features and coefficient features, the audio prediction features corresponding to each audio frame are obtained.
[0143] In some embodiments, the current sample filtering result is corrected based on historical sampling features, historical correction values, and multiple conditional features to obtain the current sample correction value; the current sample filtering result is predicted based on the current sample correction value to obtain the audio prediction feature of the current sample; the audio prediction feature corresponding to the current audio frame is obtained based on the audio prediction feature of the current sample; and the audio prediction feature corresponding to each audio frame is obtained based on the audio prediction feature corresponding to the current audio frame.
[0144] For example, the historical sampling features include the sample value S_(t) corresponding to the previous sampling point of the current sampling point. 1) Historical correction values include the residual value e_(t) corresponding to the previous sampling point of the current sampling point. 1) The current sample filtering result is corrected using multiple conditional features to obtain the current sample correction value. The current sample filtering result is then added to the current sample correction value to obtain the audio prediction feature of the current sample point. The audio prediction features corresponding to each sample point are merged according to the order of the sample points in the audio frame in the preset time series to obtain the audio prediction feature corresponding to the current audio frame. After performing the same processing on each audio frame, the audio prediction feature corresponding to each audio frame is obtained.
[0145] In step 3034, audio prediction features are synthesized to obtain decoded audio data.
[0146] For example, the feature waveform corresponding to each audio frame is determined based on the audio prediction features corresponding to each audio frame, and the feature waveform corresponding to each audio frame is synthesized to obtain the decoded audio data.
[0147] The following will describe an exemplary application of the embodiments of this application in a real-world application scenario.
[0148] The embodiments of this application can be applied to various audio signal conversion scenarios, such as high-concurrency voice conferencing, live voice broadcasting services, and some bandwidth-limited (2G baseband mode) applications. The following description uses a voice conferencing scenario as an example: In related technologies, speech encoding and decoding based on deep learning methods, such as the LPCNet vocoder, primarily inputs include the following feature values extracted during the encoding stage through relevant acoustic model operators: pitch period, cross-correlation value, Barker cepstral coefficients, etc. The output of the LPCNet network is a Pulse Code Modulation (PCM) sound signal. Because the LPCNet network is a deep integration of traditional acoustic models and deep learning networks, it extracts key and necessary speech feature model parameters through traditional acoustic models and uses deep learning networks to iteratively generate time-domain PCM sound signals from these speech features, significantly reducing the encoding bitrate. However, LPCNet is only suitable for ideal scenarios with a single human voice. When the input signal is not limited to a single human voice, the decoded sound from LPCNet will exhibit discontinuous and incoherent sound, resulting in severely degraded sound quality.
[0149] To address the aforementioned problems, this application proposes an audio processing solution. See also... Figure 15 , Figure 15 This is a schematic diagram of an encoding / decoding principle provided in an embodiment of this application.
[0150] like Figure 15As shown, the frame rate network takes multiple candidate pitches and their multiple cross-correlation values as input, and extracts speech features through multiple convolutional layers as multiple conditional features f of the subsequent sampling rate network. The sampling rate network can calculate LPC coefficients based on multi-dimensional audio features, and based on the LPC coefficients, combine the predicted sampling point values St obtained from multiple time points before the current time. 16...St 1. The linear predictive coding outputs the current sample filtering result pt corresponding to the current sampling point, and the sampling rate network outputs the sample value St corresponding to the sampling point of the previous time step. 1. The residual value et corresponding to the sampling point at the previous time step. 1. The current sample filtering result pt and multiple conditional features f output by the frame rate network are used as input. The output is the residual value et corresponding to the current sampling point. Then, the sampling rate network adds the current sample filtering result pt to the residual value et corresponding to the current sampling point to obtain the predicted value St at the current time. The sampling rate network performs the same processing on each sampling point in the multi-dimensional audio features, continuously looping until it completes the prediction of the sampling values of all sampling points. Based on the predicted values at each sampling point, the entire target audio to be synthesized is obtained.
[0151] In some embodiments, see Figure 16 , Figure 16 This is a flowchart illustrating an audio processing method provided in an embodiment of this application.
[0152] Step 401: Extract multiple candidate pitches (i.e., pitch period) and their multiple autocorrelation values (i.e., correlation features), as well as Barker cepstral coefficient features (i.e., cepstral features).
[0153] From each audio frame of the audio data, obtain the corresponding multiple candidate pitches and their multiple cross-correlation values, as well as the Barker cepstral coefficient features.
[0154] In some embodiments, see Figure 17 , Figure 17 This is a schematic diagram of a process for obtaining multiple candidate pitches, multiple cross-correlation values, and Barker cepstral coefficient features provided in an embodiment of this application. Figure 16 The illustrated step 401 can be achieved through Figure 17 Steps 4011-4015 are implemented.
[0155] Step 4011: Extract the Barker cepstral coefficient features.
[0156] The method for calculating Barker cepstral coefficient features includes obtaining the power spectrum of a speech frame, multiplying the power of multiple frequency points within a frequency band by coefficients and adding them together to obtain the band power, and then taking the logarithm of the band power and performing a Discrete Cosine Transform (DCT) to obtain the Barker cepstral coefficient features.
[0157] Step 4012: Extract multiple candidate pitches and their multiple autocorrelation values.
[0158] The input audio signal (i.e., audio data) is processed by framing. For example, a frame is 40ms in length, resulting in sample values S(i), where i takes the value [1, L], and L is the maximum length of samples in this frame. For example, if fs is the signal sampling frequency and fs = 16000, then L = 0.04. 16000 = 640, the number of samples taken per second.
[0159] Step 4013: Normalize the sample value S(i).
[0160] Find the maximum absolute value Smax of the sample point value (i.e. the sampling feature) S(i). Divide all sample point values S(i) by Smax to obtain the normalized sample point value Sn(i).
[0161] Step 4014: Calculate the autocorrelation value R(k) for the normalized sample value Sn(i).
[0162] Calculate the value of the sample point before the sampling point of this frame and the value of the sample point after a delay of k points, and multiply the cumulative value of the signal of this frame by formula 1. Formula 1 is a method for calculating the value of the sample point provided in the embodiment of this application.
[0163] Formula 1 Here, M represents the number of delay samples (k value) corresponding to the maximum fundamental frequency period (i.e., the maximum period value) (for example, if the minimum fundamental frequency is 80 Hz, it corresponds to a maximum fundamental frequency period (i.e., the maximum period value) of 12.5 ms), which is 12.5. fs / 1000, for example, if fs=16000, then M=200.
[0164] The value of k is directly proportional to the candidate pitch period value and can be used to represent the pitch period characteristics. The value of k ranges from [1, M].
[0165] Step 4015: Select multiple candidate pitch values and their corresponding autocorrelation values from multiple candidate pitches to determine multiple pitch features.
[0166] Based on the first filtering condition, select from M candidate pitches: filter out autocorrelation values R0 that are less than the threshold thrd0 from the M autocorrelation values R0. The autocorrelation values R0 that meet the first filtering condition are sorted from largest to smallest, and the corresponding k value is selected. The second filtering condition is the K0 value corresponding to the largest R value. Filter out autocorrelation values R0 with k value floor(K0 / j), where j is an integer 2, 3, 4, or 5.
[0167] The autocorrelation values R0 that meet the first and second filtering conditions are sorted from largest to smallest. The first N k values and their corresponding R(k) are the final candidate pitch period features, which are used as one of the input features of the deep learning network. When the autocorrelation values R0 that meet the above conditions 1 and 2 are less than N, for example, if Q autocorrelation values R0 do not meet conditions 1 and 2, then the remaining NQ candidate features are filled with 0. For example, if N=4 and Q=2, then the multi-pitch features input to the deep learning network are: K0 corresponding to the pitch period with the largest autocorrelation value R0 and its autocorrelation value R0, K1 corresponding to the pitch period with the second largest autocorrelation value and its autocorrelation value R1, and the next two features are 0, 0, 0, 0.
[0168] Step 402: After compressing and encoding multiple candidate pitches and their multiple cross-correlation values and Barker cepstral coefficient features, transmit them to the decoder at the other end.
[0169] Step 403: The decoder resolves multiple candidate pitches and their respective correlation values, as well as the Barker domain cepstral coefficients.
[0170] The Barker domain cepstral coefficients are subjected to Discrete Cosine Transform (IDCT), then processed by the exponential function to obtain the Barker subband power spectrum, and then linear interpolation is used to obtain the linear spectrum (power spectrum). The autocorrelation coefficients are obtained by inverse Fast Fourier Transform (IFFT), and finally the LPC coefficients (i.e., autocorrelation coefficients) are calculated by the Levinson-Durbin algorithm for subsequent linear filtering.
[0171] Step 404: Based on multiple candidate pitch values and their correlation values, and Barker domain cepstral coefficients, conditional features are obtained. The frame rate network generates multiple f-conditional features from multiple candidate pitch values and their related values, Barker domain cepstral coefficients, through two-level one-dimensional convolution (Conv1D) and two-level fully connected network (FC).
[0172] Step 405: Determine the current sample filtering result based on the linear predictive coding coefficients.
[0173] The sampling rate network uses a linear filtering module based on LPC coefficients to filter the output sample values S_(t-16)...S_(t-1) of the previous 16 historical points to obtain the current sample filtering result pt.
[0174] Step 406: Perform residual processing on the current sample filtering results and conditional features based on historical sample values and residual values to determine the residual value corresponding to the sampling point at the current time.
[0175] The sampling rate network will use the sample value S_(t) corresponding to the previous sampling point. 1) The residual value e_(t) corresponding to the previous sampling point 1) The current sample filtering result pt and the conditional features f output by the frame rate network are used as inputs to output the residual value et corresponding to the current sampling point. The sampling rate network contains two levels of GRU units and a Dual FC (containing two fully connected FC units for processing, whose output is a weighted sum of the independent outputs of the two FC units). The maximum classification probability of the residual value in the current iteration is obtained through the activation layer, and then through the sampling layer, the residual value et corresponding to the maximum classification probability is finally output.
[0176] Step 407: Based on the current sample filtering result and the residual value corresponding to the sampling point at the current time, determine the predicted value at the current time.
[0177] The sampling rate network uses the current sample filtering result pt plus the residual value et corresponding to the current sampling point to obtain the predicted value St at the current time. For example, after 160 iterations of the sampling rate network per frame, it outputs 160 sample values, which is the decoded sound signal (i.e., audio signal) of this frame.
[0178] It should be noted that the sampling rate network performs the same processing on each sampling point in the audio features, continuously looping until it completes the prediction of the sample values of all sampling points. Based on the predicted values at each sampling point, the entire target audio to be synthesized is obtained.
[0179] In this embodiment, multiple candidate pitches and their cross-correlation values and Barker cepstral coefficients are obtained from each audio frame in the audio data. These are then compressed and encoded to obtain bitstream data, which is sent to the decoder. The decoder parses multiple candidate pitches and their respective cross-correlation values, as well as Barker cepstral coefficients, from the bitstream data. Multiple f-conditional features are then determined based on the multiple candidate pitches and their respective cross-correlation values, as well as the Barker cepstral coefficients. Finally, the sampled values at each sampling point are predicted based on the multiple f-conditional features and the Barker cepstral coefficients, thereby obtaining the predicted values at each sampling point and thus the entire target audio to be synthesized.
[0180] This application embodiment obtains bitstream data by compressing and encoding cepstral features and target pitch features, without needing to transmit a large number of encoding parameters to the decoder at the other end. This reduces the bit rate of the bitstream data during transmission, thereby occupying less network bandwidth or storage space. Based on multiple candidate pitch values and their correlation values, as well as multiple conditional features determined by Barker domain cepstral coefficients, it predicts the sampled values of each sampling point corresponding to the audio data in non-single voice signal scenarios. This makes the decoded target audio playback coherent and improves the sound quality, thereby effectively making up for the shortcomings of current deep learning vocoders and effectively solving the current situation that deep learning speech encoders cannot be applied to non-single voice signal scenarios.
[0181] It is understood that, in the embodiments of this application, the audio data and other related data involved need to obtain user permission or consent when the embodiments of this application are applied to specific products and technologies, and the collection, use and processing of related data need to comply with the relevant laws, regulations and standards of relevant countries and regions.
[0182] The following description continues to illustrate the exemplary structure of the audio processing device 555 provided in the embodiments of this application as a software module. In some embodiments, such as... Figure 2 As shown, the software modules stored in the audio processing device 555 of the memory 240 may include: The data parsing module 5551 is used to parse and process the audio data to obtain multiple sampling features and multiple pitch periods of each audio frame, wherein the audio data includes multiple audio frames; the first feature calculation module 5552 is used to perform autocorrelation processing on the multiple sampling features of each audio frame based on the multiple pitch periods of each audio frame to obtain the correlation features corresponding to the pitch periods; the second feature calculation module 5553 is used to determine the target pitch feature corresponding to each audio frame based on the multiple pitch periods of each audio frame and the correlation features corresponding to the pitch periods; the cepstral processing module 5554 is used to perform cepstral processing on the audio frames to obtain the cepstral features of each audio frame; and the compression encoding processing module 5555 is used to perform compression encoding processing on the cepstral features and target pitch features of the audio frames to obtain bitstream data.
[0183] In some embodiments, the data parsing module is further configured to perform frame segmentation on the audio data to obtain multiple audio frames included in the audio data; and to perform the following processing on any one of the multiple audio frames: to perform sampling processing on the audio frame based on the sampling frequency to obtain multiple sampling features corresponding to the audio frame; and to determine multiple pitch periods of the audio frame based on the minimum fundamental frequency and sampling frequency corresponding to the audio frame.
[0184] In some embodiments, the data parsing module is further configured to take the reciprocal of the minimum fundamental frequency corresponding to the audio frame to obtain the maximum fundamental value corresponding to the audio frame; multiply the maximum fundamental value, the sampling frequency, and the fundamental period coefficient to obtain the maximum period value of the fundamental period; and obtain multiple fundamental periods from the range of values from the minimum period value to the maximum period value of the fundamental period.
[0185] In some embodiments, the first feature calculation module is further configured to sort multiple sampling features of an audio frame in ascending order based on the sampling points of the audio frame, and take the first LM sampling features in the ascending order as candidate sampling features, where L represents the number of sampling features of the audio frame and M represents the maximum period value of the pitch period; and perform the following processing on any candidate sampling feature: obtain the next sampling feature that is separated from the candidate sampling feature by a pitch period; multiply the candidate sampling feature by the next sampling feature as the relevant feature of the candidate sampling feature; and sum the relevant features of the LM candidate sampling features as the relevant feature corresponding to the pitch period.
[0186] In some embodiments, the second feature calculation module is further configured to perform a first filtering process on multiple pitch periods and related features corresponding to the pitch periods to obtain a first filtering result; perform a second filtering process on the first filtering result based on the maximum value of the pitch period in the first filtering result to obtain a second filtering result; sort the second filtering result in descending order based on the related features in the second filtering result to obtain a descending sorting result; and determine the target pitch feature corresponding to each audio frame according to the descending sorting result.
[0187] In some embodiments, the second feature calculation module is further configured to delete relevant features that do not reach the first threshold to obtain deleted relevant features, and delete the pitch period corresponding to the relevant features that do not reach the first threshold to obtain deleted pitch period; and use the deleted pitch period and the deleted relevant features as the first filtering result.
[0188] In some embodiments, the second feature calculation module is further configured to round down the ratio of the maximum value to the set value to obtain a filter value; delete the pitch period in the first filtering result that is equal to the filter value to obtain a deleted pitch period; delete the related features corresponding to the pitch period that is equal to the filter value to obtain deleted related features; and use the deleted pitch period and deleted related features as the second filtering result.
[0189] In some embodiments, the second feature calculation module is further configured to, when the number of pitch periods in the descending sorting result reaches a number threshold, use the first N pitch periods in the descending sorting result and the related features corresponding to the first N pitch periods as the target pitch feature corresponding to the audio frame, where N represents the number threshold; when the number of pitch periods in the descending sorting result does not reach the number threshold, supplement the set value with the pitch periods in the descending sorting result whose sorting order is between K and N, and the related features corresponding to the pitch periods whose sorting order is between K and N, where K represents the number of pitch periods, and use the first N pitch periods in the supplemented descending sorting result and the related features corresponding to the first N pitch periods as the target pitch feature corresponding to the audio frame.
[0190] In some embodiments, the cepstral processing module is further configured to obtain the power spectrum of the audio frame; perform weighted summation of the power at multiple frequency points in the same frequency domain in the power spectrum to obtain the frequency band power corresponding to the audio frame; perform logarithmic processing on the frequency band power to obtain the logarithmic processing result; and perform discrete cosine transform processing on the logarithmic processing result to obtain the cepstral features of each audio frame.
[0191] In some embodiments, the compression encoding processing module is further configured to aggregate the cepstral features of the audio frame and the target pitch features corresponding to the audio frame to obtain an aggregation result; quantize and encode the aggregation result to obtain bitstream data, and use the bitstream data as code stream data.
[0192] The acquisition processing module 5556 is used to acquire bitstream data; the decompression processing module 5557 is used to decompress the bitstream data to obtain the cepstral features and target pitch features of each audio frame included in the bitstream data; the decoding processing module 5558 is used to decode the cepstral features of the audio frame based on the target pitch features of the audio frame to obtain the decoded audio data.
[0193] This application provides a computer program product or computer program that includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the audio processing method described above in this application.
[0194] This application provides a computer-readable storage medium storing executable instructions, wherein the executable instructions are stored and when executed by a processor, the processor will execute the audio processing method provided in this application.
[0195] In some embodiments, the computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface memory, optical disk, or CD-ROM; or it may be a variety of devices including one or any combination of the above-mentioned memories.
[0196] In some embodiments, executable instructions may take the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a standalone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.
[0197] As an example, executable instructions may, but do not necessarily, correspond to files in a file system. They may be stored as part of a file that holds other programs or data, for example, in one or more scripts in a Hyper Text Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple collaborating files (e.g., a file that stores one or more modules, subroutines, or code sections).
[0198] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, and improvements made within the spirit and scope of this application are included within the scope of protection of this application.
Claims
1. An audio processing method, characterized in that, include: The audio data is parsed and processed to obtain multiple sampling features and multiple pitch periods for each audio frame, wherein the audio data includes multiple audio frames; the sampling features are parameters characterizing changes in sound fluctuations, and the pitch periods characterize the trajectory features of vocal cord vibration frequencies in the audio data; Based on multiple pitch periods of each audio frame, autocorrelation processing is performed on multiple sampling features of each audio frame to obtain the correlation features corresponding to the pitch periods; A first filtering process is performed on the plurality of pitch periods and the related features corresponding to the pitch periods to obtain a first filtering result; Based on the maximum value of the pitch period in the first filtering result, the first filtering result is subjected to a second filtering process to obtain a second filtering result; Based on the relevant features in the second filtering result, the second filtering result is sorted in descending order to obtain the descending sort result; Based on the descending sorting result, the target pitch feature corresponding to each audio frame is determined; The audio frames are subjected to cepstral processing to obtain the cepstral features of each audio frame; The cepstral features of the audio frame and the target pitch features are compressed and encoded to obtain the bitstream data.
2. The method according to claim 1, characterized in that, The process of parsing and processing the audio data to obtain multiple sampling features and multiple pitch periods for each audio frame includes: The audio data is segmented into frames to obtain multiple audio frames comprising the audio data; For any one of the plurality of audio frames, the following processing is performed: The audio frame is sampled based on the sampling frequency to obtain multiple sampling features corresponding to the audio frame; Based on the minimum fundamental frequency corresponding to the audio frame and the sampling frequency, multiple fundamental frequency periods of the audio frame are determined.
3. The method according to claim 2, characterized in that, Determining the plurality of pitch periods based on the minimum fundamental frequency corresponding to the audio frame and the sampling frequency includes: The maximum pitch value corresponding to the audio frame is obtained by taking the reciprocal of the minimum pitch frequency corresponding to the audio frame. The maximum pitch value, the sampling frequency, and the pitch period coefficient are multiplied to obtain the maximum period value of the pitch period. The plurality of fundamental periods are obtained from the range of values from the minimum period value to the maximum period value of the fundamental period.
4. The method according to claim 1, characterized in that, The method involves performing autocorrelation processing on multiple sampling features of each audio frame based on multiple pitch periods of each audio frame to obtain correlation features corresponding to the pitch periods, including: Based on the sampling points of the audio frame, the multiple sampling features of the audio frame are sorted in ascending order, and the first LM sampling features in the ascending order are taken as candidate sampling features, where L represents the number of sampling features of the audio frame and M represents the maximum period value of the pitch period. Perform the following processing on any of the candidate sampling features: Obtain the next sampling feature that is spaced from the candidate sampling feature by the pitch period; The product of the candidate sampling feature and the next sampling feature is used as the relevant feature of the candidate sampling feature; The sum of the relevant features of the LM candidate sampling features is taken as the relevant feature corresponding to the pitch period.
5. The method according to claim 4, characterized in that, The first filtering process, which involves performing a first filtering process on the plurality of pitch periods and the corresponding features to obtain a first filtering result, includes: The relevant features that do not reach the first threshold are deleted to obtain the deleted relevant features, and the pitch period corresponding to the relevant features that do not reach the first threshold is deleted to obtain the deleted pitch period; The deleted pitch period and the deleted related features are used as the first filtering result.
6. The method according to claim 4, characterized in that, The step of performing a second filtering process on the first filtering result based on the maximum value of the pitch period in the first filtering result to obtain a second filtering result includes: The ratio of the maximum value to the set value is rounded down to obtain the filtered value; The pitch period that is equal to the filter value in the first filtering result is deleted to obtain the deleted pitch period, and the related features corresponding to the pitch period that is equal to the filter value are deleted to obtain the deleted related features. The deleted pitch period and the deleted related features are used as the second filtering result.
7. The method according to claim 4, characterized in that, The step of determining the target pitch feature corresponding to each audio frame based on the descending sorting result includes: When the number of pitch periods in the descending sort result reaches a number threshold, the first N pitch periods in the descending sort result and the related features corresponding to the first N pitch periods are taken as the target pitch features corresponding to the audio frame, where N represents the number threshold. When the number of fundamental frequencies in the descending sorting result does not reach the number threshold, the set value is supplemented with the fundamental frequencies in the descending sorting result between the Kth and Nth frequencies, and the related features corresponding to the fundamental frequencies in the descending sorting result between the Kth and Nth frequencies, where K represents the number of fundamental frequencies. The first N fundamental frequencies in the supplemented descending sorting result and the related features corresponding to the first N fundamental frequencies are used as the target fundamental frequency features corresponding to the audio frame.
8. The method according to claim 1, characterized in that, The cepstral processing of the audio frames to obtain the cepstral features of each audio frame includes: Obtain the power spectrum of the audio frame; The power of multiple frequency points in the same frequency domain in the power spectrum is weighted and summed to obtain the frequency band power corresponding to the audio frame. Logarithmic processing is performed on the bandwidth power to obtain the logarithmic processing result; The logarithmic processing result is subjected to discrete cosine transformation to obtain the cepstral features of each audio frame.
9. The method according to claim 1, characterized in that, The compression encoding process of the cepstral features of the audio frame and the target pitch features to obtain bitstream data includes: The cepstral features of the audio frame and the target pitch features corresponding to the audio frame are aggregated to obtain the aggregation result; The aggregation result is quantized and encoded to obtain bitstream data, and the bitstream data is used as the code stream data.
10. An audio processing method, executed by an electronic device, applied to bitstream data obtained by the audio processing method according to any one of claims 1-9; The method includes: Obtain the bitstream data; The bitstream data is decompressed to obtain the cepstral features and target pitch features of each audio frame included in the bitstream data; Based on the target pitch features of the audio frame, the cepstral features of the audio frame are decoded to obtain the decoded audio data.
11. An audio processing apparatus, characterized in that, The device includes: The data parsing module is used to parse and process the audio data to obtain multiple sampling features and multiple pitch periods for each audio frame. The audio data includes multiple audio frames. The sampling features are parameters that characterize changes in sound fluctuations, and the pitch periods characterize the trajectory features of vocal cord vibration frequencies in the audio data. The first feature calculation module is used to perform autocorrelation processing on multiple sampling features of each audio frame based on multiple pitch periods of each audio frame to obtain the correlation features corresponding to the pitch periods; The second feature calculation module is used to perform a first filtering process on the plurality of pitch periods and the related features corresponding to the pitch periods to obtain a first filtering result; based on the maximum value of the pitch periods in the first filtering result, perform a second filtering process on the first filtering result to obtain a second filtering result; based on the related features in the second filtering result, sort the second filtering result in descending order to obtain a descending sort result; and determine the target pitch feature corresponding to each audio frame according to the descending sort result. The cepstral processing module is used to perform cepstral processing on the audio frames to obtain the cepstral features of each audio frame; The compression encoding processing module is used to perform compression encoding processing on the cepstral features of the audio frame and the target pitch features to obtain bitstream data.
12. An audio processing device, characterized in that, The audio processing device includes: Memory, used to store executable instructions; A processor, when executing executable instructions stored in the memory, implements the method according to any one of claims 1 to 10.
13. A computer-readable storage medium storing executable instructions, characterized in that, When the executable instructions are executed by the processor, they implement the method according to any one of claims 1 to 10.
14. A computer program product comprising a computer program or instructions, characterized in that, When the computer program or instructions are executed by a processor, they implement the method described in any one of claims 1 to 10.
Citation Information
Patent Citations
Encoding method and encoder
CN101303857A
Coding method and decoding method for voice data
CN103247293A