Audio data compression method and related products

CN116259322BActive Publication Date: 2026-05-29TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
TENCENT TECHNOLOGY (SHENZHEN) CO LTD
Filing Date
2021-12-10
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

In audio and video calls and live streaming, network jitter can cause unstable data packets, leading to playback channel congestion, buffer overflow, and consequently playback delays, playback lag, and intermittent sound.

Method used

By classifying the audio data, different categories of audio sub-data are obtained, and different compression ratios are assigned to different categories of audio sub-data to control their differentiated compression playback.

Benefits of technology

It improves the compression and playback effect of audio data, avoids unnatural sound issues, and optimizes the playback effect after audio data compression and speed adjustment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116259322B_ABST
    Figure CN116259322B_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of audio and video, and particularly relates to an audio data compression method, an audio data compression device, a computer readable medium, an electronic device and a computer program product. The method comprises the following steps: acquiring a target compression amount for data compression of audio data, wherein the target compression amount is a data amount difference before and after compression of the audio data; performing classification processing on the audio data to obtain at least two kinds of audio sub-data; respectively assigning target compression ratios to the at least two kinds of audio sub-data according to the target compression amount, wherein the target compression ratio is a ratio of a compression amount of the audio sub-data to a data amount before compression; and performing compression processing on the audio sub-data according to the target compression ratio. The method can control different categories of audio sub-data to be differentially compressed, and improve the compression and playback effect of the audio data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of audio and video technology, specifically relating to an audio data compression method, an audio data compression device, a computer-readable medium, an electronic device, and a computer program product. Background Technology

[0002] In audio and video calls, live streaming, and other applications, audio signals are collected, compressed, and encoded at the sending terminal, then transmitted or distributed over the network to the receiving terminal, where they are finally decoded and played. Under normal circumstances, the sending terminal can ensure a smooth and even transmission of voice-encoded data packets. However, due to unpredictable network jitter, the arrival time of data packets at the receiving terminal is also unstable. Sometimes a long time passes without receiving a single data packet, while at other times a large number of data packets are received in a short period. This can cause intermittent audio playback issues when directly playing these packets. When the receiving terminal receives a large number of data packets in a short period, it can easily lead to playback channel congestion or even buffer overflow, resulting in playback delays, playback lag, and intermittent audio. Summary of the Invention

[0003] The purpose of this application is to provide an audio data compression method, an audio data compression device, a computer-readable medium, an electronic device, and a computer program product, which at least to some extent overcomes the technical problem of poor audio playback stability in related technologies.

[0004] Other features and advantages of this application will become apparent from the following detailed description, or may be learned in part from practice of this application.

[0005] According to one aspect of the embodiments of this application, an audio data compression method is provided, the method comprising:

[0006] Obtain the target compression amount for compressing audio data, wherein the target compression amount is the difference in data volume of the audio data before and after compression;

[0007] The audio data is classified to obtain at least two types of audio sub-data;

[0008] Based on the target compression amount, a target compression ratio is assigned to each of the at least two types of audio sub-data, wherein the target compression ratio is the ratio of the amount of audio sub-data compressed to the amount of data before compression;

[0009] The audio sub-data is compressed according to the target compression ratio.

[0010] According to one aspect of the embodiments of this application, an audio data compression apparatus is provided, the apparatus comprising:

[0011] The acquisition module is configured to acquire a target compression amount for compressing audio data, wherein the target compression amount is the difference in the amount of audio data before and after compression.

[0012] The classification module is configured to classify the audio data to obtain at least two types of audio sub-data;

[0013] The allocation module is configured to allocate a target compression ratio to each of the at least two types of audio sub-data according to the target compression amount, wherein the target compression ratio is the ratio of the compressed amount of the audio sub-data to the amount of data before compression;

[0014] The compression module is configured to compress the audio sub-data according to the target compression ratio.

[0015] According to one aspect of the embodiments of this application, a computer-readable medium is provided having a computer program stored thereon, which, when executed by a processor, implements the audio data compression method as described above.

[0016] According to one aspect of the embodiments of this application, an electronic device is provided, the electronic device comprising: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to perform an audio data compression method as described above by executing the executable instructions.

[0017] According to one aspect of the embodiments of this application, a computer program product or computer program is provided, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the audio data compression method as described in the above technical solutions.

[0018] In the technical solution provided in the embodiments of this application, different categories of audio sub-data can be obtained by classifying audio data. By allocating compression ratios to different categories of audio sub-data, the data compression of different categories of audio sub-data can be controlled to be differentiated. Therefore, the compression and playback of different categories of audio sub-data can be adaptively controlled at an appropriate speed, thereby improving the compression and playback effect of audio data.

[0019] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Attached Figure Description

[0020] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application. It is obvious that the drawings described below are merely some embodiments of this application, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort.

[0021] Figure 1 A schematic diagram of an exemplary system architecture to which the technical solutions of the embodiments of this application can be applied is shown.

[0022] Figure 2 This illustrates the placement of audio / video encoding and decoding devices in a streaming environment.

[0023] Figure 3 This paper illustrates a method in related technologies that uses data compression playback to alleviate playback data congestion problems.

[0024] Figure 4 A flowchart illustrating the steps of an audio data compression method according to one embodiment of this application is shown.

[0025] Figure 5 This illustration shows the effect of classifying audio data according to whether it carries voice content in one embodiment of this application.

[0026] Figure 6 A flowchart illustrating the steps of speech rate estimation for speech sub-data in one embodiment of this application is shown.

[0027] Figure 7 A flowchart illustrating the steps of assigning a target compression ratio to speech sub-data and non-speech sub-data in one embodiment of this application is shown.

[0028] Figure 8 A flowchart illustrating the steps of assigning a target compression ratio to speech segments with different speech rate levels in one embodiment of this application is shown.

[0029] Figure 9 A structural block diagram of the audio data compression apparatus provided in an embodiment of this application is shown.

[0030] Figure 10 A computer system architecture block diagram suitable for implementing the embodiments of this application is shown. Detailed Implementation

[0031] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided to make this application more comprehensive and complete, and to fully convey the concept of the exemplary embodiments to those skilled in the art.

[0032] Furthermore, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. Numerous specific details are provided in the following description to give a thorough understanding of embodiments of this application. However, those skilled in the art will recognize that the technical solutions of this application can be practiced without one or more of the specific details, or other methods, components, apparatuses, steps, etc., can be employed. In other instances, well-known methods, apparatuses, implementations, or operations are not shown or described in detail to avoid obscuring various aspects of this application.

[0033] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.

[0034] The flowcharts shown in the accompanying drawings are merely illustrative and do not necessarily include all content and operations / steps, nor do they necessarily need to be performed in the described order. For example, some operations / steps can be broken down, while others can be combined or partially combined; therefore, the actual execution order may change depending on the specific circumstances.

[0035] In the specific implementation of this application, user voice, video and other related data are involved. When the various embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0036] It should be noted that "multiple" in this article refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following related objects have an "or" relationship.

[0037] Figure 1 A schematic diagram of an exemplary system architecture to which the technical solutions of the embodiments of this application can be applied is shown.

[0038] like Figure 1 As shown, system architecture 100 includes multiple terminal devices that can communicate with each other via, for example, a network 150. For instance, system architecture 100 may include a first terminal device 110 and a second terminal device 120 interconnected via network 150. Figure 1 In one embodiment, the first terminal device 110 and the second terminal device 120 perform unidirectional data transmission.

[0039] For example, the first terminal device 110 can encode audio and video data (e.g., audio and video data streams collected by the terminal device 110) to transmit to the second terminal device 120 via the network 150. The encoded audio and video data is transmitted in the form of one or more encoded audio and video streams. The second terminal device 120 can receive the encoded audio and video data from the network 150, decode the encoded audio and video data to recover the audio and video data, and play or display the content based on the recovered audio and video data.

[0040] In one embodiment of this application, system architecture 100 may include a third terminal device 130 and a fourth terminal device 140 that perform bidirectional transmission of encoded audio and video data, such as during an audio-visual conference. For bidirectional data transmission, each of the third terminal device 130 and the fourth terminal device 140 may encode audio and video data (e.g., an audio and video data stream acquired by the terminal device) for transmission over network 150 to the other terminal device. Each of the third terminal device 130 and the fourth terminal device 140 may also receive encoded audio and video data transmitted by the other terminal device, decode the encoded audio and video data to recover the audio and video data, and play or display content based on the recovered audio and video data.

[0041] exist Figure 1 In the embodiments disclosed herein, the first terminal device 110, the second terminal device 120, the third terminal device 130, and the fourth terminal device 140 may be servers, personal computers, and smartphones, but the principles disclosed herein are not limited to these. The embodiments disclosed herein are applicable to laptop computers, tablet computers, media players, and / or dedicated audio and video conferencing equipment. Network 150 refers to any number of networks that transmit encoded audio and video data between the first terminal device 110, the second terminal device 120, the third terminal device 130, and the fourth terminal device 140, including, for example, wired and / or wireless communication networks. Communication network 150 may exchange data in circuit-switched and / or packet-switched channels. This network may include telecommunications networks, local area networks, wide area networks, and / or the Internet. For the purposes of this application, unless explained below, the architecture and topology of network 150 may be irrelevant to the operation of this application.

[0042] In one embodiment of this application, Figure 2The illustration shows the placement of audio / video encoding and decoding devices in a streaming environment. The subject matter disclosed in this application is equally applicable to other audio / video supported applications, including, for example, audio / video conferencing, digital television (television), and storing compressed audio / video on digital media including CDs, DVDs, memory sticks, etc.

[0043] The streaming system may include an acquisition subsystem 213, which may include audio / video sources 201 such as microphones and cameras, which create an uncompressed audio / video data stream 202. The audio / video data stream 202 is depicted as a thick line to emphasize its high data volume compared to encoded audio / video data 204 (or encoded audio / video bitstream 204). The audio / video data stream 202 may be processed by an electronic device 220, which includes an audio / video encoding device 203 coupled to the audio / video source 201. The audio / video encoding device 203 may include hardware, software, or a combination of hardware and software to implement or enforce aspects of the disclosed subject matter as described in more detail below. Compared to the audio / video data stream 202, the encoded audio / video data 204 (or encoded audio / video bitstream 204) is depicted as thin lines to emphasize the lower data volume of the encoded audio / video data 204 (or encoded audio / video bitstream 204), which can be stored on the streaming server 205 for future use. One or more streaming client subsystems, such as... Figure 2 Client subsystems 206 and 208 can access streaming server 205 to retrieve copies 207 and 209 of encoded audio and video data 204. Client subsystem 206 may include, for example, an audio / video decoding device 210 in electronic device 230. Audio / video decoding device 210 decodes the incoming copy 207 of the encoded audio and video data and produces an output audio / video data stream 211 that can be presented on output device 212 (e.g., a speaker, a display) or another presentation device. In some streaming systems, the encoded audio / video data 204, audio / video data 207, and audio / video data 209 (e.g., audio / video streams) may be encoded according to certain audio / video encoding / compression standards.

[0044] It should be noted that electronic devices 220 and 230 may include other components not shown in the figures. For example, electronic device 220 may include an audio / video decoding device, and electronic device 230 may also include an audio / video encoding device.

[0045] Under normal circumstances, the audio and video data sender can smoothly and evenly transmit encoded data packets to the data receiver. However, when network jitter occurs, the data packet reception time at the data receiver becomes unstable, leading to delays, lags, and intermittent sound during audio and video playback. To address this issue, relevant technologies typically employ data buffering, where the data receiver temporarily stores the received audio and video data in a data buffer to mitigate the impact of network jitter. If congestion occurs in the data buffer, data compression strategies can be implemented to reduce playback lag and delay.

[0046] Figure 3 This illustrates a method in related technologies that uses data compression playback to alleviate playback data congestion problems. For example... Figure 3 As shown, the method for compressing and playing the received audio data at the data receiving end includes the following steps.

[0047] Step S301: Decode the data compressed packet transmitted from the network to obtain the sound signal.

[0048] Step S302: Store the decoded audio signal in the playback buffer. The audio signal obtained after decoding is a signal that can be directly played out. However, in order to avoid poor playback stability due to network jitter, the audio signal can be saved to the playback buffer first, and then the data can be read and played from the playback buffer according to the order in which the data is stored.

[0049] Step S303: Determine if there is congestion in the data stored in the playback buffer. Under normal network transmission conditions, the data write speed and data read speed of the playback buffer are roughly equal, so the playback buffer is always in a dynamically balanced data storage state. However, if severe network jitter occurs, the data receiving speed of the playback buffer may exceed its playback speed. If the amount of data to be played in the playback buffer exceeds the expected amount of cached data, a congestion problem will occur. If a congestion problem occurs, proceed to step S304; if no congestion problem occurs, proceed to step S305.

[0050] Step S304: Calculate the data compression ratio based on the congestion level of the playback buffer, and play the cached data at variable speed according to the data compression ratio. After compressing the cached data, the playback duration of the compressed cached data can be controlled to be lower than the original playback duration, thereby achieving the purpose of quickly consuming the cached data and alleviating the congestion problem of the playback buffer.

[0051] Step S305: If no congestion problem occurs, play and output the cached data at the normal playback speed.

[0052] Based on the above compression playback scheme, it can be seen that the buffered audio signal in the playback buffer can be played at varying speeds through data compression, thereby alleviating the data congestion problem in the playback buffer. However, after the audio signal is compressed and its speed is varied, unnatural auditory problems may easily occur, such as the sound suddenly speeding up, sounding robotic, or the sound becoming unclear after acceleration.

[0053] To address the issue of poor sound playback quality after compression and speed adjustment in related technologies, this application adopts a scheme of content analysis of cached data in the playback buffer. By classifying and processing the cached audio data, different compression ratios can be configured for different categories of audio data, thereby controlling different categories of audio data to use different compression ratios for accelerated playback and optimizing the playback effect after audio data compression and speed adjustment.

[0054] The following detailed description of the audio data compression method, audio data compression device, computer-readable medium, electronic device, and computer program product provided in this application, in conjunction with specific embodiments, provides a detailed explanation of these technical solutions.

[0055] Figure 4 This invention illustrates a flowchart of an audio data compression method according to one embodiment of the present application. The method can be performed by… Figure 1 The terminal device shown can be used for execution, or it can be executed by... Figure 2 The electronic device shown performs this action. For example... Figure 4 As shown, the audio data compression method in this application embodiment can mainly include the following steps S410 to S440.

[0056] Step S410: Obtain the target compression amount for compressing the audio data. The target compression amount is the difference in the amount of audio data before and after compression.

[0057] Step S420: Classify the audio data to obtain at least two types of audio sub-data.

[0058] Step S430: Assign target compression ratios to at least two types of audio sub-data based on the target compression amount. The target compression ratio is the ratio of the amount of audio sub-data compressed to the amount of data before compression.

[0059] Step S440: Compress the audio sub-data according to the target compression ratio.

[0060] In the audio data compression method provided in this application embodiment, different categories of audio sub-data can be obtained by classifying the audio data. The compression ratio of different categories of audio sub-data can be allocated separately, and the data compression of different categories of audio sub-data can be controlled to be differentiated. Therefore, the compression playback effect of different categories of audio sub-data can be adaptively controlled to compress and play at an appropriate speed, thereby improving the compression playback effect of audio data.

[0061] The following sections will provide a detailed explanation of each step in the audio data compression method.

[0062] In step S410, the target compression amount for compressing the audio data is obtained. The target compression amount is the difference in the amount of audio data before and after compression.

[0063] In one embodiment of this application, the target compression level for compressing the audio data can be determined based on the amount of data stored in the audio buffer in real time. For example, the target compression level can be positively correlated with the amount of data stored in the audio buffer in real time; the larger the amount of data stored in real time, the larger the target compression level.

[0064] In one embodiment of this application, the audio data in the audio buffer can be divided into multiple data frames according to a certain time interval, for example, every 20ms of audio data is divided into one data frame. The system monitors the real-time number of stored frames of audio data in the audio buffer; obtains the expected number of stored frames in the audio buffer; and when the real-time number of stored frames is greater than the expected number of stored frames, determines the target compression amount for data compression of the audio data based on the difference between the real-time number of stored frames and the expected number of stored frames.

[0065] Data packets transmitted from the data sender are decoded and written to the audio buffer for playback. Audio data that has already been played is removed from the audio buffer, thus the audio data stored in the buffer is constantly changing. By monitoring the real-time number of stored frames in the audio buffer and comparing it with the expected number of stored frames, the target compression level for compressing the audio data can be predicted. When compressing the audio data according to the target compression level, the real-time number of stored frames in the audio buffer can be kept dynamically changing around the expected number of stored frames.

[0066] In step S420, the audio data is classified to obtain at least two types of audio sub-data.

[0067] Figure 5 This illustration shows the effect of classifying audio data according to whether it carries voice content in one embodiment of this application. Figure 5As shown, speech activity detection is performed on each data frame in the audio data to determine whether each data frame is a speech frame 501 or a non-speech frame 502; the continuously distributed speech frames 501 in the audio data are marked as speech sub-data 503 carrying speech content; and the continuously distributed non-speech frames 502 in the audio data are marked as non-speech sub-data 504 without carrying speech content.

[0068] Voice Activity Detection (VAD) is a detection method that distinguishes between speech and non-speech regions. By extracting features from audio data frames, it can predict whether an audio data frame is a speech frame carrying speech content or a non-speech frame without speech content. Non-speech frames can include silence frames or noise frames.

[0069] VAD algorithms can include various types such as threshold-based VAD, classifier-based VAD, and model-based VAD. Threshold-based VAD algorithms extract features from the time domain (short-time energy, short-term zero-crossing rate, etc.) or frequency domain (MFCC, spectral entropy, etc.) and, by setting appropriate thresholds, distinguish between speech and non-speech. Classifier-based VAD treats speech detection as a binary classification problem of speech / non-speech, and then uses machine learning methods to train a classifier to detect speech. Model-based VAD algorithms utilize a complete acoustic model to distinguish between speech and non-speech segments based on global information after decoding.

[0070] In one embodiment of this application, the speech sub-data carrying speech content can be further classified to obtain speech segments with different speech rate levels. The higher the speech rate level, the faster the speech content carried by the speech sub-data has a speech rate.

[0071] In one embodiment of this application, speech rate estimation is performed on speech sub-data carrying speech content to obtain speech rate state parameters that represent the speed of speech sub-data; and the speech sub-data is marked as speech segments with different speech rate levels according to the speech rate state parameters.

[0072] This application embodiment can pre-configure one or more parameter thresholds for classifying different speech rate levels. The speech rate state parameters obtained through speech rate estimation are compared with each parameter threshold to determine the corresponding speech rate level based on the comparison results. For example, this application embodiment can classify speech rates into three levels from fastest to slowest: high-speed speech signal, medium-speed speech signal, and low-speed speech signal. A high-speed speech signal indicates that the speech rate state parameter of a speech segment is greater than a first parameter threshold; a low-speed speech signal indicates that the speech rate state parameter of a speech segment is less than a second parameter threshold; and a medium-speed speech signal indicates that the speech rate state parameter of a speech segment is between the first and second parameter thresholds. The second parameter threshold is less than the first parameter threshold.

[0073] In one embodiment of this application, a method for estimating speech rate of speech sub-data carrying speech content may include: performing pitch detection on the speech sub-data carrying audio content to obtain the pitch period of each data frame in the speech sub-data; and estimating speech rate of the speech sub-data based on the temporal variation state of the pitch period, wherein the temporal variation state is used to represent the periodic variation trend of the pitch period in the time domain.

[0074] The fundamental tone, as the name suggests, is the basis of sound. Based on the different ways the vocal cords vibrate, sound signals can be divided into voiced and unvoiced sounds. Voiced sounds require periodic vibration of the vocal cords, thus exhibiting a clear periodicity. The frequency of this vocal cord vibration is called the fundamental frequency, and the corresponding period is called the fundamental frequency period. Typically, the fundamental frequency is closely related to the structure of an individual's vocal cords, so the fundamental frequency can also be used to identify the sound source. Estimating the fundamental frequency period is called fundamental frequency detection. The ultimate goal of fundamental frequency detection is to find a trajectory curve that perfectly matches or closely approximates the vocal cord vibration frequency. As one of the important parameters describing the excitation source in speech signal processing, the fundamental frequency period has wide and important applications in speech synthesis, speech compression coding, speech recognition, and speaker identification.

[0075] Pitch detection methods can be broadly categorized into three types: 1) Time-domain methods, which directly estimate the pitch period from the speech waveform. Common methods include autocorrelation, parallel processing, average amplitude difference, and data reduction. 2) Frequency-domain methods, which transform the speech signal to the frequency domain to estimate the pitch period. First, homomorphic analysis is used to eliminate the influence of the vocal tract and obtain information belonging to the excitation part. Then, the pitch period is calculated. The most commonly used method is the cepstral method. The disadvantage of this method is that the algorithm is relatively complex, but the pitch estimation effect is very good. 3) Hybrid methods, which first extract the signal vocal tract model parameters, then use them to filter the signal to obtain the sound source sequence, and finally use the autocorrelation method or the average amplitude difference method to obtain the pitch period.

[0076] Figure 6A flowchart illustrating the steps of speech rate estimation for speech sub-data in one embodiment of this application is shown. Figure 6 As shown, based on the above embodiments, the method for estimating speech rate of speech sub-data according to the temporal variation state of the pitch period may include the following steps S610 to S640.

[0077] Step S610: Compare the pitch periods of two adjacent data frames in the time domain to obtain the period change trend and period change amplitude of the pitch period of the later data frame relative to the earlier data frame. The period change trend is used to represent the rising, falling or flat trend of the pitch period, and the period change amplitude is used to represent the pitch period difference between the later data frame and the earlier data frame.

[0078] Step S620: Determine the temporal variation state of the pitch period between two adjacent data frames based on the period variation trend and the period variation amplitude. The temporal variation state includes at least two of the following: period rising state, period decreasing state, or period remaining flat state.

[0079] In one embodiment of this application, an amplitude threshold associated with the periodic change trend is obtained. The amplitude threshold includes a first threshold associated with an upward trend in the periodicity and a second threshold associated with a downward trend in the periodicity. The first threshold is a positive number, and the second threshold is a negative number. If the periodic change amplitude is less than the first threshold and greater than the second threshold, the temporal change state of the pitch period between two adjacent data frames is marked as a periodicity flat state. If the periodic change amplitude is greater than the first threshold, the temporal change state is marked as a periodicity rising state. If the periodic change amplitude is less than the second threshold, the temporal change state is marked as a periodicity falling state.

[0080] Step S630: Count the frequency of state transitions in the time domain when state changes occur. The state transition frequency is used to represent the number of times the state changes from one state to another in the time domain.

[0081] The pitch period of adjacent frames is divided into three states: "rising," "flat," and "falling." The number of adjacent frames in the same state is counted. For example, if the pitch period states of ten adjacent frames are 0000111122, state "0" represents a "rising" state, meaning the current frame's pitch period value is greater than the previous frame's. A threshold can be set for comparison; a value greater than the threshold indicates a "rising" state. State "1" represents a "flat" state, meaning the current frame's pitch period value is equal to or very close to the previous frame's. A threshold can also be set for comparison; if the difference between the current frame's pitch period and the previous frame's pitch period value is less than the threshold, it is considered a flat state. The "flat" state; state "2" is the pitch period "falling" state, that is, the pitch period value of the current frame is smaller than the pitch period value of the previous frame. Here, a threshold can be set for comparison, that is, if the pitch period value of the current frame is smaller than the pitch period value of the previous frame and is less than the threshold, it is judged as the "falling" state; therefore, the statistical results of the 0000111122 state are that the cumulative value of the continuous "rising" state is 4 (four consecutive 0s), the cumulative value of the continuous "flat" state is 4 (four consecutive 1s), and the cumulative value of the continuous "falling" state is 2 (two consecutive 2s). By counting the number of state switching per unit time, the speech rate is approximated. For example, the series of states 0000111122 switched a total of 3 times.

[0082] Step S640: Estimate the speech rate of the speech sub-data based on the frequency of state switching. The speech rate of the speech sub-data is positively correlated with the frequency of state switching.

[0083] In one embodiment of this application, the frame number of the data frame in the speech sub-data is obtained; and a speech rate state parameter for representing the speech rate of the speech sub-data is determined based on the ratio of the state switching frequency to the frame number.

[0084] The number of data frames in the statistical speech sub-data is Cnt_V, and the switching frequency of the pitch period state in these data frames is Cnt_P. The speech rate can be approximately represented by the speech rate parameter rate_V = Cnt_P / Cnt_V. Two thresholds are set empirically, for example: 0.08 and 0.15. If rate_V is lower than 0.08, it is a low-speed speech signal; if rate_V is higher than 0.15, it is a high-speed speech signal; and if rate_V is between 0.08 and 0.15, it is a medium-speed speech signal.

[0085] In one embodiment of this application, speech rate estimation can also be performed on speech sub-data using curve fitting. In this embodiment, a periodic distribution map of the pitch period in the time domain is plotted based on the pitch period of each data frame in the speech sub-data; curve fitting is performed on the periodic distribution map to obtain a time-domain variation curve representing the time-domain change state of the pitch period; and speech rate estimation is performed on the speech sub-data based on the time-domain variation curve.

[0086] In one embodiment of this application, the method for estimating speech rate of speech sub-data based on the time-domain variation curve includes: obtaining the number of data frames in the speech sub-data; counting the number of extreme points in the time-domain variation curve; and determining a speech rate state parameter to represent the speech rate of the speech sub-data based on the ratio of the number of extreme points to the number of frames.

[0087] Using curve fitting for speech rate estimation can avoid high-frequency changes in the pitch period in a short period of time due to detection errors and other reasons, thereby improving the stability and reliability of speech rate estimation.

[0088] In step S430, a target compression ratio is assigned to at least two types of audio sub-data according to the target compression amount. The target compression ratio is the ratio of the amount of audio sub-data compressed to the amount of data before compression.

[0089] In one embodiment of this application, at least two types of audio sub-data obtained by classifying audio data include speech sub-data carrying speech content and non-speech sub-data not carrying speech content.

[0090] Figure 7 A flowchart illustrating the steps of assigning a target compression ratio to speech sub-data and non-speech sub-data in one embodiment of this application is shown. Figure 7 As shown, based on the above embodiments, step S430, which assigns a target compression ratio to at least two types of audio sub-data according to the target compression amount, may include the following steps S710 to S750.

[0091] Step S710: Determine the first compression ratio based on the target compression amount and the number of frames of non-speech sub-data.

[0092] The number of data frames in the non-speech sub-data is N. unvoice If the target compression amount is, for example, N, then the first compression ratio can be expressed as the ratio of the target compression amount N to the number of frames N. unvoice The ratio, i.e., N / N unvoice .

[0093] Step S720: Configure the target compression ratio for non-speech sub-data based on the smaller value between the first compression ratio and the first compression ratio threshold.

[0094] The first compression ratio threshold represents the maximum compression ratio for non-speech sub-data; for example, the first compression ratio threshold is 0.2. If N / N unvoice If the value is less than or equal to 0.2, the target compression ratio for non-speech sub-data can be configured as N / N. unvoice If N / N unvoice If the value is greater than 0.2, the target compression ratio for non-speech sub-data can be configured to 0.2.

[0095] Step S730: Determine the actual compression amount of the non-speech sub-data based on the target compression ratio of the non-speech sub-data, and determine the expected compression amount of the speech sub-data based on the target compression amount and the actual compression amount of the non-speech sub-data.

[0096] After determining the target compression ratio of the non-speech sub-data in step S720, the actual compression amount of the non-speech sub-data can be determined based on the target compression ratio. For example, when the target compression ratio is N / N unvoice When the target compression ratio is 0.2 (the first compression ratio threshold), the actual compression amount of the non-speech sub-data is N; for example, when the target compression ratio is 0.2, the actual compression amount of the non-speech sub-data is 0.2 * N. unvoice .

[0097] When the actual compression amount of the non-speech sub-data is equal to the target compression amount N, the expected compression amount of the speech sub-data is 0, meaning that the desired data compression effect can be achieved by compressing only the non-speech sub-data. When the actual compression amount of the non-speech sub-data is N1 (N1 is less than the target compression amount N), the expected compression amount of the speech frame can be determined as M = N - N1.

[0098] Step S740: Determine the second compression ratio based on the desired compression amount and the number of frames of the speech sub-data.

[0099] In one embodiment of this application, all voice sub-data are compressed and played using a uniform compression ratio, with the expected compression amount M being proportional to the number of frames N of the voice sub-data. voice The ratio of is the second compression ratio.

[0100] Step S750: Configure the target compression ratio for the speech sub-data based on the smaller value between the second compression ratio and the second compression ratio threshold.

[0101] To avoid excessive compression of voice data leading to sound distortion, this embodiment of the application can set a second compression ratio threshold as the maximum compression ratio. When the second compression ratio determined in step S740 is less than (or equal to) the second compression ratio threshold, the second compression ratio can be used as the target compression ratio for data compression, so that the compression amount of the voice sub-data reaches the desired compression amount. When the second compression ratio determined in step S740 is greater than the second compression ratio threshold, the second compression ratio threshold can be used as the target compression ratio.

[0102] In one embodiment of this application, speech sub-data can be classified into speech segments with different speech rate levels according to speech speed. Based on this, different target compression ratios can be assigned to each speech segment according to the speech rate level. The target compression ratio of a speech segment is negatively correlated with its speech rate level, that is, a speech segment with a relatively fast speech rate can be assigned a lower target compression ratio, while a speech segment with a relatively slow speech rate can be assigned a higher target compression ratio.

[0103] Figure 8 A flowchart illustrating the steps of assigning a target compression ratio to speech segments with different speech rate levels is shown in one embodiment of this application.

[0104] like Figure 8 As shown, based on the above embodiments, the second compression ratio is determined according to the desired compression amount and the number of frames of the voice sub-data, and the target compression ratio of the voice sub-data is configured according to the smaller value between the second compression ratio and the second compression ratio threshold. This may include the following steps S810 to S840.

[0105] Step S810: Obtain the speech frame weights that are negatively correlated with the speech rate level of the speech segment.

[0106] A higher speech rate level indicates a faster speech content delivery rate within a speech segment. Speech frame weights are assigned to speech segments at each speech rate level based on a negative correlation. For example, low-speed speech information can be assigned a first speech frame weight a1, medium-speed speech information can be assigned a second speech frame weight a2 (a smaller than the first speech frame weight a1), and high-speed speech information can be assigned a third speech frame weight a3 (a smaller than the second speech frame weight a2). For instance, the first speech frame weight a1 might be configured as 1.3, the second speech frame weight a2 as 1.15, and the third speech frame weight a3 as 1.

[0107] Step S820: Weight the number of frames of the speech segment according to the speech frame weight to obtain the weighted number of frames of the speech sub-data.

[0108] The weighted frame count of the speech sub-data can be obtained by weighting and summing the frame counts of each speech segment according to their corresponding speech frame weights. The frame counts of each speech segment at different speech rates can be counted in the audio buffer. For example, if the frame count for low-speed speech is K1, for medium-speed speech is K2, and for high-speed speech is K3, then the weighted frame count of the speech sub-data is K = a1*K1 + a2*K2 + a3*K3.

[0109] Step S830: Determine the second compression ratio based on the ratio of the desired compression amount to the weighted frame number.

[0110] By weighting speech segments at different speech rates, the equivalent number of weighted frames for the overall speech subdata can be determined based on balancing the speech rate of each speech segment. The second compression ratio, M / K, can be determined using the ratio of the expected compression amount M to the number of weighted frames K.

[0111] Step S840: The smaller value between the second compression ratio and the second compression ratio threshold is weighted according to the speech frame weight to obtain the target compression ratio of each speech segment.

[0112] The second compression ratio threshold is, for example, Cmax. The smaller value between the second compression ratio threshold and the second compression ratio is selected as the overall target compression ratio of the speech sub-data, C = min(Cmax, K / M). Then, the overall target compression ratio C is weighted according to the speech frame weights of each speech segment to obtain the target compression ratio of each speech segment. For example, the target compression ratio of low-speed speech is a1*C, the target compression ratio of medium-speed speech is a2*C, and the target compression ratio of high-speed speech is a3*C.

[0113] By assigning different compression ratios to speech segments with different speaking speeds, excessive compression of speech frames can be avoided, which can lead to sound distortion and discomfort, thus enabling low-perceptibility compressed playback.

[0114] In step S440, the audio sub-data is compressed according to the target compression ratio.

[0115] Differentiated data compression can be performed on audio sub-data based on the target compression ratio assigned to different categories of audio sub-data.

[0116] In one embodiment of this application, the audio data includes speech sub-data carrying speech content and non-speech sub-data not carrying speech content. The method for compressing the audio sub-data according to a target compression ratio includes: deleting a portion of the non-speech frames in the non-speech sub-data according to the target compression ratio to obtain compressed non-speech sub-data; and superimposing a portion of the speech frames in the speech sub-data according to waveform similarity according to the target compression ratio to obtain compressed speech sub-data.

[0117] For non-speech sub-data, directly deleting some non-speech frames can improve data compression efficiency and save computational resources. For speech sub-data, to improve the playback quality after compression, overlay processing can be performed based on the waveform similarity between speech frames. This waveform similarity-based overlay processing uses a windowed fusion method. For example, first, a speech frame is selected as the reference frame. Then, a sliding time window is used to select the speech frame with the highest waveform similarity to the reference frame. This selected speech frame is then overlaid and fused with the reference frame, and so on, until all speech sub-data is compressed. The sliding time window step size is negatively correlated with the target compression ratio. For example, if the target compression ratio is 1 / X, the sliding time window step size can be X, meaning that overlay fusion is performed once every X speech frames on average.

[0118] In one embodiment of this application, the method for compressing non-speech sub-data includes: randomly selecting a portion of non-speech frames to be deleted from the non-speech sub-data according to a target compression ratio. For example, if the target compression ratio is 1 / X, it means that one non-speech frame needs to be randomly selected from X non-speech frames for deletion. Alternatively, when the target compression ratio is 1 / X, one frame can be deleted sequentially at intervals of X non-speech frames. After deleting the selected non-speech frame from the non-speech sub-data, the adjacent frames of the deleted non-speech frame need to be fade-in / fade-out spliced ​​to avoid introducing noise after deleting the non-speech frame. In some optional embodiments, the same compression strategy as for the speech sub-data can be used for the non-speech sub-data, such as superposition processing based on the waveform similarity between non-speech frames.

[0119] It should be noted that although the steps of the method in this application are described in a specific order in the accompanying drawings, this does not require or imply that the steps must be performed in that specific order, or that all the steps shown must be performed to achieve the desired result. Additional or alternative steps may be omitted, multiple steps may be combined into one step, and / or one step may be broken down into multiple steps.

[0120] The following describes an apparatus embodiment of this application, which can be used to execute the audio data compression method in the above embodiments of this application. Figure 9 A structural block diagram of an audio data compression apparatus provided in an embodiment of this application is shown. Figure 9 As shown, the audio data compression device 900 mainly includes:

[0121] The acquisition module 910 is configured to acquire a target compression amount for compressing audio data, wherein the target compression amount is the difference in the amount of audio data before and after compression.

[0122] The classification module 920 is configured to classify the audio data to obtain at least two types of audio sub-data.

[0123] The allocation module 930 is configured to allocate a target compression ratio to the at least two types of audio sub-data according to the target compression amount, wherein the target compression ratio is the ratio of the compressed amount of the audio sub-data to the amount of data before compression;

[0124] Compression module 940 is configured to compress the audio sub-data according to the target compression ratio.

[0125] In one embodiment of this application, based on the above embodiments, the acquisition module 910 may further include:

[0126] The frame count monitoring module 911 is configured to monitor the real-time storage frame count of the audio data stored in the audio buffer;

[0127] The frame count acquisition module 912 is configured to acquire the desired number of frames stored in the audio buffer;

[0128] The compression amount determination module 913 is configured to determine the target compression amount for data compression of audio data based on the difference between the real-time storage frame number and the expected storage frame number when the real-time storage frame number is greater than the expected storage frame number.

[0129] In one embodiment of this application, based on the above embodiments, the classification module 920 may further include:

[0130] The speech detection module 921 is configured to perform speech activity detection on each data frame in the audio data to determine whether the data frame is a speech frame or a non-speech frame.

[0131] The speech tagging module 922 is configured to tag consecutively distributed speech frames in the audio data as speech sub-data carrying speech content;

[0132] The non-speech tagging module 923 is configured to tag consecutively distributed non-speech frames in the audio data as non-speech sub-data that do not carry speech content.

[0133] In one embodiment of this application, based on the above embodiments, the voice tagging module 922 may further include:

[0134] The speech rate estimation module is configured to estimate the speech rate of the speech sub-data carrying speech content to obtain a speech rate state parameter that represents the speed of speech of the speech sub-data.

[0135] The speech segment labeling module is configured to label the speech sub-data as speech segments with different speech rate levels based on the speech rate status parameter.

[0136] In one embodiment of this application, based on the above embodiments, the speech rate estimation module can be further configured to: perform pitch detection on the speech sub-data carrying audio content to obtain the pitch period of each data frame in the speech sub-data; and perform speech rate estimation on the speech sub-data according to the temporal change state of the pitch period, wherein the temporal change state is used to represent the periodic change trend of the pitch period in the time domain.

[0137] In one embodiment of this application, based on the above embodiments, the speech rate estimation module can be further configured to: compare the pitch periods of two adjacent data frames in the time domain to obtain the periodic change trend and periodic change amplitude of the pitch period of the subsequent data frame relative to the previous data frame, wherein the periodic change trend is used to represent the rising, falling, or flat trend of the pitch period, and the periodic change amplitude is used to represent the pitch period difference between the subsequent data frame and the previous data frame; determine the temporal change state of the pitch period between two adjacent data frames based on the periodic change trend and periodic change amplitude, wherein the temporal change state includes at least two of the following: a periodic rising state, a periodic decreasing state, or a periodic flat state; count the state switching frequency of the time domain change state in the time domain, wherein the state switching frequency is used to represent the number of times the time domain change state switches from one state to another; and estimate the speech rate of the speech sub-data based on the state switching frequency, wherein the speech rate of the speech sub-data is positively correlated with the state switching frequency.

[0138] In one embodiment of this application, based on the above embodiments, the speech rate estimation module can be further configured to: obtain an amplitude threshold associated with the periodic change trend, the amplitude threshold including a first threshold associated with an upward periodic trend and a second threshold associated with a downward periodic trend, the first threshold being a positive number and the second threshold being a negative number; if the periodic change amplitude is less than the first threshold and greater than the second threshold, then the temporal change state of the pitch period between two adjacent data frames is marked as a periodic flat state; if the periodic change amplitude is greater than the first threshold, then the temporal change state is marked as a periodic upward state; if the periodic change amplitude is less than the second threshold, then the temporal change state is marked as a periodic downward state.

[0139] In one embodiment of this application, based on the above embodiments, the speech rate estimation module may be further configured to: obtain the frame number of the data frames in the speech sub-data; and determine a speech rate state parameter representing the speech rate speed of the speech sub-data according to the ratio of the state switching frequency to the frame number.

[0140] In one embodiment of this application, based on the above embodiments, the speech rate estimation module can be configured to: plot a periodic distribution map of the pitch period in the time domain according to the pitch period of each data frame in the speech sub-data; perform curve fitting on the periodic distribution map to obtain a time-domain variation curve representing the time-domain variation state of the pitch period; and estimate the speech rate of the speech sub-data according to the time-domain variation curve.

[0141] In one embodiment of this application, based on the above embodiments, the speech rate estimation module may be further configured to: obtain the number of data frames in the speech sub-data; count the number of extreme points in the time-domain variation curve; and determine a speech rate state parameter to represent the speech rate speed of the speech sub-data based on the ratio of the number of extreme points to the number of frames.

[0142] In one embodiment of this application, based on the above embodiments, the at least two types of audio sub-data include voice sub-data carrying voice content and non-voice sub-data not carrying voice content; the allocation module 930 includes:

[0143] The first compression ratio determination module 931 is configured to determine a first compression ratio based on the target compression amount and the number of frames of the non-speech sub-data.

[0144] The non-speech compression ratio configuration module 932 is configured to configure a target compression ratio for the non-speech sub-data based on the smaller value between the first compression ratio and the first compression ratio threshold.

[0145] The expected compression amount determination module 933 is configured to determine the actual compression amount of the non-speech sub-data based on the target compression ratio of the non-speech sub-data, and to determine the expected compression amount of the speech sub-data based on the target compression amount and the actual compression amount of the non-speech sub-data.

[0146] The second compression ratio determination module 934 is configured to determine a second compression ratio based on the desired compression amount and the number of frames of the speech sub-data.

[0147] The voice compression ratio configuration module 935 is configured to configure a target compression ratio for the voice sub-data based on the smaller value between the second compression ratio and the second compression ratio threshold.

[0148] In one embodiment of this application, based on the above embodiments, the speech sub-data includes speech segments with different speech rate levels; the second compression ratio determination module 934 may be further configured to: obtain speech frame weights that are positively correlated with the speech rate level of the speech segment; perform weighted processing on the number of frames of the speech segment according to the speech frame weights to obtain the weighted number of frames of the speech sub-data; and determine the second compression ratio according to the ratio of the expected compression amount to the weighted number of frames.

[0149] In one embodiment of this application, based on the above embodiments, the speech sub-data includes speech segments with different speech rate levels; the speech compression ratio configuration module 935 may be further configured to: obtain speech frame weights that are positively correlated with the speech rate level of the speech segment; and perform weighted processing on the smaller value between the second compression ratio and the second compression ratio threshold according to the speech frame weights to obtain the target compression ratio of each speech segment.

[0150] In one embodiment of this application, based on the above embodiments, the audio data includes voice sub-data carrying voice content and non-voice sub-data not carrying voice content; the compression module 940 includes:

[0151] The non-speech deletion module 941 is configured to delete a portion of the non-speech frames in the non-speech sub-data according to the target compression ratio, so as to obtain compressed non-speech sub-data.

[0152] The speech overlay module 942 is configured to overlay a portion of the speech frames in the speech sub-data according to waveform similarity based on the target compression ratio, so as to obtain compressed speech sub-data.

[0153] The specific details of the audio data compression apparatus provided in the various embodiments of this application have been described in detail in the corresponding method embodiments, and will not be repeated here.

[0154] Figure 10 A schematic block diagram of a computer system architecture for implementing an electronic device according to embodiments of the present application is shown.

[0155] It should be noted that, Figure 10 The computer system 1000 of the electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.

[0156] like Figure 10As shown, the computer system 1000 includes a central processing unit (CPU) 1001, which can perform various appropriate actions and processes based on programs stored in read-only memory (ROM) 1002 or programs loaded from storage section 1008 into random access memory (RAM). The RAM 1003 also stores various programs and data required for system operation. The CPU 1001, ROM 1002, and RAM 1003 are interconnected via a bus 1004. An input / output interface 1005 (I / O interface) is also connected to the bus 1004.

[0157] The following components are connected to the input / output interface 1005: an input section 1006 including a keyboard, mouse, etc.; an output section 1007 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 1008 including a hard disk, etc.; and a communication section 1009 including a network interface card such as a local area network card, modem, etc. The communication section 1009 performs communication processing via a network such as the Internet. A drive 1010 is also connected to the input / output interface 1005 as needed. A removable medium 1011, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on the drive 1010 as needed so that computer programs read from it can be installed into the storage section 1008 as needed.

[0158] Specifically, according to embodiments of this application, the processes described in the various method flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 1009, and / or installed from removable medium 1011. When the computer program is executed by central processing unit 1001, it performs various functions defined in the system of this application.

[0159] It should be noted that the computer-readable medium shown in the embodiments of this application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disc read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this application, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this application, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such transmitted data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to wireless, wired, etc., or any suitable combination thereof.

[0160] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0161] It should be noted that although several modules or units for the device used to perform actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to the embodiments of this application, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0162] Through the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, server, touch terminal, or network device, etc.) to execute the method according to the embodiments of this application.

[0163] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein.

[0164] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.

Claims

1. An audio data compression method, characterized in that, include: Obtain the target compression amount for compressing audio data, wherein the target compression amount is the difference in data volume of the audio data before and after compression; The audio data is classified to obtain at least two types of audio sub-data, including speech sub-data carrying speech content and non-speech sub-data not carrying speech content. The frequency of state switching of the speech sub-data in the audio data is obtained. The frequency of state switching is used to represent the number of times the time domain change state of the pitch period changes from one state to another. The time domain change state of the pitch period includes at least two of the following: period rising state, period falling state, or period flat state. Based on the frequency of state switching, the speech sub-data is classified into speech segments with different speech rate levels; Based on the target compression amount, a target compression ratio is assigned to each speech segment according to the speech rate level. The target compression ratio is the ratio of the compression amount of the speech segment to the amount of data before compression. The speech segment is compressed according to the target compression ratio.

2. The audio data compression method according to claim 1, characterized in that, Obtain the target compression level for compressing the audio data, including: Monitor the real-time number of stored frames of audio data in the audio buffer; Obtain the desired number of frames to store in the audio buffer; When the number of real-time stored frames is greater than the number of expected stored frames, the target compression amount for compressing the audio data is determined based on the difference between the number of real-time stored frames and the number of expected stored frames.

3. The audio data compression method according to claim 1, characterized in that, The audio data is classified and processed, including: Speech activity detection is performed on each data frame in the audio data to determine whether the data frame is a speech frame or a non-speech frame. The continuously distributed speech frames in the audio data are marked as speech sub-data carrying speech content; The continuously distributed non-speech frames in the audio data are marked as non-speech sub-data that do not carry speech content.

4. The audio data compression method according to claim 3, characterized in that, Based on the frequency of state switching, the speech sub-data is classified into speech segments with different speech rate levels, including: Based on the state switching frequency, the speech rate of the speech sub-data carrying speech content is estimated to obtain a speech rate state parameter that represents the speed of speech of the speech sub-data. The speech sub-data is labeled as speech segments with different speech rate levels based on the speech rate status parameter.

5. The audio data compression method according to claim 4, characterized in that, Obtaining the state switching frequency of the speech sub-data in the audio data includes: Pitch detection is performed on the speech sub-data carrying speech content to obtain the pitch period of each data frame in the speech sub-data; The state switching frequency of the speech sub-data is obtained based on the time-domain change state of the pitch period.

6. The audio data compression method according to claim 5, characterized in that, The frequency of state switching of the speech sub-data is obtained based on the time-domain variation of the pitch period, including: By comparing the pitch periods of two adjacent data frames in the time domain, the periodic variation trend and periodic variation amplitude of the pitch period of the subsequent data frame relative to the previous data frame are obtained. The periodic variation trend is used to represent the rising, falling or flat trend of the pitch period, and the periodic variation amplitude is used to represent the difference in pitch period between the subsequent data frame and the previous data frame. The temporal variation state of the pitch period between two adjacent data frames is determined based on the periodic variation trend and the periodic variation amplitude. The frequency of state transitions during time-domain changes is statistically analyzed.

7. The audio data compression method according to claim 6, characterized in that, Determining the temporal variation state of the pitch period between two adjacent data frames based on the period variation trend and period variation amplitude includes: Obtain an amplitude threshold associated with the cyclical change trend, the amplitude threshold including a first threshold associated with an upward cyclical trend and a second threshold associated with a downward cyclical trend, the first threshold being a positive number and the second threshold being a negative number; If the period change amplitude is less than the first threshold and greater than the second threshold, then the temporal change state of the pitch period between two adjacent data frames is marked as a period flat state. If the periodic change amplitude is greater than the first threshold, then the time-domain change state is marked as a periodic rising state; If the periodic change amplitude is less than the second threshold, the time-domain change state is marked as a periodic decrease state.

8. The audio data compression method according to claim 6, characterized in that, Based on the frequency of state switching, the speech rate of the speech sub-data carrying the speech content is estimated, including: Obtain the frame number of the data frames in the speech sub-data; The speech rate state parameter, used to represent the speech rate speed of the speech sub-data, is determined based on the ratio of the state switching frequency to the number of frames.

9. The audio data compression method according to any one of claims 1 to 8, characterized in that, The method further includes: A first compression ratio is determined based on the target compression amount and the number of frames of the non-speech sub-data. The target compression ratio is configured for the non-speech sub-data based on the smaller value between the first compression ratio and the first compression ratio threshold. The actual compression amount of the non-speech sub-data is determined based on the target compression ratio of the non-speech sub-data, and the expected compression amount of the speech sub-data is determined based on the target compression amount and the actual compression amount of the non-speech sub-data. The second compression ratio is determined based on the desired compression amount and the number of frames of the speech sub-data. The target compression ratio for the speech subdata is configured based on the smaller value between the second compression ratio and the second compression ratio threshold.

10. The audio data compression method according to claim 9, characterized in that, The speech sub-data includes speech segments with different speech rate levels; Determining a second compression ratio based on the desired compression amount and the number of frames of the speech sub-data includes: Obtain the speech frame weights that are positively correlated with the speech rate level of the speech segment; The number of frames in the speech segment is weighted according to the speech frame weight to obtain the weighted number of frames in the speech sub-data. The second compression ratio is determined based on the ratio of the desired compression amount to the weighted number of frames.

11. The audio data compression method according to claim 9, characterized in that, Configure a target compression ratio for the speech sub-data based on the smaller value between the second compression ratio and the second compression ratio threshold, including: Obtain the speech frame weights that are positively correlated with the speech rate level of the speech segment; The smaller value between the second compression ratio and the second compression ratio threshold is weighted according to the speech frame weights to obtain the target compression ratio for each speech segment.

12. The audio data compression method according to any one of claims 1 to 8, characterized in that, The method further includes: According to the target compression ratio, some non-speech frames in the non-speech sub-data are deleted to obtain compressed non-speech sub-data. According to the target compression ratio, some speech frames in the speech sub-data are superimposed based on waveform similarity to obtain compressed speech sub-data.

13. An audio data compression device, characterized in that, include: The acquisition module is configured to acquire a target compression amount for compressing audio data, wherein the target compression amount is the difference in the amount of audio data before and after compression. A classification module is configured to classify the audio data to obtain at least two types of audio sub-data, including speech sub-data carrying speech content and non-speech sub-data without speech content; to obtain the state switching frequency of the speech sub-data in the audio data, wherein the state switching frequency is used to represent the number of times the temporal change state of the pitch period switches from one state to another, and the temporal change state of the pitch period includes at least two of the following: period rising state, period falling state, or period flat state; and to classify the speech sub-data into speech segments with different speech rate levels according to the state switching frequency. The allocation module is configured to allocate a target compression ratio to each of the speech segments according to the target compression amount and the speech rate level, wherein the target compression ratio is the ratio of the compression amount of the speech segment to the amount of data before compression; The compression module is configured to compress the speech segment according to the target compression ratio.

14. A computer-readable medium, characterized in that, The computer-readable medium stores a computer program that, when executed by a processor, implements the audio data compression method according to any one of claims 1 to 12.

15. An electronic device, characterized in that, include: processor; as well as Memory for storing the executable instructions of the processor; The processor is configured to cause the electronic device to perform the audio data compression method according to any one of claims 1 to 12 by executing the executable instructions.

16. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the audio data compression method according to any one of claims 1 to 12.