Audio transmission method, device, terminal, storage medium and program product

By performing subband decomposition and compression encoding of audio signals, error correction encoding is performed for signal energy concentration frequency bands, and redundant data is generated, which solves the problems of high transmission bandwidth and cost in the prior art, and realizes efficient audio data recovery.

CN116959458BActive Publication Date: 2025-08-15TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210405956.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-18
Publication Date
2025-08-15
Estimated Expiration
2042-04-18

AI Technical Summary

Technical Problem

In the prior art, forward error correction encoding consumes a large amount of transmission bandwidth and operational costs in audio transmission, and the packet loss resistance is positively correlated with encoding redundancy, resulting in difficulty in improving communication quality.

Method used

Subband decomposition and compression encoding methods are adopted to correct and encode subband coded data in the frequency band of signal energy concentration, redundant data is generated, and data recovery is used on the receiving end to reduce the amount of redundant data.

Benefits of technology

While improving the audio transmission quality, it reduces the consumption of transmission bandwidth and operating costs by error correction encoding, and improves the data recovery capability of the audio receiver.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116959458B_ABST
    Figure CN116959458B_ABST
Patent Text Reader

Abstract

The embodiments of the present application disclose an audio transmission method, device, terminal, storage medium and program product, which belong to the field of multimedia transmission technology. The method includes: performing sub-band decomposition and compression coding on the input signal to obtain sub-band coding data of at least two groups of signal sub-bands; determining target sub-band coding data from the sub-band coding data based on the energy distribution of the input signal; performing error correction coding on the target sub-band coding data to obtain redundant data; and sending an audio data packet containing sub-band coding data and redundant data to the audio receiving end. The present application decomposes and compresses the input signal into frequency bands, performs error correction coding on the sub-band coding data where the signal energy is concentrated, and can reduce the amount of redundant data while improving the audio transmission quality, thereby reducing the consumption of transmission bandwidth and operating costs by error correction coding.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of multimedia transmission technology, and in particular to an audio transmission method, device, terminal, storage medium, and program product. Background Art

[0002] Voice codecs play a crucial role in modern communication systems. In audio and video calls, the transmitter compresses and packages the audio signal using an encoder, then sends the data to the receiver according to the network transmission format and protocol. The receiver then decompresses and decodes the data packets to produce the audio signal.

[0003] To address packet loss during transmission, the transmitter typically uses forward error correction (FEC) technology for channel coding to generate redundant data packets. Upon determining packet loss, the receiver can recover the data based on the redundant packets to obtain the complete multimedia data.

[0004] However, FEC redundant packets consume additional transmission bandwidth, and the transmission system's ability to withstand packet loss is positively correlated with the degree of coding redundancy. To ensure communication quality, FEC coding redundancy needs to be increased, which results in a significant increase in transmission bandwidth and operating costs. Summary of the Invention

[0005] The embodiments of the present application provide an audio transmission method, apparatus, terminal, storage medium, and program product that can improve audio transmission quality while reducing the consumption of transmission bandwidth and operating costs for error correction coding. The technical solution is as follows:

[0006] In one aspect, the present application provides an audio transmission method, which is performed by an audio transmitting terminal and includes:

[0007] Performing sub-band decomposition and compression coding on the input signal to obtain sub-band coded data of at least two groups of signal sub-bands, where different signal sub-bands correspond to different audio frequency bands;

[0008] determining target sub-band coded data from the sub-band coded data based on energy distribution of the input signal, wherein the audio frequency band corresponding to the target sub-band coded data is a frequency band where signal energy is concentrated;

[0009] performing error correction coding on the target sub-band coded data to obtain redundant data;

[0010] An audio data packet including the sub-band coded data and the redundant data is sent to an audio receiving end, and the audio receiving end is configured to perform data recovery on the sub-band coded data based on the redundant data in the event of packet loss.

[0011] On the other hand, the present application provides an audio transmission method, which is performed by an audio receiving end, and the method includes:

[0012] receiving an audio data packet, the audio data packet comprising redundant data and at least two sets of sub-band coded data, the redundant data being obtained by an audio transmitter performing forward error correction coding on target sub-band coded data among the sub-band coded data, different sub-band coded data corresponding to different audio frequency bands, and the audio frequency band corresponding to the target sub-band coded data being a frequency band where signal energy is concentrated;

[0013] performing packet loss detection on the sub-band coded data;

[0014] In the event that the sub-band coded data is lost, data recovery is performed on the sub-band coded data based on the redundant data to obtain an output signal.

[0015] In another aspect, the present application provides an audio transmission device, comprising:

[0016] A sub-band coding module is used to perform sub-band decomposition and compression coding on the input signal to obtain sub-band coded data of at least two groups of signal sub-bands, where different signal sub-bands correspond to different audio frequency bands;

[0017] a determination module, configured to determine target sub-band coded data from the sub-band coded data based on energy distribution of the input signal, wherein the audio frequency band corresponding to the target sub-band coded data is a frequency band where signal energy is concentrated;

[0018] an error correction coding module, configured to perform error correction coding on the target sub-band coded data to obtain redundant data;

[0019] The data sending module is used to send an audio data packet containing the sub-band coded data and the redundant data to an audio receiving end, and the audio receiving end is used to recover the sub-band coded data based on the redundant data in the event of packet loss.

[0020] In another aspect, the present application provides an audio transmission device, comprising:

[0021] a data receiving module, configured to receive an audio data packet, the audio data packet including redundant data and at least two sets of sub-band coded data, the redundant data being obtained by performing forward error correction coding on target sub-band coded data among the sub-band coded data by an audio transmitting end, different sub-band coded data corresponding to different audio frequency bands, and the audio frequency band corresponding to the target sub-band coded data being a frequency band where signal energy is concentrated;

[0022] a packet loss detection module, configured to perform packet loss detection on the sub-band coded data;

[0023] The decoding module is used to recover the sub-band coded data based on the redundant data to obtain an output signal when the sub-band coded data is lost.

[0024] On the other hand, the present application provides a terminal comprising a processor and a memory; the memory stores at least one program, and the at least one program is loaded and executed by the processor to implement the audio transmission method as described in the above aspects.

[0025] On the other hand, the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores at least one computer program, and the computer program is loaded and executed by a processor to implement the audio transmission method as described in the above aspects.

[0026] According to one aspect of the present application, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a terminal reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the terminal to perform the audio transmission method provided in various optional implementations of the above aspects.

[0027] The technical solutions provided by the embodiments of the present application include at least the following beneficial effects:

[0028] In this embodiment, the input signal is decomposed and compressed into frequency bands to generate at least two sets of sub-band coded data. Error correction coding is then performed on the sub-band coded data where signal energy is concentrated, ensuring the audio receiver's ability to recover the primary audio data. Compared to solutions that directly perform error correction coding on the entire input signal, this approach improves audio transmission quality while reducing the amount of redundant data, thereby reducing the transmission bandwidth and operating costs consumed by error correction coding. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 It is an audio transmission flow chart of the related technical solution;

[0030] Figure 2 is a schematic diagram of an implementation environment provided by an exemplary embodiment of the present application;

[0031] Figure 3 is a flowchart of an audio transmission method provided by an exemplary embodiment of the present application;

[0032] Figure 4 is a flowchart of an audio transmission method provided by another exemplary embodiment of the present application;

[0033] Figure 5is a framework diagram of a subband coding model provided by an exemplary embodiment of the present application;

[0034] Figure 6 is a flowchart of an audio transmission method provided by another exemplary embodiment of the present application;

[0035] Figure 7 is a flowchart of an audio transmission method provided by another exemplary embodiment of the present application;

[0036] Figure 8 is a framework diagram of an audio codec system provided by an exemplary embodiment of the present application;

[0037] Figure 9 is a structural block diagram of an audio transmission device provided by an exemplary embodiment of the present application;

[0038] Figure 10 is a structural block diagram of an audio transmission device provided by another exemplary embodiment of the present application;

[0039] Figure 11 It is a structural block diagram of a terminal provided by an exemplary embodiment of the present application. DETAILED DESCRIPTION

[0040] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.

[0041] In this document, "plurality" refers to two or more. "And / or" describes a relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can mean: A exists alone, A and B exist simultaneously, or B exists alone. The character " / " generally indicates an "or" relationship between the associated objects.

[0042] Speech coding and decoding plays an important role in modern communication systems. Figure 1 As shown in the figure, in a voice call scenario, the sound signal is collected by a microphone. The terminal converts the analog sound signal into a digital sound signal through an analog-to-digital conversion circuit. The digital signal is compressed by a voice encoder and then packaged according to the communication network transmission format and protocol and sent to the receiving end. The receiving device decompresses the data packet and outputs the voice coding compressed code stream. After passing through the voice decoder, the voice digital signal is regenerated. Finally, the voice digital signal is played through the speaker. Voice encoding and decoding effectively reduces the bandwidth required for voice signal transmission, playing a crucial role in saving voice information storage and transmission costs and ensuring the integrity of voice information during transmission over the communication network.

[0043] In practical applications, the instability of the transmission network can lead to packet loss during the transmission process, causing audio stuttering and discontinuity at the receiving end, resulting in a poor listening experience for the listener. Various methods have been adopted to combat network packet loss, including forward error correction (FEC), packet loss concealment, and automatic retransmission requests. FEC anti-packet loss solutions are an effective way to perfectly recover the location of lost packets. FEC-encoded data is packaged and sent to the receiving end, which decodes the FEC code to recover the complete data at the location of the lost packet, achieving perfect recovery. FEC requires additional bandwidth, and the higher the FEC redundancy, the stronger the packet loss resistance, but this also results in an increase in bandwidth. Therefore, how to effectively control FEC redundancy, reduce bandwidth consumption, and achieve optimal end-to-end audio transmission is a topic worth studying.

[0044] This application proposes an audio transmission method, please refer to Figure 2 , which shows a schematic diagram of an implementation environment provided by an exemplary embodiment of the present application. The implementation environment includes: an audio transmitting terminal 110 and an audio receiving terminal 120.

[0045] The audio transmitter 110 combines the sub-band coding and decoding method to perform sub-band decomposition and compression coding on the input signal and perform signal classification. Based on the signal classification results, error correction coding is performed on the sub-band coded data where the energy is concentrated to generate redundant data. The audio transmitter 110 sends each group of sub-band coded data and redundant data to the audio receiver 120. The audio receiver 120 receives and parses the data to detect whether the sub-band coded data is lost. In the event of packet loss, the audio receiver 120 can recover the sub-band signal in the frequency band where the energy is concentrated based on the redundant data, and then obtain the complete output signal through sub-band prediction. By combining sub-band coding and error correction coding, the error correction code of the transmission part of the sub-band can effectively reduce the bit consumption of the error correction coding compared to the error correction coding scheme in the related art, thereby reducing the transmission bandwidth and operating costs.

[0046] It is worth mentioning that the audio transmitting end 110 shown in the figure can also be used as a receiving end to receive audio data, and the audio receiving end 120 can also be used as a transmitting end to send audio data. In addition, the figure only shows two terminals accessing the transmission network. In actual application scenarios (such as multi-person call scenarios or online conference scenarios, etc.), the number of terminals can be greater. The embodiments of this application do not limit the number of terminals or device types.

[0047] Please refer to Figure 3 , which shows a flow chart of an audio transmission method provided by an exemplary embodiment of the present application. This embodiment is described by taking the method executed by the audio sending end as an example. The method includes the following steps:

[0048] Step 301 : performing sub-band decomposition and compression coding on an input signal to obtain sub-band coding data of at least two groups of signal sub-bands, where different signal sub-bands correspond to different audio frequency bands.

[0049] The input signal is a sound signal collected by the terminal through a microphone or other device. In one possible implementation, the audio transmitter converts the input signal from the time domain to the frequency domain, decomposes the input signal into subbands in the frequency domain, and compresses and encodes the signals of each group of signal subbands to obtain subband encoded data for each group of signal subbands. Therefore, different signal subbands correspond to different audio frequency bands.

[0050] Optionally, the audio sending end obtains sub-band coding data of each signal sub-band through one sub-band decomposition and compression coding, or the audio sending end performs multiple sub-band decompositions (for example, first obtains two groups of signal sub-bands through one sub-band decomposition, and then continues to decompose part or all of the signal sub-bands), and then performs compression coding.

[0051] Schematically, in a voice call scenario, the frequency of human speech is usually distributed in the range of 500Hz to 4KHz. Therefore, for the transmission of 16KHz audio files, the audio transmitter first sub-band decomposes and compresses the input signal to obtain sub-band coded data in the 0-8KHz and 8KHz-16KHz frequency bands.

[0052] Step 302 : Based on the energy distribution of the input signal, target sub-band coded data is determined from the sub-band coded data. The audio frequency band corresponding to the target sub-band coded data is a frequency band where the signal energy is concentrated.

[0053] The audio transmitter calculates the energy distribution of the input signal to determine the frequency band where the signal energy is concentrated, and determines the sub-band coded data corresponding to the frequency band as the target sub-band coded data.

[0054] For example, for the input signal in step 301, if the audio transmitting end determines that the signal energy is concentrated in the frequency band of 0-8 kHz, the sub-band coded data of the frequency band is determined as the target sub-band coded data.

[0055] Optionally, the target sub-band coded data is a group of sub-band coded data with the highest energy percentage, or, in the case of finer frequency band division, the target sub-band coded data includes multiple groups of sub-band coded data with the highest energy percentage. This embodiment of the present application is not limited to this.

[0056] Step 303: Perform error correction coding on the target sub-band coded data to obtain redundant data.

[0057] In actual audio transmission scenarios, packet loss can occur during audio data transmission due to factors such as network instability and device hardware failures. This can cause stuttering and discontinuity in the audio playback at the receiving end, resulting in a poor listener experience. Transmission systems typically employ error correction coding to combat network packet loss. Error correction coding, also known as channel coding, primarily includes technologies such as packet loss concealment (PLC), automatic repeat-request (ARQ), forward error correction (FEC), hybrid error correction coding, bit interleaving, and BCH error correction coding. Forward error correction can be implemented using various algorithms, including Reed-Solomon code (RS code), Hamming code, or low-density parity check code (LDPC).

[0058] The audio transmitter performs error correction coding on the target sub-band coded data to generate redundant data, while not performing error correction coding on the coded data in other sub-bands. This ensures that the audio receiver can first recover the sound signal in the important frequency band based on the redundant data in the event of packet loss. It also reduces the loss of transmission bandwidth caused by redundant data.

[0059] Step 304: Send an audio data packet including the sub-band coded data and redundant data to an audio receiving end. The audio receiving end is configured to recover the sub-band coded data based on the redundant data in the event of packet loss.

[0060] The audio transmitter packages each group of sub-band coded data and redundant data corresponding to the input signal and sends them to the audio receiver, so that the audio receiver decodes the sub-band coded data and redundant data and finally outputs the sound signal.

[0061] In summary, in this embodiment, the input signal is decomposed and compressed into frequency bands to generate at least two sets of sub-band coded data. Error correction coding is then performed on the sub-band coded data where signal energy is concentrated, ensuring the audio receiver's ability to recover the primary audio data. Compared to solutions that directly perform error correction coding on the complete input signal, this approach improves audio transmission quality while reducing the amount of redundant data, thereby reducing the transmission bandwidth and operating costs of error correction coding.

[0062] In one possible implementation, developers can set fixed frequency bands for error correction coding based on actual application scenarios. For example, in voice calls, since human voices are typically low-frequency signals, the audio transmitter can be configured to use sub-band coded data in the low-frequency sub-band as the target sub-band coded data. To improve audio coding and transmission quality, the audio transmitter can also determine the target sub-band coded data by calculating the energy ratio.

[0063] Please refer to Figure 4 , which shows a flow chart of an audio transmission method provided by another exemplary embodiment of the present application. This embodiment is described by taking the method executed by the audio sending end as an example, and the method includes the following steps:

[0064] Step 401 : performing analog-to-digital conversion on the analog sound signal collected by the microphone to generate a digital sound signal.

[0065] In a voice call scenario, the sound signal is collected by a microphone. In this case, the sound signal collected by the audio transmitter is an analog signal. The audio transmitter converts the analog sound signal into a digital sound signal through an analog-to-digital conversion circuit for subsequent compression encoding, error correction encoding, and audio transmission.

[0066] Step 402: Perform Fourier transform on the digital sound signal to obtain a frequency domain signal.

[0067] Subband coding converts the original signal from the time domain into the frequency domain, then divides it into several sub-bands and digitally encodes the signal in each sub-band. Because the audio transmitter needs to decompose the input signal into sub-bands, it first converts the time domain signal into the frequency domain. The audio transmitter then performs a Fourier transform on the digital audio signal to obtain the frequency domain audio signal.

[0068] Step 403: perform sub-band decomposition and compression coding on the frequency domain signal to generate sub-band coded data.

[0069] The audio transmitter decomposes the signal into components in different frequency bands to remove signal correlation, and then samples, quantizes, and encodes each component separately, thereby obtaining multiple sets of mutually uncorrelated codewords. In one possible implementation, step 403 specifically includes the following steps 403a to 403b (not shown):

[0070] Step 403a: performing sub-band decomposition on the frequency domain signal through at least two band-pass filters corresponding to different frequency bands to obtain at least two sub-band signals, where the frequency bands of the band-pass filters are continuous.

[0071] like Figure 5As shown in the figure, the basic concept of speech subband coding is that the transmitter first decomposes the input signal into several subband signals in different frequency bands through a set of bandpass filters. These subband signals are then frequency-shifted to convert them into baseband signals, and each baseband signal is sampled separately. The sampled signals are quantized and encoded, and then combined into a total bit stream for transmission to the receiver. Subband coding optimizes the number of bits per subband based on the human auditory characteristics to achieve better auditory quality, while also saving storage resources and reducing transmission bandwidth.

[0072] In one possible implementation, the audio transmitter in the embodiment of the present application performs sub-band decomposition and compression encoding on the input signal based on the above basic concept to obtain sub-band encoded data. The audio transmitter first uses a set of bandpass filters, such as a quadrature mirror filter (QMF), to divide the frequency band of a frame of input signal into several continuous frequency bands, each of which is called a sub-band.

[0073] Step 403b: frequency shift and quantization coding are performed on the sub-band signal to obtain sub-band coded data.

[0074] The audio transmitter performs frequency shifting on each sub-band signal to a higher frequency end and quantizes and encodes the frequency-shifted sub-band signal. Optionally, the audio transmitter uses a unified encoding scheme to encode each group of sub-band signals, or uses a separate encoding scheme to encode each group of sub-band signals. This is not limited in the present embodiment.

[0075] Step 404 : Determine the low-frequency energy proportion of the low-frequency sub-band based on the sample signals of the input signal in each audio frequency band.

[0076] Optionally, after sub-band decomposition, the audio sending end performs compression encoding and calculation of the low-frequency energy ratio simultaneously, or the audio sending end calculates the low-frequency energy ratio after compression encoding the sub-band signal. This embodiment of the present application is not limited to this.

[0077] The audio transmitter calculates the energy percentage of the low-frequency sub-band to determine the sub-band signal with the most concentrated energy. If the energy percentage of the low-frequency sub-band is high, it indicates that the signal energy is concentrated in the low-frequency sub-band; if the energy percentage of the low-frequency sub-band is low, it indicates that the signal energy is concentrated in the high-frequency sub-band.

[0078] Indicatively, the formula for calculating the proportion of low-frequency energy is as follows:

[0079]

[0080] Wherein, x(k, i) is the i-th sample signal of the k-th subband after subband decomposition of a single frame signal. The larger the k value, the higher the corresponding subband frequency. k=1 represents a low-frequency subband, and M is the total number of subbands.

[0081] Optionally, when the input signal is decomposed into two groups of sub-band signals, the audio transmitter only needs to calculate the energy proportion of a low-frequency sub-band. When the total number of sub-bands is greater than 2, the audio transmitter calculates the energy proportion of a group of sub-bands with the lowest frequency, or calculates the energy proportion of multiple groups of sub-bands with the lowest frequency. Developers can set the calculation method of the low-frequency energy proportion and the determination method of the target sub-band encoding data according to factors such as the actual application scenario and the audio file format. For example, when the audio file transmitted by the terminal is 32KHz, the audio transmitter can first decompose the input signal into two frequency bands of 0-16KHz and 16-32KHz, and then decompose the 0-16KHz frequency band into two frequency bands of 0-8KHz and 8-16KHz, and calculate the low-frequency energy proportion of the 0-8KHz frequency band and the 8-16KHz frequency band. The embodiments of the present application are not limited to this.

[0082] Step 405 : Determine target sub-band coded data from the sub-band coded data based on the low-frequency energy ratio.

[0083] The audio transmitter determines the frequency band where the energy is concentrated based on the low-frequency energy ratio, thereby determining the target sub-band coded data. In one possible implementation, step 405 specifically includes the following steps 405a to 405b (not shown in the figure):

[0084] Step 405a: when the low-frequency energy ratio is higher than the threshold, the sub-band coded data corresponding to the low-frequency sub-band is determined as the target sub-band coded data.

[0085] The audio transmitter stores a threshold value. After calculating the low-frequency energy ratio, the input signal is classified by comparing the low-frequency energy ratio with the threshold value, and the target sub-band coding data is determined.

[0086] Schematically, the input signal is decomposed into sub-band signals of two frequency bands, 0-8KHz and 8-16KHz, with a threshold of 50%. If the low-frequency energy proportion of 0-8KHz is higher than 50%, the audio transmitter determines the sub-band coded data of the 0-8KHz frequency band as the target sub-band coded data.

[0087] Step 405b: when the low frequency energy ratio is lower than the threshold, the sub-band coded data corresponding to the high frequency sub-band is determined as the target sub-band coded data, and the audio frequency corresponding to the high frequency sub-band is higher than the audio frequencies corresponding to other signal sub-bands.

[0088] Optionally, when the input signal is decomposed into two sub-band signals, if the low-frequency energy ratio is lower than a threshold, the audio transmitter directly determines the sub-band coded data of the high-frequency sub-band as the target sub-band coded data. When the input signal is decomposed into three or more sub-band signals, the audio transmitter can calculate the energy ratios of the multiple sub-band signals to determine the sub-band with the most concentrated energy, and then determine the target sub-band coded data.

[0089] Step 406: Perform error correction coding on the target sub-band coded data to obtain redundant data.

[0090] The specific implementation of step 406 can refer to the above step 303, and will not be repeated here in this embodiment of the present application.

[0091] Step 407: Generate a signal type identifier based on the low-frequency energy ratio.

[0092] The signal type identifier is used to indicate whether the input signal belongs to a voiced signal or an unvoiced signal, wherein the low-frequency energy ratio of the voiced signal is higher than a threshold, and the low-frequency energy ratio of the unvoiced signal is lower than the threshold.

[0093] In one possible implementation, after calculating the low-frequency energy ratio, the terminal classifies the input signal into two types: voiced and unvoiced. Voiced signals are sound signals whose energy is concentrated in the low-frequency region, while unvoiced signals are sound signals whose energy is concentrated in the high-frequency region. Voiced and unvoiced signals have different signal type identifiers.

[0094] In another possible implementation, after calculating the low-frequency energy fraction, the audio transmitter first classifies the input signal and generates a signal type identifier. Based on the signal type identifier, the audio transmitter determines the target subband coded data and performs error correction coding. If the signal type identifier indicates a voiced sound signal, the audio transmitter performs error correction coding on the subband coded data for the low-frequency subband; if the signal type identifier indicates an unvoiced sound signal, the audio transmitter performs error correction coding on the subband coded data for the high-frequency subband.

[0095] Step 408: Pack the sub-band coded data, redundant data, and signal type identifier to generate an audio data packet.

[0096] The audio transmitter sends the signal type identifier together with the subband coded data and redundant data to the audio receiver, so that the audio receiver can determine the target subband coded data based on the signal type identifier in the event of packet loss and perform data recovery and subband prediction.

[0097] Step 409: Send the audio data packet to the signal receiving end.

[0098] The audio receiving end is used to determine the target sub-band coded data based on the signal type identifier in the event of packet loss, and to perform data recovery based on the target sub-band coded data and redundant data.

[0099] In an embodiment of the present application, the audio transmitter determines the frequency band where the energy is concentrated by calculating the proportion of low-frequency energy in the low-frequency sub-band, and then determines the target sub-band coding data, so that error correction coding can be performed on the actually important sub-band signal, avoiding the situation where a continuous signal cannot be restored when a packet is lost due to error correction coding of a fixed frequency band, thereby improving the signal transmission quality on the basis of reducing the transmission bandwidth.

[0100] The above embodiments illustrate the process of sub-band coding and error correction coding at the audio transmitter. For the audio receiver, after receiving the audio data packet, it first determines whether there is packet loss. In the event of packet loss, the audio receiver needs to perform data recovery and sub-band prediction on the sub-band coded data based on redundant data, so as to output a continuous sound signal. Please refer to Figure 6 , which shows a flowchart of an audio transmission method provided by an exemplary embodiment of the present application. This embodiment is described by taking the method executed by an audio receiving end as an example. The method includes the following steps:

[0101] Step 601: Receive an audio data packet.

[0102] The audio data packet contains redundant data and at least two sets of sub-band coded data. The redundant data is obtained by the audio transmitter performing forward error correction coding on the target sub-band coded data in the sub-band coded data. Different sub-band coded data correspond to different audio frequency bands. The audio frequency band corresponding to the target sub-band coded data is the frequency band where the signal energy is concentrated.

[0103] After receiving the audio data packet, the audio receiving end parses the data to obtain the sub-band coded data and redundant data and caches the data.

[0104] Step 602: Perform packet loss detection on the sub-band coded data.

[0105] In one possible implementation, during data encoding, the audio transmitter assigns consecutive numbers to the sub-band coded data according to the timing of signal acquisition. After parsing the data, the audio receiver checks whether the corresponding numbers for the sub-band coded data are consecutive. If the numbers are consecutive, it is determined that the sub-band coded data has not been lost. If the numbers are not consecutive, it is determined that packet loss has occurred.

[0106] Step 603: When sub-band coded data is lost, data recovery is performed on the sub-band coded data based on the redundant data to obtain an output signal.

[0107] If packet loss is not detected, the audio receiver directly performs sub-band decoding. If packet loss is detected, the audio receiver first retrieves redundant data from the data buffer and adjacent data packets for error correction decoding to obtain the sub-band coded data at the packet loss location. Sub-band decoding and sub-band prediction are then performed to obtain a continuous output signal.

[0108] In an embodiment of the present application, an audio receiving end receives an audio data packet containing redundant data and sub-band coded data, wherein the redundant data is obtained by error correction coding of data in the energy concentrated frequency band by the audio sending end. Compared with the method of directly performing error correction coding on the complete input signal, while improving the network's anti-packet loss capability, on the one hand, it can reduce the amount of redundant data and reduce the storage resources consumed by the audio receiving end to cache data, and on the other hand, it can reduce the transmission bandwidth and operating costs.

[0109] Please refer to Figure 7 , which shows a flowchart of an audio transmission method provided by another exemplary embodiment of the present application. This embodiment is described by taking the method executed by the audio receiving end as an example, and the method includes the following steps:

[0110] Step 701: Receive an audio data packet.

[0111] Step 702: Perform packet loss detection on the sub-band coded data.

[0112] The specific implementation of steps 701 to 702 can refer to the above steps 601 to 602, and will not be repeated here in this embodiment of the present application.

[0113] Step 703: Determine target sub-band coded data from the sub-band coded data based on the signal type identifier.

[0114] In one possible implementation, the audio data packet further includes a signal type identifier, which indicates whether the sound signal corresponding to the subband coded data is a voiced signal or an unvoiced signal. The target subband coded data for voiced signals is the subband coded data for the low-frequency subband, while the target subband coded data for unvoiced signals is the subband coded data for the high-frequency subband. In other words, a voiced signal is one whose signal energy is concentrated in the low-frequency region, while an unvoiced signal is one whose signal energy is concentrated in the non-low-frequency region.

[0115] When packet loss occurs, the audio receiver needs to read the relevant redundant data and adjacent data packets from the data buffer for error correction decoding. The redundant data is obtained by the audio receiver through error correction encoding of the target sub-band coded data. Therefore, the audio receiver first determines the target sub-band coded data from at least two sets of sub-band coded data based on the signal type (voiced or unvoiced) indicated by the signal type identifier.

[0116] Step 704: Perform error correction decoding on the target sub-band coded data based on the redundant data and the sub-band coded data in the adjacent audio data packets.

[0117] Based on the error correction coding algorithm of the audio transmitter, the audio receiver uses the corresponding error correction decoding algorithm to perform error correction decoding to obtain the sub-band coded data and signal classification identifier of the packet loss position.

[0118] Step 705: perform sub-band decoding on the target sub-band coded data after error correction decoding to obtain a target sub-band signal.

[0119] After the audio receiving end recovers the sub-band coded data at the packet loss position, it compresses and decodes the complete target sub-band coded data to obtain the target sub-band signal.

[0120] Step 706: Perform data recovery on other sub-band coded data based on the target sub-band signal.

[0121] Redundant data is generated by the audio transmitter through error correction coding of the target subband's coded data. The audio receiver also uses this redundant data to recover lost packets of the target subband's coded data. Audio data is transmitted over the channel in bundled packets, so packet loss means that all subbands of coded data are lost. Therefore, the audio receiver must perform subband prediction on the data of other subbands based on the recovered target subband signal and signal classification identifier to obtain a complete sound signal.

[0122] The embodiment of the present application adopts a deep learning method to perform subband prediction. In one possible implementation, when the received audio frame is a voiced frame, step 706 specifically includes the following steps 706a to 706c (not shown in the figure):

[0123] Step 706a: When the signal type identifier is a voiced sound signal identifier, feature extraction is performed on the target subband signal to obtain a first signal feature, where the first signal feature includes at least one of a logarithmic power spectrum, a gene period, and a cross-correlation value.

[0124] For voiced frames, the audio receiver predicts the high-frequency subband signal from the decoded low-frequency subband signal. First, relevant features of the low-frequency signal are extracted as input to the deep learning network, such as the logarithmic power spectrum, pitch period, and cross-correlation value.

[0125] Step 706b: input the first signal feature into the first deep learning network to obtain the high-frequency sub-band power spectrum output by the first deep learning network.

[0126] The first deep learning network is trained based on signal characteristics of a sample low-frequency signal and a power spectrum of a sample high-frequency signal. The sample low-frequency signal and the sample high-frequency signal belong to different subband signals of the same sound signal.

[0127] In one possible implementation, during the model training phase, a computer device performs subband decomposition on a sample sound signal to obtain a sample low-frequency signal and a sample high-frequency signal. The computer device then inputs the signal features of the sample low-frequency signal into a first deep learning network to obtain a high-frequency subband power spectrum predicted by the first deep learning network. Based on the power spectrum of the sample high-frequency signal and the prediction results of the first deep learning network, the computer device performs backpropagation training on the first deep learning network.

[0128] The first deep learning network can be a combination of multi-layer convolutional neural networks (CNN) and multi-layer long short-term memory networks (LSTM).

[0129] Step 706c: Perform inverse Fourier transform based on the high-frequency sub-band power spectrum and the random phase value to obtain a high-frequency sub-band signal.

[0130] The high-frequency power spectrum value predicted by the first deep learning network is combined with the random phase value and subjected to inverse Fourier transform to obtain the time domain high-frequency sub-band signal.

[0131] In a possible implementation, when the received audio frame is an unvoiced frame, step 706 specifically includes the following steps 706d to 706f (not shown in the figure):

[0132] Step 706d: When the signal type identifier is a non-voiced signal identifier, feature extraction is performed on the target sub-band signal to obtain a second signal feature, where the second signal feature includes a logarithmic power spectrum.

[0133] For unvoiced frames, the audio receiver predicts the low-frequency subband signal from the decoded high-frequency subband signal. First, relevant features of the high-frequency signal are extracted as input to the deep learning network, such as the logarithmic power spectrum.

[0134] Step 706e: Input the second signal feature into the second deep learning network to obtain the low-frequency sub-band power spectrum output by the second deep learning network.

[0135] The second deep learning network is trained based on the signal characteristics of the sample high-frequency signal and the power spectrum of the sample low-frequency signal. The sample low-frequency signal and the sample high-frequency signal belong to different sub-band signals of the same sound signal.

[0136] In one possible implementation, during the model training phase, the computer device performs subband decomposition on the sample sound signal to obtain a sample low-frequency signal and a sample high-frequency signal. The computer device then inputs the signal features of the sample high-frequency signal into a second deep learning network to obtain a low-frequency subband power spectrum predicted by the second deep learning network. Based on the power spectrum of the sample low-frequency signal and the prediction results of the second deep learning network, the computer device performs backpropagation training on the second deep learning network.

[0137] The second deep learning network can be a combination of multi-layer CNN and multi-layer LSTM.

[0138] Step 706f: Perform inverse Fourier transform based on the low-frequency sub-band power spectrum and the random phase value to obtain a low-frequency sub-band signal.

[0139] The low-frequency signal power spectrum value predicted by the second deep learning network, combined with the random phase value, and then subjected to inverse Fourier transform, can obtain the time domain low-frequency sub-band signal.

[0140] Step 707: Perform sub-band synthesis based on the sub-band signals to obtain an output signal.

[0141] After performing sub-band prediction and recovery, the audio receiver obtains complete sub-band signals for all sub-bands. Sub-band synthesis, such as the QMF sub-band synthesis method, is then performed to synthesize multiple groups of sub-band signals into a complete sub-band signal for output.

[0142] The above steps 703 to 707 are the process in which the audio receiving end performs error correction decoding and sub-band prediction to obtain a complete sound signal when the sub-band coded data is lost. In one possible implementation, the following steps are also included after step 702 (not shown in the figure):

[0143] In the case that there is no packet loss of the sub-band coded data, sub-band decoding and sub-band synthesis are performed on the sub-band coded data to obtain an output signal.

[0144] If there is no packet loss, the audio receiver can directly compress and decode each group of sub-band coded data to obtain the sub-band signal. Then, through the inverse Fourier transform and sub-band synthesis process, the output signal is obtained.

[0145] In an embodiment of the present application, in the event of packet loss, the audio receiving end can recover the complete target sub-band coded data based on redundant data, and then predict other sub-band signals through a deep learning network, thereby further improving the audio transmission network's ability to resist packet loss.

[0146] like Figure 8As shown, it shows the process of the audio sending end collecting and sending audio and the audio receiving end receiving and outputting audio. The audio transmitter encodes the input signal by first performing sub-band decomposition and sub-band coding on the input signal and determining the input signal type, which includes voiced and unvoiced signals. For voiced signals, the audio transmitter extracts the low-frequency sub-band coded bitstream and signal type identifier for error correction coding. For unvoiced signals, the audio transmitter extracts the high-frequency sub-band coded bitstream and signal classification identifier for error correction coding. Finally, the sub-band coded data and error-correction coded redundant data are bundled and transmitted to the audio receiver. The audio receiver decodes the received signal by first receiving and buffering the data. If packet loss has not occurred, the sub-band decoding process is performed. The decoded sub-band signals are synthesized to obtain the complete output signal. If packet loss has occurred, the relevant redundant data and adjacent data packets are retrieved from the data buffer for error correction decoding. The sub-band coded data and signal type identifier at the location of the packet loss are obtained through error correction decoding. Based on the sub-band bitstream obtained through error correction decoding, the remaining sub-bands are predicted and recovered to obtain all sub-band signals. The complete output signal is then synthesized through sub-band synthesis.

[0147] Figure 9 This is a structural block diagram of an audio transmission device provided by an exemplary embodiment of the present application, which includes the following structure:

[0148] The sub-band coding module 901 is configured to perform sub-band decomposition and compression coding on the input signal to obtain sub-band coded data of at least two groups of signal sub-bands, where different signal sub-bands correspond to different audio frequency bands;

[0149] A determination module 902 is configured to determine target sub-band coded data from the sub-band coded data based on energy distribution of the input signal, wherein the audio frequency band corresponding to the target sub-band coded data is a frequency band where signal energy is concentrated;

[0150] an error correction coding module 903, configured to perform error correction coding on the target sub-band coded data to obtain redundant data;

[0151] The data sending module 904 is configured to send an audio data packet including the sub-band coded data and the redundant data to an audio receiving end, and the audio receiving end is configured to recover the sub-band coded data based on the redundant data in case of packet loss.

[0152] Optionally, the determining module 902 is further configured to:

[0153] determining a low-frequency energy proportion of a low-frequency sub-band based on sample signals of the input signal in each audio frequency band, wherein the audio frequency corresponding to the low-frequency sub-band is lower than the audio frequencies corresponding to other signal sub-bands;

[0154] The target sub-band coded data is determined from the sub-band coded data based on the low-frequency energy proportion.

[0155] Optionally, the determining module 902 is further configured to:

[0156] When the low-frequency energy proportion is higher than a threshold, determining the sub-band coded data corresponding to the low-frequency sub-band as the target sub-band coded data;

[0157] When the low-frequency energy ratio is lower than the threshold, the sub-band coding data corresponding to the high-frequency sub-band is determined as the target sub-band coding data, and the audio frequency corresponding to the high-frequency sub-band is higher than the audio frequencies corresponding to other signal sub-bands.

[0158] Optionally, the device further includes:

[0159] an identifier generating module, configured to generate a signal type identifier based on the low-frequency energy ratio, the signal type identifier being used to indicate whether the input signal belongs to a voiced sound signal or an unvoiced sound signal, wherein the low-frequency energy ratio of the voiced sound signal is higher than the threshold, and the low-frequency energy ratio of the unvoiced sound signal is lower than the threshold;

[0160] The data sending module 904 is further configured to:

[0161] Packing the sub-band coded data, the redundant data, and the signal type identifier to generate the audio data packet;

[0162] The audio data packet is sent to the signal receiving end, and the audio receiving end is used to determine the target sub-band coded data based on the signal type identifier in the event of packet loss, and perform data recovery based on the target sub-band coded data and the redundant data.

[0163] Optionally, the subband coding module 901 is further configured to:

[0164] Perform analog-to-digital conversion on the analog sound signal collected by the microphone to generate a digital sound signal;

[0165] Performing Fourier transform on the digital sound signal to obtain a frequency domain signal;

[0166] The frequency domain signal is subjected to sub-band decomposition and compression coding to generate the sub-band coded data.

[0167] Optionally, the subband coding module 901 is further configured to:

[0168] Performing sub-band decomposition on the frequency domain signal through at least two band-pass filters corresponding to different frequency bands to obtain at least two sub-band signals, wherein the frequency bands of the band-pass filters are continuous;

[0169] The sub-band signal is subjected to frequency shifting and quantization encoding to obtain the sub-band encoded data.

[0170] Figure 10 : is a structural block diagram of an audio transmission device provided by another exemplary embodiment of the present application, the device includes the following structure:

[0171] A data receiving module 1001 is configured to receive an audio data packet, wherein the audio data packet includes redundant data and at least two sets of sub-band coded data. The redundant data is obtained by performing forward error correction coding on target sub-band coded data among the sub-band coded data by an audio transmitting end. Different sub-band coded data correspond to different audio frequency bands. The audio frequency band corresponding to the target sub-band coded data is a frequency band where signal energy is concentrated.

[0172] a packet loss detection module 1002, configured to perform packet loss detection on the sub-band coded data;

[0173] The decoding module 1003 is configured to recover the sub-band coded data based on the redundant data to obtain an output signal when the sub-band coded data is lost.

[0174] Optionally, the audio data packet further includes a signal type identifier, where the signal type identifier is used to indicate whether the sound signal corresponding to the sub-band coded data is a voiced sound signal or an unvoiced sound signal, wherein the target sub-band coded data for the voiced sound signal is the sub-band coded data for a low-frequency sub-band, and the target sub-band coded data for the unvoiced sound signal is the sub-band coded data for a high-frequency sub-band;

[0175] The decoding module 1003 is further configured to:

[0176] determining the target sub-band coded data from the sub-band coded data based on the signal type identifier;

[0177] performing error correction decoding on the target sub-band coded data based on the redundant data and the sub-band coded data in the adjacent audio data packets;

[0178] performing sub-band decoding on the target sub-band coded data after error correction decoding to obtain a target sub-band signal;

[0179] Performing data recovery on other sub-band coded data based on the target sub-band signal;

[0180] Sub-band synthesis is performed based on the respective sub-band signals to obtain the output signal.

[0181] Optionally, the decoding module 1003 is further configured to:

[0182] When the signal type identifier is a voiced sound signal identifier, performing feature extraction on the target subband signal to obtain a first signal feature, where the first signal feature includes at least one of a logarithmic power spectrum, a gene period, and a cross-correlation value;

[0183] Inputting the first signal feature into a first deep learning network to obtain a high-frequency sub-band power spectrum output by the first deep learning network, wherein the first deep learning network is trained based on the signal feature of a sample low-frequency signal and the power spectrum of a sample high-frequency signal, wherein the sample low-frequency signal and the sample high-frequency signal belong to different sub-band signals of the same sound signal;

[0184] An inverse Fourier transform is performed based on the high frequency sub-band power spectrum and the random phase value to obtain a high frequency sub-band signal.

[0185] Optionally, the decoding module 1003 is further configured to:

[0186] When the signal type identifier is a non-voiced signal identifier, performing feature extraction on the target subband signal to obtain a second signal feature, where the second signal feature includes a logarithmic power spectrum;

[0187] Inputting the second signal feature into a second deep learning network to obtain a low-frequency sub-band power spectrum output by the second deep learning network, wherein the second deep learning network is trained based on the signal feature of the sample high-frequency signal and the power spectrum of the sample low-frequency signal, wherein the sample low-frequency signal and the sample high-frequency signal belong to different sub-band signals of the same sound signal;

[0188] An inverse Fourier transform is performed based on the low-frequency sub-band power spectrum and the random phase value to obtain a low-frequency sub-band signal.

[0189] Optionally, the decoding module 1003 is further configured to:

[0190] In the case that there is no packet loss of the sub-band coded data, sub-band decoding and sub-band synthesis are performed on the sub-band coded data to obtain the output signal.

[0191] In summary, in this embodiment, the input signal is decomposed and compressed into frequency bands to generate at least two sets of sub-band coded data. Error correction coding is then performed on the sub-band coded data where signal energy is concentrated, ensuring the audio receiver's ability to recover the primary audio data. Compared to solutions that directly perform error correction coding on the complete input signal, this approach improves audio transmission quality while reducing the amount of redundant data, thereby reducing the transmission bandwidth and operating costs of error correction coding.

[0192] Please refer to Figure 11, which shows a block diagram of the structure of a terminal 1100 provided by an exemplary embodiment of the present application. Terminal 1100 may be a portable mobile terminal, such as a smartphone, a tablet computer, a Moving Picture Experts Group Audio Layer III (MP3) player, or a Moving Picture Experts Group Audio Layer IV (MP4) player. Terminal 1100 may also be referred to as user equipment, a portable terminal, or other names.

[0193] Typically, the terminal 1100 includes a processor 1101 and a memory 1102 .

[0194] The processor 1101 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 1101 may be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), and programmable logic array (PLA). The processor 1101 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state, also known as a central processing unit (CPU); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 1101 may be integrated with a graphics processing unit (GPU), which is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 1101 may also include an artificial intelligence (AI) processor, which is used to process computing operations related to machine learning.

[0195] The memory 1102 may include one or more computer-readable storage media, which may be tangible and non-transitory. The memory 1102 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash memory storage devices. In some embodiments, the non-transitory computer-readable storage medium in the memory 1102 is used to store at least one instruction, which is used to be executed by the processor 1101 to implement the method provided in the embodiment of the present application.

[0196] In some embodiments, the terminal 1100 may optionally further include: a peripheral device interface 1103 .

[0197] The peripheral device interface 1103 can be used to connect at least one input / output (I / O)-related peripheral device to the processor 1101 and the memory 1102. In some embodiments, the processor 1101, the memory 1102, and the peripheral device interface 1103 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 1101, the memory 1102, and the peripheral device interface 1103 can be implemented on separate chips or circuit boards, which is not limited in this embodiment.

[0198] An embodiment of the present application further provides a computer-readable storage medium, which stores at least one instruction, and the at least one instruction is loaded and executed by a processor to implement the audio transmission method described in the above embodiments.

[0199] According to one aspect of the present application, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the audio transmission method provided in various optional implementations of the above aspects.

[0200] Those skilled in the art will appreciate that in one or more of the above examples, the functions described in the embodiments of the present application can be implemented using hardware, software, firmware, or any combination thereof. When implemented using software, these functions can be stored in a computer-readable storage medium or transmitted as one or more instructions or codes on a computer-readable storage medium. Computer-readable storage media include computer storage media and communication media, wherein communication media include any media that facilitates the transmission of computer programs from one place to another. The storage medium can be any available medium that can be accessed by a general-purpose or special-purpose computer.

[0201] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, storage, display, etc.), and signals involved in this application are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the input signals and audio data involved in this application are all obtained with full authorization.

[0202] The above description is merely an optional embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.

Claims

1. An audio transmission method, characterized in that: The method is performed by an audio sending end, and the method includes: Performing sub-band decomposition and compression coding on the input signal to obtain sub-band coded data of at least two groups of signal sub-bands, where different signal sub-bands correspond to different audio frequency bands; determining target sub-band coded data from the sub-band coded data based on energy distribution of the input signal, wherein the audio frequency band corresponding to the target sub-band coded data is a frequency band where signal energy is concentrated; performing error correction coding on the target sub-band coded data to obtain redundant data; An audio data packet including the sub-band coded data and the redundant data is sent to an audio receiving end, and the audio receiving end is configured to perform data recovery on the sub-band coded data based on the redundant data in the event of packet loss.

2. The method according to claim 1, characterized in that The determining target sub-band coded data from the sub-band coded data based on the energy distribution of the input signal includes: determining a low-frequency energy proportion of a low-frequency sub-band based on sample signals of the input signal in each audio frequency band, wherein the audio frequency corresponding to the low-frequency sub-band is lower than the audio frequencies corresponding to other signal sub-bands; The target sub-band coded data is determined from the sub-band coded data based on the low-frequency energy proportion.

3. The method according to claim 2, characterized in that The determining the target sub-band coded data from the sub-band coded data based on the low-frequency energy proportion includes: When the low-frequency energy proportion is higher than a threshold, determining the sub-band coded data corresponding to the low-frequency sub-band as the target sub-band coded data; When the low-frequency energy ratio is lower than the threshold, the sub-band coding data corresponding to the high-frequency sub-band is determined as the target sub-band coding data, and the audio frequency corresponding to the high-frequency sub-band is higher than the audio frequencies corresponding to other signal sub-bands.

4. The method according to claim 3, characterized in that After determining the low-frequency energy proportion of the low-frequency sub-band based on the sample signals of the input signal in each audio frequency band, the method includes: generating a signal type identifier based on the low-frequency energy ratio, the signal type identifier being used to indicate whether the input signal is a voiced sound signal or an unvoiced sound signal, wherein the low-frequency energy ratio of the voiced sound signal is higher than the threshold, and the low-frequency energy ratio of the unvoiced sound signal is lower than the threshold; The sending of the audio data packet including the sub-band coded data and the redundant data to the receiving end comprises: Packing the sub-band coded data, the redundant data, and the signal type identifier to generate the audio data packet; The audio data packet is sent to the signal receiving end, and the audio receiving end is used to determine the target sub-band coded data based on the signal type identifier in the event of packet loss, and perform data recovery based on the target sub-band coded data and the redundant data.

5. The method according to any one of claims 1 to 4, characterized in that: The step of performing sub-band decomposition and compression coding on the input signal to obtain sub-band coded data of at least two groups of signal sub-bands includes: Perform analog-to-digital conversion on the analog sound signal collected by the microphone to generate a digital sound signal; Performing Fourier transform on the digital sound signal to obtain a frequency domain signal; The frequency domain signal is subjected to sub-band decomposition and compression coding to generate the sub-band coded data.

6. The method according to claim 5, characterized in that The performing sub-band decomposition and compression coding on the frequency domain signal to generate the sub-band coded data includes: Performing sub-band decomposition on the frequency domain signal through at least two band-pass filters corresponding to different frequency bands to obtain at least two sub-band signals, wherein the frequency bands of the band-pass filters are continuous; The sub-band signal is subjected to frequency shifting and quantization encoding to obtain the sub-band encoded data.

7. An audio transmission method, characterized in that: The method is performed by an audio receiving end, and the method includes: receiving an audio data packet, the audio data packet comprising redundant data and at least two sets of sub-band coded data, the redundant data being obtained by an audio transmitter performing forward error correction coding on target sub-band coded data among the sub-band coded data, different sub-band coded data corresponding to different audio frequency bands, and the audio frequency band corresponding to the target sub-band coded data being a frequency band where signal energy is concentrated; performing packet loss detection on the sub-band coded data; In the event that the sub-band coded data is lost, data recovery is performed on the sub-band coded data based on the redundant data to obtain an output signal.

8. The method according to claim 7, characterized in that The audio data packet further includes a signal type identifier, wherein the signal type identifier is used to indicate whether the sound signal corresponding to the sub-band coded data is a voiced sound signal or an unvoiced sound signal, wherein the target sub-band coded data for the voiced sound signal is the sub-band coded data of the low-frequency sub-band, and the target sub-band coded data for the unvoiced sound signal is the sub-band coded data of the high-frequency sub-band; When the sub-band coded data is lost, performing data recovery on the sub-band coded data based on the redundant data to obtain an output signal includes: determining the target sub-band coded data from the sub-band coded data based on the signal type identifier; performing error correction decoding on the target sub-band coded data based on the redundant data and the sub-band coded data in the adjacent audio data packets; performing sub-band decoding on the target sub-band coded data after error correction decoding to obtain a target sub-band signal; Performing data recovery on other sub-band coded data based on the target sub-band signal; Sub-band synthesis is performed based on the respective sub-band signals to obtain the output signal.

9. The method according to claim 8, characterized in that The performing data recovery on other sub-band coded data based on the target sub-band signal includes: When the signal type identifier is a voiced sound signal identifier, performing feature extraction on the target subband signal to obtain a first signal feature, where the first signal feature includes at least one of a logarithmic power spectrum, a gene period, and a cross-correlation value; Inputting the first signal feature into a first deep learning network to obtain a high-frequency sub-band power spectrum output by the first deep learning network, wherein the first deep learning network is trained based on the signal feature of a sample low-frequency signal and the power spectrum of a sample high-frequency signal, wherein the sample low-frequency signal and the sample high-frequency signal belong to different sub-band signals of the same sound signal; An inverse Fourier transform is performed based on the high frequency sub-band power spectrum and the random phase value to obtain a high frequency sub-band signal.

10. The method according to claim 8, characterized in that The performing data recovery on other sub-band coded data based on the target sub-band signal includes: When the signal type identifier is a non-voiced signal identifier, performing feature extraction on the target subband signal to obtain a second signal feature, where the second signal feature includes a logarithmic power spectrum; Inputting the second signal feature into a second deep learning network to obtain a low-frequency sub-band power spectrum output by the second deep learning network, wherein the second deep learning network is trained based on the signal feature of the sample high-frequency signal and the power spectrum of the sample low-frequency signal, wherein the sample low-frequency signal and the sample high-frequency signal belong to different sub-band signals of the same sound signal; An inverse Fourier transform is performed based on the low-frequency sub-band power spectrum and the random phase value to obtain a low-frequency sub-band signal.

11. The method according to any one of claims 7 to 10, characterized in that: After performing packet loss detection on the sub-band coded data, the method further includes: In the case that there is no packet loss of the sub-band coded data, sub-band decoding and sub-band synthesis are performed on the sub-band coded data to obtain the output signal.

12. An audio transmission device, characterized in that: The device comprises: A sub-band coding module is used to perform sub-band decomposition and compression coding on the input signal to obtain sub-band coded data of at least two groups of signal sub-bands, where different signal sub-bands correspond to different audio frequency bands; a determination module, configured to determine target sub-band coded data from the sub-band coded data based on energy distribution of the input signal, wherein the audio frequency band corresponding to the target sub-band coded data is a frequency band where signal energy is concentrated; an error correction coding module, configured to perform error correction coding on the target sub-band coded data to obtain redundant data; The data sending module is used to send an audio data packet containing the sub-band coded data and the redundant data to an audio receiving end, and the audio receiving end is used to recover the sub-band coded data based on the redundant data in the event of packet loss.

13. An audio transmission device, characterized in that: The device comprises: a data receiving module, configured to receive an audio data packet, the audio data packet including redundant data and at least two sets of sub-band coded data, the redundant data being obtained by performing forward error correction coding on target sub-band coded data among the sub-band coded data by an audio transmitting end, different sub-band coded data corresponding to different audio frequency bands, and the audio frequency band corresponding to the target sub-band coded data being a frequency band where signal energy is concentrated; a packet loss detection module, configured to perform packet loss detection on the sub-band coded data; The decoding module is configured to recover the sub-band coded data based on the redundant data to obtain an output signal when the sub-band coded data is lost.

14. A terminal, characterized in that: The terminal includes a processor and a memory; the memory stores at least one program, and the at least one program is loaded and executed by the processor to implement the audio transmission method according to any one of claims 1 to 6 or the audio transmission method according to any one of claims 7 to 11.

15. A computer-readable storage medium, characterized in that The computer-readable storage medium stores at least one computer program, which is loaded and executed by a processor to implement the audio transmission method according to any one of claims 1 to 6 or the audio transmission method according to any one of claims 7 to 11.

16. A computer program product, characterized in that The computer program product includes computer instructions, which are stored in a computer-readable storage medium; the processor of the terminal reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the terminal executes the audio transmission method according to any one of claims 1 to 6 or the audio transmission method according to any one of claims 7 to 11.

Citation Information

Patent Citations

  • System and method for perform replacement to considered loss part of audio signal

    CN101136201A

  • Data transmission method, system and device, computer readable storage medium and equipment

    CN113936669A