Data transmission method and system for audio stream
By using adaptive bit allocation and dynamic bandwidth estimation, combined with time-domain and frequency-domain characteristics, the problem of high-frequency detail loss and low-frequency resource waste in existing audio streaming transmission is solved, achieving high-quality audio transmission and improving network resource utilization efficiency and system adaptability.
Patent Information
- Application Number
- CN202511453841.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-13
- Publication Date
- 2026-01-20
AI Technical Summary
Existing audio streaming technologies use a single-mode bit allocation based on a uniform masking threshold, ignoring frequency band energy differences. This results in the loss of high-frequency details and waste of low-frequency resources, making it difficult to adapt to mixed signal requirements and unable to meet high-quality transmission requirements.
By using adaptive bit allocation and dynamic bandwidth estimation, the bit allocation is dynamically adjusted according to the time-domain energy difference and frequency-domain characteristics of audio frames to optimize audio stream data transmission. Fourier transform is used to obtain the energy spectrum and frequency band division. By combining the time-domain masking degree and the frequency-domain exposure degree, the perceptual importance of the frequency band and the bit division weight are calculated to achieve adaptive bit allocation of the frequency band.
It significantly improves audio quality, avoids loss of high-frequency details and waste of low-frequency resources, improves network resource utilization efficiency, enhances the system's adaptability to different audio signals and network conditions, and ensures high-quality transmission under limited bandwidth.
Smart Images

Figure CN121366580A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of data processing. In particular, it relates to a data transmission method and system for audio stream. BACKGROUND
[0002] Audio stream refers to a technology that transmits audio signals in the form of continuous data stream. It divides audio data into a series of data packets and transmits them in real time over the network to achieve continuous playback of audio content. Audio stream is widely used in online music, video conferencing, network broadcasting and other fields. Its core is to efficiently transmit high-quality audio signals under limited network bandwidth, while ensuring smoothness and real-time performance of playback.
[0003] In the process of audio stream data transmission, the existing technology generally adopts a method based on a unified masking threshold, relying on a single audio mode of music or speech for bit allocation, ignoring the energy differences between different frequency bands, resulting in loss of high-frequency details due to insufficient bits and waste of resources in the low-frequency part due to excessive bit allocation.
[0004] In addition, this single-mode bit allocation method is difficult to adapt to the complexity and diversity of audio signals, especially when dealing with audio streams containing mixed signals of music and speech, the limitations of existing technology are more prominent, and it cannot meet the demand of high-quality audio transmission. SUMMARY
[0005] To solve the problem that the existing audio stream transmission technology adopts a single-mode bit allocation based on a unified masking threshold, ignores the energy differences between different frequency bands, results in loss of high-frequency details, waste of low-frequency resources, is difficult to adapt to mixed signal requirements, and cannot meet the requirements of high-quality transmission, the present application provides solutions in the following aspects.
[0006] In a first aspect, a data transmission method for audio stream includes: obtaining an analog audio signal of an audio input device, converting the analog audio signal into a digital signal to obtain PCM data, dividing the PCM data into a plurality of audio frames with equal time length according to the collection time of the data, wherein each audio frame contains audio data within a period of time; determining the time domain masking degree of each audio frame and the corresponding bit rate according to the difference in time domain average energy between the audio frames; performing Fourier transform on each audio frame to obtain the energy spectrum and frequency band division of each audio frame, taking the ratio between the energy of a frequency and the hearing threshold of the frequency as the frequency domain exposure degree, analyzing the perceptual importance of each frequency band according to the frequency domain exposure degree and the time domain masking degree, and calculating the bit division weight of the frequency band according to the energy difference and the perceptual importance between the frequency bands; performing adaptive allocation of bits in the frequency bands within the audio frame according to the bit division weight of each frequency band and the target bit rate corresponding to each audio frame, sequentially performing quantization and entropy coding processing on the audio data after the bit allocation is completed, obtaining compressed audio data frames, and performing protocol packaging to generate encoded data frames suitable for network transmission, thereby completing the data transmission of the audio stream.
[0007] Through adaptive bit allocation and dynamic bandwidth estimation, the efficiency and quality of audio stream data transmission are optimized, the bit allocation is dynamically adjusted according to the time domain energy difference and frequency domain characteristics of the audio frame, the problems of high frequency detail loss and low frequency resource waste in the prior art are avoided, and the audio quality is significantly improved. At the same time, the utilization efficiency of network resources is improved, the adaptability of the system to different audio signals and network conditions is enhanced, and high-quality audio transmission is ensured under limited bandwidth.
[0008] Preferably, the step of obtaining the time domain average energy comprises: Taking any audio frame as a target audio frame, the sum of squares of all time point data in the target audio frame is calculated and divided by the number of time points to obtain the time domain average energy of the target audio frame.
[0009] Preferably, the step of calculating the time domain masking degree of each audio frame comprises: Taking any audio frame as a target audio frame, the difference between the time domain average energy of the target audio frame and the time domain average energy of the forward audio frame of the target audio frame is calculated, and normalized to obtain a forward energy difference value; The difference between the time domain average energy of the target audio frame and the time domain average energy of the backward audio frame of the target audio frame is calculated, and normalized to obtain a backward energy difference value; The natural exponential function is taken for the forward energy difference value and the backward energy difference value respectively, and the sum is obtained to obtain the time domain masking degree of the target audio frame.
[0010] Preferably, the step of calculating the bit rate comprises: taking any audio frame as a target audio frame, determining an estimated bandwidth of the current transmission network based on the dynamic bandwidth estimation method, taking the estimated bandwidth as a maximum transmission bit rate, and calculating a target bit rate of the target audio frame according to the maximum transmission bit rate of the current transmission network and the normalized time-domain masking degree of the audio frame. The target bit rate of the target audio frame is calculated by multiplying the maximum transmission bit rate of the current transmission network by 1 minus the difference between the normalized time-domain masking degree.
[0011] Preferably, the step of calculating the perceptual importance of each frequency band comprises: taking any audio frame as a target audio frame, taking any frequency band in the target audio frame as a target frequency band, and taking any frequency in the target frequency band as a target frequency, calculating a ratio between the time-domain average energy of the target audio frame and the hearing threshold of the target frequency in the target frequency band to obtain an energy proportion of the target audio frame, and taking the product between the energy proportion and the normalized time-domain masking degree of the target audio frame as the time-domain masking degree of the target frequency. The average of the sum of the time-domain masking degree and the frequency-domain revealing degree is taken as the perceptual importance of the target frequency band.
[0012] Preferably, the step of calculating the bit allocation weight of each frequency band comprises: taking any frequency band in the target audio frame as a target frequency band, calculating the average energy of all frequencies in the target frequency band, calculating the ratio between the absolute difference between the average energy and the maximum average energy of all frequency bands in the target audio frame and the absolute difference between the average energy and the minimum average energy of all frequency bands in the target audio frame to obtain an energy difference, and taking the product between the mapping result obtained by nonlinearly mapping the energy difference by using a negative exponential function and the perceptual importance of the target frequency band as the bit allocation weight of the target frequency band.
[0013] Preferably, the step of adaptively allocating bits to the frequency bands in the audio frame comprises: taking any frequency band in the target audio frame as a target frequency band, and taking the product between the target frequency band and the corresponding bit allocation weight as the bit size of the target frequency band.
[0014] In a second aspect, a data transmission system for an audio stream comprises a processor and a memory, and the memory stores computer program instructions which, when executed by the processor, implement the above-mentioned data transmission method for an audio stream.
[0015] The present application has the following effects: 1、The present application can dynamically adjust the bit allocation strategy according to the actual characteristics of the audio signal through adaptive bit allocation. Not only the energy difference between frequency bands is considered, but also the time domain masking effect and the frequency domain exposure degree are combined, thereby effectively avoiding the problems of high frequency detail loss and low frequency resource waste in the prior art, and significantly improving the sound quality of audio transmission.
[0016] 2、The present application realizes bit adaptive allocation of different frequency bands through fine bit division weight calculation, can allocate bits according to the actual importance of the frequency band, avoids the waste phenomenon of bit allocation, improves the utilization efficiency of network resources, and is especially suitable for limited bandwidth network environment. BRIEF DESCRIPTION OF DRAWINGS
[0017] Figure 1 is a method flowchart of steps S1-S4 in a data transmission method for audio stream.
[0018] Figure 2 is a structural block diagram of a data transmission system for audio stream. DETAILED DESCRIPTION
[0019] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments of the present application.
[0020] Referring to Figure 1 A data transmission method for audio stream includes steps S1-S4, specifically as follows: S1: Obtain the analog audio signal of the audio input device, convert the analog audio signal into a digital signal to obtain PCM data, and divide the PCM data into a plurality of audio frames with equal time length according to the data collection time, wherein each audio frame contains audio data in a time period.
[0021] In an embodiment of the present application, the audio input device is used to collect the analog audio signal of the audio output object to generate the analog audio signal. The analog audio signal is converted into a digital signal by an analog-to-digital converter (ADC) to obtain PCM (Pulse Code Modulation) data. Further, the PCM data is divided into a plurality of audio frames with equal time length according to the data collection time. The sound data contained in each audio frame is a sampling set of the audio signal in a specific time interval, rather than only the instantaneous value from a single time point. In other words, the audio frame is a discrete representation of the audio signal in a time window, reflecting the continuous change of the audio signal in the time window, rather than the audio state at a certain moment.
[0022] The time length of an audio frame can be calculated according to a sampling rate and a number of samples. For example, the time length of an audio frame can be 20 ms, and the amount of data contained in the audio frame is 20 ms of PCM data.
[0023] In the present application, the term "analog signal" does not refer to a fictitious or experimentally constructed signal, but is a term of art in the field of audio signal processing, referring to a continuously varying signal. The term "audio frame" is also a term of art in the prior art, and is also referred to as a "sample frame" in some literature, and is used to describe a set of audio data collected within a certain time interval.
[0024] Further analysis shows that, considering that an audio stream usually contains a mixed signal of music and speech, the distribution of such a mixed signal in the time-frequency domain is significantly different from that of a single signal. Existing audio encoders mainly rely on a psychoacoustic model to analyze the importance of each frequency band and allocate bits based on a fixed time-frequency domain masking threshold according to a single audio mode of music or speech. However, this method ignores the differences between different frequency bands, resulting in loss of audio data or waste of resources.
[0025] For example, in a music scenario, the energy of high-frequency overtones can be lower than the masking threshold, so fewer bits are allocated, resulting in loss of high-frequency details and fuzzy sound quality; and in a speech scenario, too many bits are allocated to the low-frequency part, causing resource waste and reducing transmission efficiency. In order to solve these problems, the specific operation steps are as follows: S2: determining the time-domain masking degree of each audio frame and the corresponding bit rate according to the difference in time-domain average energy between the audio frames.
[0026] The step of obtaining the time-domain average energy includes: Taking any audio frame as a target audio frame, the sum of squares of data at all time points in the target audio frame is calculated and divided by the number of time points to obtain the time-domain average energy of the target audio frame.
[0027] It should be noted that a target audio frame has a forward audio frame and a backward audio frame, but when the first audio frame is selected as the target audio frame, it does not have a forward audio frame; and when the last audio frame is selected as the target audio frame, it does not have a backward audio frame.
[0028] The step of calculating the masking degree includes: Taking any audio frame as a target audio frame, the difference between the time-domain average energy of the target audio frame and the time-domain average energy of the forward audio frame of the target audio frame is calculated, and normalized to obtain a forward energy difference value; The difference between the time-domain average energy of the target audio frame and the time-domain average energy of the backward audio frame of the target audio frame is calculated, and normalized to obtain a backward energy difference value; The natural exponential function is taken for the forward energy difference and the backward energy difference respectively, and the sum is obtained as the time-domain masking degree of the target audio frame.
[0029] Specifically, the time-domain masking degree satisfies the following relationship: ; In the formula, the time-domain masking degree of the target audio frame, the time-domain average energy of the forward audio frame of the target audio frame, the time-domain average energy of the backward audio frame of the target audio frame, the time-domain average energy of the target audio frame, the exponential function with a natural number as the base.
[0030] That is, the forward masking degree of the target audio frame, the backward masking degree of the target audio frame, and the time-domain masking degree is normalized to obtain .
[0031] It should be noted that in human ear hearing perception, when a strong sound signal appears near a weak sound signal, the weak sound signal will be masked and become inaudible, which is called masking effect. Therefore, when the energy of a target audio frame is significantly lower than the energy of its forward audio frame and backward audio frame, it indicates that the target audio frame is more likely to be masked.
[0032] The steps of the bit rate calculation method include: Taking any audio frame as a target audio frame, the estimated bandwidth of the current transmission network is determined based on the dynamic bandwidth estimation method, the estimated bandwidth is taken as the maximum transmission bit rate, and the target bit rate of the target audio frame is calculated according to the maximum transmission bit rate of the current transmission network and the normalized time-domain masking degree of the audio frame; Wherein, the calculation method of the bit rate of the target audio frame includes: taking the product of the maximum transmission bit rate of the current transmission network and 1 minus the difference value of the normalized time-domain masking degree, to obtain the target bit rate of the target audio frame.
[0033] For example, in addition to using the GCC (Google Congestion Control) algorithm to estimate the bandwidth of the current transmission network, the bandwidth can also be estimated by measuring the RTT (Round-Trip Time) of the data packet. The longer the RTT, the higher the network delay, which may be due to network congestion. By analyzing the changes of RTT, the available bandwidth of the network can be inferred.
[0034] Specifically, the target bit rate satisfies the following relationship: ; In the formula, denotes the target bit rate of the target audio frame, denotes the maximum transmission bit rate, denotes the normalized time-domain masking degree of the target audio frame.
[0035] That is, the smaller the time-domain masking degree of an audio frame, the greater the energy of the audio frame, the more key sound data it contains, and the larger the bit rate should be to ensure the transmission quality of the audio; otherwise, the greater the time-domain masking degree, the smaller the probability of capturing the sound by the human ear, and the smaller the bit rate should be to avoid resource waste.
[0036] Bandwidth refers to the maximum amount of data that can be transmitted per unit time in network transmission, rather than the frequency bandwidth concept in filters. Bit rate refers to the amount of data transmitted per unit time by an audio encoder. Both are measured in bits per second (bit / s) and satisfy the relationship that the bit rate is less than or equal to the bandwidth, which is known to those skilled in the art and will not be described in detail.
[0037] S3: Fourier transform is performed on each audio frame to obtain the energy spectrum and frequency band division of each audio frame, the ratio between the energy of a frequency and the hearing threshold of the frequency is taken as the frequency domain exposure degree, the perceptual importance of each frequency band is analyzed according to the frequency domain exposure degree and the time-domain masking degree, and the bit division weight of the frequency band is calculated according to the energy difference and the perceptual importance between the frequency bands.
[0038] It should be noted that Fourier transform is performed on each audio frame to obtain the energy spectrum of the audio frame, and MDCT (Modified Discrete Cosine Transform) technology is used to divide each audio frame into a plurality of non-uniform frequency bands with different lengths. Among them, the non-uniform frequency band means that the width of each frequency band is not consistent, and the frequency band means a certain range of frequency interval.
[0039] The calculation steps of the perceptual importance of each frequency band include: Taking any audio frame as a target audio frame, taking any frequency band in the target audio frame as a target frequency band, and taking any frequency in the target frequency band as a target frequency, the ratio between the time-domain average energy of the target audio frame and the hearing threshold of the target frequency in the target frequency band is calculated to obtain the energy proportion of the target audio frame. The product of the energy proportion and the normalized time-domain masking degree of the target audio frame is taken as the time-domain masking degree of the target frequency; The average value of the sum of the time-domain masking degree and the frequency domain exposure degree is taken as the perceptual importance of the target frequency band.
[0040] Further, an audio frame represents a segment in time domain, containing audio data in a time period. A frequency band represents a range in frequency domain, containing signal components in a frequency interval. A frequency represents a specific component in frequency domain, representing a specific audio signal component.
[0041] Specifically, the perceptual importance satisfies the following polynomial: ; ; In the formula, represents the perceptual importance of the target frequency band, represents the number of frequencies contained in the target frequency band, represents the frequency domain exposure of the th frequency, represents the energy of the th frequency in the target frequency band in the energy spectrum, represents the hearing threshold of the th frequency in the target frequency band, represents the time domain average energy of the target audio frame, represents the time domain masking degree of the normalized target audio frame.
[0042] That is, represents the frequency domain exposure of the th frequency, represents the time domain masking influence degree of the th frequency.
[0043] Further, in audio signal processing, theoretically, the human ear cannot perceive the sound of a frequency when the energy of the frequency is lower than the hearing threshold. However, the hearing threshold is measured in an ideal environment, and in actual situations, even if the energy of the frequency is lower than the hearing threshold, it is still possible to be captured by the human ear. Therefore, the frequency domain exposure of the frequency is determined by calculating the ratio of the energy of the frequency to the hearing threshold. The smaller the ratio, the lower the possibility that the sound of the frequency is perceived by the human ear, and the smaller the perceptual importance in the audio signal. In addition, since the energies of different frequencies are different, the masking influences they receive in the time domain are also different. Specifically, the greater the energy of the frequency, the smaller the masking influence it receives, and accordingly, the greater its perceptual importance.
[0044] By carefully considering the differences in the frequency domain exposure of different frequencies in the same frequency band and the degree of influence of time domain masking on each of them, the masking details in the frequency band are accurately refined, thereby significantly improving the accuracy of the perceptual importance calculation. In contrast, the prior art only evaluates the perceptual importance according to the average energy of the frequency band and the overall time domain masking degree, assuming that all frequencies in the same frequency band are uniformly affected by masking, ignoring the differences in the energies of the frequencies in the frequency band.
[0045] The calculating step of the bit allocation weight comprises: Taking any frequency band in the target audio frame as a target frequency band, calculating the average energy of all frequencies in the target frequency band, calculating the ratio between the absolute difference between the average energy and the maximum average energy of all frequency bands in the target audio frame and the absolute difference between the average energy and the minimum average energy, obtaining the energy difference, performing nonlinear mapping on the energy difference by using a negative exponential function, and taking the product between the mapping result and the perceptual importance of the target frequency band as the bit allocation weight of the target frequency band.
[0046] Specifically, the bit allocation weight satisfies the following relationship: ; In the formula, denotes the bit allocation weight of the target frequency band, denotes the perceptual importance of the target frequency band, denotes the average energy of all frequencies in the target frequency band, denotes the maximum average energy value of all frequency bands in the target audio frame, denotes the minimum average energy value of all frequency bands in the target audio frame, denotes the exponential function with a natural number as the base.
[0047] That is, denotes the frequency band importance of the target frequency band. The higher the perceptual importance of the frequency band, the richer the information amount and the degree of details contained in the frequency band, and therefore more bits need to be allocated to ensure that the key information in the audio signal is accurately retained. At the same time, when the difference between the average energy value of a frequency band and the maximum average energy value of all frequency bands in the target audio frame is smaller, and the difference between the average energy value and the minimum average energy value is larger, it indicates that the frequency band belongs to the main component in the audio signal, and more bits should be allocated to ensure high-quality presentation in the process of audio encoding and transmission. The bit allocation weight is normalized to obtain the bit allocation weight of each frequency band as the target frequency band after normalization.
[0048] Further analysis, the energy of the audio stream data at different frequencies is significantly different, resulting in extremely uneven distribution of signal energy in the same type of frequency band. The prior art usually simply divides the frequency band into two types of high and low frequencies, and adopts a distribution strategy of equally allocating bits to the two types of frequency bands. This strategy ignores the energy difference of part of the frequency bands in the same type of frequency band, resulting in that part of the frequency bands with small energy and being masked are still allocated with too many bits.
[0049] For example, in the frequency band range of 0-4kHz, 50 frequency bands are contained and 75% of bits are allocated; while in the frequency band range of 4-8kHz, 32 frequency bands are contained and only 25% of bits are allocated. This allocation is obviously unreasonable, causing waste of resources. In order to solve this problem, the specific steps are as follows: S4: According to the bit division weight of each frequency band and the target bit rate corresponding to each audio frame, the bits of the frequency bands in the audio frame are adaptively allocated, the audio data after bit allocation is sequentially quantized and entropy coded, the compressed audio data frame is obtained, and protocol encapsulation is performed to generate the encoded data frame suitable for network transmission, and the data transmission of the audio stream is completed.
[0050] Taking any frequency band in the target audio frame as a target frequency band, the product of the target frequency band and the corresponding bit division weight is taken as the bit size of the target frequency band.
[0051] Finally, the quantization and entropy coding operation is performed on the audio data after bit allocation, so as to generate the compressed audio data frame, thereby realizing efficient coding and transmission.
[0052] The application also provides a data transmission system for an audio stream. As shown in the figure, Figure 2 The system includes a processor and a memory, and the memory stores computer program instructions which realize the data transmission method for an audio stream according to the first aspect of the application when executed by the processor. The system also includes a communication bus and a communication interface and other components familiar to those skilled in the art, the settings and functions of which are known in the art, so they will not be described here.
[0053] It should be pointed out that for those skilled in the art, without departing from the concept of the present application, a number of modifications and improvements can be made, which all belong to the protection scope of the present application. Therefore, the protection scope of the present application patent should be subject to the appended claims.
Claims
1. A data transmission method for audio streams, characterized in that, include: The analog audio signal from the audio input device is acquired, converted into a digital signal to obtain PCM data, and divided into multiple audio frames of equal length according to the data acquisition time. Each audio frame contains audio data within a certain time period. Based on the difference in average energy in the time domain between audio frames, the degree of time domain masking of each audio frame and the corresponding bit rate are determined. Fourier transform is performed on each audio frame to obtain the energy spectrum and frequency band division of each audio frame. The ratio between the energy of the frequency and the hearing threshold of the frequency is used as the frequency domain exposure. Based on the frequency domain exposure and the time domain masking, the perceptual importance of each frequency band is analyzed. Based on the energy difference and perceptual importance between frequency bands, the bit division weight of the frequency band is calculated. Based on the bit division weights of each frequency band and the target bit rate corresponding to each audio frame, the frequency bands within the audio frame are adaptively allocated. The audio data after bit allocation is sequentially quantized and entropy encoded to obtain compressed audio data frames. Protocol encapsulation is then performed to generate encoded data frames suitable for network transmission, thus completing the data transmission of the audio stream.
2. The data transmission method for audio streams according to claim 1, characterized in that, The steps for obtaining the time-domain average energy include: Using any audio frame as the target audio frame, calculate the sum of squares of all time points within the target audio frame and divide by the number of time points to obtain the temporal average energy of the target audio frame.
3. The data transmission method for audio streams according to claim 1, characterized in that, The steps for calculating the temporal masking degree of each audio frame include: Taking any audio frame as the target audio frame, calculate the difference between the temporal average energy of the target audio frame and the temporal average energy of the preceding audio frames of the target audio frame, and perform normalization processing to obtain the forward energy difference. Calculate the difference between the temporal average energy of the target audio frame and the temporal average energy of the backward audio frame, and perform normalization to obtain the backward energy difference. The natural exponential function is applied to the forward energy difference and the backward energy difference respectively, and the summation is performed to obtain the degree of temporal masking of the target audio frame.
4. The data transmission method for audio streams according to claim 1, characterized in that, The steps for calculating the bit rate include: Using any audio frame as the target audio frame, the estimated bandwidth of the current transmission network is determined based on the dynamic bandwidth estimation method. The estimated bandwidth is used as the maximum transmission bit rate. Based on the maximum transmission bit rate of the current transmission network and the normalized temporal masking degree of the audio frame, the target bit rate of the target audio frame is calculated. The calculation method for the bit rate of the target audio frame includes: multiplying the maximum transmission bit rate of the current transmission network by 1 and then subtracting the difference in the normalized time-domain masking degree to obtain the target bit rate of the target audio frame.
5. A data transmission method for audio streams according to claim 1, characterized in that, The calculation steps for the perceived importance of each frequency band include: Using any audio frame as the target audio frame, any frequency band in the target audio frame as the target frequency band, and any frequency in the target frequency band as the target frequency, calculate the ratio between the time-domain average energy of the target audio frame and the hearing threshold of the target frequency in the target frequency band to obtain the energy ratio of the target audio frame. The product of the energy ratio and the normalized time-domain masking degree of the target audio frame is taken as the time-domain masking degree of the target frequency. The average value of the sum of the product of the degree of time-domain masking and the degree of frequency-domain exposure is taken as the perceived importance of the target frequency band.
6. A data transmission method for audio streams according to claim 1, characterized in that, The calculation steps for the bit division weight of the frequency band include: Taking any frequency band in the target audio frame as the target frequency band, calculate the average energy of all frequencies within the target frequency band. Calculate the ratio between the absolute difference between the average energy and the maximum average energy of all frequency bands in the target audio frame, and between the absolute difference between the maximum and minimum average energy, to obtain the energy difference. Use a negative exponential function to perform a nonlinear mapping on the energy difference. The product of the mapping result and the perceived importance of the target frequency band is used as the target frequency band bit partitioning weight.
7. A data transmission method for audio streams according to claim 1, characterized in that, The step of adaptively allocating bits within the frequency band of the audio frame includes: The target frequency band is defined as any frequency band in the target audio frame, and the product of the target frequency band and the corresponding bit division weight is used as the bit size of the target frequency band.
8. A data transmission system for audio streams, characterized in that, include: A processor and a memory, the memory storing computer program instructions that, when executed by the processor, implement the data transmission method for an audio stream according to any one of claims 1-7.