Audio data real-time compression transmission method, device, equipment, medium and product

By extracting and encoding data on the transmitter side, decoding and interpolation on the receiver side, the problems of delay and sound quality in existing audio data transmission are solved, and the audio data transmission effect with low latency and high sound quality is achieved.

CN120452458APending Publication Date: 2025-08-08SHENZHEN JIAYZ PHOTO IND LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510687770.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

In the existing audio data transmission methods, the buffer causes large delays, long-term transmission increases the risk of data overflow and loss, large lossless compression delay, low lossy compressed audio quality, and cannot take into account both low latency and high sound quality requirements.

Method used

Data extraction, linear prediction and encoding are performed on the transmitter side to reduce the amount of data transmitted in real time; decoding, prediction recovery and data interpolation are performed on the receiver side to realize low-latency and high-sound quality audio data transmission.

Benefits of technology

It realizes low-latency data transmission while maintaining high sound quality, improving the performance of the audio data transmission system, avoiding the delay risk brought by the buffer and loss of audio quality loss due to lossy compression.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120452458A_ABST
    Figure CN120452458A_ABST
Patent Text Reader

Abstract

The invention discloses an audio data real-time compression transmission method and device, equipment, a medium and a product, and relates to the technical field of audio data processing, the method comprises the following steps: determining target transmission information, the target transmission information being audio data obtained by processing to-be-transmitted audio data by using a transmitter; the control information transmission channel transmits the target transmission information to the receiver; the receiver is controlled to solve the target transmission information to obtain audio data to be played, and the audio data to be played is audio data obtained after the target transmission information is processed. According to the method, data extraction, linear prediction and coding are carried out on the transmitter side, the real-time transmission data volume is reduced, low-delay data transmission is achieved, decoding, prediction recovery and data interpolation are carried out on the receiver side, the high-tone-quality data transmission effect is achieved, the low-delay and high-tone-quality requirements of the data transmission process are considered, and the data transmission efficiency is improved. And the performance of the audio data transmission system is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of audio data processing, and in particular to a transmission method, device, equipment, medium and product for real-time compression of audio data. Background Art

[0002] The digital audio transmission system can transmit sound signals to a player for playback. For example, the sound signal is sampled and quantized by an audio analog-to-digital converter (ADC), processed by a transmitter processor, output by a channel digital-to-analog converter (DAC), transmitted through a channel, sampled and quantized by a channel ADC, processed by a receiver processor, and output by an audio DAC before being transmitted to the player for playback.

[0003] As audio devices become more widely and intensively used, the sampling rate and quantization depth of audio ADCs are increasing, leading to an ever-increasing amount of data required to be transmitted. Consequently, the method for transmitting audio data has become a key concern. Without changing the channel bandwidth, methods for transmitting audio data include: 1) adding a data buffer to the transmitter to temporarily cache data that is not yet transmitted, waiting for the previous data to be transmitted before transmitting the remaining data; 2) compressing the sampled data at the transmitter to reduce the amount of data required to be transmitted through the channel. The receiver then decompresses the compressed data to recover the sampled data.

[0004] However, the existence of the buffer will cause a large delay in audio data, and long-term data transmission will increase the risk of data overflow and loss; data compression includes lossless and lossy methods. Lossless compression has a large delay, and lossy compression has low audio quality. Data compression methods cannot guarantee the transmission effect of audio data. Summary of the Invention

[0005] The present invention provides a method, device, equipment, medium and product for transmitting audio data in real-time compression, which performs data extraction, linear prediction and encoding on the transmitter side to reduce the amount of real-time transmission data and realize low-latency data transmission. It performs decoding, prediction recovery and data interpolation on the receiver side to achieve high-quality data transmission, taking into account the low latency and high sound quality requirements of the data transmission process, and improving the performance of the audio data transmission system.

[0006] According to one aspect of the present invention, a method for transmitting audio data in real-time compression is provided, the method comprising:

[0007] Determining target transmission information, wherein the target transmission information is audio data obtained after downsampling, error prediction, and encoding of the audio data to be transmitted by the transmitter;

[0008] Control the information transmission channel to transmit the target transmission information to the receiver;

[0009] The receiver is controlled to decode the target transmission information to obtain audio data to be played, wherein the audio data to be played is audio data obtained after decoding, predicting and restoring the target transmission information and performing data interpolation processing.

[0010] According to another aspect of the present invention, a device for transmitting audio data in real time with real-time compression is provided. The device for transmitting audio data in real time with real-time compression is used to implement the method for transmitting audio data in real time with real-time compression in any embodiment of the present invention. The device comprises:

[0011] An information acquisition module is used to determine target transmission information, wherein the target transmission information is audio data obtained by downsampling, error prediction and encoding the audio data to be transmitted using a transmitter;

[0012] An information sending module is used to control the information transmission channel and transmit the target transmission information to the receiver;

[0013] The information decoding module is used to control the receiver to decode the target transmission information to obtain the audio data to be played, wherein the audio data to be played is the audio data obtained after decoding, predicting and recovering the target transmission information and performing data interpolation processing.

[0014] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:

[0015] at least one processor; and a memory communicatively coupled to the at least one processor;

[0016] The memory stores a computer program that can be executed by at least one processor, and the computer program is executed by at least one processor so that the at least one processor can execute the method for transmitting audio data in real-time compression in any embodiment of the present invention.

[0017] According to another aspect of the present invention, a computer-readable storage medium is provided, which stores computer instructions, which are used to enable a processor to implement the real-time compressed audio data transmission method in any embodiment of the present invention when executed.

[0018] According to another aspect of the present invention, a computer program product is provided. The computer program product includes a computer program. When the computer program is executed by a processor, it implements the method for transmitting audio data in real-time compression according to any embodiment of the present invention.

[0019] The present invention discloses a method for transmitting audio data in real-time compression, comprising: determining target transmission information, wherein the target transmission information is the audio data obtained by downsampling, error prediction, and encoding the audio data to be transmitted by a transmitter; controlling an information transmission channel to transmit the target transmission information to a receiver; and controlling the receiver to decode the target transmission information to obtain audio data to be played, wherein the audio data to be played is the audio data obtained by decoding, predicting, recovering, and interpolating the target transmission information. The technical solution of the present invention reduces the amount of real-time transmission data and achieves low-latency data transmission by performing data extraction, linear prediction, and encoding on the transmitter side, and performs decoding, predicting, recovering, and interpolating data on the receiver side to achieve high-quality data transmission, balancing the low latency and high sound quality requirements of the data transmission process and improving the performance of the audio data transmission system. The method solves the problems of existing transmission methods, such as buffers causing large audio data delays, long data transmission processes increasing the risk of data overflow and loss, large delays in lossless compression methods, low audio quality in lossy compression, and the inability of data compression methods to guarantee audio data transmission.

[0020] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present invention, nor is it intended to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0022] Figure 1 This is a flow chart of a method for transmitting audio data in real-time compression provided by the present invention;

[0023] Figure 2 is a schematic diagram of an audio transmission process provided by the present invention;

[0024] Figure 3 It is a schematic diagram of a data processing process provided by the present invention;

[0025] Figure 4It is a flow chart of another method for transmitting audio data in real-time compression provided by the present invention;

[0026] Figure 5 This is a structural diagram of a transmission device for real-time compression of audio data provided by the present invention;

[0027] Figure 6 It is a structural schematic diagram of an electronic device provided by the present invention. DETAILED DESCRIPTION

[0028] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.

[0029] It should be noted that the terms "first", "second", "candidate", "initial", "intermediate", "target", "alternative", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0030] Traditional audio data transmission systems use relatively low sampling rates and quantization depths for audio ADCs. For example, a 22.5 kHz, 16 bit / s audio ADC yields a data rate equal to the product of 22.5 kHz and 16 bit / s, or 360 kbit / s. Hz and bit / s represent the sampling frequency in communications, while bit / s represents the quantized digital signal. When using phase-based modulation (Quadrature Phase Shift Keying, QPSK) for channel modulation and demodulation, the symbol rate of the transmitted signal in the channel is 180 ksps, meaning that less than 200 kHz of bandwidth is required for lossless, real-time transmission of all data. However, with the increasing widespread and widespread use of audio devices, the sampling rates and quantization depths of audio ADCs are also increasing, leading to an increasing amount of data being transmitted through the channels. Without compression of the sampled data, the existing channel bandwidth will no longer be sufficient for real-time data transmission.

[0031] Figure 1 This is a flow chart of a method for transmitting audio data in real time with compression provided by the present invention. This embodiment is applicable to situations such as real-time transmission of audio data, high efficiency and high quality transmission of audio data, etc. The method can be executed by the device for transmitting audio data in real time with compression provided by the present invention. The device can be implemented in the form of hardware and / or software. In a specific embodiment, the device can be integrated into an electronic device. The following embodiments will be described using the device integrated into an electronic device as an example. Figure 1 , the method specifically comprises the following steps:

[0032] S101: Determine target transmission information.

[0033] From audio source acquisition to final playback, multiple stages are required, each working in concert to ensure efficient and accurate audio signal reception and reproduction. The audio signal processing process primarily includes audio signal acquisition, preprocessing, digitization, data compression, encapsulation, transmission, reception, decapsulation, decompression, decoding, post-processing, and audio signal playback. The purpose of audio signal acquisition is to convert sound into an electrical signal (i.e., analog audio signal) or a digital signal (i.e., digital audio signal) using devices such as microphones and audio interfaces. For example, a microphone converts the mechanical vibrations of sound into an electrical signal, while an audio interface converts analog signals into digital signals. Preprocessing aims to optimize audio signal quality and reduce interference. Examples include gain adjustment, noise cancellation, and pre-emphasis. Gain adjustment adjusts the signal amplitude to ensure moderate signal strength, noise cancellation removes background noise and improves the signal-to-noise ratio, and pre-emphasis boosts the energy of high-frequency signals to reduce high-frequency attenuation during transmission. Digitization converts analog audio signals into digital signals for subsequent processing and transmission. Examples include sampling, quantization, and encoding. Sampling converts a continuous analog signal into discrete sampling points, while quantization converts the amplitude values of the sampling points into discrete digital codes. Encoding involves subjecting the digitized audio signal to encoding methods such as Pulse Code Modulation (PCM), Moving Picture Experts Group Audio Layer III (MP3), and Advanced Audio Coding (AAC). Data compression aims to reduce data volume and improve transmission efficiency. Ideally, it can reduce file size by removing audio information imperceptible to the human ear while ensuring that the audio signal can be fully restored after compression. Encapsulation encapsulates audio data into a format suitable for transmission. For example, audio data and metadata (such as sampling rate and encoding format) are encapsulated into a specific file format (e.g., MP4 or WAV). The audio data is then transmitted via a wired or wireless network to a receiver, which receives the audio data via an antenna or other receiving device. Decapsulation extracts audio data and metadata. Decompression restores compressed audio data to its original size. Decoding converts the encoded audio signal into a digital audio signal. The goal of decompression and decoding is to decompress and decode the compressed audio data into a playable audio signal. Post-processing further optimizes the quality of the audio signal to suit the playback device. Examples include dynamic range compression and filtering. Dynamic range compression ensures consistent volume across different devices, while filtering tailors the audio signal to the specific characteristics of the playback device.Audio signal playback refers to converting audio signals into sounds and playing them through devices such as speakers and headphones for users to hear. For example, digital audio is converted into analog signals through a digital-to-analog converter (DAC), and then used to drive speakers or headphones to produce sound.

[0034] The target transmission information can be understood as audio information that can be transmitted after digital processing, data compression and encapsulation. For example, the audio data to be transmitted is obtained after the transmitter performs downsampling processing, error prediction processing and encoding processing. The audio data to be transmitted can be understood as the sound signal collected by a microphone or audio interface. The downsampling processing is to convert the continuous analog signal into discrete sampling points to reduce the data volume. The error prediction processing can be understood as data calibration processing to improve the sampling accuracy to ensure the audio playback effect. The encoding processing is to encode and encapsulate the calibrated audio signal so that the audio data can be transmitted to the receiving end through a wired or wireless network. The receiving end can be understood as a receiver.

[0035] Figure 2 This is a schematic diagram of an audio transmission process provided by the present invention. The raw sampled data in the figure can be understood as the collected, unprocessed audio signal, the encoded data is the packaged audio signal to be transmitted, and the interpolated data can be understood as the audio information to be played. As can be seen from the figure, the transmitter performs decimation (downsampling), linear prediction, and encoding, while the receiver performs decoding, prediction recovery, and interpolation (upsampling). Both sides cooperate to ensure the quality of audio information transmission while improving transmission efficiency. Among them, decoding and encoding correspond to each other, prediction recovery corresponds to linear prediction, and interpolation corresponds to decimation.

[0036] In one embodiment, S101 may specifically include: determining the audio data to be transmitted, where the audio data to be transmitted is the audio data received by a receiving device (microphone, audio interface, etc.); using the data extraction unit of the transmitter to downsample the audio data to be transmitted based on preset data extraction rules to obtain first intermediate audio data, where the preset data extraction rules include an equal-interval quarter extraction rule; performing error prediction processing on the first intermediate audio data to obtain a target prediction error; encoding the first intermediate audio data based on preset encoding rules to obtain candidate transmission information; and determining the target transmission information based on the candidate transmission information and the target prediction error.

[0037] The data extraction unit can be understood as a module in the transmitter that performs data extraction or downsampling, the first intermediate audio data can be understood as audio data that has been sampled, the target prediction error can be understood as the first intermediate audio data that has been error-corrected, the preset encoding rule can be understood as a method for converting the first intermediate audio data into a binary stream, and the candidate transmission information can be understood as the first intermediate audio data converted into a binary stream. The target transmission information is determined based on the candidate transmission information and the target prediction error in order to obtain encapsulated data with a small data volume and corrected errors, thereby improving audio playback quality.

[0038] Specifically, Figure 3 is a schematic diagram of a data processing process provided by the present invention, Figure 3 The horizontal axis is time, indicating the order, and the vertical axis is the collected data value, indicating the size relationship. Taking the equal-interval quarter decimation rule as an example, the audio data obtained by sampling is downsampled by using the equal-interval 1 / 4 decimation method. Specifically, the sampling points are processed, and one sampling point is discarded for every four sampling points. The discarded sampling points are separated by three retained sampling points. The sampling method and sampling results are shown in the figure. Figure 3 As shown in the figure, the solid arrows correspond to the retained sampling points and the data values of the retained sampling points, and the dotted arrows correspond to the discarded sampling points and the data values corresponding to the discarded sampling points.

[0039] The advantages of this downsampling method are: 1) The data at the undecimated points (i.e., the retained sampling points) is the original sampled data, with a zero error. 2) No redundant data calculations are required; only the data at the sampling points to be decremented need to be discarded, resulting in high data processing efficiency and low data processing effort. 3) The fitting function calculation requires caching multiple sampling points, while direct decrement does not, resulting in lower transmission latency.

[0040] When the signal is interpolated and reconstructed at the receiving end, the error between the reconstructed signal and the original signal mainly comes from the high-frequency signal. The fitting function calculation of the high-frequency signal is often less accurate, and the overall error is higher than the signal error obtained by directly reconstructing the original data obtained by extraction. Therefore, the present invention adopts a third-order fixed coefficient linear prediction algorithm, which no longer calculates coefficients on the sampled data and no longer calculates the Toeplitz matrix. The overall transmission delay of the audio signal is lower and the amount of calculation is reduced. By adding an error correction function to the prediction unit, the prediction error during non-floating point calculations is effectively controlled and kept within the range of ±1LSB.

[0041] The first intermediate audio data includes at least three sub-audio data and at least three sampling points, and the sub-audio data and the sampling points have a one-to-one correspondence. For the i-th sub-audio data, where i is an integer greater than 2, error prediction processing is performed on the first intermediate audio data to obtain a target prediction error, including: determining initial prediction data for the current sampling point based on the sub-audio data of the current sampling point, the sub-audio data of the previous sampling point, and the sub-audio data of the two previous sampling points; determining a prediction error for the current sampling point based on the sub-audio data of the current sampling point and the initial prediction data of the current sampling point; determining candidate prediction data for the current sampling point based on the prediction error of the current sampling point, candidate prediction data for the previous sampling point, and candidate prediction data for the two previous sampling points; and determining a target prediction error based on the candidate prediction data for the current sampling point, the sub-audio data of the current sampling point, and the prediction error of the current sampling point.

[0042] The sub-audio data is the original sampling data, combined with Figure 3 , assuming that the current sampling point is the sampling point corresponding to the dotted arrow between x1 and x2, the sub-audio data of the current sampling point is the data value corresponding to the dotted arrow between x1 and x2, the previous sampling point is x1, the sub-audio data of the previous sampling point is f1, the previous two sampling points are x0, the sub-audio data of the previous two sampling points is f0, and the prediction error of the current sampling point is d0 represents the original data of the current sampling point, d1 represents the original data of the previous sampling point, and d2 represents the original data of the two previous sampling points. The prediction error of the current sampling point is e0 = d0-n0 = d1-0.5d2-0.5d0, the candidate prediction data of the current sampling point is r0 = 2r1-r2-2e0, and the target prediction error is e c =e0-(d0-r0), r1 represents the candidate prediction data of the previous sampling point, r2 represents the candidate prediction data of the two previous sampling points, and the candidate prediction data of the previous sampling point and the candidate prediction data of the two previous sampling points are determined in the same way as the candidate prediction data of the current sampling point, which is not repeated here.

[0043] The preset encoding rules include data conversion rules, data partitioning rules, and data encoding rules. Based on the preset encoding rules, encoding the first intermediate audio data to obtain candidate transmission information includes: using the data conversion rules to process the first intermediate audio data to obtain second intermediate audio data, the data conversion rules being used to convert the first intermediate audio data into natural numbers; using the data partitioning rules to partition the second intermediate audio data and determine the quotient and remainder of each partition result; and using the data encoding rules to encode the quotient and remainder of each partition result to obtain candidate transmission information.

[0044] The data conversion rule is used to adjust the first intermediate audio data to a unified format for subsequent processing. The second intermediate audio data is the converted first intermediate audio data. For example, the positive and negative data values in the first intermediate audio data are converted to natural numbers. The conversion method is: d a Indicates the converted data value, d i Represents the data value before conversion. The data division rule is the basis for the division of the second intermediate audio data. For example, natural numbers are grouped in chronological order, with 16 data in each group, natural numbers are grouped in chronological order, with 8 data in each group, etc. Taking each group of 16 data as an example, after division, it is necessary to find the maximum value ma of each group of data and calculate the m value. The m value is used to assist in the encoding work, m=log2ma-2. After obtaining the m value, the quotient q and remainder r of all natural numbers will also be calculated, q=d a / 2 m , r=mod(d a ,2 m The data encoding rules are used to process audio information into binary information output. For example, the remainder is output using binary encoding, and the quotient is output using unary encoding. Taking the value of m equal to 6 as an example, assuming that the quotient of a certain point is 2 and the remainder r is 5, the encoded output is 001000101. Parsing the code from front to back, the number of consecutive 0s is q, q=2, indicating two 0s, and the third 1 is the interval between the quotient and the remainder. m=6 represents 6 bits, and r=5 indicates that the binary number of the bit is 5, that is, 000101.

[0045] S102: Control the information transmission channel to transmit the target transmission information to the receiver.

[0046] The information transmission channel can be understood as the transmission path of the audio signal, including wired and wireless transmission. The receiver can be understood as the device that receives the packaged audio signal and parses and processes it to obtain a playable audio signal. Specifically, the information transmission channel may be related to information such as user settings, the audio signal source, the audio signal transmission requirements, the audio signal collection location, and the playback location, and the present invention is not limited to this.

[0047] Furthermore, the information transmission channel can also be understood as the transmission medium, including but not limited to optical fiber transmission, coaxial transmission, analog transmission, interface transmission, Bluetooth transmission, network port transmission and other 2.4 non-universal protocol wireless transmission. Different media are suitable for different types of audio transmission requirements and application scenarios, and can be limited according to information such as transmission distance, anti-interference ability, and sound quality requirements.

[0048] S103: Control the receiver to decode the target transmission information to obtain the audio data to be played.

[0049] The audio data to be played is the audio data obtained after decoding, predictive recovery, and data interpolation of the target transmission information, and can be played by a playback device. The decoding process corresponds to the encoding process, the predictive recovery process corresponds to the error prediction process, and the data interpolation process corresponds to the downsampling process. Specifically, the receiver performs the inverse of the processing tasks at the transmitter to restore the audio data and achieve high-quality audio playback.

[0050] In one embodiment, S103 may specifically include: using preset decoding rules to decode the target transmission information to obtain third intermediate audio data; using prediction recovery rules to predict and recover the third intermediate audio data to obtain fourth intermediate audio data, the fourth intermediate audio data including audio data of two sampling points before the missing sampling point; processing the fourth intermediate audio data to obtain audio data to be completed for the missing sampling point; using the audio data to be completed to supplement the audio data of the missing sampling point in the fourth intermediate audio data to obtain audio data to be played.

[0051] Among them, the preset decoding rule is the reverse algorithm of the preset encoding rule, the third intermediate audio data is the decoded audio data, the prediction recovery processing is the reverse work of the linear prediction processing, the fourth intermediate audio data is the audio data after data recovery, and the audio data to be completed is the audio data after upsampling, that is, the restored extracted audio data, and the audio data to be played can be understood as the complete audio data that can be played after restoration.

[0052] The decoding process includes: 1) counting the number of consecutive '0's input until the input binary data becomes '1', and recording the number of consecutive '0's as the quotient q, and the binary data value corresponding to the preset number (i.e. m) bits after '1' as r. 2) calculating the natural number data, the natural number data d b =q*2 m +r. 3) Calculate the decoded output data, decode the output data

[0053] Audio data to be completed r a =2r1-r2-2e c. Furthermore, the present invention adopts a 4th-order interpolation algorithm, that is, the interpolation point data is restored by calculating the interpolation polynomial of 4 points, two points on both sides of the interpolation point. For any missing sampling point, the fourth intermediate audio data is processed to obtain the audio data to be completed for the missing sampling point, including: sorting the audio data in the fourth intermediate audio data according to the time information of the sampling point; determining the first-order quotient difference of two adjacent audio data, the second-order quotient difference of two adjacent first-order quotient differences, and the third-order quotient difference of two adjacent second-order quotient differences; fitting the first-order quotient difference, the second-order quotient difference, and the third-order quotient difference to obtain the audio data to be completed.

[0054] The purpose of sorting the audio data in the fourth intermediate audio data according to the time information of the sampling points is to ensure temporal coherence and continuity. Figure 3 , assuming that the missing sampling point is the sampling point between x1 and x2, then the adjacent sampling points include x0, x1, x2 and x3, and the adjacent audio data include f0, f1, f2 and f3. Specifically, the order of the sampling points is x0=1, x1=2, x2=4, x3=5, f0=r-2, indicating two sampling points before the missing sampling point, f1=r-1, indicating one sampling point before the missing sampling point, f2=r1, indicating one sampling point after the missing sampling point, and f3=r2, indicating two sampling points after the missing sampling point. The first-order difference quotients of two adjacent audio data are f01, f12 and f23, The second-order quotient difference of two adjacent first-order quotient differences is f012 and f123, The third-order quotient difference of two adjacent second-order quotient differences The audio data to be completed can be understood as the interpolation value of the missing sampling points, and the audio data to be completed y=f0+f 01 *(3-x0)+f 012 *(3-x0)*(3-x1)+f 0123 *(3-x0)*(3-x1)*(3-x2).

[0055] Using the audio data to be completed, the audio data of the missing sampling points in the fourth intermediate audio data are supplemented to obtain the audio data to be played. This can be understood as adding the audio data to be completed to the gaps in the fourth intermediate audio data in a time series, supplementing the fourth intermediate audio data, and obtaining complete and continuous audio information to ensure the audio playback effect.

[0056] The technical solution of the present invention reduces the amount of real-time data transmission and achieves low-latency data transmission by performing data extraction, linear prediction, and encoding on the transmitter side. Decoding, prediction recovery, and data interpolation are performed on the receiver side to achieve high-quality data transmission, balancing the requirements of low latency and high sound quality during the data transmission process and improving the performance of the audio data transmission system. It addresses existing transmission methods, such as buffers causing significant audio data delays, long data transmission times increasing the risk of data overflow and loss, large delays in lossless compression methods, low audio quality in lossy compression, and the inability of data compression methods to guarantee audio data transmission.

[0057] Figure 4 This is a flow chart of another method for transmitting audio data in real-time compression provided by the present invention. This embodiment provides a preferred method for transmitting audio data based on the above embodiment. Specifically, Figure 4 As shown, the method includes:

[0058] S201: Determine audio data to be transmitted.

[0059] The audio data to be transmitted is the audio data received by the audio receiving device.

[0060] S202 : Using a data extraction unit of the transmitter, based on a preset data extraction rule, downsample the audio data to be transmitted to obtain first intermediate audio data.

[0061] Among them, the preset data extraction rules include the equal-interval quarter extraction rule.

[0062] S203: Perform error prediction processing on the first intermediate audio data to obtain a target prediction error.

[0063] S204. Encode the first intermediate audio data based on a preset encoding rule to obtain candidate transmission information.

[0064] S205: Determine target transmission information based on the candidate transmission information and the target prediction error.

[0065] The target transmission information is audio data obtained by performing downsampling processing, error prediction processing and encoding processing on the audio data to be transmitted by the transmitter.

[0066] S206: Control the information transmission channel to transmit the target transmission information to the receiver.

[0067] S207: Decode the target transmission information using a preset decoding rule to obtain third intermediate audio data.

[0068] S208. Use the prediction recovery rule to perform prediction recovery processing on the third intermediate audio data to obtain fourth intermediate audio data.

[0069] The fourth intermediate audio data includes audio data of two sampling points before the vacant sampling point.

[0070] S209: Process the fourth intermediate audio data to obtain audio data of the missing sampling points to be completed.

[0071] S210. Use the audio data to be completed to supplement the missing sampling points in the fourth intermediate audio data to obtain the audio data to be played.

[0072] The audio data to be played is audio data obtained after decoding, predicting and restoring the target transmission information, and performing data interpolation processing.

[0073] Figure 5 This is a structural diagram of a transmission device for real-time compression of audio data provided by the present invention. Figure 5 As shown, the device includes: an information acquisition module 301, an information sending module 302 and an information calculation module 303.

[0074] The information acquisition module 301 is used to determine target transmission information, wherein the target transmission information is audio data obtained by downsampling, error prediction and encoding the audio data to be transmitted by the transmitter.

[0075] The information sending module 302 is used to control the information transmission channel and transmit the target transmission information to the receiver.

[0076] The information decoding module 303 is used to control the receiver to decode the target transmission information to obtain the audio data to be played, wherein the audio data to be played is the audio data obtained after decoding, predicting and recovering the target transmission information and performing data interpolation processing.

[0077] Optionally, the information acquisition module 301 is specifically used to determine the audio data to be transmitted, which is the audio data received by the receiving device; using the data extraction unit of the transmitter, based on the preset data extraction rules, down-sample the audio data to be transmitted to obtain first intermediate audio data, and the preset data extraction rules include the equal-interval quarter extraction rule; error prediction processing is performed on the first intermediate audio data to obtain a target prediction error; based on the preset encoding rules, the first intermediate audio data is encoded to obtain candidate transmission information; based on the candidate transmission information and the target prediction error, the target transmission information is determined.

[0078] Optionally, the first intermediate audio data includes at least three sub-audio data and at least three sampling points, and the sub-audio data and the sampling points have a one-to-one correspondence. The information acquisition module 301 is specifically configured to determine initial prediction data for the current sampling point based on the sub-audio data for the current sampling point, the sub-audio data for the previous sampling point, and the sub-audio data for the two previous sampling points; determine a prediction error for the current sampling point based on the sub-audio data for the current sampling point and the initial prediction data for the current sampling point; determine candidate prediction data for the current sampling point based on the prediction error for the current sampling point, candidate prediction data for the previous sampling point, and candidate prediction data for the two previous sampling points; and determine a target prediction error based on the candidate prediction data for the current sampling point, the sub-audio data for the current sampling point, and the prediction error for the current sampling point.

[0079] Optionally, the preset encoding rules include data conversion rules, data partitioning rules, and data encoding rules. The information acquisition module 301 is specifically configured to process the first intermediate audio data using the data conversion rules to obtain second intermediate audio data, where the data conversion rules are configured to convert the first intermediate audio data into natural numbers; partition the second intermediate audio data using the data partitioning rules and determine the quotient and remainder of each partition result; and encode the quotient and remainder of each partition result using the data encoding rules to obtain candidate transmission information.

[0080] Optionally, the information solution module 303 is specifically used to use preset decoding rules to decode the target transmission information to obtain third intermediate audio data; use prediction recovery rules to predict and recover the third intermediate audio data to obtain fourth intermediate audio data, and the fourth intermediate audio data includes audio data of two sampling points before the missing sampling point; process the fourth intermediate audio data to obtain audio data to be completed for the missing sampling point; use the audio data to be completed to supplement the audio data of the missing sampling points in the fourth intermediate audio data to obtain audio data to be played.

[0081] Optionally, for any missing sampling point, the information solution module 303 is specifically used to sort the audio data in the fourth intermediate audio data according to the time information of the sampling point; determine the first-order quotient difference of two adjacent audio data, the second-order quotient difference of two adjacent first-order quotient differences, and the third-order quotient difference of two adjacent second-order quotient differences; fit the first-order quotient difference, the second-order quotient difference, and the third-order quotient difference to obtain the audio data to be completed.

[0082] The audio data real-time compressed transmission device provided in the above embodiments can execute the audio data real-time compressed transmission method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.

[0083] Figure 6: is a structural diagram of an electronic device provided by the present invention. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0084] like Figure 6 As shown, the electronic device 10 includes at least one processor 11 and a memory connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (also known as random access memory, RAM) 13, etc., wherein the memory stores a computer program that can be executed by the at least one processor, and the processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 to the random access memory (RAM) 13. Various programs and data required for the operation of the electronic device 10 can also be stored in the RAM 13. The processor 11, ROM 12 and RAM 13 are connected to each other via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0085] Multiple components in the electronic device 10 are connected to the I / O interface 15, including an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0086] The processor 11 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the method for transmitting audio data in real-time compression.

[0087] In some embodiments, the method for transmitting audio data in real time with compression can be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the method for transmitting audio data in real time with compression can be performed. Alternatively, in other embodiments, the processor 11 can be configured to perform the method for transmitting audio data in real time with compression by any other appropriate means (e.g., by means of firmware).

[0088] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0089] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer program is executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer program may be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0090] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in combination with an instruction execution system, device or equipment. A computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium can be a machine-readable signal medium. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device or any suitable combination of the foregoing.

[0091] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device that has: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0092] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0093] A computing system may include clients and servers. The clients and servers are typically remote from each other and typically interact via a communication network. This client-server relationship arises through computer programs running on the respective computers, creating a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host. This server is a hosting product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosting and VPS services.

[0094] In one embodiment, the present invention further includes a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements the method for transmitting audio data in real-time compression according to any embodiment of the present invention.

[0095] The computer program product may be implemented in a computer program code for performing the operations of the present invention written in one or more programming languages, or a combination thereof, including object-oriented programming languages and conventional procedural programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0096] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in the present invention can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved. This is not limited herein.

[0097] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.

Claims

1. A method for transmitting audio data in real-time compression, characterized in that: include: Determining target transmission information, wherein the target transmission information is audio data obtained by downsampling, error prediction, and encoding the audio data to be transmitted using a transmitter; Controlling the information transmission channel to transmit the target transmission information to the receiver; The receiver is controlled to decode the target transmission information to obtain audio data to be played, wherein the audio data to be played is audio data obtained after decoding, predicting and restoring, and interpolating the target transmission information.

2. The method for transmitting audio data in real-time compression according to claim 1, wherein: The determining of target transmission information includes: Determining the audio data to be transmitted, wherein the audio data to be transmitted is audio data received by a sound receiving device; Using the data extraction unit of the transmitter, based on a preset data extraction rule, the downsampling process is performed on the audio data to be transmitted to obtain first intermediate audio data, wherein the preset data extraction rule includes an equal-interval quarter extraction rule; performing the error prediction process on the first intermediate audio data to obtain a target prediction error; performing the encoding process on the first intermediate audio data based on a preset encoding rule to obtain candidate transmission information; The target transmission information is determined based on the candidate transmission information and the target prediction error.

3. The method for transmitting audio data in real-time compression according to claim 2, wherein: The first intermediate audio data includes at least three sub-audio data and at least three sampling points, and the sub-audio data and the sampling points correspond one to one; For the i-th sub-audio data, where i is an integer greater than 2, performing the error prediction process on the first intermediate audio data to obtain a target prediction error includes: Determining initial prediction data for the current sampling point based on the sub-audio data of the current sampling point, the sub-audio data of the previous sampling point, and the sub-audio data of the previous two sampling points; determining a prediction error of the current sampling point based on the sub-audio data of the current sampling point and the initial prediction data of the current sampling point; Determining candidate prediction data for the current sampling point based on the prediction error of the current sampling point, the candidate prediction data for the previous sampling point, and the candidate prediction data for the two previous sampling points; The target prediction error is determined based on the candidate prediction data of the current sampling point, the sub-audio data of the current sampling point, and the prediction error of the current sampling point.

4. The method for transmitting audio data in real-time compression according to claim 2, wherein: The preset coding rules include data conversion rules, data division rules and data coding rules; The encoding process is performed on the first intermediate audio data based on a preset encoding rule to obtain candidate transmission information, including: Processing the first intermediate audio data using the data conversion rule to obtain second intermediate audio data, wherein the data conversion rule is used to convert the first intermediate audio data into a natural number; Dividing the second intermediate audio data using the data division rule, and determining a quotient and a remainder of each division result; The data encoding rule is used to encode the quotient and remainder of each division result to obtain the candidate transmission information.

5. The method for transmitting audio data in real-time compression according to claim 1, wherein: The control receiver solves the target transmission information to obtain the audio data to be played, including: Decoding the target transmission information using a preset decoding rule to obtain third intermediate audio data; performing predictive recovery processing on the third intermediate audio data using a predictive recovery rule to obtain fourth intermediate audio data, wherein the fourth intermediate audio data includes audio data of two sampling points before the missing sampling point; Processing the fourth intermediate audio data to obtain audio data to be completed for the missing sampling points; The audio data to be completed is used to supplement the missing sampling points in the fourth intermediate audio data to obtain the audio data to be played.

6. The method for transmitting audio data in real-time compression according to claim 5, characterized in that: For any missing sampling point, the processing of the fourth intermediate audio data to obtain the audio data to be completed for the missing sampling point includes: sorting the audio data in the fourth intermediate audio data according to time information of the sampling points; Determine a first-order difference quotient of two adjacent audio data, a second-order quotient difference of two adjacent first-order quotient differences, and a third-order quotient difference of two adjacent second-order quotient differences; The first-order quotient difference, the second-order quotient difference, and the third-order quotient difference are fitted to obtain the audio data to be completed.

7. A transmission device for real-time compression of audio data, characterized in that: A method for transmitting audio data in real-time compression according to any one of claims 1 to 6, wherein the transmission device for transmitting audio data in real-time compression comprises: an information acquisition module, configured to determine target transmission information, wherein the target transmission information is audio data obtained by downsampling, error prediction, and encoding the audio data to be transmitted using a transmitter; An information sending module is used to control the information transmission channel and transmit the target transmission information to the receiver; The information decoding module is used to control the receiver to decode the target transmission information to obtain audio data to be played, wherein the audio data to be played is the audio data obtained after decoding, prediction recovery and data interpolation processing of the target transmission information.

8. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively coupled to the at least one processor; The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the real-time compressed audio data transmission method described in any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the audio data real-time compression and transmission method according to any one of claims 1 to 6 when executed.

10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the computer program implements the method for transmitting audio data in real-time compression according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Backward block adaptive Golomb-Rice coding and decoding method and apparatus thereof

    CN102368385A

  • Quasi lossless compression algorithm for correcting subcarriers based on assistance of a part of sampling points

    CN107018107A

  • Audio signal coding compression and transmission method and electronic equipment

    CN113129911A

  • Signal processing method and device, computer equipment, storage medium and program product

    CN117334204A

  • Method and system for reducing transmission delay of Bluetooth headset

    CN117459510A