Audio processing method, electronic device and audio transmission system
By processing the original audio data, separating and compressing the main sound data and the residual sound data, the problems of high resource consumption and poor real-time performance of existing audio compression algorithms are solved. The algorithm is simplified while maintaining audio quality, improving the user experience.
Patent Information
- Application Number
- CN202510940703.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-09
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2045-07-09
AI Technical Summary
Existing audio compression algorithms have problems of high resource consumption and poor real-time performance when processing signals with a large dynamic range, and the audio quality is easily affected.
The original audio data is processed to separate the main sound data and the residual sound data. The main sound data and the residual sound data are compressed using a simple audio compression algorithm to generate first compressed data and second compressed data. These data are then sent to a pre-set second device for decompression and reassembly to generate the target audio data.
While simplifying the audio compression algorithm, it reduces resource consumption and improves real-time performance, avoids a significant decline in audio quality, and improves user experience.
Smart Images

Figure CN120472917B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of audio technology, and in particular to an audio processing method, electronic equipment, and audio transmission system. Background Art
[0002] In the recording industry, dynamic range is a key metric for measuring recording quality. It directly determines the sound reproduction and listening experience, especially in vocal and instrument recordings. Insufficient dynamic range can easily lead to sound clipping and distortion, seriously affecting audio quality. However, while existing audio compression algorithms can achieve a large dynamic range input, they place high demands on algorithmic complexity. While simple audio compression algorithms are easy to implement, they have many limitations when processing signals with a large dynamic range.
[0003] Therefore, how to simplify the audio compression algorithm to reduce resource consumption and improve real-time performance while ensuring that the sound quality is not seriously affected is a technical problem that needs to be solved urgently. Summary of the Invention
[0004] The present application provides an audio processing method, an electronic device, and an audio transmission system, which can simplify the audio compression algorithm to reduce resource consumption and improve real-time performance while avoiding a significant decline in audio quality.
[0005] In a first aspect, the present application provides an audio processing method, which is applied to a first device, the method comprising:
[0006] Processing the original audio data to obtain main sound data and residual sound data of the original audio;
[0007] compressing the main sound data to obtain first compressed data, and compressing the residual sound data to obtain second compressed data;
[0008] The first compressed data and the second compressed data are sent to a preset second device for decompression and reassembly, so as to generate target audio data in the second device.
[0009] In a second aspect, the present application further provides an audio processing method, which is applied to a second device, the method comprising:
[0010] receiving first compressed data and second compressed data transmitted by a preset first device;
[0011] Decompressing the first compressed data and the second compressed data respectively to obtain first data and second data;
[0012] The first data and the second data are recombined to obtain target audio data.
[0013] In a third aspect, the present application further provides an audio processing apparatus, which is applied to a first device, and the apparatus includes:
[0014] A processing unit, configured to process the original audio data to obtain main sound data and residual sound data of the original audio;
[0015] a compression unit, configured to compress the main sound data to obtain first compressed data, and compress the residual sound data to obtain second compressed data;
[0016] The sending unit is used to send the first compressed data and the second compressed data to a preset second device for decompression and reassembly, so as to generate target audio data in the second device.
[0017] In a fourth aspect, the present application further provides an audio processing apparatus, which is applied to a second device, and the apparatus includes:
[0018] A receiving unit, configured to receive first compressed data and second compressed data transmitted by a preset first device;
[0019] a decompression unit, configured to decompress the first compressed data and the second compressed data respectively to obtain first data and second data;
[0020] The recombining unit is used to recombining the first data and the second data to obtain target audio data.
[0021] In a third aspect, an embodiment of the present application provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the audio processing method provided in the first or second aspect above is implemented.
[0022] In a fourth aspect, an embodiment of the present application also provides an audio transmission system, which includes a first device and a second device, and audio data is transmitted between the first device and the second device. The first device is configured to execute the audio processing method provided by the first aspect, and the second device is configured to execute the audio processing method provided by the second aspect.
[0023] In a fifth aspect, an embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the processor executes the audio processing method provided in the first or second aspect above.
[0024] In a sixth aspect, an embodiment of the present application further provides a computer program product, including a computer program or instructions, where the computer program or instructions are executed by a processor to perform the audio processing method provided in the first aspect or the second aspect.
[0025] The audio processing method provided in the present application can process the original audio data to obtain the main sound data and residual sound data of the original audio, and can apply a relatively simple audio compression algorithm to compress the main sound data and the residual sound data respectively to obtain first compressed data and second compressed data. Finally, the first compressed data and the second compressed data can be sent to a preset second device for decompression and reorganization to generate target audio data in the second device. In this way, while simplifying the audio compression algorithm to reduce resource consumption and improve real-time performance, a significant decline in audio quality can be avoided, thereby greatly improving the user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0027] Figure 1 A framework diagram of the audio transmission system provided in an embodiment of the present application;
[0028] Figure 2 A schematic diagram of a first flow chart of the audio processing method provided in an embodiment of the present application;
[0029] Figure 3 A second flow chart of the audio processing method provided in an embodiment of the present application;
[0030] Figure 4 A schematic block diagram of an audio processing device provided in an embodiment of the present application;
[0031] Figure 5 A schematic block diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0032] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0033] It will be understood that when used in this specification and the appended claims, the terms including and comprising indicate the presence of described features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.
[0034] It should also be understood that the terms used in this specification are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in this specification and the appended claims, the singular forms "a", "an", and "the" are intended to include the plural forms unless the context clearly indicates otherwise.
[0035] It should be further understood that the terms and / or used in the specification and appended claims refer to and include any and all possible combinations of one or more of the associated listed items.
[0036] In addition, in this application, unless otherwise clearly specified or limited in the embodiments, the terms "installation", "connection", "connection" and "fixation" appearing in the embodiments should be understood in a broad sense. For example, "connection" can be a fixed connection, a detachable connection, or an integral connection. It can also be a mechanical connection, an electrical connection, etc.; of course, it can also be a direct connection, or an indirect connection through an intermediate medium, or it can be the internal communication of two elements, or the interaction between two elements. For those skilled in the art, the specific meanings of the above terms in this application can be understood based on the specific implementation.
[0037] The present application provides an audio processing method, an electronic device, and an audio transmission system.
[0038] To facilitate understanding, the audio transmission system is first introduced, and then the audio processing method is introduced in detail based on this system.
[0039] See also Figure 1 , Figure 1 This is a framework diagram of an audio transmission system provided in an embodiment of the present application. The audio transmission system includes a first device and a second device, and audio data can be transmitted between the first and second devices. Furthermore, the audio data transmission between the first and second devices can be implemented using a wireless network. The first device can be a recording device, and the second device can be an audio playback device. Furthermore, the first and second devices can be integrated or separated, depending on the actual application and is not specifically limited in this application.
[0040] It should be noted that the application scenarios described in the following embodiments of the present application are intended to more clearly illustrate the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Ordinary technicians in this field can know that with the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are also applicable to similar technical problems.
[0041] The audio processing method provided in this application is described in detail below.
[0042] like Figure 2 As shown, the method includes the following steps S210~S230.
[0043] S210, processing the original audio data to obtain main sound data and residual sound data of the original audio;
[0044] S220: compress the main sound data to obtain first compressed data, and compress the residual sound data to obtain second compressed data;
[0045] S230: Send the first compressed data and the second compressed data to a preset second device for decompression and reassembly, so as to generate target audio data in the second device.
[0046] Specifically, in the process of processing the original audio data, the present application can use a preset audio processing algorithm to process the original audio data. Specifically, the original audio data can be compressed or normalized through the dynamic range to obtain the main sound information within the preset first dynamic range, that is, the main sound data, and the main sound data is subtracted from the original audio data to obtain residual sound data. The residual sound data can restore more dynamic range information after subsequent processing, which not only provides a guarantee for the accuracy and effectiveness of subsequent processing, but also lays the foundation for improving audio quality. At the same time, in the data separation process, the algorithm accurately identifies and extracts the components of the main sound, while retaining the remaining part as residual sound data for processing in subsequent links.
[0047] The primary sound data represents the main sound in the original audio data. As the core of the original audio, it carries the main information and features. The residual sound data represents the difference between the main sound and the original audio, and can be understood as the secondary sound. For example, if a drum beat is loud at a certain moment in the original audio, and the main sound is scaled, the drum beat becomes quieter, then the residual sound data may contain the information about the drum beat that was lost due to the scaling.
[0048] At the same time, in the process of compressing the main sound data of the present application, the main sound can be scaled to a first dynamic range. The first dynamic range is smaller than the dynamic range of the main sound in the original audio data. Then, a simple small dynamic range audio compression algorithm can be used to compress the main sound data to obtain first compressed data. Among them, the small dynamic range audio compression algorithm has a lower computational complexity, which can more easily control parameters such as the compression ratio and can better adapt to the characteristics of the main sound. For example, small amplitude changes in the main sound can be effectively compressed. For example, pydub can be used to compress the main sound data into MP3 format data.
[0049] During the compression process of the residual sound data, the residual sound data also contains some audio information, and its dynamic range is different from that of the main sound. The present application can select a corresponding compression strategy to compress the residual sound data according to its own characteristics. For example, FLAC or WAV can be used to compress the residual sound data to preserve the detailed information.
[0050] In this application, the first compressed data and the second compressed data can be sent to the second device via a stable and efficient transmission channel. At the same time, the second device can utilize specially developed decompression and reconstruction technology to restore the compressed data to high-quality audio data, i.e., target audio data. This ensures the fidelity and integrity of the audio during transmission and processing, thereby providing users with a high-quality listening experience, allowing users to experience the rich details and high-quality effects of the original audio when enjoying the audio content.
[0051] In some embodiments, processing the original audio data to obtain the main sound data and the residual sound data of the original audio includes: processing the original audio data to obtain the main sound data; and generating the residual sound data based on the main sound data.
[0052] Specifically, in the process of processing the original audio data, the present application can directly process the original audio data to use the processed main data as the main sound data, which can be obtained through dynamic range compression, normalization or other processing methods, and can be based on the main sound data. The main sound data can be subtracted from the original audio data to obtain the residual sound data, so as to effectively process audio signals with different dynamic ranges while maintaining high precision and low distortion.
[0053] In some embodiments, processing the original audio data to obtain the main sound data includes: mapping the original audio data into a preset integer space to generate audio data in an integer format; and generating the main sound data based on the audio data in the integer format.
[0054] In the present application, in the process of processing the original audio data, the file of the original audio data can be read, and the original audio data can be read as 32-bit integer data. The 32-bit integer data can then be converted from a floating-point format to a preset integer format, and the main sound data can be generated based on the audio data in the integer format.
[0055] Specifically, in the process of converting the original audio data from floating-point format to integer format, the floating-point signal can be multiplied by a scaling factor and rounded to convert the floating-point audio data into a preset integer format. At the same time, in the process of generating the main sound data based on the integer format audio data, the present application can process the integer format audio data to generate the main sound data, such as performing dynamic range compression or normalization on the integer format audio data.
[0056] In this embodiment, the integer space can be a 64-bit integer space. The present application can construct a full integer processing pipeline based on int64_t and amplify the 32-bit floating-point signal to the 64-bit integer space so as to achieve zero overflow error during subsequent multiplication and addition operations, thereby breaking through the dynamic limitations of traditional int32_t processors.
[0057] Specifically, 32-bit floating point numbers The formula for scaling up to 64-bit integer space can be:
[0058]
[0059] in, It can be understood as audio data in integer format. It can be understood as raw audio data.
[0060] In some embodiments, generating the main sound data based on the audio data in the integer format includes: performing a limiting process on the audio data in the integer format within a preset first dynamic range to obtain the main sound data.
[0061] In this application, in the process of determining the first dynamic range, it is necessary to determine a suitable dynamic range to include the main components of the audio signal while avoiding signal distortion. For example, the first dynamic range can be , that is [-32768, 32767].
[0062] At the same time, in the process of limiting the audio data in integer format, the present application needs to limit the audio data in integer format to a first dynamic range to generate main sound data. Specifically, a hard limiting or soft limiting function, such as tanh or clip, can be used to limit the signal to the first dynamic range, and the processed main sound data can be saved as a new audio file for subsequent use.
[0063] Among them, hard limiting is a simple limiting method, which can directly truncate signal values that exceed the specified range to the boundary value of the range. That is, hard limiting can be understood as directly truncating signal values that exceed the dynamic range to the boundary value. It may introduce obvious distortion, especially when the signal amplitude is close to the boundary of the limiting range. At the same time, the signal transition is not smooth, which may cause "clicking" or "clipping" in the audio.
[0064] Soft limiting is a smoother limiting method that gradually reduces the amplitude of the signal when it approaches the limit range boundary, rather than directly truncating it. That is, soft limiting can be understood as using a smooth function (such as the tanh function) to limit the signal to the dynamic range to reduce distortion. However, soft limiting is more complex than hard limiting and has slightly lower computational efficiency. At the same time, in some cases, soft limiting may slightly compress the dynamic range of the signal.
[0065] In this embodiment, since it is necessary to simplify the audio compression algorithm to reduce resource consumption, and the present application separates the residual audio data from the original audio data, and the data volume of the audio data in integer format is large, the present application can directly perform hard limiting processing on the audio data in integer format within the first dynamic range to reduce the complexity of the algorithm, thereby achieving the purpose of reducing resource consumption.
[0066] Specifically, the formula for hard limiting processing of audio data in integer format can be:
[0067]
[0068] in, It can be understood as the main sound data. It can be understood as audio data in integer format.
[0069] In some embodiments, generating residual sound data of the original audio based on the main sound data includes: obtaining audio data and the main sound data in integer format; and generating residual sound data according to the audio data and the main sound data in integer format.
[0070] In this application, in the process of generating residual sound data, the main sound data can be subtracted from the integer-format audio data to obtain the residual sound data, thereby effectively processing audio signals of different dynamic ranges while maintaining high precision and low distortion. The main sound data is the main component extracted from the integer-format audio data, which is used to represent the main part of the audio signal; the residual sound data is the difference between the original audio data and the main sound data.
[0071] In some embodiments, residual sound data is generated based on the audio data in integer format and the main sound data, including: generating initial parameter sound data based on the audio data in integer format and the main sound data; and limiting the initial parameter sound data within a preset second dynamic range to obtain residual sound data.
[0072] In this application, the initial parameter sound data can be understood as the difference between the audio data in integer format and the main sound data, which can be used to represent the detailed information of the audio signal. The initial parameter sound data can be obtained by subtracting the main sound data from the audio data in integer format.
[0073] Specifically, the second dynamic range can be used to limit the amplitude of the initial parameter sound data. The present application uses the second dynamic range to limit the initial parameter sound data, which can ensure that the amplitude of the residual sound data is within a reasonable range and avoid distortion caused by an excessively large dynamic range. In the process of limiting the initial parameter sound data within the second dynamic range, a hard limiting or soft limiting function, such as tanh or clip, can be used to limit the signal to the second dynamic range, and the obtained residual sound data can be saved as a new audio file for subsequent use. The residual sound data can be saved as a new audio file using To represent it, save it as a 24-bit PCM format file.
[0074] In this embodiment, the amount of data occupied by the initial parameter sound data is relatively small, and this application needs to ensure the high precision and low distortion of the final audio data. Therefore, in the process of limiting the initial parameter sound data within the second dynamic range, the initial parameter sound data can be soft-limited. Specifically, a nonlinear saturation function can be used to avoid extreme value overflow, and residual sound data can be obtained.
[0075] Specifically, the formula for generating the initial parameter sound data can be:
[0076]
[0077] in, It can be understood as the initial parameter sound data. It can be understood as audio data in integer format. It can be understood as the main sound data.
[0078] The formula for soft limiting the initial parameter sound data can be:
[0079]
[0080] in, It can be understood as residual sound data. It can be understood as the initial parameter sound data.
[0081] In some embodiments, compressing the residual sound data to obtain second compressed data includes: compressing the residual sound data using a preset residual dynamic mapping model to obtain the second compressed data.
[0082] Specifically, in the process of compressing the residual sound data, the present application can specifically use a residual dynamic mapping model for compression processing to achieve lossless encapsulation of the residual sound data, so that the lost information of the main sound can be restored to a large extent in the second device later. The residual dynamic mapping model can be expressed by the following formula:
[0083]
[0084] in, Residual sound data saved in 24-bit integer format, It can be understood as the second compressed data, that is, the residual sound data saved in the 24-bit integer format.
[0085] In some embodiments, a preset residual dynamic mapping model is used to compress the residual sound data to obtain second compressed data, including: performing dithering processing on the residual sound data to obtain dithered residual sound data; and performing shifting processing on the dithered residual sound data to obtain second compressed data.
[0086] Specifically, in the process of compressing the residual sound data using the residual dynamic mapping model, the residual sound data is subjected to dithering processing, specifically using the residual dynamic mapping model in the residual sound data. Dithering the residual sound data can generate a random integer between 0 and 255, that is, a 24-bit integer data, and use the random integer as a dithering signal. After addition, the generated new 24-bit integer data is shifted, specifically by right shifting 8 bits (equivalent to dividing by 256), and then the 24-bit data can be converted to the range of 16-bit data. Among them, the value range of 16-bit integer data is arrive , and the value range of 24-bit integer data is arrive Therefore, the present application can right-shift a newly generated 24-bit integer data by 8 bits, thereby converting the 24-bit integer data into the range of 16-bit integer data.
[0087] In this application, by performing dithering processing on the residual sound data, the quantization error can be reduced and the sound quality can be improved, and the residual sound data after dithering processing is shifted, thereby achieving lossless conversion of 24-bit residual sound data into 16-bit residual sound data.
[0088] In the audio processing method provided in the embodiment of the present application, the original audio data is processed to obtain the main sound data and residual sound data of the original audio, and a relatively simple audio compression algorithm can be applied to compress the main sound data and the residual sound data respectively to obtain first compressed data and second compressed data. Finally, the first compressed data and the second compressed data can be sent to a preset second device for decompression and reorganization to generate target audio data in the second device. This can achieve the goal of simplifying the audio compression algorithm to reduce resource consumption and improve real-time performance while avoiding a significant decline in audio quality and greatly improving user experience.
[0089] In some embodiments, as Figure 3 As shown, the present application also provides an audio processing method, which is applied to a second device. The audio processing method includes steps S310, S320 and S330.
[0090] S310, receiving first compressed data and second compressed data transmitted by a preset first device;
[0091] S320, decompressing the first compressed data and the second compressed data respectively to obtain first data and second data;
[0092] S330: Recombining the first data and the second data to obtain target audio data.
[0093] In this application, after the second device receives the first compressed data and the second compressed data transmitted by the first device, it can decompress the first compressed data (main sound) and the second compressed data (secondary sound) respectively and reorganize them, thereby ensuring the fidelity and integrity of the audio during transmission and processing.
[0094] During the compression process, decompression can be performed using the corresponding decompression algorithm or library (such as libmp3lame or libflac) based on the compression format (e.g., MP3, FLAC, AAC, etc.). Furthermore, because the secondary compressed data contains the difference information between the original sound source and the primary sound, the combination can restore a wider dynamic range than the primary sound alone after decompression. For example, if the original primary sound has a smaller dynamic range after scaling, the combination of the decompressed primary sound and the residual sound can restore the originally reduced dynamic range of parts such as drum beats to a near-original larger dynamic range.
[0095] In some embodiments, the first data and the second data are recombined to obtain the target audio data, including: using a preset accumulator to recombined the first data and the second data to obtain the target audio data.
[0096] In the present application, the first data corresponds to the main sound data, and the second data corresponds to the residual sound data. After decompressing the first compressed data and the second compressed data respectively to obtain the first data and the second data, the present application can use a preset accumulator to reorganize the first data and the second data, thereby achieving sub-bit compensation of floating-point errors in the signal restoration stage, achieving higher-precision accumulation operations, and thus realizing suitability for audio signal processing scenarios requiring high precision.
[0097] In this embodiment, the accumulator may be a Kahan accumulator. The core concept of the Kahan accumulator is to track and compensate for the error caused by rounding in floating-point addition. That is, by retaining an error term (err) and updating the error term after each addition operation, the error is compensated in subsequent calculations. The accumulator can be characterized as follows:
[0098]
[0099] in, It can be understood as the output value of the accumulator initialization. It can be understood as the subsequent output value of the accumulator. It can be understood as residual sound data. It can be understood as the main sound data. It can be understood as an error term.
[0100] Specifically, in the process of using a preset accumulator to reorganize the first data and the second data, the present application can initialize the accumulator and the error term. Specifically, the output value of the accumulator can be initialized to 0, and the error term can be initialized to 0. Then, the Kahan accumulator is used to reorganize the data. Specifically, for each sample, the sum of the new data and the current error term can be calculated first, and then added to the accumulator, and the error term is updated to the rounding error of this addition. Finally, the target audio data can be output. Specifically, the final value of the accumulator can be used as the reorganized target audio data.
[0101] It should be noted that when using a pre-set accumulator to reconstruct the first and second data, the present application must ensure that the input data is floating-point numbers, as integer types do not support rounding error compensation in floating-point operations. Furthermore, correctly updating the error term after each addition is crucial for the Kahan accumulator. Furthermore, while the Kahan accumulator improves accuracy, it may slightly reduce computational speed, requiring additional computation to track and compensate for errors.
[0102] In this embodiment, the original audio data can be traditional 24-bit packaged audio data, with a dynamic range of 144 dB and a THD (total harmonic distortion) + N (noise) less than -110 dB, making it suitable for high-precision measurement. If the original audio data is directly truncated and transmitted to the second device without adding residual audio data, the resulting target audio data can only achieve a dynamic range of 96 dB, with a THD (total harmonic distortion) + N (noise) less than -60 dB, making it suitable only for voice communication.
[0103] This application generates target audio data by combining main sound data with residual sound data, and its dynamic range can reach 120 dB, while THD (total harmonic distortion) + N (noise) is less than -100 dB, which can realize professional audio production.
[0104] In the audio processing method provided in the embodiment of the present application, first compressed data and second compressed data transmitted by a preset first device are received, and the first compressed data and the second compressed data are decompressed respectively to obtain first data and second data, and finally the first data and the second data are recombined to obtain target audio data, thereby ensuring the fidelity and integrity of the audio during the transmission and processing process.
[0105] The embodiment of the present application further provides an audio processing device 400, which is used to execute any embodiment of the aforementioned audio processing method.
[0106] Specifically, see Figure 4 , Figure 4 4 is a schematic block diagram of an audio processing device 400 provided in an embodiment of the present application.
[0107] like Figure 4 As shown, the audio processing device 400 provided in the present application is applied in a first device, and the device includes a processing unit 410, a compression unit 420 and a sending unit 430.
[0108] The processing unit 410 is used to process the original audio data to obtain the main sound data and residual sound data of the original audio; the compression unit 420 is used to compress the main sound data to obtain the first compressed data, and compress the residual sound data to obtain the second compressed data; the sending unit 430 is used to send the first compressed data and the second compressed data to a preset second device for decompression and reorganization to generate the target audio data in the second device.
[0109] The audio processing device 400 provided in the embodiment of the present application can process the original audio data to obtain the main sound data and residual sound data of the original audio, and can apply a relatively simple audio compression algorithm to compress the main sound data and the residual sound data respectively to obtain first compressed data and second compressed data. Finally, the first compressed data and the second compressed data can be sent to a preset second device for decompression and reorganization to generate target audio data in the second device. This can achieve the goal of simplifying the audio compression algorithm to reduce resource consumption and improve real-time performance while avoiding a significant decline in audio quality and greatly improving user experience.
[0110] In some embodiments, the audio processing apparatus 400 may also be applied to a second device, and the apparatus further includes a receiving unit, a decompression unit, and a reassembly unit.
[0111] The receiving unit is used to receive the first compressed data and the second compressed data transmitted by the preset first device; the decompression unit is used to decompress the first compressed data and the second compressed data respectively to obtain the first data and the second data; the recombining unit is used to recombine the first data and the second data to obtain the target audio data.
[0112] The audio processing device 400 provided in the embodiment of the present application can receive first compressed data and second compressed data transmitted by a preset first device, and decompress the first compressed data and the second compressed data respectively to obtain first data and second data, and finally reorganize the first data and the second data to obtain target audio data, thereby ensuring the fidelity and integrity of the audio during the transmission and processing process.
[0113] It should be noted that those skilled in the art can clearly understand that the specific implementation process of the above-mentioned audio processing device 400 and each unit can refer to the corresponding description in the aforementioned method embodiment. For the convenience and brevity of description, it will not be repeated here.
[0114] The above-mentioned audio processing device 400 can be implemented in the form of a computer program. The computer program can be used in Figure 5 Runs on the electronic devices shown.
[0115] See also Figure 5 , Figure 5 It is a schematic block diagram of an electronic device 500 provided in an embodiment of the present application.
[0116] See Figure 5 The electronic device 500 includes a processor 502 , a memory, and a network interface 505 connected via a system bus 501 , wherein the memory may include a storage medium 503 and an internal memory 504 .
[0117] The storage medium 503 may store an operating system 5031 and a computer program 5032. When the computer program 5032 is executed, the processor 502 may execute an audio processing method.
[0118] The processor 502 is used to provide computing and control capabilities to support the operation of the entire electronic device 500.
[0119] The internal memory 504 provides an environment for the operation of the computer program 5032 in the non-volatile storage medium 503. When the computer program 5032 is executed by the processor 502, the processor 502 can execute the audio processing method.
[0120] The network interface 505 is used for network communication, such as providing data information transmission. Those skilled in the art will understand that Figure 5 The structure shown in the figure is merely a block diagram of a portion of the structure related to the solution of the present application, and does not constitute a limitation on the device 500 to which the solution of the present application is applied. The specific device 500 may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0121] Among them, the processor 502 is used to run the computer program 5032 stored in the memory to achieve the following functions: processing the original audio data to obtain the main sound data and residual sound data of the original audio; compressing the main sound data to obtain first compressed data, and compressing the residual sound data to obtain second compressed data; sending the first compressed data and the second compressed data to a preset second device for decompression and reorganization to generate target audio data in the second device.
[0122] The processor 502 is used to run the computer program 5032 stored in the memory, and can also implement the following functions: receiving first compressed data and second compressed data transmitted by a preset first device; decompressing the first compressed data and the second compressed data respectively to obtain first data and second data; and recombining the first data and the second data to obtain target audio data.
[0123] Those skilled in the art will understand that Figure 5 The embodiment of the electronic device 500 shown in the figure does not constitute a limitation on the specific structure of the electronic device 500. In other embodiments, the electronic device 500 may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently. For example, in some embodiments, the electronic device 500 may only include a memory and a processor 502. In such an embodiment, the structure and function of the memory and processor 502 are the same as those in the figure. Figure 5 The embodiments shown are consistent and will not be described again here.
[0124] It should be understood that in the embodiments of the present application, the processor 502 may be a central processing unit (CPU), another general-purpose processor 502, a digital signal processor 502 (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor 502 may be a microprocessor 502 or any conventional processor 502.
[0125] In some embodiments, the embodiments of the present application also provide an audio transmission system, which includes a first device and a second device, and audio data is transmitted between the first device and the second device. The first device is configured to perform: processing the original audio data to obtain the main sound data and residual sound data of the original audio; compressing the main sound data to obtain first compressed data, and compressing the residual sound data to obtain second compressed data; sending the first compressed data and the second compressed data to a preset second device for decompression and reorganization to generate target audio data in the second device.
[0126] The second device is configured to perform: receiving first compressed data and second compressed data transmitted by a preset first device; decompressing the first compressed data and the second compressed data respectively to obtain first data and second data; and recombining the first data and the second data to obtain target audio data.
[0127] According to one aspect of the present application, a computer program product or computer program is also provided, the computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, causing the electronic device to implement the following steps: processing original audio data to obtain main sound data and residual sound data of the original audio; compressing the main sound data to obtain first compressed data, and compressing the residual sound data to obtain second compressed data; and sending the first compressed data and the second compressed data to a preset second device for decompression and reassembly, so as to generate target audio data in the second device.
[0128] The processor of the electronic device reads the computer instructions from a computer-readable storage medium, and the processor executes the computer instructions, so that the electronic device can also implement the following steps: receiving first compressed data and second compressed data transmitted by a preset first device; decompressing the first compressed data and the second compressed data respectively to obtain first data and second data; and recombining the first data and the second data to obtain target audio data.
[0129] Those skilled in the art will appreciate that all or part of the steps in the method of the above-described embodiment can be implemented by instructing the relevant hardware through a computer program. The computer program includes program instructions, which can be stored in a storage medium that is computer-readable. The program instructions are executed by at least one processor in the computer system to implement the steps in the method of the above-described embodiment.
[0130] In another embodiment of the present application, a computer storage medium is provided. The storage medium may be a non-volatile computer-readable storage medium or a volatile storage medium. The storage medium stores a computer program 5032, wherein when the computer program 5032 is executed by the processor 502, the following steps are implemented: processing original audio data to obtain primary sound data and residual sound data of the original audio; compressing the primary sound data to obtain first compressed data, and compressing the residual sound data to obtain second compressed data; and sending the first compressed data and the second compressed data to a pre-set second device for decompression and reassembly, thereby generating target audio data in the second device.
[0131] When the computer program 5032 is executed by the processor 502, the following steps can also be implemented: receiving first compressed data and second compressed data transmitted by a preset first device; decompressing the first compressed data and the second compressed data respectively to obtain first data and second data; and recombining the first data and the second data to obtain target audio data.
[0132] The storage medium can be any computer-readable storage medium that can store program code, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a magnetic disk, or an optical disk.
[0133] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in terms of function in the above description. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0134] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of each unit is merely a logical functional division, and other division methods may be used in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not implemented.
[0135] The steps in the method of the embodiment of the present application can be adjusted in order, combined, and deleted according to actual needs. The units in the device of the embodiment of the present application can be combined, divided, and deleted according to actual needs. In addition, the functional units in the various embodiments of the present application can be integrated into a processing unit, or each unit can exist physically separately, or two or more units can be integrated into a single unit.
[0136] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the existing technology, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for enabling an electronic device (which can be a personal computer, terminal, or network device, etc.) to perform all or part of the steps of the method provided in each embodiment of the present application.
[0137] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present application, and such modifications or substitutions should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. An audio processing method, characterized in that: Applied to a first device, the method includes: Processing the original audio data to obtain main sound data and residual sound data of the original audio; compressing the main sound data to obtain first compressed data, and compressing the residual sound data to obtain second compressed data; Sending the first compressed data and the second compressed data to a preset second device for decompression and reassembly, so as to generate target audio data in the second device; The compressing the residual sound data to obtain second compressed data includes: compressing the residual sound data using a preset residual dynamic mapping model to obtain the second compressed data; The method of using a preset residual dynamic mapping model to compress the residual sound data to obtain the second compressed data includes: performing dithering processing on the residual sound data to obtain dithered residual sound data; and performing shifting processing on the dithered residual sound data to obtain the second compressed data.
2. The audio processing method according to claim 1, wherein: The processing of the original audio data to obtain the main sound data and the residual sound data of the original audio includes: Processing the original audio data to obtain the main sound data; The residual audio data is generated based on the main audio data.
3. The audio processing method according to claim 2, wherein: The processing of the original audio data to obtain the main sound data includes: Mapping the original audio data into a preset integer space to generate audio data in an integer format; The main sound data is generated based on the audio data in the integer format.
4. The audio processing method according to claim 3, characterized in that: The step of generating the main sound data based on the audio data in the integer format includes: The audio data in the integer format is subjected to amplitude limiting processing within a preset first dynamic range to obtain the main sound data.
5. The audio processing method according to claim 3, wherein: The generating the residual sound data based on the main sound data includes: Acquire the audio data in integer format and the main sound data; The residual sound data is generated based on the audio data in the integer format and the main sound data.
6. The audio processing method according to claim 5, characterized in that: The step of generating the residual sound data according to the audio data in integer format and the main sound data includes: generating initial parameter sound data according to the audio data in integer format and the main sound data; The initial parameter sound data is subjected to a limiting process within a preset second dynamic range to obtain the residual sound data.
7. An audio processing method, characterized in that: Applied to the second device, the method includes: Receiving first compressed data and second compressed data transmitted by a preset first device; the first compressed data and the second compressed data are obtained by processing the audio processing method according to any one of claims 1 to 6; Decompressing the first compressed data and the second compressed data respectively to obtain first data and second data; The first data and the second data are recombined to obtain target audio data.
8. The audio processing method according to claim 7, characterized in that: The recombining the first data and the second data to obtain target audio data includes: The first data and the second data are reassembled using a preset accumulator to obtain the target audio data.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the audio processing method according to any one of claims 1 to 8 is implemented.
10. An audio transmission system, characterized in that: The invention comprises a first device and a second device, wherein audio data is transmitted between the first device and the second device, the first device is configured to execute the audio processing method according to any one of claims 1 to 6, and the second device is configured to execute the audio processing method according to any one of claims 7 to 8.
Citation Information
Patent Citations
Perceptually weighted digital audio level compression
US20090116664A1
Audio compression device, audio compression system, and audio compression method
US20240112688A1