Method and system for full link low distortion original sound restoration from digital read to analog output
Patent Information
- Application Number
- CN202610987739.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-03
- Publication Date
- 2026-08-04
AI Technical Summary
[0005]本发明提供了从数字读取到模拟输出的全链路低失真原声还原方法及系统,用于解决数字媒体播放器中因采用固定信号处理流水线而导致高规格音源本征品质被后续处理环节额外劣化且系统无法在直通路径与补偿路径之间依据实时失真水平进行闭环动态选择的问题
[0022] The technical solution of this invention first parses the encoded feature code in the audio file header and the bit depth, sampling rate, number of channels and other fields in the frame structure header information, and matches them with pre-built feature templates and threshold rules to automatically classify the audio source into four format categories. Clear classification labels can be obtained at the output end of the playback engine, providing a basis for differentiated decision-making for subsequent processing paths.
Smart Images

Figure CN122511265A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of electroacoustic conversion and audio signal processing, specifically to a method and system for end-to-end low-distortion original sound restoration from digital reading to analog output. Background Technology
[0002] In the field of high-end digital media players, after the audio signal is read from the storage medium, it must sequentially pass through the decoding and encapsulation of the playback engine, the sampling rate conversion and room acoustic calibration in the digital signal processor, the digital-to-analog domain conversion in the digital-to-analog converter, the amplitude control of the R2R resistor network or analog volume chip, and the buffer drive of the analog preamplifier output stage, before finally being sent to the power amplifier or directly driving the transducer load. In the above complete chain, the transmission characteristics of each link will deviate from the original digital audio data, and the amount of deviation will show a differential cumulative effect depending on the audio source format, bit depth, and sampling rate.
[0003] Current digital media players generally employ a fixed-sequence signal processing pipeline in their system architecture. After the playback engine decodes the audio file, the data unconditionally enters the digital signal processor (DSP), which sequentially converts the input sampling rate to a fixed rate under the DAC's master clock domain via an asynchronous sampling rate converter. Then, it passes through multiple stages of digital filters for oversampling and out-of-band noise suppression, while simultaneously adding frequency domain compensation coefficients for room acoustic calibration. The digital-to-analog converter (DAC) typically operates at its highest supported bit width and fixed modulation mode. The attenuation of the R2R resistor network or analog volume chip is uniquely determined by the user's volume setting, and the analog preamplifier output stage maintains a constant bias current and gain structure. This fixed pipeline, when processing high-quality audio sources, cannot determine whether the current pass-through truly achieves the low-distortion target, nor can it autonomously switch back to a compensation protection path when distortion worsens. This forces the player to operate with a uniform compromise path when facing audio sources of varying quality, making it difficult to achieve both high fidelity and reliability simultaneously.
[0004] In summary, within the actual product chain layout of digital media players, how to construct a low-distortion original sound restoration mechanism that dynamically selects a direct or compensated path based on the audio source format category and achieves closed-loop path back-cutting, around the complete chain of digital reading, playback engine, clock synchronization, DSP and room calibration, DAC, R2R analog volume, and analog preamp output, is an urgent problem to be solved in this field. Summary of the Invention
[0005] This invention provides a method and system for end-to-end low-distortion original sound restoration from digital reading to analog output, which solves the problem that the inherent quality of high-specification audio sources is additionally degraded by subsequent processing stages due to the use of a fixed signal processing pipeline in digital media players, and that the system cannot make a closed-loop dynamic selection between the direct path and the compensation path based on the real-time distortion level.
[0006] In view of the above problems, the present invention provides a method and system for end-to-end low-distortion original sound restoration from digital reading to analog output.
[0007] In a first aspect, the present invention provides a method for end-to-end low-distortion original sound restoration from digital reading to analog output, including:
[0008] The encoding format and original bit depth of the target lossless audio file are analyzed, and the audio source format category is identified.
[0009] Based on the audio source format category, the processing mode of the digital signal processing link is determined, wherein the processing mode includes at least a pass-through mode and a compensation mode;
[0010] When configured in pass-through mode, the native bit-width pass-through state of the digital-to-analog converter and the zero-attenuation pass-through state of the analog output stage are activated simultaneously, and the data path connection between the sampling rate converter and the digital filter in the digital signal processor is disconnected.
[0011] The analog output signal processed by the analog output stage is acquired in real time, and the end-to-end distortion parameters are determined based on the analog output signal.
[0012] Based on the original bit depth, the intrinsic distortion limit is determined, and the distortion deviation is obtained by combining the end-to-end distortion parameters.
[0013] When the distortion deviation exceeds the dynamic low distortion tolerance threshold, a path back-switching process is triggered to dynamically switch the signal link from the direct mode to the compensation mode, and the working status of the digital-to-analog converter and the analog output stage are adjusted accordingly.
[0014] Secondly, this invention provides a low-distortion original sound restoration system covering the entire chain from digital reading to analog output, including:
[0015] The audio source analysis and classification module is used to analyze the encoding format and original bit depth of the target lossless audio file and identify its audio source format category;
[0016] The processing mode decision module is used to determine the processing mode of the digital signal processing link according to the audio source format category, wherein the processing mode includes at least a pass-through mode and a compensation mode.
[0017] The pass-through mode link configuration module is used to simultaneously activate the native bit-width direct drive state of the digital-to-analog converter and the zero-attenuation pass-through state of the analog output stage when configured as pass-through mode, and disconnect the data path connection between the sampling rate converter and the digital filter in the digital signal processor.
[0018] The end-to-end distortion monitoring module is used to acquire the analog output signal after processing by the analog output stage in real time, and determine the end-to-end distortion parameters based on the analog output signal.
[0019] The distortion deviation evaluation module is used to determine the intrinsic distortion limit based on the original bit depth, and to obtain the distortion deviation by combining the end-to-end distortion parameters.
[0020] The dynamic path back-cutting and linkage adjustment module is used to trigger the path back-cutting process when the distortion deviation exceeds the dynamic low distortion tolerance threshold, dynamically switch the signal link from the direct mode to the compensation mode, and adjust the working state of the digital-to-analog converter and the analog output stage in linkage.
[0021] One or more technical solutions provided in this invention have at least the following technical effects or advantages:
[0022] The technical solution of this invention first parses the encoded feature code in the audio file header and the bit depth, sampling rate, number of channels and other fields in the frame structure header information, and matches them with pre-built feature templates and threshold rules to automatically classify the audio source into four format categories. Clear classification labels can be obtained at the output end of the playback engine, providing a basis for differentiated decision-making for subsequent processing paths.
[0023] Furthermore, differentiated processing is performed based on the audio source format category. A signal processing path that precisely matches the intrinsic quality of the audio source is established between the digital playback engine and the DAC, allowing high-specification audio sources to retain their original waveform information, while audio sources requiring system assistance receive appropriate compensation processing.
[0024] Furthermore, in pass-through mode, writing the original bit depth to the DAC bit width configuration register achieves hardware-level locking of the quantization bit depth; writing the sampling frequency to the clock divider register achieves phase-locking between the master clock and the audio clock at integer multiples; writing 0 to the analog output stage volume attenuation register and bypassing the output buffer achieves zero-attenuation unity-gain pass-through of the analog link; and switching the data selector to the first path and disabling the digital filter enable register within the DSP achieves complete bypassing and silencing of the digital domain processing module. This constructs a bit-transparent, unprocessed pass-through link directly from the playback engine to the analog preamp output, eliminating the combined effects of multiple distortion sources such as DSP interpolation errors, filter ringing, R2R resistor network attenuation noise, and buffer amplifier nonlinearity.
[0025] Furthermore, at the analog output stage backend, the actual analog signal is synchronously sampled with high precision at the original sampling frequency to obtain the monitoring sampling sequence. The original data for the corresponding time period is extracted from the playback engine cache and a reference sampling sequence is obtained through the same downmixing rules. The alignment offset is calculated using a cross-correlation function to complete the temporal alignment. The two sequences are subtracted point by point to obtain the residual sequence. The ratio of the root mean square value of the residual to the root mean square value of the reference is used as the end-to-end distortion parameter. This allows for real-time quantization of the actual distortion level of the link from the analog domain output, eliminating temporal misalignment pseudo-distortion caused by link delay and providing an accurate and reliable measurement benchmark for subsequent decision-making.
[0026] Furthermore, the original bit depth is substituted into the theoretical quantization noise model to calculate the root mean square value of the quantization noise. This value is then compared with the root mean square value of the full-amplitude sine wave to obtain the intrinsic distortion limit. The distortion deviation is then compared with the intrinsic distortion limit to obtain the distortion deviation. The degree of link fidelity degradation is intuitively evaluated by the multiple of the deviation from the intrinsic distortion limit, providing a clear trigger criterion for dynamic threshold determination.
[0027] Finally, using the intrinsic distortion threshold as the baseline tolerance, three correction factors are generated by logarithmically scaling the deviations of sampling frequency, bit depth, and number of channels from their respective thresholds. These factors are then merged into a comprehensive correction value, dynamically calculating a low distortion tolerance threshold suitable for the current audio source specifications. When the distortion deviation exceeds this threshold, a path back-switching is triggered, simultaneously executing the following actions: switching the DSP from bypass to compensated data selector, converting the DAC from native bit-width direct drive to compensated modulation mode, and restoring the analog output stage from zero-attenuation pass-through to controlled gain state. By incorporating the intrinsic quality of the audio source, measured distortion in the link, and engineering feasibility into a closed-loop decision framework, the system can autonomously backtrack to a compensated path with correction capabilities when distortion deteriorates in the pass-through path.
[0028] In summary, the technical solution of this invention, under the actual product link layout of a digital media player, enables high-specification audio sources to be transmitted to the subsequent stage with minimal loss in the direct path, while standard-specification and multi-channel audio sources obtain appropriate digital domain pre-correction in the compensation path. At the same time, closed-loop monitoring ensures that link distortion is controllable and recoverable under any operating condition, achieving a dynamic optimal balance between fidelity and reliability across the entire range of audio source types. Attached Figure Description
[0029] Figure 1 This is a flowchart illustrating the end-to-end low-distortion original sound restoration method provided by the present invention, which involves digital reading and analog output.
[0030] Figure 2 This is a schematic diagram of the trigger path back-switching logic in the end-to-end low-distortion original sound restoration method from digital reading to analog output provided by the present invention.
[0031] Figure 3This is a schematic diagram of the structure of the end-to-end low-distortion original sound restoration system from digital reading to analog output provided by the present invention.
[0032] In the attached diagram, the labels representing each component are as follows:
[0033] The module includes: 11 Sound source analysis and classification module, 12 Processing mode decision module, 13 Direct mode link configuration module, 14 Full link distortion monitoring module, 15 Distortion deviation evaluation module, and 16 Dynamic path back-cut and linkage adjustment module. Detailed Implementation
[0034] This invention provides a method and system for low-distortion original sound restoration across the entire chain from digital reading to analog output. It solves the problem that the inherent quality of high-specification audio sources is additionally degraded by subsequent processing stages due to the use of a fixed signal processing pipeline in digital media players, and that the system cannot make a closed-loop dynamic selection between the direct path and the compensation path based on the real-time distortion level.
[0035] It should be noted that the terms "comprising" and "having" are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or server that includes a series of steps or units is not necessarily limited to those steps or units that are explicitly listed, but may include other steps or modules that are not explicitly listed or that are inherent to these processes, methods, products, or devices.
[0036] Example 1, as Figure 1 As shown, this invention provides a method for end-to-end low-distortion original sound restoration from digital reading to analog output, the method comprising:
[0037] S100: Analyzes the encoding format and original bit depth of the target lossless audio file, and identifies the audio source format category to which it belongs.
[0038] After the playback engine completes the audio file decoding, this step extracts the encoding format identifier and the original bit depth field from the file header information, metadata, and data frame structure, and determines the audio source format category to which the file belongs based on the preset audio source classification rules, providing a basis for subsequent processing mode decisions.
[0039] Step S100 in the method provided by the present invention includes:
[0040] Read the file header extension identifier of the target lossless audio file and extract the encoding feature code from the encoding type field;
[0041] The encoding format is determined by matching the encoded feature code with the first feature template of the pulse code modulation format and the second feature template of the direct stream digital encoding format.
[0042] The value of the sampling bit depth field is extracted from the frame structure header information of the target lossless audio file and used as the original bit depth;
[0043] When the encoding format is pulse code modulation, the channel number tag and sampling rate tag in the frame structure header information are parsed to obtain the number of channels and the sampling frequency.
[0044] When the sampling frequency exceeds the first frequency threshold or the original bit depth exceeds the first depth threshold, the audio source format category is determined to be high-magnification sampling lossless type;
[0045] When the sampling frequency does not exceed the first frequency threshold and the original bit depth does not exceed the first bit depth threshold, if the number of channels exceeds the first channel threshold, the audio source format category is determined to be multi-channel interleaved lossless; otherwise, the audio source format category is determined to be standard bit depth lossless.
[0046] When the encoding format is a direct stream digital encoding format, the audio source format category is determined to be a one-bit stream native lossless class, and the original bit depth is set to 1.
[0047] In this step, the header extension identifier of the target lossless audio file is first read, and the encoding feature code is extracted from the encoding type field. For example, when reading the header of a FLAC format audio file, the encoding type field is located in the Vorbis Comment extension identifier area, and the encoding feature code 0x0001 is extracted. FLAC (Free Lossless Audio Codec) is a lossless audio compression encoding format that can compress audio files to 50% to 60% of their original size without loss.
[0048] Next, the encoded feature codes are matched with the first feature template of the Pulse Code Modulation (PCM) format and the second feature template of the Direct Stream Digital (DSD) encoding format to determine the encoding format. The first feature template refers to a pre-stored set of feature codes corresponding to PCM encoding; for example, the PCM feature code for FLAC packages is 0x0001, for WAV packages it is 0x0001, and for APE packages it is 0x1000, etc. The second feature template refers to a pre-stored set of feature codes corresponding to DSD encoding; for example, the DSD feature code for DSF packages is 0x0000, and for DFF packages it is 0x0200, etc. These templates are pre-built and stored in non-volatile memory before the system leaves the factory.
[0049] Among them, PCM (Pulse Code Modulation) is the most basic encoding method for converting analog signals into digital signals, representing continuous waveforms as discrete numerical sequences through equally spaced sampling and quantization. DSD (Direct Stream Digital) is an audio encoding method that samples at a high rate of 1 bit, encoding analog waveforms into a single-bit data stream with an ultra-high sampling rate through delta-sigma modulation. DSF (DSD Stream File) is a file encapsulation format used to store DSD audio data. DFF (DSD Interchange File Format) is a file format used for exchanging DSD audio data.
[0050] For example, the feature code 0x0001 is compared one by one with the PCM feature code entries in the first feature template, and a match is found for the corresponding record of FLAC encapsulated PCM; however, a match is not found with the DSD feature code entries in the second feature template. Therefore, the encoding format is determined to be pulse code modulation.
[0051] Next, the sample bit depth field value is extracted from the frame structure header information of the target lossless audio file as the raw bit depth. The frame structure header information refers to a set of structured fields carried at the beginning of each audio data frame, used to describe the encoding parameters of the audio data within that frame. For FLAC files, it includes the bits per sample (sample bit depth) field in the frame header and the channels (number of channels) and sample rate (sample rate) fields in the stream information block; for WAV files, it corresponds to the wBitsPerSample, nChannels, and nSamplesPerSec fields in the PCM format block; for DSF / DFF files, it corresponds to the channel number and sample rate information directly recorded in the file header. By parsing these fields at fixed offset positions frame by frame, key parameters such as the raw bit depth, number of channels, and sampling frequency of the current audio stream are obtained.
[0052] For example, in the frame structure header of a FLAC file, the bits per sample field indicates that the sample bit depth is 24, meaning the original bit depth of the audio file is 24 bits.
[0053] Next, when the encoding format is Pulse Code Modulation (PCM), the channel number tag and sample rate tag in the frame structure header are parsed to obtain the number of channels and the sampling frequency. For example, further parsing the channels tag in the frame structure header yields a channel number of 2, and the sample rate tag yields a sampling frequency of 192000Hz, meaning the audio is stereo with a 192kHz sampling rate.
[0054] Furthermore, based on the extracted parameters, the audio source format category is determined step by step according to thresholds. Specifically, there are three thresholds: the first deep threshold, determined based on the maximum linear quantization bit width natively supported by current mainstream DAC hardware. When the R2R DAC natively supports 24-bit linear quantization, the first deep threshold is set to 24 bits; if the DAC natively supports 20 bits, the threshold is set to 20 bits.
[0055] Among them, the R2R type DAC (Resistor-to-Resistor Digital-to-Analog Converter) is a digital-to-analog converter architecture that uses a precision resistor ladder network to directly convert digital codes into analog voltages, and has high linearity and native bit-width direct drive capability.
[0056] The first frequency threshold is determined based on the Nyquist sampling requirement of the upper limit of human hearing and the general standard for high-resolution audio. The upper limit of human hearing is approximately 20kHz, corresponding to a Nyquist sampling frequency of 40kHz; the general standard for high-resolution audio uses 88.2kHz as the starting point. Since 40kHz only satisfies the theoretical audible range reconstruction, while 88.2kHz represents the actual starting point for high-resolution audio products, 88.2kHz is chosen as the first frequency threshold to distinguish between standard resolution and high-resolution audio.
[0057] The first channel threshold is determined based on the standard channel configuration of a general stereo playback device, with a typical value of 2. When the number of channels exceeds 2, it is determined to be a multi-channel audio source.
[0058] It should be noted that the above three thresholds are preset and stored in non-volatile memory before the device leaves the factory, based on the actual DAC hardware capabilities, and serve as preset parameters during device power-on initialization.
[0059] In this step, the specific judgment logic is as follows: when the sampling frequency exceeds the first frequency threshold or the original bit depth exceeds the first depth threshold, the audio source format category is determined to be high-sampling lossless. For example, using the FLAC audio file from the previous example, the sampling frequency is 192000Hz, which is greater than the first frequency threshold of 88200Hz. The condition is met, therefore the audio source format category of the FLAC audio file is determined to be high-sampling lossless.
[0060] If the sampling frequency does not exceed the first frequency threshold and the original bit depth does not exceed the first depth threshold, then the number of channels is further determined: if the number of channels exceeds the first channel threshold, it is determined to be a multi-channel interleaved lossless type; otherwise, it is determined to be a standard bit depth lossless type.
[0061] In another possible embodiment, if an audio file has a sampling frequency of 44100Hz, a bit depth of 16 bits, and 6 channels, it does not meet the conditions for high-sampling lossless audio, and the number of channels (6) exceeds the first channel threshold of 2, thus it is classified as multi-channel interleaved lossless audio. If the number of channels is 2, i.e., it does not exceed the first channel threshold, it is classified as standard bit depth lossless audio.
[0062] Furthermore, when the encoding format is a direct stream digital encoding format, the audio source format category is determined to be a one-bit stream native lossless class, and the original bit depth is set to 1 bit. For example, when reading an audio file in DSF container format, the encoded feature code matches the second feature template, which determines that it is a direct stream digital encoding format. The audio source format category is determined to be a one-bit stream native lossless class, and the original bit depth is forcibly set to 1 bit. Subsequent processing links will adapt to the DSD native pass-through path accordingly.
[0063] It should be noted that the above values are for illustrative purposes only and do not constitute a limitation on the present invention.
[0064] In summary, this step can automatically distinguish between four types of audio source formats: high-sampling lossless, standard bit-depth lossless, multi-channel interleaved lossless, and bit-stream native lossless. This provides a clear and reusable classification basis for the subsequent digital signal processing link to accurately select between pass-through and compensation modes, thereby avoiding unnecessary numerical calculation distortion caused by sending high-specification audio sources into the DSP processing pipeline in a unified manner.
[0065] S200: Determine the processing mode of the digital signal processing link according to the audio source format category, wherein the processing mode includes at least a pass-through mode and a compensation mode.
[0066] This step enables the digital signal processing link to adopt differentiated processing strategies based on the intrinsic quality differences of the audio source. For audio sources with high fidelity potential, a pass-through mode is used to completely preserve their original waveform information, while for audio sources that need to be improved in fidelity, a compensation mode is used to apply digital domain pre-correction, thereby achieving a reasonable match between processing depth and audio source quality at the entire link level.
[0067] Step S200 in the method provided by the present invention includes:
[0068] When the audio source format category is standard bit-depth lossless, the processing mode is determined to be pass-through mode, and the bit width locking value corresponding to the native bit width direct drive state of the digital-to-analog converter is verified to be consistent with the original bit depth.
[0069] When the audio source format category is high-sampling lossless, the processing mode is determined to be compensation mode, and the first ratio of the sampling frequency to the first frequency threshold and the second ratio of the original bit depth to the first bit depth threshold are calculated.
[0070] The larger of the first ratio and the second ratio is selected, and the sampling rate converter is controlled to perform resampling according to the larger value;
[0071] When the audio source format category is multi-channel interleaved lossless, the processing mode is determined to be the compensation mode, and the third ratio of the number of channels to the first channel threshold is calculated.
[0072] The value obtained by rounding up the third ratio is used to perform downmixing on the number of channels to obtain the sampled values of the left output channel and the right output channel.
[0073] When the audio source format category is a one-bit stream native lossless type, check whether the digital-to-analog converter supports native pass-through decoding of direct stream digital encoding format;
[0074] If supported, the processing mode is determined to be the pass-through mode; otherwise, the processing mode is determined to be the compensation mode, and the order of the digital filter is configured to be twice the original bit depth.
[0075] In this step, the determined audio source format category is first obtained to determine the processing mode of the digital signal processing link. When the audio source format category is standard bit-depth lossless, the processing mode is determined to be pass-through mode, and the bit-width lock value corresponding to the DAC's native bit-width direct drive state is verified to be consistent with the original bit depth.
[0076] Among them, DAC (Digital-to-Analog Converter) is an integrated circuit chip that converts discrete digital codes into continuous analog voltage or current signals.
[0077] Native bit-width direct drive mode is the operating mode in which the exponential analog-to-digital converter (DAC) directly receives the input digital code and performs digital-to-analog conversion using its original physical bit width, after bypassing the internal digital filter and sampling rate conversion unit. In this mode, the DAC no longer performs internal truncation or padding on the input data; each bit of digital input corresponds to a one-bit physical switch drive signal.
[0078] The bit-width lock value refers to the bit-width parameter written to the DAC configuration register when the native bit-width direct drive mode is activated, forcing the DAC's input interface to align to a specified number of bits, such as 16 bits, 20 bits, or 24 bits. After receiving the bit-width lock value, the valid bit segments of the internal shift register and latch are fixed, and the high-bit padding logic above the lock value and the low-bit truncation logic below the lock value are both disabled.
[0079] In one possible embodiment, if the audio file has a sampling frequency of 44.1kHz, a bit depth of 16 bits, and 2 channels, it is determined to be a standard bit depth lossless type. The processing mode is determined to be pass-through mode. The lock value of the DAC's native bit width direct drive is read as 16 bits, which is consistent with the original bit depth. After verification, the pass-through path is maintained in a ready state.
[0080] Furthermore, when the audio source format is of the high-sampling lossless type, the processing mode is determined to be the compensation mode. A first ratio of the sampling frequency to a first frequency threshold and a second ratio of the original bit depth to a first bit depth threshold are calculated. The larger of the first ratio and the second ratio is taken, and the integer obtained by rounding it up is used as the resampling factor. The sampling rate converter is then controlled to perform resampling according to this factor.
[0081] The sample rate converter (SRC) is a digital signal processing unit that changes the sampling rate of a digital audio signal using interpolation or decimation algorithms. It is a functional module within the digital signal processor used to convert the sampling rate of the input digital audio signal to another sampling rate. The sample rate converter is located between the playback engine output and the digital filter, receiving resampling ratio control parameters from the mode decision module.
[0082] For example, using the FLAC audio file scenario from the previous steps, the first frequency threshold is 88.2kHz, and the sampling frequency is 192kHz. Therefore, the first ratio = 192 / 88.2 ≈ 2.18. The first bit depth threshold is 24 bits, and the original bit depth is 24 bits, so the second ratio = 24 / 24 = 1. Taking the larger value 2.18, rounding it up to 3, controls the sample rate converter to perform resampling at 3 times the original sampling rate, i.e., resampling to 576kHz before sending it to the subsequent digital filter cascade in the compensation path.
[0083] Furthermore, when the audio source format is a multi-channel interleaved lossless type, the processing mode is determined to be the compensation mode, and the third ratio of the number of channels to the first channel threshold is calculated. The value of the third ratio is rounded up, and downmixing is performed on the number of channels to obtain the sampled values of the left output channel and the right output channel.
[0084] Specifically, the third ratio, rounded up, is used to perform downmixing on the number of channels, including:
[0085] The target number of downmixing channels is fixed at the threshold of the first channel.
[0086] Parse the channel configuration field in the frame structure header information of the target lossless audio file to obtain the type identifier of each channel. The type identifier includes at least the left channel, right channel, center channel, left surround channel, right surround channel and low frequency effect channel.
[0087] When the input channel type is identified as the left channel, the sampled value of the input channel is multiplied by the value 1 and then accumulated to the left output channel;
[0088] When the input channel type is identified as left surround channel, the sampled value of the input channel is multiplied by the left allocation coefficient and then accumulated to the left output channel. The left allocation coefficient is determined based on the 3dB sound pressure level attenuation principle.
[0089] When the input channel type is identified as the right channel, the sampled value of the input channel is multiplied by the value 1 and then accumulated to the right output channel;
[0090] When the input channel type is identified as right surround channel, the sampled value of the input channel is multiplied by the right allocation coefficient and then accumulated to the right output channel. The right allocation coefficient is determined based on the 3dB sound pressure level attenuation principle.
[0091] When the input channel type is identified as center channel, the sampled value of the input channel is multiplied by the center allocation coefficient and then accumulated to the left output channel and the right output channel respectively. The center allocation coefficient is determined based on the principle of equal distribution of dual channel energy.
[0092] When the input channel type is identified as a low-frequency effects channel and the current audio source format is a multi-channel music recording type, skip the input channel.
[0093] The value obtained by rounding up the third ratio is used as the amplitude normalization divisor. The accumulated value of the left output channel is divided by the amplitude normalization divisor to obtain the sampled value of the left output channel. The accumulated value of the right output channel is divided by the amplitude normalization divisor to obtain the sampled value of the right output channel.
[0094] Specifically, the number of target channels for downmixing is first fixed at the first channel threshold. The channel configuration field in the frame structure header information is parsed to obtain the type identifier of each channel. The type identifier includes at least the left channel, right channel, center channel, left surround channel, right surround channel, and low-frequency effects channel.
[0095] In one possible embodiment, a multi-channel FLAC file has a sampling frequency of 48kHz, a bit depth of 24 bits, and 6 channels. The third ratio = 6 / 2 = 3, rounded up to 3. The parsed channel configuration field results are: Channel 1 - Left channel, Channel 2 - Right channel, Channel 3 - Center channel, Channel 4 - Low-frequency effects channel, Channel 5 - Left surround channel, Channel 6 - Right surround channel.
[0096] Furthermore, weighted accumulation is performed according to channel type. For the left channel, the sampled value is multiplied by 1 and then accumulated to the left output channel. For the right channel, the sampled value is multiplied by 1 and then accumulated to the right output channel.
[0097] For the left surround channel, the sampled value is multiplied by the left allocation factor and then summed to the left output channel. The left allocation factor is determined based on a 3dB sound pressure level attenuation principle, i.e., left allocation factor = 10. (dB / 20) Substituting the -3dB sound pressure level attenuation, we get the left distribution coefficient = 10. (-3 / 20) The 3dB sound pressure level attenuation principle means that when surround channels are mixed in stereo, their sound pressure level contribution should be reduced by 3dB to simulate the difference in human perception between side and front sound sources. For the right surround channel, the sampled value is multiplied by the right allocation coefficient and then added to the right output channel. The right allocation coefficient has the same value as the left allocation coefficient, approximately 0.708.
[0098] For the center channel, the sampled value is multiplied by the center allocation factor and then accumulated to the left and right output channels respectively. The center allocation factor is determined based on the principle of equal energy distribution between the two channels, i.e., center allocation factor = 1 / 2≈0.707. The principle of equal energy distribution in dual-channel audio refers to distributing the energy of the center channel equally to the left and right channels, so that each channel receives a relative level of -3dB, thereby keeping the total sound power constant.
[0099] In addition, for low-frequency effect channels, they are skipped and not included in the downmixing accumulation in multi-channel audio sources for music recording, while in multi-channel audio sources for movie sound effects, they are accumulated to the left and right output channels respectively after being filtered by a 120Hz low-pass filter with a distribution coefficient of 0.1 to 0.3, in order to preserve the ultra-low frequency energy in movie special effects.
[0100] For example, at a certain sampling moment, the sampled values of the six channels are: left channel 0.5, right channel -0.3, center channel 0.2, low-frequency effects channel 0.1, left surround channel 0.15, and right surround channel -0.1. The accumulation process is as follows: left output accumulated value = 0.5 × 1 + 0.15 × 0.708 + 0.2 × 0.707 = 0.7476, right output accumulated value = (-0.3) × 1 + (-0.1) × 0.708 + 0.2 × 0.707 = -0.2294. In multi-channel audio sources for music recording, the low-frequency effects channel 0.1 is skipped and does not participate in the accumulation.
[0101] Furthermore, the value obtained by rounding up the third ratio is used as the amplitude normalization divisor. The left output channel accumulated value is divided by the amplitude normalization divisor to obtain the left output channel sample value, and the right output channel accumulated value is divided by the amplitude normalization divisor to obtain the right output channel sample value.
[0102] For example, continuing from the previous example of output accumulation value, with the amplitude normalization divisor being 3, the sample value of the left output channel = 0.7476 / 3 = 0.2492, and the sample value of the right output channel = -0.2294 / 3 = -0.0765. Thus, the 6-channel interleaved audio stream is downmixed into a 2-channel standard PCM stream.
[0103] Furthermore, when the audio source format is a native lossless bitstream, check if the DAC supports native DSD pass-through decoding. If it does, set the processing mode to pass-through mode; if it does not, set the processing mode to compensation mode and configure the digital filter order to twice the original bit depth.
[0104] Among them, DSD native pass-through decoding means that the exponential analog-to-digital converter has a dedicated DSD signal path inside. This path can directly receive a 1-bit DSD modulated data stream without any decimation filtering, bit width conversion or PCM recoding processing, and directly drive the back-end switched capacitor network or current rudder array to complete analog domain reconstruction.
[0105] Specifically, the method involves reading the DAC chip's device ID register or feature descriptor and checking if the native DSD support flag is set. If the flag is true, native DSD pass-through decoding is supported; if the flag is false or the DAC datasheet does not declare this feature, it is determined that it is not supported.
[0106] In another possible embodiment, a DSF file is determined to be a native lossless one-bit stream. The DAC capability register is queried to confirm that the current DAC supports native DSD pass-through decoding, and the processing mode is set to pass-through mode. Subsequent DSD data streams will be directly sent to the DAC's dedicated DSD path.
[0107] In another possible embodiment, if the DAC does not support native DSD pass-through, the processing mode is set to compensation mode, the original bit depth is 1, the digital filter order is configured as 1×2=2, and the 1-bit DSD stream is lightly filtered and converted into PCM before being sent to the DAC.
[0108] Finally, based on the processing mode determined by each branch, the mode selection signal is written to the DSP link configuration register, completing the unified setting of the bypass or enable states of all functional modules from the playback engine output to the DAC input. Taking high-sampling lossless as an example, the mode register is written with the compensation mode identifier, the sampling rate converter is enabled and locked at 3x resampling ratio, the digital filter is enabled, and the DAC is configured in standard PCM input mode. Taking standard bit-depth lossless as an example, the mode register is written with the pass-through mode identifier, the sampling rate converter is bypassed, the digital filter is bypassed, and the DAC is configured in native bit-width direct drive mode and locked at 16 bits.
[0109] It should be noted that the above values are for illustrative purposes only and do not constitute a limitation on the present invention.
[0110] In summary, this step establishes a signal processing path between the digital playback engine and the DAC that precisely matches the intrinsic quality of the audio source. This means that the data from high-specification audio sources goes directly to the analog conversion stage without undergoing digital domain operations, while audio sources requiring system assistance receive appropriate compensation processing, thus avoiding the additional degradation of high-fidelity audio sources caused by fixed pipelines at the source.
[0111] S300: When configured in pass-through mode, it simultaneously activates the native bit-width pass-through state of the digital-to-analog converter and the zero-attenuation pass-through state of the analog output stage, and disconnects the data path connection between the sampling rate converter and the digital filter in the digital signal processor.
[0112] This step constructs a transparent, unprocessed path from the digital playback engine to the analog preamp output in pass-through mode. This eliminates numerical truncation errors introduced by DSP interpolation, the broadening and contamination of the time-domain waveform caused by the ringing effect of digital filters, and the noise floor rise caused by the additional gain / attenuation of the analog volume network and buffer stage. This allows the original sampling point information of the high-specification audio source to be transmitted to the subsequent equipment with minimal loss.
[0113] Step S300 in the method provided by the present invention includes:
[0114] Configure the digital audio input format of the digital-to-analog converter according to the original bit depth, and set the internal processing precision of the digital-to-analog converter to be no less than the original bit depth;
[0115] After reading the status register of the digital-to-analog converter and confirming that the digital audio input format and internal processing precision configuration are effective, the sampling frequency is written to the clock divider register of the digital-to-analog converter, and the main clock of the digital-to-analog converter is configured to be an integer multiple of the sampling frequency.
[0116] Write the value 0 to the volume attenuation register of the analog output stage, and configure the output buffer of the analog output stage to bypass state;
[0117] The data selector connected to the sampling rate converter in the digital signal processor is switched to the first path, wherein the first path directly sends the input sampled value to the digital-to-analog converter, and at the same time writes the value 0 to the enable register of the digital filter.
[0118] In this step, the digital audio input format of the digital-to-analog converter is first configured according to the original bit depth (e.g., I...). 2 (Data word length in formats such as C, left-aligned, and right-aligned), and set the internal processing precision of the digital-to-analog converter to be no less than the original bit depth to ensure that the working bit width of the DAC's internal digital filter and Δ-Σ modulator can fully preserve the quantization precision of the original audio data.
[0119] Specifically, configuration is achieved by writing corresponding parameters to the DAC chip's digital audio interface control register and precision configuration register, which are typically mapped to the DAC chip's I / O pins. 2 In the C or SPI control address space.
[0120] Among them, I 2 C (Inter-Integrated Circuit) is a serial communication bus protocol, and SPI (Serial Peripheral Interface) is a full-duplex serial communication protocol. Both are commonly used for the exchange of control instructions and data between chips.
[0121] For example, the currently playing audio source is a standard bit-depth lossless audio source with an original bit depth of 16 bits. (This is achieved through I...) 2 The C control interface configures the DAC chip's digital audio input format to the I²S standard 16-bit data word length and sets the internal processing precision to 24 bits (not less than the original bit depth), enabling the DAC's internal digital filters and Δ-Σ modulator to operate with 24-bit precision, fully preserving the quantization information of the original 16-bit audio data.
[0122] Further, the DAC's status register is read to confirm that the digital audio input format and internal processing precision configuration are effective. Then, the sampling frequency is written to the DAC's clock divider register, configuring the DAC's main clock to be an integer multiple of the sampling frequency. The clock divider register is a programmable control register within the digital-to-analog converter's internal phase-locked loop or clock management unit, used to set the integer division ratio between the DAC's main clock and the input audio sampling frequency. It is typically mapped to the I / O pin of the DAC chip. 2 In the C or SPI control address space.
[0123] For example, after detecting that the digital audio input format and internal processing precision configuration are effective, the sampling frequency value 44100 is written to the clock divider register. Based on this, the DAC's internal phase-locked loop divides the externally provided 45.1584MHz master clock to 1024 times 44100Hz, so that the DAC's master clock and audio sampling frequency maintain a strict integer multiple phase-locked relationship, eliminating the periodic sampling jitter introduced by asynchronous clock domain conversion.
[0124] Furthermore, the volume attenuation register of the analog output stage is written with the value 0, and the output buffer of the analog output stage is configured to bypass mode. The volume attenuation register is a programmable control register inside the volume control chip of the R2R resistor network in the analog output stage, used to set the attenuation of the resistor network in the signal path. It is typically accessed via SPI or I / O. 2 The C interface allows for direct control of the conduction state combination of each switch in the R2R network, thereby changing the amplitude attenuation ratio of the signal after passing through the resistor divider network.
[0125] For example, the analog output stage uses an R2R resistor network for volume control, followed by an analog preamplifier. Writing a value of 0 to the attenuation register of the R2R volume control chip via the SPI interface switches all switches within the R2R network to the no-attenuation pass-through mode, meaning the signal passes directly without any resistor voltage division. Simultaneously, writing a bypass enable bit of 1 to the configuration register of the analog preamplifier shorts the unity-gain amplifier of the buffer via a relay, creating a direct copper trace between the input and output pins, achieving zero-attenuation unity-gain pass-through.
[0126] Furthermore, the data selector connected to the sampling rate converter in the digital signal processor (DSP) is switched to the first path. The first path directly sends the input sampled value to the DAC, while the enable register of the digital filter is written with the value 0.
[0127] The first path is the data selector inside the digital signal processor, which is one of the input channels of the 2-to-1 multiplexer. It directly connects the raw audio sample values output by the playback engine to the data input interface of the DAC without going through any digital domain processing modules such as the sampling rate converter and digital filter.
[0128] Enable registers are switch control registers corresponding to various functional modules within a digital signal processor (DSP). They are used to independently control the clock gating and data path activation states of modules such as the sampling rate converter and digital filters. They are typically mapped in the DSP's memory-mapped I / O address space, with each module corresponding to an independent enable bit.
[0129] For example, the data selector inside the DSP is a 2-to-1 multiplexer. The first path connects the playback engine output to the DAC input, and the second path connects the cascaded output of the sampling rate converter and digital filter to the DAC input. Writing a selection signal of 0 to the control register of the data selector enables the first path, allowing the 16-bit / 44.1kHz PCM data stream output from the playback engine to bypass the sampling rate converter and digital filter and be directly sent to the parallel data interface of the DAC. Simultaneously, writing a value of 0 to the enable register of the digital filter disables the clock gate of the digital filter, stopping the flipping of its internal multiply-accumulate units and completely eliminating the coupling interference of dynamic power consumption and switching noise introduced by digital domain operations to the analog reference ground.
[0130] It should be noted that the above values are for illustrative purposes only and do not constitute a limitation on the present invention.
[0131] In summary, this step eliminates numerical errors and ringing distortion introduced by DSP interpolation and filtering, and avoids additional noise and distortion caused by R2R resistor network attenuation and buffer amplifier nonlinearity, so that the original sampling information of high-specification audio source is transmitted to the subsequent equipment with minimal loss.
[0132] S400: Real-time acquisition of the analog output signal after processing by the analog output stage, and determination of the end-to-end distortion parameters based on the analog output signal.
[0133] During operation in either pass-through or compensation mode, this step continuously acquires real waveform data from the analog domain output and quantifies its distortion level, providing real-time and accurate measurement data for assessing the actual deviation of the current link from the intrinsic distortion limit in subsequent steps.
[0134] Step S400 in the method provided by the present invention includes:
[0135] When the encoding format is pulse code modulation, the analog output signal is sampled at the sampling frequency to obtain a monitoring sampling sequence;
[0136] Read the original audio sampling sequence that is the same as the monitoring sampling sequence time period, and obtain the reference sampling sequence after the same decoding and downmixing process as in the current signal link. If the number of channels is greater than the first channel threshold, the original audio sampling sequence is downmixed according to the same channel type allocation coefficient and amplitude normalization method; otherwise, the read original audio sampling sequence is used as the reference sampling sequence.
[0137] Calculate the cross-correlation function between the monitoring sampling sequence and the reference sampling sequence, and use the delay value corresponding to the peak value of the cross-correlation function as the alignment offset;
[0138] The monitoring sampling sequence is shifted by the alignment offset to align the monitoring sampling sequence with the reference sampling sequence in the time domain.
[0139] The aligned monitoring sampling sequence is subtracted from the reference sampling sequence point by point to obtain the residual sequence;
[0140] Calculate the root mean square value of the residual sequence and the root mean square value of the reference sampling sequence, and obtain the basic distortion ratio based on the ratio of the root mean square value of the residual to the root mean square value of the reference.
[0141] The base distortion ratio is used as the end-to-end distortion parameter.
[0142] In this step, when the encoding format is pulse code modulation, the analog output signal after processing by the analog output stage is sampled at the sampling frequency to obtain the monitoring sampling sequence.
[0143] For example, the currently playing audio source is a standard bit-depth lossless PCM file with a sampling frequency of 44.1kHz, a bit depth of 16 bits, and 2 channels, and is in pass-through mode. The high-precision analog-to-digital converter (ADC) at the back end of the analog output stage synchronously samples the analog output waveform at a sampling frequency of 44.1kHz, with a sampling bit depth of 32 bits to retain sufficient margin, continuously acquiring 1024 sampling points to obtain the left channel monitoring sampling sequence M.
[0144] Among them, ADC (Analog-to-Digital Converter) is an integrated circuit chip that converts continuous analog signals into discrete digital codes.
[0145] Furthermore, the original audio sample sequence, identical to the one used in the monitoring sample sequence, is read from the playback engine cache. This original audio sample sequence undergoes the same decoding process as in the current signal link, with an added delay consistent with the inherent delay of the current signal link (the delay is composed of DAC conversion delay, analog buffer propagation delay, and ADC conversion delay). If the number of channels exceeds the first channel threshold, the decoded audio sample sequence is downmixed using the same channel type allocation coefficients and amplitude normalization methods as in the previous steps; otherwise, the decoded audio sample sequence is directly used as the reference sample sequence. This process ensures that the reference sample sequence and the monitoring sample sequence are comparable in the time domain.
[0146] For example, the number of channels in the aforementioned standard bit-depth lossless PCM file is 2, which does not exceed the first channel threshold of 2, so no downmixing is required. The original PCM data corresponding to the re-sampling period is read from the playback engine's cache, and after being decoded in the same way as the current signal link, 1024 sampling points of the left channel are extracted to obtain the reference sampling sequence R. This reference sampling sequence represents the expected waveform of the original audio data after the link decoding and processing.
[0147] Furthermore, the cross-correlation function between the monitored sampling sequence and the reference sampling sequence is calculated, and the delay value corresponding to the peak of the cross-correlation function is used as the alignment offset. The cross-correlation value at delay m is equal to the sum of the product of the monitored sampling sequence M and the reference sampling sequence R at the corresponding offset positions. The final value of delay m is the alignment offset.
[0148] For example, the cross-correlation sequence of M and R within the delay range [-100, +100] is calculated. A maximum peak occurs at m=2, indicating that the monitoring sequence lags the reference sequence by two sampling points. This delay is jointly introduced by the DAC conversion delay, the analog buffer propagation delay, and the ADC conversion delay. The alignment offset is determined to be 2.
[0149] Furthermore, the monitoring sampling sequence is shifted by an alignment offset to align it with the reference sampling sequence in the time domain. For example, applying an alignment offset of 2 to the monitoring sequence shifts M forward by 2 sampling points. After the shift, the first 1022 valid sampling points of the monitoring sequence are aligned with the reference sequence R. The aligned valid interval is then truncated to obtain the aligned monitoring sequence M. a With reference sequence R a .
[0150] Furthermore, the aligned monitoring sampling sequence is subtracted from the reference sampling sequence point by point to obtain the residual sequence. For example, the subtraction operation is performed point by point on 1022 aligned sampling points to obtain the residual sequence E. Each value in the residual sequence represents the instantaneous deviation between the actual output and the ideal output of the link at that sampling moment, including the contributions of all non-ideal factors such as DAC nonlinearity, thermal noise, buffer distortion, and power supply ripple coupling.
[0151] Finally, the root mean square (RMS) values of the residual sequence and the reference sample sequence are calculated. Based on the ratio of the residual RMS value to the reference RMS value, the base distortion ratio is obtained. The resulting base distortion ratio is the end-to-end distortion parameter.
[0152] For example, the calculated root mean square value of the residual sequence is 0.00031, the root mean square value of the reference sample sequence is 0.124, and the base distortion ratio is approximately 0.00031 / 0.124 ≈ 0.0025. Converted to decibels, this is approximately -52 dB, which is the end-to-end distortion parameter, representing the overall distortion level. The end-to-end distortion parameter of 0.0025 is written to the distortion parameter field of the system status register.
[0153] It should be noted that the above values are for illustrative purposes only and do not constitute a limitation on the present invention.
[0154] In summary, this step obtains the actual distortion level of the current link in real time and quantitatively from the analog domain output, eliminating the pseudo-distortion component caused by the time-domain misalignment between the monitoring and reference sequences due to link delay, and providing an accurate and reliable measurement benchmark for subsequent distortion deviation assessment and path back-cutting decision.
[0155] S500: Based on the original bit depth, determine the intrinsic distortion limit, and in conjunction with the end-to-end distortion parameters, obtain the distortion deviation.
[0156] This step takes the original bit depth and end-to-end distortion parameters as input, calculates the theoretical minimum distortion level that can be achieved based on quantization theory as the intrinsic distortion limit, and then compares the end-to-end distortion parameters with the intrinsic distortion limit to quantify the degree of deviation of the current link's actual distortion from the theoretical limit, thus obtaining the distortion deviation.
[0157] Step S500 in the method provided by the present invention includes:
[0158] Substitute the original bit depth into the theoretical quantization noise model to calculate the root mean square value of quantization noise corresponding to the ideal quantization step size, and use the ratio of the root mean square value of quantization noise to the root mean square value of the full-amplitude sine wave as the intrinsic distortion limit.
[0159] The theoretical quantization noise model takes the original bit depth as input and outputs the root mean square value of the quantization noise based on the uniform quantization theory. The full-amplitude sine wave is a sine wave signal whose amplitude range covers the full scale of the digital domain.
[0160] Read the end-to-end distortion parameters and calculate the ratio of the end-to-end distortion parameters to the intrinsic distortion limit. Record the ratio as the distortion deviation.
[0161] In this step, the original bit depth is first substituted into the theoretical quantization noise model to calculate the root mean square value of the quantization noise corresponding to the ideal quantization step size. The theoretical quantization noise model is a statistical model of the quantization error under uniform quantization conditions, assuming that the quantization error follows a uniform distribution within the interval [-Δ / 2, +Δ / 2], where Δ is the quantization step size and Δ = full-scale voltage swing / 2. NN is the original bit depth. The formula for calculating the root mean square value of quantization noise is: Root mean square value of quantization noise = Δ / 12. For example, normalizing the full-scale voltage swing to 1, the original bit depth of the output is N=16. Substituting this into the model, the root mean square value of the quantization noise is calculated as 1 / (2 16 ×3.464)=4.41×10 -6 .
[0162] Furthermore, the root mean square (RMS) value of the sine wave signal whose amplitude range covers the full scale of the digital domain is calculated. The peak amplitude of the full-scale sine wave is 1, and its RMS value is the peak value divided by √2, that is, the RMS value of the full-scale sine wave is 0.707.
[0163] Furthermore, the ratio of the root mean square value of quantization noise to the root mean square value of the full-amplitude sine wave is used as the intrinsic distortion limit, representing the minimum possible distortion level introduced by finite-bit quantization under ideal digital-to-analog conversion conditions. Intrinsic distortion limit = root mean square value of quantization noise / root mean square value of the full-amplitude sine wave.
[0164] For example, when the original bit depth N=16, the intrinsic distortion limit = 1 / (2 16 × 6)≈6.24×10 -6 This translates to -98.08 dB, representing the theoretical minimum distortion level of a 16-bit PCM audio source under ideal conditions.
[0165] Finally, the output end-to-end distortion parameters are read, and the ratio of the end-to-end distortion parameters to the intrinsic distortion floor is calculated. This ratio is recorded as the distortion deviation. The ratio obtained above is dimensionless. A value of 1 indicates that the link distortion is exactly equal to the theoretical quantization noise level; the larger the value, the more the additional distortion introduced by the link exceeds the intrinsic distortion floor.
[0166] For example, in the previous example, the output end-to-end distortion parameter is 0.0025, and the intrinsic distortion floor is 6.24 × 10⁻⁶. -6 Therefore, the distortion deviation = 0.0025 / 6.24 × 10 -6 ≈400.6 indicates that the actual distortion of the current link is about 400 times greater than the theoretical limit of 16 bits, and there is a serious source of additional distortion in the link.
[0167] It should be noted that the above values are for illustrative purposes only and do not constitute a limitation on the present invention.
[0168] In summary, this step establishes a quantifiable comparative relationship between the actual distortion level of the link and the theoretical limit determined by the bit depth of the audio source itself, and can judge the degree of fidelity degradation of the current link by the multiple of deviation from the intrinsic limit.
[0169] S600: When the distortion deviation exceeds the dynamic low distortion tolerance threshold, a path back-cut process is triggered to dynamically switch the signal link from the through mode to the compensation mode, and the working status of the digital-to-analog converter and the analog output stage are adjusted accordingly.
[0170] like Figure 2 As shown, in this step, when the distortion deviation exceeds the dynamic low distortion tolerance threshold, the path reversal process is triggered. The calculation steps for the dynamic low distortion tolerance threshold include:
[0171] Obtain the intrinsic distortion limit and use the intrinsic distortion limit as the benchmark tolerance value;
[0172] Calculate the frequency deviation ratio between the sampling frequency and the first frequency threshold, and perform logarithmic scaling on the frequency deviation ratio to obtain the frequency correction factor;
[0173] Calculate the depth deviation ratio between the original bit depth and the first depth threshold, and perform logarithmic scaling on the depth deviation ratio to obtain the depth correction factor;
[0174] Calculate the channel deviation ratio between the number of channels and the first channel threshold, and perform logarithmic scaling on the channel deviation ratio to obtain the channel correction factor;
[0175] The logarithmic scaling is defined as follows: when the input deviation ratio is greater than 1, the output is the logarithm of the deviation ratio with base 2; when the input deviation ratio is less than or equal to 1, the output is 0.
[0176] The frequency correction factor, the depth correction factor, and the channel correction factor are fused to obtain a comprehensive correction amount, and combined with the reference tolerance value, a dynamic low distortion tolerance threshold is obtained.
[0177] In this step, the dynamic low distortion tolerance threshold is first determined. Specifically, the calculated intrinsic distortion floor is first obtained and used as the baseline tolerance value. The baseline tolerance value represents the theoretical distortion floor of the sound source under ideal conditions.
[0178] For example, the currently playing audio source is a standard bit-depth lossless PCM file with a sampling frequency of 44.1kHz, a bit depth of 16 bits, and 2 channels. The output intrinsic distortion threshold is 6.24 × 10⁻⁶. -6 The baseline tolerance value is 6.24 × 10⁻⁶. -6 .
[0179] Furthermore, the frequency deviation ratio between the sampling frequency and the first frequency threshold is calculated, and the ratio is logarithmically scaled to obtain the frequency correction factor. The logarithmic scaling rule is as follows: when the deviation ratio > 1, log2(deviation ratio) is output; when the deviation ratio ≤ 1, 0 is output.
[0180] For example, if the sampling frequency is 44100Hz, the first frequency threshold is 88200Hz, the frequency deviation ratio is 44100 / 88200=0.5, which is less than 1, so the logarithmic scaling output is 0 and the frequency correction factor is 0.
[0181] In another possible embodiment, if the sound source is a high-sampling lossless type with a sampling frequency of 192000Hz, the frequency deviation ratio = 192000 / 88200≈2.18, and the ratio is greater than 1, then the frequency correction factor = log2(2.18)≈1.12.
[0182] Furthermore, the depth deviation ratio between the original bit depth and the first depth threshold is calculated, and the ratio is logarithmically scaled to obtain the depth correction factor. For example, if the original bit depth is 16 bits and the first depth threshold is 24 bits, the depth deviation ratio = 16 / 24 ≈ 0.67, which is less than 1, so the logarithmic scaling output is 0, and the depth correction factor = 0.
[0183] In another possible embodiment, if the source bit depth is 32 bits, the depth deviation ratio = 32 / 24 ≈ 1.33, and the ratio is greater than 1, then the depth correction factor = log2(1.33) ≈ 0.42.
[0184] Furthermore, the channel deviation ratio between the number of channels and the first channel threshold is calculated, and the ratio is logarithmically scaled to obtain the channel correction factor. For example, if the number of channels is 2, the first channel threshold is 2, the channel deviation ratio = 2 / 2 = 1, the ratio is equal to 1, then the logarithmic scaling output is 0, and the channel correction factor = 0.
[0185] In another possible embodiment, if the sound source is a multi-channel interleaved lossless type with 6 channels, the channel deviation ratio = 6 / 2 = 3, and the ratio is greater than 1, then the channel correction factor = log2(3)≈1.58.
[0186] Finally, the frequency correction factor, depth correction factor, and channel correction factor are fused to obtain the overall correction amount, which is then combined with the baseline tolerance value to obtain the dynamic low distortion tolerance threshold. Wherein, the overall correction amount = frequency correction factor + depth correction factor + channel correction factor. The dynamic low distortion tolerance threshold = baseline tolerance value × 2 raised to the power of the overall correction amount.
[0187] For example, if the audio file is a standard bit-deep lossless file, all three correction factors are 0, the overall correction amount is 0, and the dynamic low distortion tolerance threshold is equal to the intrinsic distortion limit. The strictest distortion tolerance standard is applied to standard specification audio sources, without any additional relaxation.
[0188] In another possible embodiment, for a high-sampling multi-channel mixing scenario, a 32-bit / 192kHz / 6-channel audio source has a frequency correction factor of 1.12, a depth correction factor of 0.42, and a channel correction factor of 1.58. Therefore, the total correction amount = 1.12 + 0.42 + 1.58 = 3.12. The intrinsic distortion floor calculated for 32 bits is approximately 1.54 × 10⁻⁶. -9 Dynamic low distortion tolerance threshold = 1.54 × 10 -9 ×2 3.12 ≈1.34×10 -8 The tolerance threshold is relaxed by approximately 8.7 times compared to the intrinsic lower limit, in order to accommodate the engineering reality that high-specification audio sources may not be able to fully reach their theoretical limits in the link.
[0189] After obtaining the dynamic low-distortion tolerance threshold, the output distortion deviation is read and compared with the dynamic low-distortion tolerance threshold. When the distortion deviation exceeds the dynamic low-distortion tolerance threshold, a path reversal process is triggered. For example, when the audio file is a standard bit-depth lossless file, the output distortion deviation is 400.6, which means the total distortion parameter is approximately -52dB, and the dynamic low-distortion tolerance threshold is 6.24 × 10⁻⁶. -6 If the distortion deviation exceeds the dynamic low distortion tolerance threshold, the trigger condition is met, and the execution path is reversed.
[0190] Before triggering the path back-cut process, the distortion deviation must continuously exceed the dynamic low distortion tolerance threshold for a preset stable confirmation time, and the dynamic low distortion tolerance threshold includes the back-cut lag amount to avoid frequent switching between modes.
[0191] The preset stability confirmation duration is a time parameter pre-built before the system leaves the factory to prevent accidental triggering of the back-off due to transient interference. Timing starts when the distortion deviation first exceeds the dynamic low distortion tolerance threshold. If the deviation remains above the threshold within this duration, the distortion condition is deemed met and a back-off is triggered. If the deviation falls below the threshold during the timing period, the timer is reset and no back-off is triggered. The preset stability confirmation duration is a default value of 3 seconds.
[0192] The cutback hysteresis is a bias value added to the dynamic low distortion tolerance threshold. Specifically, after cutting back from pass-through mode to compensated mode, the distortion deviation must decrease to below the dynamic low distortion tolerance threshold minus the cutback hysteresis and maintain a stable acknowledgment time before allowing a return from compensated mode to pass-through mode. This creates a hysteresis window between the cutback threshold and the recovery threshold, preventing repeated mode switching near the critical point. The cutback hysteresis is typically set to half of the dynamic low distortion tolerance threshold.
[0193] For example, the dynamic low distortion tolerance threshold for a 16-bit / 44.1kHz standard bit-depth lossless audio source is calculated to be 6.24 × 10⁻⁶. -6The preset stability confirmation time is set to 3 seconds, and the back-cut hysteresis is set to half of the threshold, i.e., 3.12 × 10. -6 In pass-through mode, the distortion deviation exceeded 6.24 × 10⁻⁶ at a certain moment. -6 A 3-second timer is initiated. If the deviation remains above the threshold after this time, the timer completes and the system switches back to compensation mode. If, 1.5 seconds later, the deviation falls below the threshold due to a brief power disturbance, the timer is reset, and the switchback does not trigger. After switching back to compensation mode, when the deviation drops to 3.12 × 10⁻⁶, the system will switch back to compensation mode. -6 If the dynamic low distortion tolerance threshold is below half of the threshold for 3 seconds, the link is considered to have recovered stably and will automatically switch back to pass-through mode.
[0194] After the path back-cut is triggered, in order to completely switch the signal link from the direct-through mode to the compensation mode, the following linkage adjustments need to be performed simultaneously: switch the DAC from the native bit-width direct-drive state to the modulation state corresponding to the compensation mode, and re-enable the internal digital filter and oversampling unit; restore the analog output stage from the zero-attenuation direct-through state to the controlled gain state, restore the volume attenuation register to the user-set value, and cancel the bypass of the output buffer; select the second path on the DSP side, reconnect the sampling rate converter and digital filter, and load the pre-compensation coefficient.
[0195] Among them, the DSP (Digital Signal Processor) is an integrated circuit dedicated to high-speed digital signal processing. It is located between the playback engine and the DAC and is responsible for performing operations such as sampling rate conversion, digital filtering, and room calibration.
[0196] For example, a link distortion of -52dB was detected in the 16-bit audio source under the direct path, far exceeding the intrinsic limit of -98dB, triggering a path reversal. The DAC exits native bit-width direct drive and switches to 8x oversampling modulation mode; the R2R volume network exits zero-attenuation direct drive and restores to the user-set -20dB attenuation value; the output buffer is de-bypassed and reconnected; the DSP data selector switches to the second path, the sampling rate converter and digital filter are re-enabled, and the link continues to operate in compensated mode.
[0197] It should be noted that the above values are for illustrative purposes only and do not constitute a limitation on the present invention.
[0198] In summary, this process maintains a stringent distortion tolerance standard for standard audio sources to maximize fidelity, while appropriately relaxing the threshold for high-specification audio sources to take into account engineering feasibility. At the same time, it automatically reverts to a compensation path with correction capabilities when distortion deteriorates in the direct path, enabling the digital media player to dynamically maintain the optimal balance between fidelity and reliability under different audio sources and operating conditions.
[0199] In summary, this invention enables high-specification audio sources to be transmitted to the next stage with minimal loss in a direct path, while standard-specification and multi-channel audio sources obtain appropriate digital domain pre-correction in a compensated path. At the same time, closed-loop monitoring ensures that link distortion is controllable and recoverable under any operating condition, achieving a dynamic optimal balance between fidelity and reliability across the entire range of audio source types.
[0200] Example 2, as Figure 3 As shown, this invention provides a low-distortion original sound restoration system with an end-to-end link from digital readout to analog output, the system comprising:
[0201] The audio source parsing and classification module 11 is used to parse the encoding format and original bit depth of the target lossless audio file and identify its audio source format category.
[0202] This includes parsing the encoding format and original bit depth of the target lossless audio file, and identifying its audio source format category, including:
[0203] Read the file header extension identifier of the target lossless audio file and extract the encoding feature code from the encoding type field;
[0204] The encoding format is determined by matching the encoded feature code with the first feature template of the pulse code modulation format and the second feature template of the direct stream digital encoding format.
[0205] The value of the sampling bit depth field is extracted from the frame structure header information of the target lossless audio file and used as the original bit depth;
[0206] When the encoding format is pulse code modulation, the channel number tag and sampling rate tag in the frame structure header information are parsed to obtain the number of channels and the sampling frequency.
[0207] When the sampling frequency exceeds the first frequency threshold or the original bit depth exceeds the first depth threshold, the audio source format category is determined to be high-magnification sampling lossless type;
[0208] When the sampling frequency does not exceed the first frequency threshold and the original bit depth does not exceed the first bit depth threshold, if the number of channels exceeds the first channel threshold, the audio source format category is determined to be multi-channel interleaved lossless; otherwise, the audio source format category is determined to be standard bit depth lossless.
[0209] When the encoding format is a direct stream digital encoding format, the audio source format category is determined to be a one-bit stream native lossless class, and the original bit depth is set to 1.
[0210] The first bit depth threshold is determined based on the maximum linear quantization bit width natively supported by the current mainstream digital-to-analog converter hardware; the first frequency threshold is determined based on the Nyquist sampling requirement of the upper limit of human hearing frequency and the general standard of high-resolution audio; and the first channel threshold is determined based on the standard channel configuration of a general stereo playback device.
[0211] The processing mode decision module 12 is used to determine the processing mode of the digital signal processing link according to the audio source format category, wherein the processing mode includes at least a pass-through mode and a compensation mode.
[0212] The process mode of the digital signal processing link is determined based on the audio source format category, including:
[0213] When the audio source format category is standard bit-depth lossless, the processing mode is determined to be pass-through mode, and the bit width locking value corresponding to the native bit width direct drive state of the digital-to-analog converter is verified to be consistent with the original bit depth.
[0214] When the audio source format category is high-sampling lossless, the processing mode is determined to be compensation mode, and the first ratio of the sampling frequency to the first frequency threshold and the second ratio of the original bit depth to the first bit depth threshold are calculated.
[0215] The larger of the first ratio and the second ratio is selected, and the sampling rate converter is controlled to perform resampling according to the larger value;
[0216] When the audio source format category is multi-channel interleaved lossless, the processing mode is determined to be the compensation mode, and the third ratio of the number of channels to the first channel threshold is calculated.
[0217] The value obtained by rounding up the third ratio is used to perform downmixing on the number of channels to obtain the sampled values of the left output channel and the right output channel.
[0218] When the audio source format category is a one-bit stream native lossless type, check whether the digital-to-analog converter supports native pass-through decoding of direct stream digital encoding format;
[0219] If supported, the processing mode is determined to be the pass-through mode; otherwise, the processing mode is determined to be the compensation mode, and the order of the digital filter is configured to be twice the original bit depth.
[0220] Specifically, the third ratio, rounded up, is used to perform downmixing on the number of channels, including:
[0221] The target number of downmixing channels is fixed at the threshold of the first channel.
[0222] Parse the channel configuration field in the frame structure header information of the target lossless audio file to obtain the type identifier of each channel. The type identifier includes at least the left channel, right channel, center channel, left surround channel, right surround channel and low frequency effect channel.
[0223] When the input channel type is identified as the left channel, the sampled value of the input channel is multiplied by the value 1 and then accumulated to the left output channel;
[0224] When the input channel type is identified as left surround channel, the sampled value of the input channel is multiplied by the left allocation coefficient and then accumulated to the left output channel. The left allocation coefficient is determined based on the 3dB sound pressure level attenuation principle.
[0225] When the input channel type is identified as the right channel, the sampled value of the input channel is multiplied by the value 1 and then accumulated to the right output channel;
[0226] When the input channel type is identified as right surround channel, the sampled value of the input channel is multiplied by the right allocation coefficient and then accumulated to the right output channel. The right allocation coefficient is determined based on the 3dB sound pressure level attenuation principle.
[0227] When the input channel type is identified as center channel, the sampled value of the input channel is multiplied by the center allocation coefficient and then accumulated to the left output channel and the right output channel respectively. The center allocation coefficient is determined based on the principle of equal distribution of dual channel energy.
[0228] When the input channel type is identified as a low-frequency effects channel and the current audio source format is a multi-channel music recording type, skip the input channel.
[0229] The value obtained by rounding up the third ratio is used as the amplitude normalization divisor. The accumulated value of the left output channel is divided by the amplitude normalization divisor to obtain the sampled value of the left output channel. The accumulated value of the right output channel is divided by the amplitude normalization divisor to obtain the sampled value of the right output channel.
[0230] The pass-through mode link configuration module 13 is used to simultaneously activate the native bit-width direct drive state of the digital-to-analog converter and the zero-attenuation pass-through state of the analog output stage when configured as pass-through mode, and disconnect the data path connection between the sampling rate converter and the digital filter in the digital signal processor.
[0231] When configured in pass-through mode, the native bit-width direct-drive state of the digital-to-analog converter and the zero-attenuation pass-through state of the analog output stage are activated simultaneously, and the data path connection between the sampling rate converter and the digital filter in the digital signal processor is disconnected, including:
[0232] Configure the digital audio input format of the digital-to-analog converter according to the original bit depth, and set the internal processing precision of the digital-to-analog converter to be no less than the original bit depth;
[0233] After reading the status register of the digital-to-analog converter and confirming that the digital audio input format and internal processing precision configuration are effective, the sampling frequency is written to the clock divider register of the digital-to-analog converter, and the main clock of the digital-to-analog converter is configured to be an integer multiple of the sampling frequency.
[0234] Write the value 0 to the volume attenuation register of the analog output stage, and configure the output buffer of the analog output stage to bypass state;
[0235] The data selector connected to the sampling rate converter in the digital signal processor is switched to the first path, wherein the first path directly sends the input sampled value to the digital-to-analog converter, and at the same time writes the value 0 to the enable register of the digital filter.
[0236] The end-to-end distortion monitoring module 14 is used to acquire the analog output signal after processing by the analog output stage in real time, and to determine the end-to-end distortion parameters based on the analog output signal.
[0237] This includes real-time acquisition of the analog output signal after processing by the analog output stage, and determination of end-to-end distortion parameters based on the analog output signal, including:
[0238] When the encoding format is pulse code modulation, the analog output signal is sampled at the sampling frequency to obtain a monitoring sampling sequence;
[0239] Read the original audio sampling sequence that is the same as the monitoring sampling sequence time period, and obtain the reference sampling sequence after the same decoding and downmixing process as in the current signal link. If the number of channels is greater than the first channel threshold, the original audio sampling sequence is downmixed according to the same channel type allocation coefficient and amplitude normalization method; otherwise, the read original audio sampling sequence is used as the reference sampling sequence.
[0240] Calculate the cross-correlation function between the monitoring sampling sequence and the reference sampling sequence, and use the delay value corresponding to the peak value of the cross-correlation function as the alignment offset;
[0241] The monitoring sampling sequence is shifted by the alignment offset to align the monitoring sampling sequence with the reference sampling sequence in the time domain.
[0242] The aligned monitoring sampling sequence is subtracted from the reference sampling sequence point by point to obtain the residual sequence;
[0243] Calculate the root mean square value of the residual sequence and the root mean square value of the reference sampling sequence, and obtain the basic distortion ratio based on the ratio of the root mean square value of the residual to the root mean square value of the reference.
[0244] The base distortion ratio is used as the end-to-end distortion parameter.
[0245] The distortion deviation evaluation module 15 is used to determine the intrinsic distortion limit based on the original bit depth and to obtain the distortion deviation by combining the end-to-end distortion parameters.
[0246] Specifically, based on the original bit depth, the intrinsic distortion threshold is determined, and combined with the end-to-end distortion parameters, the distortion deviation is obtained, including:
[0247] Substitute the original bit depth into the theoretical quantization noise model to calculate the root mean square value of quantization noise corresponding to the ideal quantization step size, and use the ratio of the root mean square value of quantization noise to the root mean square value of the full-amplitude sine wave as the intrinsic distortion limit.
[0248] The theoretical quantization noise model takes the original bit depth as input and outputs the root mean square value of the quantization noise based on the uniform quantization theory. The full-amplitude sine wave is a sine wave signal whose amplitude range covers the full scale of the digital domain.
[0249] Read the end-to-end distortion parameters and calculate the ratio of the end-to-end distortion parameters to the intrinsic distortion limit. Record the ratio as the distortion deviation.
[0250] The dynamic path back-cutting and linkage adjustment module 16 is used to trigger the path back-cutting process when the distortion deviation exceeds the dynamic low distortion tolerance threshold, dynamically switch the signal link from the direct mode to the compensation mode, and adjust the working state of the digital-to-analog converter and the analog output stage in linkage.
[0251] The calculation steps for the dynamic low distortion tolerance threshold include:
[0252] Obtain the intrinsic distortion limit and use the intrinsic distortion limit as the benchmark tolerance value;
[0253] Calculate the frequency deviation ratio between the sampling frequency and the first frequency threshold, and perform logarithmic scaling on the frequency deviation ratio to obtain the frequency correction factor;
[0254] Calculate the depth deviation ratio between the original bit depth and the first depth threshold, and perform logarithmic scaling on the depth deviation ratio to obtain the depth correction factor;
[0255] Calculate the channel deviation ratio between the number of channels and the first channel threshold, and perform logarithmic scaling on the channel deviation ratio to obtain the channel correction factor;
[0256] The logarithmic scaling is defined as follows: when the input deviation ratio is greater than 1, the output is the logarithm of the deviation ratio with base 2; when the input deviation ratio is less than or equal to 1, the output is 0.
[0257] The frequency correction factor, the depth correction factor, and the channel correction factor are fused to obtain a comprehensive correction amount, and combined with the reference tolerance value, a dynamic low distortion tolerance threshold is obtained.
[0258] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
[0259] This specification and accompanying drawings are merely illustrative examples of the invention and are intended to cover any and all modifications, variations, combinations, or equivalents within the scope of the invention. Clearly, those skilled in the art can make various alterations and modifications to the invention without departing from its scope. Therefore, if such modifications and modifications fall within the scope of the invention and its equivalents, the invention is intended to include these modifications and modifications.
Claims
1. A method for end-to-end low-distortion original sound restoration from digital reading to analog output, characterized in that, The method includes: The encoding format and original bit depth of the target lossless audio file are analyzed, and the audio source format category is identified. Based on the audio source format category, the processing mode of the digital signal processing link is determined, wherein the processing mode includes at least a pass-through mode and a compensation mode; When configured in pass-through mode, the native bit-width pass-through state of the digital-to-analog converter and the zero-attenuation pass-through state of the analog output stage are activated simultaneously, and the data path connection between the sampling rate converter and the digital filter in the digital signal processor is disconnected. The analog output signal processed by the analog output stage is acquired in real time, and the end-to-end distortion parameters are determined based on the analog output signal. Based on the original bit depth, the intrinsic distortion limit is determined, and the distortion deviation is obtained by combining the end-to-end distortion parameters. When the distortion deviation exceeds the dynamic low distortion tolerance threshold, a path back-switching process is triggered to dynamically switch the signal link from the direct mode to the compensation mode, and the working status of the digital-to-analog converter and the analog output stage are adjusted accordingly.
2. The end-to-end low-distortion original sound restoration method from digital reading to analog output according to claim 1, characterized in that, The encoding format and original bit depth of the target lossless audio file are analyzed, and the audio source format category is identified, including: Read the file header extension identifier of the target lossless audio file and extract the encoding feature code from the encoding type field; The encoding format is determined by matching the encoded feature code with the first feature template of the pulse code modulation format and the second feature template of the direct stream digital encoding format. The value of the sampling bit depth field is extracted from the frame structure header information of the target lossless audio file and used as the original bit depth; When the encoding format is pulse code modulation, the channel number tag and sampling rate tag in the frame structure header information are parsed to obtain the number of channels and the sampling frequency. When the sampling frequency exceeds the first frequency threshold or the original bit depth exceeds the first depth threshold, the audio source format category is determined to be high-magnification sampling lossless type; When the sampling frequency does not exceed the first frequency threshold and the original bit depth does not exceed the first bit depth threshold, if the number of channels exceeds the first channel threshold, the audio source format category is determined to be multi-channel interleaved lossless; otherwise, the audio source format category is determined to be standard bit depth lossless. When the encoding format is a direct stream digital encoding format, the audio source format category is determined to be a one-bit stream native lossless class, and the original bit depth is set to 1.
3. The end-to-end low-distortion original sound restoration method from digital reading to analog output as described in claim 2, characterized in that, The first bit depth threshold is determined based on the maximum linear quantization bit width natively supported by the current mainstream digital-to-analog converter hardware. The first frequency threshold is determined based on the Nyquist sampling requirement of the upper limit of human hearing frequency and the general standard of high-resolution audio. The first channel threshold is determined based on the standard channel configuration of a general stereo playback device.
4. The end-to-end low-distortion original sound restoration method from digital reading to analog output according to claim 1, characterized in that, Based on the audio source format category, the processing mode of the digital signal processing link is determined, including: When the audio source format category is standard bit-depth lossless, the processing mode is determined to be pass-through mode, and the bit width locking value corresponding to the native bit width direct drive state of the digital-to-analog converter is verified to be consistent with the original bit depth. When the audio source format category is high-sampling lossless, the processing mode is determined to be compensation mode, and the first ratio of the sampling frequency to the first frequency threshold and the second ratio of the original bit depth to the first bit depth threshold are calculated. The larger of the first ratio and the second ratio is selected, and the sampling rate converter is controlled to perform resampling according to the larger value; When the audio source format category is multi-channel interleaved lossless, the processing mode is determined to be the compensation mode, and the third ratio of the number of channels to the first channel threshold is calculated. The value obtained by rounding up the third ratio is used to perform downmixing on the number of channels to obtain the sampled values of the left output channel and the right output channel. When the audio source format category is a one-bit stream native lossless type, check whether the digital-to-analog converter supports native pass-through decoding of direct stream digital encoding format; If supported, the processing mode is determined to be the pass-through mode; otherwise, the processing mode is determined to be the compensation mode, and the order of the digital filter is configured to be twice the original bit depth.
5. The end-to-end low-distortion original sound restoration method from digital reading to analog output according to claim 4, characterized in that, The value obtained by rounding up the third ratio is used to perform downmixing on the number of channels, including: The target number of downmixing channels is fixed at the threshold of the first channel. Parse the channel configuration field in the frame structure header information of the target lossless audio file to obtain the type identifier of each channel. The type identifier includes at least the left channel, right channel, center channel, left surround channel, right surround channel and low frequency effect channel. When the input channel type is identified as the left channel, the sampled value of the input channel is multiplied by the value 1 and then accumulated to the left output channel; When the input channel type is identified as left surround channel, the sampled value of the input channel is multiplied by the left allocation coefficient and then accumulated to the left output channel. The left allocation coefficient is determined based on the 3dB sound pressure level attenuation principle. When the input channel type is identified as the right channel, the sampled value of the input channel is multiplied by the value 1 and then accumulated to the right output channel; When the input channel type is identified as right surround channel, the sampled value of the input channel is multiplied by the right allocation coefficient and then accumulated to the right output channel. The right allocation coefficient is determined based on the 3dB sound pressure level attenuation principle. When the input channel type is identified as center channel, the sampled value of the input channel is multiplied by the center allocation coefficient and then accumulated to the left output channel and the right output channel respectively. The center allocation coefficient is determined based on the principle of equal distribution of dual channel energy. When the input channel type is identified as a low-frequency effects channel and the current audio source format is a multi-channel music recording type, skip the input channel. The value obtained by rounding up the third ratio is used as the amplitude normalization divisor. The accumulated value of the left output channel is divided by the amplitude normalization divisor to obtain the sampled value of the left output channel. The accumulated value of the right output channel is divided by the amplitude normalization divisor to obtain the sampled value of the right output channel.
6. The end-to-end low-distortion original sound restoration method from digital reading to analog output according to claim 1, characterized in that, When configured in pass-through mode, the native bit-width direct-drive state of the digital-to-analog converter and the zero-attenuation pass-through state of the analog output stage are activated simultaneously, and the data path connection between the sampling rate converter and the digital filter in the digital signal processor is disconnected, including: Configure the digital audio input format of the digital-to-analog converter according to the original bit depth, and set the internal processing precision of the digital-to-analog converter to be no less than the original bit depth; After reading the status register of the digital-to-analog converter and confirming that the digital audio input format and internal processing precision configuration are effective, the sampling frequency is written to the clock divider register of the digital-to-analog converter, and the main clock of the digital-to-analog converter is configured to be an integer multiple of the sampling frequency. Write the value 0 to the volume attenuation register of the analog output stage, and configure the output buffer of the analog output stage to bypass state; The data selector connected to the sampling rate converter in the digital signal processor is switched to the first path, wherein the first path directly sends the input sampled value to the digital-to-analog converter, and at the same time writes the value 0 to the enable register of the digital filter.
7. The end-to-end low-distortion original sound restoration method from digital reading to analog output according to claim 1, characterized in that, Real-time acquisition of the analog output signal after processing by the analog output stage, and determination of end-to-end distortion parameters based on the analog output signal, including: When the encoding format is pulse code modulation, the analog output signal is sampled at the sampling frequency to obtain a monitoring sampling sequence; Read the original audio sampling sequence that is the same as the monitoring sampling sequence time period, and obtain the reference sampling sequence after the same decoding and downmixing process as in the current signal link. If the number of channels is greater than the first channel threshold, the original audio sampling sequence is downmixed according to the same channel type allocation coefficient and amplitude normalization method; otherwise, the read original audio sampling sequence is used as the reference sampling sequence. Calculate the cross-correlation function between the monitoring sampling sequence and the reference sampling sequence, and use the delay value corresponding to the peak value of the cross-correlation function as the alignment offset; The monitoring sampling sequence is shifted by the alignment offset to align the monitoring sampling sequence with the reference sampling sequence in the time domain. The aligned monitoring sampling sequence is subtracted from the reference sampling sequence point by point to obtain the residual sequence; Calculate the root mean square value of the residual sequence and the root mean square value of the reference sampling sequence, and obtain the basic distortion ratio based on the ratio of the root mean square value of the residual to the root mean square value of the reference. The base distortion ratio is used as the end-to-end distortion parameter.
8. The end-to-end low-distortion original sound restoration method from digital reading to analog output according to claim 1, characterized in that, Based on the original bit depth, the intrinsic distortion limit is determined, and combined with the end-to-end distortion parameters, the distortion deviation is obtained, including: Substitute the original bit depth into the theoretical quantization noise model to calculate the root mean square value of quantization noise corresponding to the ideal quantization step size, and use the ratio of the root mean square value of quantization noise to the root mean square value of the full-amplitude sine wave as the intrinsic distortion limit. The theoretical quantization noise model takes the original bit depth as input and outputs the root mean square value of the quantization noise based on the uniform quantization theory. The full-amplitude sine wave is a sine wave signal whose amplitude range covers the full scale of the digital domain. Read the end-to-end distortion parameters and calculate the ratio of the end-to-end distortion parameters to the intrinsic distortion limit. Record the ratio as the distortion deviation.
9. The end-to-end low-distortion original sound restoration method from digital reading to analog output according to claim 1, characterized in that, The calculation steps for the dynamic low distortion tolerance threshold include: Obtain the intrinsic distortion limit and use the intrinsic distortion limit as the benchmark tolerance value; Calculate the frequency deviation ratio between the sampling frequency and the first frequency threshold, and perform logarithmic scaling on the frequency deviation ratio to obtain the frequency correction factor; Calculate the depth deviation ratio between the original bit depth and the first depth threshold, and perform logarithmic scaling on the depth deviation ratio to obtain the depth correction factor; Calculate the channel deviation ratio between the number of channels and the first channel threshold, and perform logarithmic scaling on the channel deviation ratio to obtain the channel correction factor; The logarithmic scaling is defined as follows: when the input deviation ratio is greater than 1, the output is the logarithm of the deviation ratio with base 2; when the input deviation ratio is less than or equal to 1, the output is 0. The frequency correction factor, the depth correction factor, and the channel correction factor are fused to obtain a comprehensive correction amount, and combined with the reference tolerance value, a dynamic low distortion tolerance threshold is obtained.
10. A low-distortion original sound restoration system with an end-to-end connection from digital reading to analog output, characterized in that: The system is used to perform the end-to-end low-distortion original sound restoration method from digital readout to analog output as described in any one of claims 1-9, the system comprising: The audio source analysis and classification module is used to analyze the encoding format and original bit depth of the target lossless audio file and identify its audio source format category; The processing mode decision module is used to determine the processing mode of the digital signal processing link according to the audio source format category, wherein the processing mode includes at least a pass-through mode and a compensation mode. The pass-through mode link configuration module is used to simultaneously activate the native bit-width direct drive state of the digital-to-analog converter and the zero-attenuation pass-through state of the analog output stage when configured as pass-through mode, and disconnect the data path connection between the sampling rate converter and the digital filter in the digital signal processor. The end-to-end distortion monitoring module is used to acquire the analog output signal after processing by the analog output stage in real time, and determine the end-to-end distortion parameters based on the analog output signal. The distortion deviation evaluation module is used to determine the intrinsic distortion limit based on the original bit depth, and to obtain the distortion deviation by combining the end-to-end distortion parameters. The dynamic path back-cutting and linkage adjustment module is used to trigger the path back-cutting process when the distortion deviation exceeds the dynamic low distortion tolerance threshold, dynamically switch the signal link from the direct mode to the compensation mode, and adjust the working state of the digital-to-analog converter and the analog output stage in linkage.