High-sampling-rate audio analysis optimization method and system, storage medium and equipment
By upgrading the FFTW library and combining technologies such as adaptive frame processing, multi-level caching and clock synchronization optimization, the repeated calculation and resource waste in high-sampling rate audio processing are solved, efficient audio analysis and resource optimization are achieved, and the performance and user experience of audio equipment are improved.
Patent Information
- Application Number
- CN202510340043.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-07-01
AI Technical Summary
The prior art has problems of repeated calculations, resource waste and mismatch in high sampling rate audio processing, especially on portable audio devices that affect device performance and user experience.
By upgrading the FFTW fast Fourier transform library, an optimized memory management mechanism is established, adaptive framing processing and multi-level caching system is adopted, FFT calculation accuracy is dynamically adjusted, clock synchronization optimization and jitter cancellation, and a multi-dimensional result cache matrix and intelligent multiplexing decision tree are established, combining subjective and objective evaluation mechanisms to monitor audio quality and system stability.
Significantly reduces the FFT repetitive calculation rate, improves memory utilization, reduces CPU peak usage, reduces clock deviation, improves system robustness and audio quality, and reduces system power consumption.
Smart Images

Figure CN120233977A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of audio signal processing, and in particular, to a method, system, storage medium and device for optimizing high-sampling-rate audio parsing. Background Art
[0002] In the field of digital audio processing, the Fast Fourier Transform (FFT) is the core technology for realizing spectrum analysis. As an efficient FFT calculation library, FFTW is widely used in various audio processing devices. However, with the popularization of high-resolution audio, especially the wide application of 192kHz sampling rate audio, traditional audio parsing solutions are facing severe challenges. High sampling rates mean larger amounts of data and more complex computational requirements, resulting in increased processing delays and excessive consumption of system resources.
[0003] In the prior art, parsing algorithms often need to perform multiple repeated calculations on the same data segment, lacking an effective data reuse mechanism. There are a large number of repeated operations in traditional parsing solutions, resulting in waste of system resources and affecting the overall performance of the device and the user experience. At the same time, using a unified computational precision for scenarios with different precision requirements also causes waste of computational resources. These problems are particularly prominent in portable audio devices such as WiiM, seriously affecting the product performance and user experience. Therefore, a new solution is urgently needed. Summary of the Invention
[0004] The purpose of the present invention is to solve the above-mentioned disadvantages existing in the prior art, and to provide a method, system, storage medium and device for optimizing high-sampling-rate audio parsing. By comprehensively upgrading and optimizing the FFTW library and combining multi-level performance optimization strategies, the key technical pain points in high-sampling-rate audio processing are successfully solved.
[0005] On the one hand, a method for optimizing high-sampling-rate audio parsing is provided, including the following steps: S1: Upgrade the FFTW Fast Fourier Transform library and establish an optimized memory management mechanism; S2: After completing the optimization of the core framework in step S1, perform adaptive frame processing on the input audio data and optimize the data access efficiency using a multi-level cache system; S3: Dynamically adjust the computational precision of the FFT Fast Fourier Transform and optimize the FFT Fast Fourier Transform calculation by pre-computing and caching FFTW plans; S4: Optimize the clock synchronization and eliminate jitter for COAX digital audio signals; S5: Establish a multi-dimensional result cache matrix and an intelligent reuse decision tree, and reduce repeated operations through intelligent cache prediction and similarity calculation; S6: Monitor the system performance metrics in real time and implement adaptive parameter adjustment based on the optimization objective function. Combine subjective and objective evaluation mechanisms to monitor the audio quality and system stability.
[0006] Further, in step S1, the upgrade of the FFTW (Fast Fourier Transform in the West) library includes: S11: Replace the old version library file and deeply optimize the SIMD (Single Instruction Multiple Data) instruction set. Automatically detect and enable the supported SIMD instruction set during the initialization of FFTW. S12: Dynamically adjust the number of threads based on the CPU load to configure dynamic thread management and avoid resource contention. The establishment of the optimized memory management mechanism includes: S13: Implement a zero-copy data transfer mechanism through memory mapping technology to reduce the data movement overhead. Allocate continuous physical memory blocks, map them to the user space address, write the audio device directly to the memory block, and the FFTW library reads data from the same address. S14: Establish a memory pool management system based on the maximum sampling rate, the number of buffer blocks, the number of samples per block, and the number of bytes per sample to optimize the memory allocation strategy.
[0007] Further, in step S2, the adaptive frame segmentation processing of the input audio data includes: S201: Set the base sampling rate and base frame length, and establish a mapping table for the sampling rate and frame length at the same time. S202: Use a sliding window to detect the actual sampling rate of the current audio stream, and dynamically calculate the optimal frame length for frame segmentation according to the detected sampling rate change. S203: Input the data into the buffer to receive the COAX audio stream, divide the data into blocks according to the optimal frame length, add an overlapping area to each data frame, and apply a Hanning window for windowing processing. S204: Start a dynamic degradation strategy to calculate the temporary frame length for frame segmentation according to the CPU load, and implement an adaptive compensation mechanism.
[0008] Preferably, in step S2, the further optimization of the data access efficiency by using a multi-level cache system includes: S211: Build a multi-level cache architecture, which is used to store the original sampling samples, the preprocessed frequency domain data, and the common calculation result templates respectively. S212: Build a data dependence graph, define the node and edge relationships of the data dependence graph, and use an incremental graph update algorithm to dynamically update the data dependence graph. S213: Optimize the cache replacement algorithm, conduct short-term prediction and long-term pattern mining based on a machine learning-based hybrid replacement strategy, evaluate the cache value through a cache value evaluation model, and replace the cache according to the set replacement decision. S214: Optimize cache hit rate control, monitor the cache hit rate in real time through a sliding window and implement an adaptive adjustment strategy, and implement a prefetching algorithm using an improved chain prediction model.
[0009] Further, in step S3, the dynamic adjustment of the FFT (Fast Fourier Transform) calculation precision includes: S301: Customize the precision level and the floating point used for each level. The precision levels include single precision and double precision, and dynamically select the precision level according to the input signal-to-noise ratio; S302: Establish a precision adaptive mechanism, set the initial mode and the decision tree for precision level switching during runtime, calculate the required precision based on the base precision and the signal-to-noise ratio factor. The base precision defaults to single precision, and the signal-to-noise ratio factor is calculated by converting decibels to amplitude ratio. Select the appropriate data type by comparing the calculated required precision with the customized precision levels; S303: Configure automatic type promotion rules to handle the situation where the input data does not match the planned precision, and use SIMD (Single Instruction, Multiple Data) instructions to batch process precision conversion.
[0010] Preferably, in step S3, the precomputation and caching of the FFTW (Fastest Fourier Transform in the West) plan include implementing the FFTW plan precomputation and establishing a plan cache pool. Among them, The implementation of the FFTW plan precomputation includes: S311: Determine the frequently used FFT sizes through historical data analysis and establish a size priority queue; S312: Divide the stages according to the size priority, quickly generate the basic FFTW plan for the highly popular sizes, and perform fine optimization on the core sizes through multi-threaded parallel processing; S313: Convert the optimized plan object into a memory binary image and store it, and record the hardware fingerprint in the plan metadata at the same time; The establishment of the plan cache pool includes: S314: Design the cache structure, divide the cache pool into memory pages of a fixed size, and each page stores FFT plans of the same size; S315: Eliminate the rarely used plans that have not been used for a long time through the LRU-K (Least Recently Used with Lookahead) replacement algorithm and set a threshold to retain the high-value plans. At the same time, automatically adjust the size of the cache pool according to the system memory pressure.
[0011] Further, in step S4, the clock synchronization optimization for the COAX digital audio signal includes: S401: Use a digital phase-locked loop (DPLL) to capture the actual sampling rate of the COAX interface in real time, compare it with the nominal value, and calculate the deviation value; S402: Implement dynamic adjustment of the phase-locked loop (PLL) parameters with the deviation value as the input, where the PLL parameters include the loop bandwidth and damping coefficient; S403: For frames with deviation values exceeding the set threshold, use a Farrow structure fractional delay filter to achieve fine adjustment of the sampling rate with minimum phase distortion.
[0012] Preferably, in step S4, the jitter elimination further includes: S411: Extract the clock deviation sequence of several consecutive sampling points in the time domain, calculate the jitter amplitude of the acquisition part, dynamically set the threshold based on the signal-to-noise ratio of the signal, and adaptively detect jitter and trigger jitter elimination according to the jitter amplitude and the threshold; S412: Jointly process the time domain and the frequency domain to optimize the jitter algorithm, where For the time domain, apply a filter to predict the trend of the clock deviation and correct the position of the sampling points; For the frequency domain, perform sub-band energy analysis on the spectrum after FFT to suppress the high-frequency noise generated by jitter.
[0013] Furthermore, in step S5, the establishment of the multi-dimensional result cache matrix includes: S501: Construct a three-dimensional tensor cache result, with dimensions including frequency binning, time window, and signal type, and construct a fast indexing mechanism using hash mapping + skip list indexing; S502: Use the historical access sequence, time interval, and signal type transition probability as input features to construct a cache prediction model, predict and output the set of cache key values that may be accessed within a certain period in the future, and implement a preloading strategy for the key values in the prediction result; S503: Establish an efficiency threshold strategy to dynamically eliminate inefficient cache entries.
[0014] Preferably, in step S5, the establishment of the intelligent reuse decision tree further includes: S511: Perform joint similarity evaluation on the input signal to extract features, where the extracted features include frequency domain features, time domain features, and energy distribution, and dynamically calculate the similarity threshold according to the signal amplitude and accuracy requirements; S512: Establish a reuse decision tree based on multi-layer decision logic, perform linear phase compensation on the reused spectrum, and perform weighted fusion of the new and old results according to the signal-to-noise ratio to obtain the final reuse decision.
[0015] Furthermore, in step S6, the real-time monitoring of the system performance indicators and the implementation of adaptive parameter adjustment based on the optimization objective function include: S601: Periodically sample the CPU usage rate through the performance interface provided by the operating system, monitor the resident memory set of the process, and use the memory water level marking strategy to count the memory occupancy. Calculate the average performance metrics through a sliding window, and measure and record the processing delay. S602: Establish a performance evaluation model based on a multi-objective optimization function, and the objective weights of the multi-objective optimization function are automatically adjusted according to the system state.
[0016] Preferably, in step S6, the monitoring of the audio quality and system stability by combining the subjective and objective evaluation mechanisms further includes: S611: Conduct an objective quality evaluation through the PEAQ audio quality evaluation algorithm. Input the time-frequency signals of the reference audio and the audio under test to calculate each MOV parameter. The MOV parameters include noise, distortion, and modulation difference. Map multiple MOV parameters to the audio quality score ODG through a non-linear regression model. S612: Collect subjective scoring standard data, including listener screening, test environment, and scoring method, and set a confidence interval for data analysis. S613: Monitor the system stability through stress testing and long-term stability verification.
[0017] On the other hand, a high sampling rate audio parsing optimization system is provided, including: A core framework optimization module for upgrading the FFTW fast Fourier transform library and establishing an optimized memory management mechanism. A data preprocessing module for adaptively frame the input audio data and optimizing the data access efficiency using a multi-level cache system. An FFT calculation optimization module for dynamically adjusting the FFT fast Fourier transform calculation accuracy and optimizing the FFT fast Fourier transform calculation by pre-calculating and caching FFTW plans. A signal processing module for clock synchronization optimization and jitter elimination for COAX digital audio signals. A dynamic multiplexing module for establishing a multi-dimensional result cache matrix and an intelligent multiplexing decision tree, and reducing repeated operations through intelligent cache prediction and similarity calculation. A quality and performance monitoring module for real-time monitoring of system performance metrics and realizing adaptive parameter adjustment based on the optimization objective function, and monitoring the audio quality and system stability by combining subjective and objective evaluation mechanisms.
[0018] In addition, a computer-readable storage medium is provided, on which a computer program is stored. When the program is executed by a processor, it implements the high sampling rate audio parsing optimization method described in any one of the above.
[0019] Meanwhile, an electronic device is provided, including: one or more processors; a storage device for storing one or more programs, which, when executed by the one or more processors, cause the one or more processors to implement the high-sampling-rate audio parsing optimization method described in any one of the above.
[0020] Compared with the prior art, the beneficial effects of the present invention are as follows: Through FFTW library optimization, plan pre-computation, and dynamic precision adjustment, the present invention reduces the time-consuming of FFT calculation, achieves less single-frame processing time in 192kHz / 24bit audio processing, and the multi-dimensional result caching and intelligent reuse strategy reduce the FFT repeated calculation rate, improving the throughput under the same hardware; The present invention reduces memory fragmentation through an intelligent memory allocator, improves memory utilization, and the adaptive parameter adjustment module dynamically allocates thread resources, reducing the CPU peak usage rate while maintaining latency stability; Through COAX clock synchronization optimization, the present invention significantly reduces clock deviation, and the jitter elimination algorithm improves the signal-to-noise ratio; The present invention monitors audio quality and system stability through schemes such as subjective and objective joint evaluation and stress testing, improving system robustness; Through intelligent cache prediction and calculation reuse, the present invention reduces the overall power consumption of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention, but do not constitute a limitation to the present invention. In the drawings: Figure 1 is a schematic flow chart of a high-sampling-rate audio parsing optimization method of the present invention; Figure 2 is a schematic flow chart of establishing an optimized memory management mechanism of the present invention; Figure 3 is a schematic flow chart of an adaptive frame processing of the present invention; Figure 4 is a schematic flow chart of optimizing data access efficiency by adopting a multi-level cache system of the present invention; Figure 5 is a schematic flow chart of dynamically adjusting the calculation precision of FFT (Fast Fourier Transform) of the present invention; Figure 6 is a schematic flow chart of pre-computing and caching FFTW plans of the present invention; Figure 7 is a schematic flow chart of clock synchronization optimization for COAX digital audio signals of the present invention; Figure 8 is a schematic flow chart of jitter elimination of the present invention; Figure 9 It is a schematic diagram of the process for establishing a multi-dimensional result cache matrix according to the present invention; Figure 10 It is a schematic diagram of the process for establishing the intelligent reuse decision tree according to the present invention; Figure 11 It is a block diagram of the structure of a high sampling rate audio parsing optimization system according to the present invention; Figure 12 It is a schematic diagram of an embodiment of an electronic device according to the present invention. Detailed implementation manners
[0022] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0023] The present invention starts with underlying optimization. By upgrading the FFTW library and enabling the SIMD instruction set, an efficient memory management mechanism is established, laying a foundation for high-performance computing. At the data processing level, an adaptive frame processing and a multi-level cache system are innovatively introduced. Especially when processing 192kHz COAX audio, through the implementation of dedicated clock synchronization optimization and jitter elimination algorithms, accurate restoration of high-fidelity audio is ensured. A complete calculation result reuse mechanism is established. Through intelligent cache prediction and similarity calculation, repeated operations are significantly reduced. At the same time, an adaptive parameter adjustment mechanism is adopted to achieve dynamic optimization allocation of computing resources. On the premise of ensuring audio quality, a comprehensive quality monitoring is carried out through the PEAQ scoring system to ensure the reliability and stability of the optimization effect.
[0024] The following illustrates the detailed implementation manners of the present invention in conjunction with the accompanying drawings and embodiments.
[0025] Embodiment 1 Please refer to Figure 1 , which is the technical solution of a high sampling rate audio parsing optimization method provided for this embodiment, including the following steps: S1: Upgrade the FFTW fast Fourier transform library and establish an optimized memory management mechanism; S2: After optimizing the core framework in step S1, perform adaptive frame processing on the input audio data and optimize the data access efficiency by using a multi-level cache system; S3: Dynamically adjust the calculation accuracy of the FFT fast Fourier transform, and optimize the FFT fast Fourier transform calculation by pre-calculation and caching the FFTW plan; S4: Optimize the clock synchronization and eliminate the jitter for the COAX digital audio signal; S5: Establish a multi-dimensional result cache matrix and an intelligent multiplexing decision tree, and reduce the repeated operations through intelligent cache prediction and similarity calculation; S6: Monitor the system performance metrics in real time and implement the adaptive parameter adjustment based on the optimization objective function, and monitor the audio quality and system stability by combining the subjective and objective evaluation mechanisms.
[0026] Among them, in step S1, upgrade the FFTW fast Fourier transform library to version v3.3.10, and directly utilize the improvements of the official algorithm and parallel computing, which belongs to the in-depth adaptation and configuration optimization of the third-party library, supports more efficient parallel computing and vectorized operations, and further includes: First, replace the old version of the FTW fast Fourier transform library file, then enable the SIMD single instruction multiple data instruction set optimization, utilize the single instruction multiple data (SIMD) instructions (such as AVX2, SSE4) to process multiple sample data in parallel, improve the FFT calculation throughput, call fftwf_init_threads() and set the FFTW_PATIENT flag during the initialization of FFTW, and automatically detect and enable the supported SIMD single instruction multiple data instruction set; Next, dynamically adjust the number of threads based on the CPU load to configure the dynamic thread management and avoid resource contention. In this embodiment, for a quad-core processor, the dynamic adjustment range of the number of threads is 1 - 4, and it is automatically adjusted according to the current CPU load, that is, input the real-time CPU usage rate , output the number of threads , where 4 is the maximum number of threads of the quad-core processor. If the CPU load is 70% ( ), then ; Secondly, establish an optimized memory management mechanism, as shown in Figure 2 , including: S13: Implement a zero-copy data transfer mechanism through the memory mapping (Memory-Mapped I / O) technology, avoid the replication of data between the user mode and the kernel mode, allocate continuous physical memory blocks, map them to the user space address, directly write the audio device into the memory block, and the FFTW fast Fourier transform library reads the data from the same address to reduce the memory copy overhead; S14: Establish a memory pool management system based on the maximum sampling rate (unit: Hz), the number of buffer blocks , the number of samples per block and the number of bytes per sample to optimize the memory allocation strategy, and the specific formula is expressed as: . In this embodiment, for a sampling rate of 192 kHz, a block size of 4096 samples, and 8 buffer blocks, the memory pool size is: 192000 × 4096 × 8 × 4 (bytes) = 25,165,824,000 bytes ≈ 24,000 MB. In this embodiment, we adopt the LRU cache policy to eliminate old data blocks according to the recent usage frequency.
[0027] Further, in step S2, adaptive frame splitting is performed on the input audio data, as Figure 3 shown, including: S201: Set the base sampling rate and the base frame length. In this embodiment, set the base sampling rate (Base_SR) = 48 kHz and the base frame length (Base_Frame) = 4096. At the same time, establish a mapping table (SR_Frame_Map) of the sampling rate and the frame length; S202: Use a sliding window to detect the actual sampling rate of the current audio stream. The window size in this embodiment is set to 5 audio data packets. The actual sampling rate of the current audio stream is obtained by calculating the number of samples per unit time: Current_SR = (number of received samples × clock frequency) / timestamp difference. Then, dynamically calculate the optimal frame length for frame splitting according to the detected sampling rate change. In this embodiment, it is set that when the detected sampling rate change exceeds ±5%, trigger the recalculation of the frame length. The formula for the optimal frame length Optimal_Frame is expressed as: Optimal_Frame = Base_Frame × (Current_SR / Base_SR), where the calculation result takes the closest 2^n value (FFT optimization requirement). Therefore, for the base frame length of 4096 in this embodiment, the optimal frame length of 192 kHz relative to 48 kHz is: 4096 × (192000 / 48000) = 16384; S203: Input the data into the buffer to receive the COAX audio stream. Perform data chunking according to the optimal frame length Optimal_Frame, add a 50% overlapping area (8192 samples) to each data frame, and perform windowing processing using the Hann window, expressed as: W[n] = 0.5(1 - cos(2πn / (N - 1))), where N is the current frame length; S204: Start the dynamic degradation policy according to the CPU load to calculate the temporary frame length for frame splitting, and implement the adaptive compensation mechanism.
[0028] On this basis, a multi-level cache system is adopted to optimize the data access efficiency, as Figure 4 shown, including: S211: Build a three - level cache architecture, including L1 Cache (SRAM), L2 Cache (DDR), and L3 Cache (NVMe), which are used to store original sampling samples, pre - processed frequency - domain data, and common calculation result templates respectively; S212: Build a data - dependence graph, define the node and edge relationships of the data - dependence graph. Specifically, the defined nodes include data fingerprints, spectral feature vectors, and timestamp information, and the defined edge relationships include temporal dependence, frequency - domain association, and computational reuse. Use an incremental graph update algorithm to dynamically update the data - dependence graph, which is expressed as follows: where α = 0.7, β = 0.3, represents the current frame, represents the historical graph; S213: Optimize the cache replacement algorithm, perform short - term prediction and long - term pattern mining based on a machine - learning - based hybrid replacement strategy. In this implementation, for short - term prediction, we use LSTM to predict the data requirements for the next 3 frames, and for long - term patterns, we use the FP - Growth algorithm to mine frequent access patterns. Then, evaluate the cache value through a cache value evaluation model , and the model is expressed as: , where 、 and assign weights according to the actual situation, represents the access frequency, represents the most recent access time, represents the pre - calculation gain, , where is the predicted frame, is the current frame. Finally, perform cache replacement according to the set replacement decision; S214: Optimize the cache hit - rate control. Monitor the cache hit - rate in real - time through a sliding window. The cache hit - rate calculation is expressed as: Hit Rate = (Cache Hits) / (Cache Hits + Cache Misses), and based on this, implement an adaptive adjustment strategy to keep the cache hit - rate above 95%. In this embodiment, when Hit_Rate < 95%, increase the L1 cache capacity by 10% and increase the look - ahead step of the pre - fetching algorithm. When Hit_Rate > 98%, reduce the cache scale by 5% and reduce the pre - fetching frequency to save power. Then, use an improved Markov chain prediction model to implement the pre - fetching algorithm.
[0029] Then proceed to dynamically adjust the FFT (Fast Fourier Transform) calculation accuracy in step S3, as Figure 5 shown, including: S301: Customize the precision levels and the floating points used for each level. The precision levels include single precision and double precision. Dynamically select the precision level according to the input signal-to-noise ratio. For single precision, when the signal-to-interference-plus-noise ratio SNR ≥ 60 dB, use the floating point float32; for double precision, when the signal-to-interference-plus-noise ratio SNR < 60 dB, use the floating point float64. Adopt the sliding window energy detection method for noise floor estimation, calculate the signal power, and combine the noise floor estimation and the signal power for real-time signal-to-noise ratio evaluation to achieve mixed precision calculation; S302: Establish a precision adaptive mechanism, set the initial mode and the decision tree for precision level switching during runtime. Calculate the required precision based on the base precision and the signal-to-noise ratio factor. The base precision defaults to single precision, and the signal-to-noise ratio factor is calculated by converting decibels to amplitude ratio. Select the appropriate data type by comparing the calculated required precision with the customized precision levels. Specifically, the calculation formula is: Required Precision = Base Precision × Signal-to-Noise Ratio Factor. The initial mode is set to select the baseline precision according to the hardware capabilities (e.g., GPU prefers single precision), and the decision tree during runtime is: if (SNR_Factor < 60 || has_overflow) → Switch to double precision elif (SNR_Factor > 80 && CPU_Util < 70%) → Allow downgrade to single precision; S303: Configure the automatic type promotion rules to handle the situation where the input data does not match the planned precision. In this embodiment, - float32 → float64: Linearly extend the mantissa, - float64 → float32: Retain the high 23 significant bits, Optimize for the conversion overhead and use SIMD instructions to batch process precision conversions.
[0030] Furthermore, in step S3, the pre-computation and caching of the FFTW plan include implementing FFTW plan pre-computation and establishing a plan cache pool, as Figure 6 shown, where The implementation of FFTW plan pre-computation includes: S311: Determine the frequently used FFT sizes through historical data analysis and establish a size priority queue; S312: Divide the stages according to the size priority, quickly generate the basic FFTW plans for the highly popular sizes, and perform fine optimization on the core sizes through multi-threaded parallel processing; S313: Convert the optimized plan objects into in-memory binary images and store them, and record the hardware fingerprint in the plan metadata at the same time; The establishment of the planned cache pool includes: S314: Design the cache structure, divide the cache pool into memory pages of a fixed size, and each page stores FFT plans of the same size; S315: Eliminate the less frequently used cold plans that have not been used for a long time through the LRU-K elimination algorithm and set a threshold to retain the high-value plans. At the same time, automatically adjust the size of the cache pool according to the system memory pressure.
[0031] In this embodiment, for the commonly used transform sizes (4096, 8192, 16384 points), we pre-calculate and cache the corresponding FFTW plans.
[0032] In step S4, the clock synchronization optimization for the COAX digital audio signal is as Figure 7 shown, and includes: S401: Use the digital phase-locked loop DPLL to capture the actual sampling rate of the COAX interface in real time, compare it with the nominal value (192 kHz), and calculate the deviation value , and the calculation formula is as follows: where, is the nominal sampling rate, is the actual sampling rate; S402: Use the deviation value as the input to achieve dynamic adjustment of the phase-locked loop PLL parameters, including dynamic adjustment and optimization of the loop bandwidth (such as adaptively expanding from 1 kHz to 5 kHz) and damping coefficient of the phase-locked loop PLL; S403: For frames with a deviation value exceeding ±100 ppm, use the Farrow structure fractional delay filter to achieve fine adjustment of the sampling rate with minimum phase distortion. In this embodiment, for the deviation, the number of compensation clock cycles inserted is expressed as follows: For take the integer part, use linear interpolation to compensate for this integer number of samples, and accumulate the remaining sample errors to the next frame for processing.
[0033] On this basis, jitter elimination is performed, as Figure 8 shown, and includes: S411: Extract the clock deviation sequence {x_i} of N consecutive sampling points in the time domain (such as N = 1024), and calculate the jitter amplitude of the acquisition part , and the calculation formula is as follows: where, is the mean value. In this embodiment, in the 92 kHz signal, μ of {x_i} is measured to be 0 ps, and the jitter amplitude is 45 ps. If it is lower than the 50 ps threshold, no processing is triggered. The threshold is dynamically set based on the signal-to-noise ratio (SNR), and the jitter is adaptively detected and the jitter elimination is triggered according to the jitter amplitude and the threshold; S412: Jointly process the time domain and the frequency domain to optimize the jitter algorithm, where, For the time domain, apply the Kalman filter to predict the clock deviation trend and correct the sampling point position; For the frequency domain, perform sub-band energy analysis on the spectrum after FFT to suppress the high-frequency noise generated by jitter.
[0034] In step S5, the establishment of the multi-dimensional result cache matrix is as follows Figure 9 shown, including: S501: Construct a three-dimensional tensor cache result, and the dimensions include frequency binning, time window, and signal type. Adopt a hash map + skip list index to construct a fast indexing mechanism; S502: Use the historical access sequence, time interval, and signal type transition probability as input features to construct a cache prediction model, predict and output the set of cache key values that may be accessed within a certain period of time in the future, and implement a preloading strategy for the key values in the prediction result; S503: Establish an efficiency threshold strategy to dynamically eliminate inefficient cache entries according to the following formula: where, is the time-consuming (ms) for word calculation results, is the cache memory occupancy (MB), is the memory cost coefficient (default 0.1 / ms·MB). In this embodiment, a certain spectrum result is reused 5 times, the calculation time-consuming is 20 ms, and the memory occupancy is 2 MB: , exceeding the threshold by 200%, retain the cache; if the efficiency drops to 180%, it is marked as to be eliminated.
[0035] Next, establish the intelligent reuse decision tree, as Figure 10 shown, including: S511: Perform joint similarity evaluation on the input signal to extract features. The extracted features include frequency domain features (cosine similarity of Mel-frequency cepstral coefficients (MFCC)), time domain features (relative error of zero-crossing rate (ZCR)), and energy distribution (KL divergence measures the energy distribution difference). Dynamically calculate the similarity threshold according to the signal amplitude and accuracy requirements. The formula is as follows: where, represents the allowable error, Denote the signal amplitude. In this embodiment, for the accuracy requirement of -96 dB ( ), normalize the signal amplitude ( ), and the similarity threshold ; S512: Establish a multiplexing decision tree based on multi-layer decision logic, perform linear phase compensation on the multiplexed spectrum, and perform weighted fusion of the new and old results according to the signal-to-noise ratio to obtain the final multiplexing decision.
[0036] In step S6, monitor the performance indicators of the real-time monitoring system respectively, implement adaptive parameter adjustment based on the optimization objective function, and monitor the audio quality and system stability by combining subjective and objective evaluation mechanisms, including: S601: Periodically sample the CPU usage rate through the performance interface provided by the operating system, monitor the resident memory set of the process, and use the memory water level marking strategy to count the memory occupancy. Calculate the average performance indicators through a sliding window. The calculation formula is as follows: Among them, Denote the window size, weighing real-time performance and stability, Denote the real-time CPU utilization rate, measure and record the processing delay. When it is detected that the index volatility (standard deviation / mean) > 30%, automatically reduce the window to W = 20 to improve the sensitivity; S602: Establish a performance evaluation model based on the multi-objective optimization function, which is expressed as follows: Among them, 、 、 Are the target weight coefficients, Denote the ratio of the current delay to the SLA delay threshold, Denote the CPU usage rate, Denote the memory occupancy. The target weights of the multi-objective optimization function are automatically adjusted according to the system state. In this embodiment, the weight coefficients are assigned as 、 、 .
[0037] S611: Perform objective quality evaluation through the PEAQ audio quality evaluation algorithm. Input the time-frequency signals of the reference audio and the measured audio to calculate each MOV parameter. The MOV parameters include noise, distortion, and modulation difference. Map multiple MOV parameters to the audio quality score ODG through a non-linear regression model. The formula is expressed as: PEAQ score = f(ODG, MOV), so as to perform objective quality evaluation. In this embodiment, the lowest ODG threshold is set to -0.5; S612: Collect subjective scoring standard data, including listener screening, test environment, and scoring method, and set a confidence interval for data analysis; S613: Monitor system stability through stress testing and long-term stability verification. Continuously process 192 kHz audio for 24 hours to monitor system stability.
[0038] By implementing the above technical solutions, the high sampling rate audio processing performance can be significantly improved, especially the efficient parsing of 192 kHz COAX audio on WiiMPlus / Pro devices. Each step in the solution cooperates with each other to form a complete optimization system, which not only ensures the quality of audio processing but also realizes the efficient utilization of system resources.
[0039] By comprehensively upgrading and optimizing the FFTW library and combining multi-level performance optimization strategies, the key technical pain points in high sampling rate audio processing have been successfully solved. First, starting from the underlying optimization, by upgrading the FFTW library and enabling the SIMD instruction set, an efficient memory management mechanism has been established, laying a foundation for high-performance computing. At the data processing level, adaptive frame processing and a multi-level cache system have been introduced, significantly improving the data access efficiency. Especially when processing 192 kHz COAX audio, by implementing special clock synchronization optimization and jitter elimination algorithms, the accurate restoration of high-fidelity audio has been ensured. The core highlight of the present invention is the establishment of a complete computational result reuse mechanism. Through intelligent cache prediction and similarity calculation, the repeated operations have been greatly reduced, and the CPU occupancy rate has been reduced by an average of 40%. At the same time, an adaptive parameter adjustment mechanism has been adopted to realize the dynamic optimization allocation of computing resources, reducing the processing delay by more than 50%. On the premise of ensuring audio quality, through the PEAQ scoring system for all-round quality monitoring, the reliability and stability of the optimization effect have been ensured. This complete optimization solution not only significantly improves the performance of the WiiM series products in processing high-quality audio but also effectively reduces the device power consumption, providing users with a better audio experience.
[0040] This embodiment also provides a high sampling rate audio parsing optimization system, as Figure 11 shown, including: A core framework optimization module for upgrading the FFTW fast Fourier transform library and establishing an optimized memory management mechanism; A data preprocessing module for performing adaptive frame processing on the input audio data and optimizing the data access efficiency using a multi-level cache system; An FFT calculation optimization module for dynamically adjusting the FFT fast Fourier transform calculation accuracy and optimizing the FFT fast Fourier transform calculation by pre-calculating and caching FFTW plans; A signal processing module for clock synchronization optimization and jitter elimination of COAX digital audio signals; A dynamic multiplexing module for establishing a multi-dimensional result cache matrix and an intelligent multiplexing decision tree to reduce repeated operations through intelligent cache prediction and similarity calculation; A quality and performance monitoring module for real-time monitoring of system performance metrics and adaptive parameter adjustment based on an optimization objective function, and monitoring of audio quality and system stability by combining subjective and objective evaluation mechanisms.
[0041] Among them, Figure 11 The functional implementation of each module and unit in the high and medium sampling rate audio parsing optimization system corresponds to each step in the above-mentioned embodiment of the high sampling rate audio parsing optimization method, and its functions and implementation processes will not be elaborated here one by one.
[0042] In this embodiment, an electronic device is also provided, such as Figure 12 shown. The electronic device includes a processor 14 and a memory 13. The memory 13 stores machine-executable instructions that can be executed by the processor 14, and the processor 14 executes the machine-executable instructions to implement the above audio control method.
[0043] Furthermore, Figure 12 the electronic device shown also includes a bus 12 and a communication interface 11. The processor 14, the communication interface 11, and the memory 13 are connected through the bus 12.
[0044] Among them, the memory 13 may include a high-speed random access memory (RAM, Random Access Memory), and may also include a non-volatile memory, such as at least one disk memory. Through at least one communication interface 11 (which can be wired or wireless), a communication connection is established between this system network element and at least one other network element, and the Internet, wide area network, local area network, metropolitan area network, etc. can be used. The bus 12 can be an ISA bus, a PCI bus, an EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity of representation, Figure 12 only a bidirectional arrow is used in the figure, but it does not mean that there is only one bus or one type of bus.
[0045] The processor 14 may be an integrated circuit chip with the ability to process signals. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in the processor 14 or the instructions in the form of software. The above-mentioned processor 14 may be a general-purpose processor, including a central processing unit (CPU for short), a network processor (NP for short), etc.; it may also be a digital signal processor (DSP for short), an application specific integrated circuit (ASIC for short), a field-programmable gate array (FPGA for short), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in this embodiment. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with this embodiment can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by the combination of the hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 1001, and the processor 1000 reads the information in the memory 1001 and combines its hardware to complete the steps of the audio control method.
[0046] The present disclosure also provides a computer-readable storage medium. The computer-readable storage medium may be a non-volatile computer-readable storage medium, or the computer-readable storage medium may also be a volatile computer-readable storage medium. A computer program is stored in the computer-readable storage medium. When the computer program runs on a computer, the computer is enabled to execute the steps of the audio control method.
[0047] Finally, it should be noted that the above description is only a preferred embodiment of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions falling within the idea of the present invention belong to the protection scope of the present invention. It should be pointed out that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements should also be regarded as the protection scope of the present invention.
[0048] The various technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the various technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
Claims
1. A high sampling rate audio analysis optimization method, characterized in that: The steps include: S1: Upgrade the FFTW fast Fourier transform library and establish an optimized memory management mechanism; S2: After completing the core framework optimization of step S1, adaptively frame the input audio data and use a multi-level cache system to optimize data access efficiency; S3: Dynamically adjust the FFT calculation accuracy and optimize the FFT calculation by pre-calculating and caching the FFTW plan; S4: Clock synchronization optimization and jitter elimination for COAX digital audio signals; S5: Establish a multi-dimensional result cache matrix and intelligent reuse decision tree to reduce repeated operations through intelligent cache prediction and similarity calculation; S6: Monitor system performance indicators in real time and implement adaptive parameter adjustment based on the optimization objective function, combining subjective and objective evaluation mechanisms to monitor audio quality and system stability.
2. The high sampling rate audio analysis optimization method according to claim 1, characterized in that: In step S1, the upgrading of the FFTW fast Fourier transform library includes: S11: Replace the old version library files, deeply optimize the SIMD instruction set, and automatically detect and enable the supported SIMD instruction set when FFTW is initialized; S12: Dynamically adjust the number of threads based on CPU load to configure dynamic thread management and avoid resource contention; The establishment of an optimized memory management mechanism includes: S13: Implement a zero-copy data transfer mechanism through memory mapping technology to reduce data movement overhead, allocate continuous physical memory blocks, map them to user space addresses, write audio devices directly into the memory blocks, and the FFTW fast Fourier transform library reads data from the same address; S14: Establish a memory pool management system based on the maximum sampling rate, the number of buffer blocks, the number of samples in a single block, and the number of bytes in a single sample to optimize the memory allocation strategy.
3. The high sampling rate audio analysis optimization method according to claim 1, characterized in that: In step S2, the adaptive framing process of the input audio data further includes: S201: Setting a basic sampling rate and a basic frame length, and establishing a mapping table between sampling rate and frame length; S202: using a sliding window to detect the actual sampling rate of the current audio stream, and dynamically calculating the optimal frame length of the frame according to the detected sampling rate change; S203: receiving the COAX audio stream by inputting the data into the buffer, dividing the data into blocks according to the optimal frame length, adding an overlapping area to each data frame, and applying a Hanning window to perform windowing processing; S204: Initiate a dynamic degradation strategy according to the CPU load to calculate the temporary frame length of the divided frames, and implement an adaptive compensation mechanism.
4. The high sampling rate audio analysis optimization method according to claim 1, characterized in that: In step S2, the use of a multi-level cache system to optimize data access efficiency further includes: S211: construct a multi-level cache architecture, which is used to store original sampled samples, pre-processed frequency domain data and common calculation result templates; S212: construct a data dependency graph, define the node and edge relationships of the data dependency graph, and dynamically update the data dependency graph using an incremental graph update algorithm; S213: Optimize the cache replacement algorithm, perform short-term prediction and long-term regularity mining based on a hybrid replacement strategy based on machine learning, evaluate the cache value through a cache value evaluation model, and replace the cache according to the set replacement decision; S214: Optimize cache hit rate control, monitor the cache hit rate in real time through a sliding window and implement an adaptive adjustment strategy, and use an improved chain prediction model to implement a prefetch algorithm.
5. The high sampling rate audio analysis optimization method according to claim 1, characterized in that: In step S3, the dynamically adjusting the FFT fast Fourier transform calculation accuracy further includes: S301: Customize the precision level and the floating point used for each level, the precision level includes single precision and double precision, and the precision level is dynamically selected according to the input signal-to-noise ratio; S302: Establishing an accuracy adaptive mechanism, setting a decision tree for switching the initial mode and the accuracy level at runtime, calculating the required accuracy according to the basic accuracy and the signal-to-noise ratio factor, the Jingchu accuracy is single precision by default, the signal-to-noise ratio factor is calculated by converting decibels into amplitude ratios, and selecting a suitable data type by comparing the required accuracy with the customized accuracy level; S303: Automatic type promotion rules are configured to deal with the situation where input data does not match the planned precision, and SIMD instructions are used to batch process precision conversion.
6. The high sampling rate audio analysis optimization method according to claim 1, characterized in that: In step S3, the pre-calculation and caching of FFTW plans includes implementing FFTW plan pre-calculation and establishing a plan cache pool, wherein: The implementation of the FFTW plan pre-calculation includes: S311: Determine the FFT size used frequently through historical data analysis and establish a size priority queue; S312: According to the size priority division stage, the basic FFTW plan is quickly generated for the high-heat size, and the core size is finely optimized through multi-threaded parallel processing; S313: Convert the optimized plan object into a memory binary image and store it, and record the hardware fingerprint in the plan metadata; The establishment of the plan cache pool includes: S314: Design a cache structure, divide the cache pool into memory pages of fixed size, and each page stores FFT plans of the same size; S315: Use the LRU-K elimination algorithm to eliminate unpopular plans that have not been used for a long time and set a threshold to retain high-value plans. At the same time, the cache pool size is automatically adjusted according to the system memory pressure.
7. The high sampling rate audio analysis optimization method according to claim 1, characterized in that: In step S4, the clock synchronization optimization for the COAX digital audio signal further includes: S401: Capture the actual sampling rate of the COAX interface in real time through the digital phase-locked loop DPLL, compare it with the nominal value, and calculate the deviation value; S402: using the deviation value as input to dynamically adjust the parameters of a phase-locked loop (PLL), wherein the parameters of the phase-locked loop (PLL) include a loop bandwidth and a damping coefficient; S403: For frames whose deviation values exceed a set threshold, a Farrow structure fractional delay filter is used to achieve sampling rate fine-tuning with minimum phase distortion.
8. The high sampling rate audio analysis optimization method according to claim 7, characterized in that: In step S4, the jitter elimination further comprises: S411: extracting a clock deviation sequence of a plurality of consecutive sampling points in the time domain, calculating the jitter amplitude of the acquisition part, dynamically setting a threshold based on the signal-to-noise ratio, adaptively detecting jitter according to the jitter amplitude and the threshold, and triggering jitter elimination; S412: Jointly process the time domain and the frequency domain to optimize the jitter algorithm, wherein: For the time domain, a filter is applied to predict the clock deviation trend and correct the sampling point position; In the frequency domain, sub-band energy analysis is performed on the spectrum after FFT to suppress high-frequency noise caused by jitter.
9. The high sampling rate audio analysis optimization method according to claim 1, characterized in that: In step S5, the establishing of the multi-dimensional result cache matrix further includes: S501: construct a three-dimensional tensor cache result, the dimensions of which include frequency binning, time window, and signal type, and use hash mapping + skip table index to construct a fast indexing mechanism; S502: Build a cache prediction model as input features, including historical access sequence, time interval, and signal type transition probability, predict and output a cache key value set that may be accessed in the future, and implement a preloading strategy for the key values in the prediction results; S503: Establish an efficiency threshold strategy to dynamically eliminate inefficient cache entries.
10. The high sampling rate audio analysis optimization method according to claim 1, characterized in that: In step S5, establishing the intelligent reuse decision tree further includes: S511: performing joint similarity evaluation on features extracted from the input signal, the extracted features including frequency domain features, time domain features and energy distribution, and dynamically calculating a similarity threshold according to signal amplitude and accuracy requirements; S512: A multiplexing decision tree is established based on the multi-layer decision logic, linear phase compensation is performed on the multiplexing spectrum, and new and old results are weightedly integrated according to the signal-to-noise ratio to obtain a final multiplexing decision.
11. The high sampling rate audio analysis optimization method according to claim 1, characterized in that: In step S6, the real-time monitoring of system performance indicators and implementing adaptive parameter adjustment based on the optimization objective function further includes: S601: periodically sample the CPU usage through the performance interface provided by the operating system, monitor the resident memory set of the process and use the memory watermark strategy to count the memory usage, calculate the average performance index through the sliding window, and measure and record the processing delay; S602: Establishing a performance evaluation model based on a multi-objective optimization function, wherein the objective weights of the multi-objective optimization function are automatically adjusted according to the system status.
12. The high sampling rate audio analysis optimization method according to claim 1, characterized in that: In step S6, the combining of subjective and objective evaluation mechanisms to monitor audio quality and system stability further comprises: S611: Perform objective quality assessment through PEAQ audio quality assessment algorithm, input reference audio and time-frequency signals of the tested audio to calculate each MOV parameter, wherein the MOV parameter includes noise, distortion and modulation difference, and map multiple MOV parameters to the audio quality score ODG through a nonlinear regression model; S612: Collect subjective scoring criteria data, including audience screening, test environment and scoring method, and set confidence intervals for data analysis; S613: Monitor system stability through stress testing and long-term stability verification.
13. A high sampling rate audio analysis optimization system, characterized in that: include: Core framework optimization module, used to upgrade the FFTW fast Fourier transform library and establish an optimized memory management mechanism; The data preprocessing module is used to perform adaptive framing of input audio data and optimize data access efficiency using a multi-level cache system; FFT calculation optimization module, used to dynamically adjust the FFT fast Fourier transform calculation accuracy and optimize the FFT fast Fourier transform calculation by pre-calculating and caching the FFTW plan; Signal processing module for clock synchronization optimization and jitter elimination of COAX digital audio signals; Dynamic reuse module, used to establish a multi-dimensional result cache matrix and intelligent reuse decision tree, reducing repeated operations through intelligent cache prediction and similarity calculation; The quality and performance monitoring module is used to monitor system performance indicators in real time and implement adaptive parameter adjustment based on the optimization objective function, combining subjective and objective evaluation mechanisms to monitor audio quality and system stability.
14. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the high sampling rate audio analysis optimization method as described in any one of claims 1-12 is implemented.
15. An electronic device, characterized in that: include: one or more processors; A storage device for storing one or more programs, which, when executed by the one or more processors, enables the one or more processors to implement the high sampling rate audio analysis optimization method as described in any one of claims 1-12.
Citation Information
Cited By
Method and system for realizing Device mode of USB audio equipment, storage medium and equipment
CN120872885A
Self-adaptive real-time reverberation processing method and system based on intelligent perception
CN121028580A
Intelligent sensing-based adaptive real-time reverberation processing method and system
CN121028580B
Large model dynamic loading reasoning method and system for terminal equipment
CN121365741A
A terminal device-oriented large model dynamic loading inference method and system
CN121365741B