Method and system for controlling rising and falling of PCM (Pulse Code Modulation) audio sampling rate

Through the deep learning model, the target sampling rate is predicted and resampled is performed, which solves the problem of lack of intelligent decision-making in sampling rate adjustment in the prior art, and the optimization balance between sound quality and computing resources is achieved.

CN120148531AInactive Publication Date: 2025-06-13SHENZHEN HAILINGWEI ELECTRONICS CO LTD
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510622573.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-15
Publication Date
2025-06-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

PCM audio sampling rate adjustment in existing Bluetooth audio devices mainly relies on fixed rules or static parameter configurations, lacks intelligent decision-making capabilities, and cannot fully utilize the characteristics of audio signals, resulting in poor balance between sound quality and computing load.

Method used

By obtaining audio signal characteristics and hardware state parameters, the target sampling rate is dynamically predicted using deep learning models, and resampling is performed using filtering processing methods and interpolation algorithms to optimize the sampling rate of audio data.

Benefits of technology

It realizes the optimization of computing resources and transmission efficiency while ensuring high-quality playback, reduces computing load, reduces sound quality loss, and improves the playback stability of Bluetooth audio devices in complex wireless environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120148531A_ABST
    Figure CN120148531A_ABST
Patent Text Reader

Abstract

The invention provides a rising and falling control method and system for a PCM audio sampling rate, and belongs to the technical field of audio processing, and the method comprises the following steps: obtaining audio signal features of pulse code modulation PCM audio data, the audio signal features comprising spectrum energy distribution, instantaneous power, a signal-to-noise ratio and a peak-to-average ratio; monitoring hardware state parameters of the Bluetooth audio equipment in real time, wherein the hardware state parameters comprise a current representative processor load rate, a bandwidth occupancy rate and a current coding and decoding mode parameter; according to the predicted value of the target sampling rate, performing resampling processing on the PCM audio data by adopting a filtering processing method and an interpolation algorithm to obtain PCM audio data matched with the target sampling rate; the resampling processing comprises up-sampling processing or down-sampling processing; and inputting the PCM audio data matched with the target sampling rate to an audio output module of the Bluetooth audio equipment for playing. The audio quality, the computing resources and the transmission efficiency are optimized through self-adaptive resampling processing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of audio processing, and particularly to a method and system for controlling the up and down conversion of PCM audio sampling rate. Background Art

[0002] With the rapid development of wireless audio technology, Bluetooth audio devices (such as Bluetooth headsets, wireless speakers) have been widely used in the consumer electronics market. Users' demands for the sound quality, transmission stability, and low latency of Bluetooth audio are increasing. However, due to the limited hardware resources of Bluetooth devices and the complex wireless transmission environment, the processing of audio data needs to strike a balance among sound quality, computing resources, and bandwidth adaptation. The sampling rate control of PCM (Pulse Code Modulation) audio data is a key link in audio signal processing. Reasonable sampling rate adjustment can optimize audio quality, improve transmission efficiency, and reduce the computing load of the device. Therefore, during the Bluetooth audio transmission and playback process, the dynamic sampling rate control technology for PCM audio data has received increasing attention.

[0003] Currently, the adjustment of the PCM audio sampling rate in Bluetooth audio devices mainly relies on fixed rules or static parameter configurations. Common methods include presetting a fixed sampling rate, simple bandwidth adaptive adjustment, and interpolation and filtering based on fixed algorithms. For example, some devices use a fixed sampling rate of 44.1kHz or 48kHz without considering the audio signal characteristics or device status; some systems adjust the sampling rate by detecting the Bluetooth bandwidth occupancy and using a simple threshold judgment strategy, such as reducing the sampling rate when the bandwidth is low, but this method ignores the dynamic characteristics of the audio signal. In addition, traditional resampling methods usually use fixed algorithms such as linear interpolation and window function filtering, which lack pertinence for different types of audio signals and are difficult to achieve an optimal balance among sound quality, latency, and computing resource consumption.

[0004] Although the existing audio sampling rate adjustment technologies can meet the requirements of Bluetooth audio devices to a certain extent, there are still many problems. First, the existing methods lack intelligent decision-making capabilities and cannot make full use of the characteristics of audio signals, resulting in a poor balance between sound quality and computing load. For example, in low-complexity audio scenarios, a high sampling rate is still used, causing waste of computing resources, while in high-complexity audio scenarios, the fixed sampling rate cannot meet the sound quality requirements. Second, the existing technologies have poor adaptability to the status of Bluetooth devices and are difficult to perceive the processor load, Bluetooth bandwidth fluctuations, and codec mode changes in real time, resulting in audio loss, increased latency, or excessive power consumption. Finally, the traditional resampling methods based on fixed filters and interpolation algorithms cannot be optimized for different audio signal characteristics, introducing additional frequency distortion or aliasing and affecting the final audio quality. Summary of the Invention

[0005] In view of this, the purpose of the embodiments of the present invention is to provide a method and system for controlling the increase and decrease of the PCM audio sampling rate to solve at least one of the above technical problems.

[0006] To achieve the above object, in a first aspect, the present invention provides a method for controlling the increase and decrease of the PCM audio sampling rate, the method comprising the following steps: Obtain the audio signal characteristics of the pulse code modulation (PCM) audio data, where the audio signal characteristics include spectral energy distribution, instantaneous power, signal-to-noise ratio, and peak-to-average ratio; Real-time monitor the hardware state parameters of the Bluetooth audio device, where the hardware state parameters include the current representative processor load rate, bandwidth occupancy rate, and current codec mode parameter; Input the audio signal characteristics and the hardware state parameters into a deep learning model, and dynamically output a predicted value of the target sampling rate through the deep learning model; According to the predicted value of the target sampling rate, perform resampling processing on the PCM audio data using a filtering processing method and an interpolation algorithm to obtain PCM audio data matching the target sampling rate; the resampling processing includes upsampling processing or downsampling processing; Input the PCM audio data matching the target sampling rate into the audio output module of the Bluetooth audio device for playback.

[0007] On the other hand, the present invention provides a system for controlling the increase and decrease of the PCM audio sampling rate, the system comprising: An audio signal characteristic acquisition module for acquiring the audio signal characteristics of the pulse code modulation audio data, where the audio signal characteristics include spectral energy distribution, instantaneous power, signal-to-noise ratio, and peak-to-average ratio; A hardware state monitoring module for real-time monitoring the hardware state parameters of the Bluetooth audio device, where the hardware state parameters include the processor computing load, Bluetooth bandwidth occupancy status, and current codec mode; A deep learning dynamic prediction module for inputting the audio signal characteristics and the hardware state parameters into a deep learning model, and dynamically outputting a predicted value of the target sampling rate through the deep learning model; A resampling processing module for performing resampling processing on the PCM audio data using a filtering processing method and an interpolation algorithm according to the predicted value of the target sampling rate to obtain PCM audio data matching the target sampling rate, where the resampling processing includes upsampling processing or downsampling processing; An audio output module for inputting the PCM audio data matching the target sampling rate into the audio output module of the Bluetooth audio device for playback.

[0008] The above technical solution has the following beneficial effects: A method for controlling the increase and decrease of the PCM audio sampling rate provided by the present invention combines audio signal characteristics (spectrum energy distribution, instantaneous power, signal-to-noise ratio, and peak-to-average ratio) with the hardware state parameters of a Bluetooth audio device (processor load rate, bandwidth occupancy rate, and codec mode parameters), uses a deep learning model to dynamically predict the optimal target sampling rate, and performs precise resampling processing through a filtering method and an interpolation algorithm, enabling audio data to optimize computing resources and transmission efficiency to the greatest extent while ensuring high-quality playback. Compared with the existing sampling rate adjustment schemes with fixed rules or static parameter configurations, this method can intelligently adapt to different audio signals and device states, effectively reduce the computing load, reduce sound quality loss, and improve the playback stability of Bluetooth audio devices in complex wireless environments, avoiding audio stuttering, distortion, or delay problems, and providing users with a better audio experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.

[0010] Figure 1 is a flowchart of a method for controlling the increase and decrease of the PCM audio sampling rate according to an embodiment of the present invention; Figure 2 is a flowchart of step S10 according to an embodiment of the present invention; Figure 3 is a flowchart of step S20 according to an embodiment of the present invention; Figure 4 is a flowchart of step S30 according to an embodiment of the present invention; Figure 5 is a flowchart of step S40 according to an embodiment of the present invention; Figure 6 is a flowchart of step S50 according to an embodiment of the present invention; Figure 7 is a functional block diagram of a system for controlling the increase and decrease of the PCM audio sampling rate according to an embodiment of the present invention; Figure 8 is a functional block diagram of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0011] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0012] In audio devices such as Bluetooth headsets or smart speakers, different sound sources, codec formats, and the highest audio specifications supported by the devices are different, resulting in the need for sample rate conversion during transmission and playback. The existing sample rate conversion methods mainly use fixed-ratio transformation or simple interpolation algorithms, which are prone to introducing quantization noise, nonlinear distortion, and delay, affecting the sound quality and user experience. The purpose of the embodiments of the present invention is to provide a method for controlling the increase and decrease of the PCM audio sample rate, which dynamically adjusts the sample rate according to the audio signal characteristics, system load, bandwidth conditions, and decoding capabilities of the playback device, improves the sound quality, reduces power consumption at the same time, and improves the playback adaptability of Bluetooth headsets or smart speakers.

[0013] The embodiments of the present invention provide an intelligent dynamic sample rate control method that combines audio signal characteristics and device status, which can predict the optimal sample rate based on deep learning technology and perform resampling processing adaptively to improve audio quality, optimize the utilization of computing resources, and ensure stable playback of Bluetooth audio devices in different environments.

[0014] Embodiment 1 As Figure 1 shown, the embodiments of the present invention provide a method for controlling the increase and decrease of the PCM (Pulse Code Modulation) audio sample rate, and the method includes the following steps: S10: Obtain the audio signal characteristics of the pulse code modulation PCM audio data, where the audio signal characteristics include spectral energy distribution, instantaneous power, signal-to-noise ratio, and peak-to-average ratio.

[0015] Specifically, the audio signal characteristics include: spectral characteristics, which are the spectral energy distribution obtained by analyzing audio data through short-time Fourier transform; instantaneous power, which is used to judge the dynamic range of the audio signal; signal-to-noise ratio, which is used to evaluate the signal quality; and peak-to-average ratio, which is used to detect the nonlinear distortion of the signal.

[0016] Specifically, in the scenario of playing music on a Bluetooth headset, the system first extracts audio signal features from the input PCM audio stream. For example, when the user plays a piece of music containing high-frequency instruments, the system analyzes the spectral energy distribution of the audio signal to identify the energy concentration regions in the high-frequency band; at the same time, it calculates the instantaneous power of the audio frame to quantify the dynamic range, detects the silent segments in the background noise region to determine the signal-to-noise ratio, and evaluates the amplitude fluctuation characteristics of the audio waveform through the peak-to-average ratio. These features together characterize the complexity of the current audio content and the demand for processing resources.

[0017] S20: Real-time monitor the hardware state parameters of the Bluetooth audio device, where the hardware state parameters include the current representative processor load rate, bandwidth occupancy rate, and the current codec mode parameters.

[0018] Specifically, based on the computing resources and transmission capabilities of the Bluetooth headset or smart speaker, determine whether it can perform high-quality resampling. High-quality means ensuring the minimum quantization error, aliasing distortion, and frequency distortion under an acceptable computational overhead, thereby improving audio fidelity.

[0019] Specifically, during the audio playback process, the system real-time monitors the hardware state of the Bluetooth headset. For example, when the processor load rate of the headset chipset rises to 60% due to multitasking (such as running a noise reduction algorithm simultaneously), the system records this load rate; at the same time, it obtains the current link bandwidth occupancy rate of 75% through the Bluetooth protocol stack interface and detects that the current SBC (Subband Coding) codec mode is used with a transmission bit rate of 328 kbps. These parameters reflect the real-time processing ability and transmission resource margin of the Bluetooth headset.

[0020] S30: Input the audio signal features and the hardware state parameters into a deep learning model, and dynamically output the predicted value of the target sampling rate through the deep learning model.

[0021] Specifically, based on the correlation analysis of audio features and hardware states, the system calls a pre-trained deep learning model for dynamic decision-making. For example, when the music contains high-frequency details and the processor load rate is high, the deep learning model predicts that the sampling rate needs to be reduced to relieve the computational pressure. For example, input the spectral energy distribution of the current audio (such as the proportion of the frequency band above 2 kHz is 30%), instantaneous power (0.8 dBFS), signal-to-noise ratio (45 dB), peak-to-average ratio (8:1), and hardware parameters (processor load rate 60%, bandwidth occupancy rate 75%, SBC coding mode) into the deep learning model, and the output target sampling rate is dynamically adjusted from the original 48 kHz to 32 kHz.

[0022] Specifically, when a low-complexity voice signal (such as a call) is detected, the prediction result is that downsampling to 8 kHz or 16 kHz is required to reduce the computational burden and power consumption; while for high-dynamic-range music content, the prediction result is to preferentially maintain the original sampling rate or increase it to 48 kHz or even 96 kHz to enhance the sound quality performance.

[0023] S40: According to the predicted value of the target sampling rate, perform resampling processing on the PCM audio data by using a filtering processing method and an interpolation algorithm to obtain PCM audio data that matches the target sampling rate; the resampling processing includes upsampling processing or downsampling processing.

[0024] Specifically, the system performs a resampling operation according to the target sampling rate. For example, when the target sampling rate of 32 kHz is lower than the original 48 kHz, downsampling processing is used: first, the high-frequency components above 16 kHz are filtered out by an anti-aliasing filter, and then the sampling points are thinned from 48 kHz to 32 kHz by using a linear interpolation algorithm to generate PCM data adapted to the target sampling rate. This process ensures that the audio after downsampling retains the effective frequency band information and avoids aliasing distortion.

[0025] S50: Input the PCM audio data that matches the target sampling rate into the audio output module of the Bluetooth audio device for playback.

[0026] Specifically, the resampled PCM audio data is transmitted to the audio output module via Bluetooth. For example, the system encapsulates the 32 kHz PCM (Pulse Code Modulation) stream in the SBC encoding format and transmits it via the Bluetooth link to the digital-to-analog converter (DAC) of the headset, which is converted into an analog signal and then drives the speaker to produce sound. This process reduces the processor load while maintaining the continuity of audio playback and avoiding stuttering or interruption due to resource overload.

[0027] Specifically, the audio output module may include: a decoding unit, a digital-to-analog conversion unit, and a speaker driving unit or a headset driving unit of a Bluetooth headset or a smart speaker. The decoding unit performs decoding processing on the PCM audio data that matches the target sampling rate to obtain the decoded PCM audio data. The digital-to-analog conversion unit receives the decoded PCM audio data and outputs an analog audio signal through digital-to-analog conversion to the speaker driving unit or the headset driving unit for playback.

[0028] In this embodiment, by dynamically analyzing the spectral characteristics of the audio signal and multi-dimensional parameters of the device hardware state, and using a deep learning model to adaptively predict the optimal sampling rate and perform precise resampling, while ensuring the effective retention of the audio core frequency band, the computational load of the processor and the transmission pressure of the Bluetooth link are reduced. At the same time, it avoids playback stuttering or sound quality deterioration caused by hardware resource overload, realizing the collaborative optimization of audio processing efficiency, device operation stability and auditory experience, and is especially suitable for high-load scenarios for processing complex audio content.

[0029] The intelligent sampling rate up / down control method of the embodiment of the present invention has technical advantages compared with traditional fixed-ratio transformation or simple interpolation schemes. By jointly analyzing in the time domain and frequency domain to optimize the sampling rate selection, the audio signal is played at the optimal resolution, reducing frequency distortion and aliasing, thereby improving the clarity and detail performance of the sound. Moreover, this method can intelligently and dynamically adjust the sampling rate, enabling Bluetooth headsets or smart speakers to still provide satisfactory sound quality in the low-power mode, effectively extending the battery life. In addition, the system can adjust the sampling rate in real time according to the Bluetooth transmission bandwidth, the computing power of the processor and the audio content, ensuring that the Bluetooth audio device automatically optimizes the playback effect in different environments, enhancing the adaptability and user experience. Compared with traditional high-order resampling algorithms (such as FIR filter interpolation), the embodiment of the present invention adopts an adaptive filtering and interpolation algorithm to optimize the calculation process, reducing the computational complexity, thereby reducing the delay and ensuring the synchronization and high stability of Bluetooth audio.

[0030] Further, the method may further include step S60: Based on the built-in sensors and system parameters of the playback end (speaker or headset), continuously monitor key audio performance indicators, such as output volume, current fluctuation, and temperature change, etc., for indirectly estimating the speaker load state. At the same time, combining known data such as the user's volume adjustment behavior and equalizer preference in different scenarios, and adopting a preset sampling rate adjustment template to perform a limited-level automatic switch during audio playback to adapt to different user needs and usage environments, thereby improving the overall audio experience without significantly increasing the system burden. The sampling rate adjustment template refers to several preset fixed audio parameter combination schemes, which are used to switch and use in different audio playback situations according to the device state or user behavior. Each template contains specific sampling rate (such as 44.1kHz, 48kHz), encoding format (such as SBC, AAC, LDAC), and bit rate and other parameters. For example, when the device load is low and the user prefers high-fidelity sound quality, enable the high-quality template and use a 48kHz sampling rate and AAC encoding; while when the system resources are tense or the user is in a commuting environment, automatically switch to the energy-saving template and use a 44.1kHz sampling rate and SBC encoding. Through template switching, a balance can be achieved between ensuring sound quality and system performance without complex real-time fine-tuning.

[0031] AsFigure 2 As shown, step S10 specifically includes the following steps: S11: Analyze the PCM audio data to obtain the number of channels of the PCM audio data. If the number of channels of the PCM audio data is greater than 1, convert the PCM audio data into single-channel audio data by using the method of weighted average or main channel extraction.

[0032] First, analyze the number of channels of the PCM audio data; if the number of channels of the PCM audio data is greater than 1, convert the PCM audio data into single-channel audio data by using the method of weighted average or main channel extraction; perform amplitude normalization processing on the single-channel audio data to limit the amplitude of the single-channel audio data between [-1, 1] to reduce the influence of amplitude overflow and quantization error.

[0033] S12: Perform time-frequency transformation processing on the single-channel audio data to extract the spectral energy distribution of the single-channel audio data.

[0034] After the preprocessing is completed, the system extracts the spectral features of the single-channel audio data through the Short-Time Fourier Transform (STFT) or Fast Fourier Transform (FFT) method. Specifically, the system frames the single-channel audio data with an overlapping sliding window (for example, a 25ms window, a 10ms step) to obtain multiple framed audio data. For each framed audio data, the system applies windowing (such as Hanning window or Hamming window) processing to reduce spectral leakage, obtaining multiple windowed framed audio data. Then, the system performs FFT operation on each windowed framed audio data to convert the time-domain signal into a frequency-domain complex spectrum. Calculate the corresponding Power Spectral Density (PSD) according to the frequency-domain complex spectrum, and then obtain the spectral energy distribution characteristics of different frequency bands by integrating multiple power spectral densities PSD over different frequency bands (low frequency, medium frequency, and high frequency).

[0035] Among them, the formula for calculating the Power Spectral Density (PSD) is as follows: ; Among them, X(f) is the frequency-domain signal, N is the length of the fast Fourier transform, that is, the number of sampling points used in the fast Fourier transform calculation, and P(f) is the power spectral density at a certain frequency f, that is, the distribution of signal power within a unit bandwidth near the frequency f.

[0036] S13: Calculate the instantaneous power of the single-channel audio data based on the framing process.

[0037] While extracting spectral features, the system calculates the instantaneous power based on the power variation of the single-channel audio data within a short period. The system frames the audio signal within a short time window (e.g., 10 ms) and calculates the instantaneous power within each time window. The specific method is to take the average of the squares of each sampling point within the window.

[0038] Among them, to calculate the instantaneous power P inst within each time window, the formula is as follows: ; where x[n] is the discrete sampling value of the audio signal, and N is the window length.

[0039] S14: Detect the silent section of the background noise and calculate the background noise power, calculate the total signal power of the single-channel audio data, and determine the signal-to-noise ratio of the single-channel audio data based on the background noise power and the total signal power.

[0040] During the audio feature extraction process, the system calculates the signal-to-noise ratio (SNR) by analyzing the background noise level in the single-channel audio data. First, the system identifies the silent section of the background noise through a silent section detection algorithm (such as the energy threshold method or the detection method based on spectral subtraction). In the silent section, the system calculates the root mean square (RMS) value of the background noise. The specific formula is as follows: ; where x silence [m] is the detected silent section signal, and M represents the number of sampling points in the silent section (background noise section), that is, the total number of signal samples used to calculate the background noise power.

[0041] Subsequently, the system calculates the total power of the audio signal within each time window and combines the power estimation of the background noise to calculate the signal-to-noise ratio (SNR) of the current window. The system smooths the signal-to-noise ratio through the moving average method to eliminate the influence of short-term fluctuations.

[0042] Among them, the formula for calculating the signal-to-noise ratio is as follows: ; S15: Determine the peak-to-average ratio based on the ratio of the peak value to the root mean square value of the single-channel audio data.

[0043] During the signal analysis process, the system extracts the crest factor by calculating the ratio of the peak value to the root mean square (RMS) value of the single-channel audio data. First, the system detects the maximum peak value of the audio signal within each window. Then, the system calculates the root mean square (RMS) value of the audio signal within the same window. Next, the system takes the ratio of the maximum peak value to the RMS value as the crest factor. The crest factor reflects the instantaneous dynamic range of the audio signal. The higher the peak value, the larger the dynamic range.

[0044] In some embodiments, step S12 may specifically include: S121: Performing frame segmentation on the single-channel audio data using a sliding window with a preset duration to obtain multiple segmented audio data; S122: Performing windowing processing on each segmented audio data to obtain multiple windowed segmented audio data; S123: Calculating the frequency-domain complex spectrum of each windowed segmented audio data, and calculating the corresponding power spectral density according to the frequency-domain complex spectrum; S124: Obtaining the spectral energy distribution of the single-channel audio data by integrating the multiple power spectral densities over different frequency bands.

[0045] In some embodiments, step S14 specifically includes: identifying the silent section by the energy threshold method; calculating the root mean square value of the background noise in the silent section; calculating the signal-to-noise ratio based on the total signal power within the sliding window and the background noise power.

[0046] As Figure 3 shown, in some embodiments, step S20 specifically includes: S21: Periodically obtaining the processor load rate of the Bluetooth audio device, establishing a time series of the processor load rate, and performing smoothing processing on the time series of the processor load rate to extract the current representative processor load rate.

[0047] The system first obtains the running state information of the current processor at the operating system level of the audio device. The system extracts the current processor load rate by calling the performance monitoring interface or the hardware monitoring module of the operating system, expressed as the CPU usage rate (i.e., the ratio of the occupied computing cycles to the total computing cycles). The calculation formula for the processor load rate is: ; where T active is the active time of the CPU within a time window, and T total is the total duration of the time window.

[0048] The system obtains the CPU load rate periodically (e.g., every 10 ms), establishes a time series of the CPU load rate, and performs smoothing processing through the moving average method to eliminate short-term jitter and obtain the representative processor load rate at the current moment. After obtaining the representative processor load rate, the system uses it as one of the input features of the model to dynamically reflect the current computing resource situation of the system. When the processor load rate exceeds a certain threshold (e.g., 80%), the system reduces the audio sampling rate to reduce the processor operation pressure.

[0049] S22: Monitor the bandwidth occupancy rate of the current Bluetooth link in real time through the Bluetooth protocol stack.

[0050] The system monitors the bandwidth occupancy of the current Bluetooth link in real time through the Bluetooth protocol stack (e.g., A2DP protocol).

[0051] The system first queries the maximum bandwidth B supported by the current Bluetooth device max . The system calculates the actually used bandwidth or the actual bandwidth usage within the current time window by monitoring the data transmission rate on the current Bluetooth link: ; where D tx represents the amount of data transmitted (in bits or bytes) within the time window length T win , and B used is the actually occupied bandwidth.

[0052] The bandwidth occupancy rate is expressed as: ; The system monitors the bandwidth occupancy rate periodically (e.g., every 10 ms) and records the historical trend of bandwidth occupancy. If the current bandwidth occupancy rate is close to the maximum value or jitters occur, the system will actively reduce the audio sampling rate to reduce the transmission pressure on the Bluetooth link and avoid audio stuttering or packet loss.

[0053] S23: Obtain the current codec mode parameters, where the current codec mode parameters include the codec format, the current codec bit rate, and the current total transmission delay.

[0054] In some embodiments, step S23 may specifically include: S231: Obtain the currently used audio codec type through the Bluetooth protocol stack interface.

[0055] The system obtains the currently used audio codec type in real time through the Bluetooth protocol stack interface (such as sub-band coding SBC, Advanced Audio Coding AAC, audio processing technology aptX, or low-latency audio codec LDAC). This step reads the codec scheme adopted by the current audio stream from the Bluetooth protocol stack, providing basic information for subsequent data extraction and analysis. The system accesses the A2DP (Advanced Audio Distribution Profile) protocol stack or the Bluetooth driver interface to determine the current audio codec type of the audio device.

[0056] S232: According to the audio codec type, extract the codec format, the current codec bit rate, and the inherent delay of the current codec, and obtain the Bluetooth protocol stack processing delay and the packet transmission delay through the Bluetooth protocol stack interface.

[0057] Based on the audio codec type obtained in step S231, the system further determines the current codec format (such as SBC, AAC, aptX, LDAC), the current codec bit rate (such as 256 kbps, 320 kbps), and the inherent delay of the current codec (such as 40 ms, 100 ms). For example, the maximum codec bit rate of SBC is 328 kbps, the maximum supported codec bit rate of AAC (Advanced Audio Coding) is 512 kbps, and the maximum supported codec bit rate of LDAC in high-quality mode is 990 kbps. This information will be used to calculate the maximum throughput and the initial transmission delay of the Bluetooth link.

[0058] S233: Calculate the maximum throughput of the current Bluetooth link according to the current codec bit rate, and calculate the total transmission delay according to the inherent delay of the codec, the Bluetooth protocol stack processing delay, and the packet transmission delay.

[0059] The system calculates the maximum throughput and delay characteristics of the current Bluetooth link based on the codec bit rate and the inherent delay of the codec extracted in step S232. The maximum throughput refers to the maximum data transmission rate that the Bluetooth link can support under the current codec format and codec bit rate, and the total transmission delay refers to the cumulative transmission time of the audio data from input to output, including the inherent delay of the codec, the Bluetooth protocol stack processing delay, and the packet transmission delay. The system calculates the throughput of the Bluetooth link using the current codec bit rate, and combines the inherent delay of the codec, the Bluetooth protocol stack processing delay, and the data transmission characteristics to determine the total transmission delay of the current audio data, so as to optimize the audio transmission strategy.

[0060] As Figure 4 shown, in some embodiments, step S30 specifically includes the following steps: S31: Normalize the audio signal features and the hardware state parameters respectively, encapsulate them into an input feature tensor, and serialize them according to time steps to form a multi-dimensional input tensor.

[0061] The system normalizes features such as spectral energy, instantaneous power, signal-to-noise ratio, and crest factor, scales them to between [0, 1], and reduces the sensitivity of the model to feature scale changes. Then, the system encapsulates the above features into a specific data structure (such as a tensor or vector) according to the time sequence order and adds timestamp information to ensure the alignment of audio signal features and time sequence information during dynamic modeling.

[0062] After the system obtains the processor load rate, Bluetooth bandwidth occupancy status, and current codec mode, it performs time alignment and synchronization processing on the above hardware state parameters. Since the sampling frequencies of different state parameters are different, the system uses interpolation and time alignment algorithms to align all parameters to a unified time axis. Linear interpolation is used for time alignment, and the formula is: ; where t 1 , t 2 are two nearest sampling time points, and P(t 1 ), P(t 2 ) are the corresponding state parameter values. Through time alignment, the system can provide consistent time step feature inputs for subsequent deep learning models.

[0063] After the system completes the extraction of the above hardware state parameters and time synchronization, it normalizes and encapsulates each hardware state parameter to form an input tensor. The system performs min-max normalization on each parameter, and the formula is: ; where X is the original parameter value, and X norm is the value after normalization.

[0064] After the system completes the extraction of audio signal features in step S10 and the acquisition of hardware state parameters in step S20, it first normalizes and encapsulates these data into a complete input feature vector. The input feature vector includes two parts: audio signal features and hardware state parameters. The audio signal features include spectral energy distribution (E low , E mid , E high ), instantaneous power (P inst ), signal-to-noise ratio (SNR), and crest factor (CF). E low , E mid and E highRepresent the spectral energy distribution of different frequency bands, which is used to describe the energy characteristics of audio signals in the low, middle, and high frequency ranges. They are obtained by calculating the power spectral density (PSD) after performing short-time Fourier transform (STFT) or fast Fourier transform (FFT) on the audio signal and integrating it over different frequency intervals. Specifically, the system first performs frame processing on the single-channel audio data (e.g., 25ms window, 10ms step), applies a window (e.g., Hanning window or Hamming window) to each frame of audio data to reduce spectral leakage, and then performs the FFT transform to convert the time-domain signal into a frequency-domain complex spectrum. Next, the power spectral density (PSD) of the spectrum is calculated, and then, according to the preset frequency partition, the PSD is numerically integrated separately in the low frequency (Elow), middle frequency (Emid), and high frequency (Ehigh) ranges to obtain the energy values of the three frequency bands. These features reflect the energy distribution of the audio signal in different frequency bands and are helpful for analyzing the signal characteristics. The hardware state parameters include the processor load rate (L CPU ), the Bluetooth bandwidth occupancy rate (B rate ), the current codec format (C codec ), the codec bit rate (R bitrate ), and the current total transmission delay (D latency ). The system serializes the above feature data by time step to form a multi-dimensional input tensor T: T = [E low ,E mid ,E high ,P inst ,SNR,CF,L CPU ,B rate ,C codec ,R bitrate ,D latency ; Among them, the data of each time step is a vector with a length of 11. The system encapsulates the features of multiple time steps within a period of time into a complete input sequence for input into the deep learning model.

[0065] S32: Input the multi-dimensional input tensor into a pre-trained deep learning model to obtain the predicted value of the target sampling rate.

[0066] The structure of the deep learning model is designed as follows: After receiving the input feature vector, the system uses the deep learning model to dynamically predict the target sampling rate. The deep learning model adopts time series modeling and multi-feature joint modeling strategies. The model structure includes an input layer, a convolutional layer (optional), a recurrent layer, a fully connected layer, and an output layer. The input layer accepts the encapsulated feature tensor with a shape of (n, 11), where n is the number of time steps. If a Convolutional Neural Networks (CNN) is used, the system can add a one-dimensional convolutional layer after the input layer to extract the local patterns of the time series features. The expression is as follows: ; In the recurrent layer, a Long-Short Term Memory Networks (LSTM) or a Gate Recurrent Unit (GRU) is used to extract the global time dependencies. The modeling formula for this layer is: ; where, h t is the current hidden state, carrying the sequence information from the past to the current. W h and W x are the weight parameters of the model, and b is the bias term. is the hidden state at the previous moment, providing the memory of historical information; W h is responsible for adjusting the influence of the hidden state at the previous moment on the current state. W x is responsible for the influence of the current input T t on the hidden state; T t is the input data at the current moment, providing new information input; b is the bias term used to adjust the output of the model to make it more flexible; f( ) is a non-linear activation function (such as tanh or ReLU), used to increase the expressive power of the model so that it can learn more complex time series features.

[0067] The time series features output by the recurrent layer are non-linearly mapped through the fully connected layer to obtain the predicted value of the target sampling rate. The formula is as follows: ; where, W out and b out are the parameters of the output layer, and σ is the activation function (such as ReLU). The deep learning model optimizes the parameters through a loss function (such as Mean-Square Error, MSE for short) to ensure that the error between the predicted sampling rate result and the real demand is minimized.

[0068] The training and optimization process of the deep learning model is as follows: The system trains a deep learning model with historical data under different audio scenarios and hardware configurations. The training process includes data collection and label generation, data augmentation and normalization, loss function definition, and backpropagation and parameter update. First, audio features and hardware state features are extracted from historical data, and the actual sampling rate used is used as the label. Then, the feature data is normalized and randomly perturbed to enhance the generalization ability of the deep learning model. The loss function uses the mean square error, which is defined as: ; where y i is the true sampling rate, is the sampling rate predicted by the model, and N is the number of samples. The parameters of the deep learning model are backpropagated and gradient-updated through an adaptive optimization method (such as adaptive moment estimation) to ensure that the deep learning model can converge quickly and has good generalization performance.

[0069] S33: Quantize and map the predicted value of the target sampling rate to the closest value in the set of discrete sampling rates supported by the Bluetooth audio device to obtain the final target sampling rate.

[0070] The target sampling rate output by the model is a continuous value, but audio devices usually support a limited set of discrete sampling rates (such as 44.1 kHz, 48 kHz, 96 kHz, etc.). Therefore, the system maps the continuous prediction value to the closest available sampling rate through a quantization mechanism: ; where F is the set of sampling rates supported by the Bluetooth audio device, and f i represents a specific sampling rate in the set F. Through this quantization mapping, the system ensures that the final output sampling rate is compatible with the device. f q represents the quantized sampling rate of the final output. f t represents the target sampling rate predicted by the deep learning model, which is a continuous value. Through this formula, the system selects the discrete sampling rate f t closest to the target sampling rate f i in the set F as the final output sampling rate f q , ensuring that the output result is compatible with the device.

[0071] Through time - series modeling and multi - feature joint modeling, the system can predict and dynamically adjust the target sampling rate in real - time according to audio features and device status, ensuring the quality of audio output and the stability of system operation. By simultaneously inputting audio features and hardware status features, the model can dynamically balance between audio quality and device load. During operation, the system adjusts the model output through a dynamic feedback mechanism to enhance the stability of the system in complex environments. In addition, through a quantization mechanism, the system achieves a seamless mapping between continuous sampling rate prediction and discrete sampling rate output, ensuring model - hardware compatibility.

[0072] In some embodiments, in step S31, the time - step serialization includes arranging the input feature vectors of multiple time steps in chronological order to construct a multi - dimensional input tensor with n time steps and m - dimensional (e.g., 11 - dimensional) feature vectors, where n represents the number of time steps and m represents the feature dimension corresponding to each time step.

[0073] As Figure 5 shown, in some embodiments, step S40 specifically includes: S41: Determine the ratio of the input sampling rate to the target sampling rate, where the input sampling rate is obtained from the format header information of the PCM audio data or the audio decoder interface, and the target sampling rate is predicted by a deep - learning model.

[0074] Before starting the resampling process, the system first determines the original input sampling rate f in of the input PCM audio data and the target sampling rate f target . The input sampling rate f in can be obtained by reading the format header information of the audio stream or the audio decoder interface. The target sampling rate f target comes from the result predicted by the deep - learning model in step S30. The system determines the resampling magnification by calculating the ratio R between the input sampling rate and the target sampling rate: ; When R > 1, it means that the audio needs to be upsampled; when R < 1, it means that the audio needs to be downsampled.

[0075] S42: Based on the ratio, pre - process the input PCM audio data to obtain pre - processed floating - point format audio data. The pre - processing includes converting the PCM audio data to floating - point format and selecting a corresponding band - pass filter according to the ratio to filter the floating - point format audio data.

[0076] Before resampling, the system preprocesses the input PCM audio data, including format conversion and data alignment. The system converts the audio data to a standard floating-point format (such as 32-bit floating-point numbers) according to the bit depth of the input audio (e.g., 16 bits, 24 bits, or 32 bits) to reduce quantization errors. During upsampling or downsampling, to avoid aliasing and spectral folding, the system needs to perform band-pass filtering (low-pass filtering or high-pass filtering) on the signal according to the Nyquist theorem.

[0077] If upsampling (R > 1) is performed, the system uses a low-pass filter to remove the high-frequency components in the input audio signal to prevent aliasing caused by the high-frequency components brought by the new interpolation points. If downsampling (R < 1) is performed, the system uses an anti-aliasing filter to limit the bandwidth of the input audio signal so that it does not exceed the Nyquist frequency at the new sampling rate: ; The low-pass filter used can be an FIR (Finite Impulse Response) or an IIR (Infinite Impulse Response) filter. The typical design method for the FIR filter is the window function method, and the window functions include the Hann window, Hamming window, and Blackman window. The frequency components of the input signal are limited within the target Nyquist frequency range through the filter to prevent frequency folding.

[0078] S43: Determine the interpolation algorithm according to the ratio, and resample the filtered floating-point format audio data to obtain floating-point format audio data matching the target sampling rate.

[0079] After completing the band-pass filtering, the system resamples the audio signal using an interpolation algorithm according to the ratio R of the input sampling rate to the target sampling rate. The interpolation algorithms can include: linear interpolation, polynomial interpolation, band-pass interpolation, fractional delay interpolation, etc. Polynomial interpolation fits adjacent points by constructing a high-order polynomial to obtain a smoother interpolation result. The methods include Lagrange interpolation and Newton interpolation. Band-pass interpolation is an interpolation method based on the Fourier transform that uses the finite bandwidth characteristic of the band-pass signal in the frequency domain to generate new sampling points. Fractional Delay Interpolation performs interpolation using a method based on a fractional delay filter at a non-integer interpolation ratio (i.e., fractional delay) to ensure the consistency of the time domain and the frequency domain. The system dynamically selects an appropriate interpolation method according to the sampling rate ratio R and the characteristics of the audio content.

[0080] S44: Convert the floating-point format audio data matching the target sampling rate to the target bit depth format data.

[0081] After the interpolation resampling process is completed, the system converts the audio data in floating-point format into a standard integer bit-depth format supported by the Bluetooth audio device, such as 16-bit or 24-bit integer PCM data. Specifically, in this step, the system scales and clips the boundaries of the floating-point format audio data according to the dynamic range of the target bit depth, multiplies the normalized floating-point value by the maximum integer amplitude value corresponding to the target bit depth (such as 32767 or 8388607), and then converts it into integer PCM data format through rounding operations and performing overflow protection, and finally outputs audio data that matches the target sampling rate and bit-depth format.

[0082] Furthermore, to avoid amplitude overflow, clipping, or dynamic range imbalance that may be introduced during the floating-point processing, the system can adopt volume normalization (Normalization) or automatic gain control (AGC) technology during the conversion process. According to the instantaneous amplitude characteristics and dynamic range of the audio signal, it adaptively adjusts the output gain to ensure that the audio signal maintains the best signal-to-noise ratio within the target bit depth range and maintains a reasonable crest factor, thereby ensuring the quality of audio playback.

[0083] S45: Encapsulate the data in the target bit-depth format into PCM audio data that matches the target sampling rate, and provide it to the audio output module for playback.

[0084] After completing resampling, filtering, and format conversion, the system encapsulates the new PCM audio data into standard audio data (such as WAV or RAW format), and transfers it to the audio output module.

[0085] The system can dynamically select the upsampling or downsampling mode according to the device operating state and the characteristics of the audio content to ensure the audio output quality and the stability of system operation. Through adaptive interpolation, the system can dynamically select the optimal interpolation method according to the audio content and real-time state, taking into account both sound quality and computational complexity. Through format conversion and gain optimization, the system ensures the dynamic range of the output signal during the resampling process, preventing distortion and clipping.

[0086] As Figure 6 shown, in some embodiments, step S50 specifically includes: S51: Store the PCM audio data that matches the target sampling rate in the dynamic audio buffer, and dynamically adjust the buffer size based on the Bluetooth link state, processor load rate, and audio stream rate.

[0087] Before data transmission, the system first stores PCM audio data that matches the target sampling rate in the dynamic audio buffer. The buffer size is dynamically adjusted according to the Bluetooth link status, processor load rate, and audio stream rate to ensure smooth audio playback and avoid audio stuttering or loss caused by data overflow or buffer depletion. Traditional systems use fixed buffers, while this method adopts an intelligent dynamic adjustment strategy, combining Bluetooth bandwidth, processor load rate, and latency requirements to optimize the buffer size and reduce power consumption and data transmission instability. The audio stream rate is calculated in the prior art by statistically counting the amount of PCM audio data transmitted or processed per unit time. The status of the Bluetooth link is the weighted sum result of the current Bluetooth link's bandwidth occupancy rate and the maximum throughput of the previous Bluetooth link.

[0088] The calculation formula for the buffer size is as follows: ; where, B buffer represents the current size of the dynamic buffer; B 0 is the basic buffer size set by the system, which is used to represent the minimum buffer capacity required in an ideal state; BLQ represents the status index of the Bluetooth link, which is a normalized value, and the numerical range is [0,1]. The closer its value is to 1, the more stable the link status; L CPU represents the load rate of the current processor, and the value range is [0,1]. The higher its value, the busier the processor; represents the change rate of the audio stream rate (ASR), which is used to reflect the stability of the audio data stream. The greater the change, the more obvious the data stream jitter; α, β, and γ are the weight adjustment coefficients of the corresponding factors, which can be set according to the actual system operation characteristics through experience or model training to achieve dynamic optimization and adjustment of the buffer size.

[0089] Specifically, when playing high-bitrate music on a Bluetooth headset, the system dynamically manages the audio buffer according to the real-time Bluetooth link quality and processor load rate. For example, when the user is in an environment with strong wireless interference such as the subway, the Bluetooth link status monitoring shows that the data transmission stability decreases (such as periodic packet loss), and at the same time, the processor load rate of the headphone chipset increases due to running the environmental noise reduction algorithm. At this time, the system automatically reduces the capacity of the dynamic audio buffer (for example, reducing from the original 500ms cache to 200ms) to reduce the audio stream transmission delay and avoid playback stuttering caused by data accumulation. On the contrary, when it is detected that the Bluetooth link is stable and the processor load rate is low, the system increases the buffer capacity to enhance the anti-burst interference ability.

[0090] S52: Read the PCM audio data that matches the target sampling rate from the dynamic audio buffer, and decode it according to the target sampling rate and codec format to obtain the decoded PCM audio data.

[0091] Specifically, the decoding unit reads the PCM audio data and decodes it according to the target sampling rate and codec format, adapting to different Bluetooth audio standards (such as SBC, AAC, aptX, LDAC). An adaptive decoding optimization strategy is adopted to dynamically adjust the decoding parameters based on the support capabilities of the Bluetooth device, the bandwidth status, and the target sampling rate, ensuring the best balance between sound quality and transmission efficiency. Traditional decoding methods use fixed-parameter decoding, while in this embodiment, intelligent adjustment is performed in combination with the target sampling rate and the device load condition to improve the decoding quality.

[0092] Specifically, the system reads the resampled PCM audio data from the dynamic buffer and performs decoding adaptation. For example, when the target sampling rate is adjusted to 32 kHz and the codec mode is AAC, the decoding sub-module calls the AAC decoder interface to reconstruct the compressed audio stream into PCM audio data at a sampling rate of 32 kHz. During this process, the decoder automatically adapts the frame length parameter of the target sampling rate (for example, the AAC frame length is adjusted from 1024 samples to 960 samples) to ensure that the decoded PCM audio data is strictly aligned with the target sampling rate.

[0093] S53: Perform dynamic gain control or volume normalization processing on the decoded PCM audio data to obtain PCM audio data in floating-point format.

[0094] Specifically, to ensure that the dynamic range of the audio output is suitable for the playback device, the system performs dynamic gain control and / or volume normalization processing on the decoded PCM audio data to prevent clipping and volume imbalance.

[0095] Specifically, during audio playback, the system performs dynamic gain control and volume optimization on the decoded PCM audio data. For example, when the user switches from a symphony with a large dynamic range to an audio program with a lower average power, the gain control sub-module based on the energy distribution characteristics and dynamic range parameters of the audio frame, real-time identifies the content type difference: for the transient peak energy of the music segment, automatically applies gain attenuation to suppress the risk of signal overload; for the average energy level of the voice segment, dynamically increases the gain coefficient to enhance voice intelligibility. At the same time, by integrating a standardized loudness algorithm, statistical analysis of the long-term loudness of the output audio is performed, and progressive gain compensation is applied in the time domain to keep the final output volume stable within a preset comfortable listening range, avoiding volume sudden changes or auditory fatigue caused by content switching.

[0096] S54: Convert the PCM audio data in floating-point format to PCM audio data in integer format.

[0097] The digital-to-analog conversion unit DAC of the audio device only supports 16-bit or 24-bit integer format, so the system needs to convert the PCM audio data in floating-point format to integer format.

[0098] Specifically, this step converts the floating-point format PCM audio data into integer format PCM audio data. For example, the system quantizes the floating-point PCM audio data (range [-1.0, 1.0]) after dynamic gain processing into a 16-bit signed integer format: by multiplying by 32767 and taking the integer part, integer data with a value range of [-32768, 32767] is generated. During this process, dithering technology is used to add low-level noise to avoid introducing harmonic distortion during the quantization process, especially to retain detailed information in low-volume audio segments (such as the soft sound of a piano).

[0099] S55: Input the integer format PCM audio data into a digital-to-analog conversion unit to generate an analog audio signal.

[0100] The converted PCM audio data is transmitted to the digital-to-analog conversion unit DAC and converted into an analog audio signal. The digital-to-analog conversion unit converts the digital signal into an analog signal. For example, the DAC chip of the headset receives 16-bit 32kHz PCM data and triggers the conversion operation at a fixed time according to the target sampling rate: each sample is converted into a 1-bit high-speed bit stream through a multi-stage modulator, and then the high-frequency quantization noise is filtered out through a low-pass filter to output a smooth analog voltage signal. This process strictly follows the timing requirements of the target sampling rate to ensure the time-domain continuity of signal reconstruction.

[0101] S56: Transmit the analog audio signal to a speaker driving unit or a headset driving unit for playback.

[0102] Specifically, the analog signal drives the speaker unit to produce sound. For example, the analog signal output by the DAC is amplified by an earphone amplifier circuit and then transmitted to the dynamic coil speaker unit. When playing a bass, the earphone amplifier dynamically adjusts the driving current according to the signal amplitude (for example, providing 20mA current at the peak), pushing the diaphragm to generate mechanical vibrations corresponding to the sound pressure. At the same time, the system monitors the temperature and load impedance of the earphone amplifier. If an abnormality is detected (such as an overload caused by a sudden drop in impedance), the current limiting protection is immediately triggered to maintain the output stability.

[0103] In some embodiments, the training process of the deep learning model includes: Extracting audio signal features and hardware state parameters from historical data under different audio scenarios and hardware configurations, and using the actual sampling rate used as a label; Normalizing the extracted audio signal features and hardware state parameters, and introducing random perturbations to obtain the normalized data as the input data for training; Using the mean square error as a loss function, calculating the error between the predicted sampling rate and the true sampling rate, and optimizing the weight parameters and bias parameters of the deep learning model based on the error; An adaptive optimization method is adopted for backpropagation and gradient update to adjust the weight parameters of the deep learning model.

[0104] In some embodiments, the quantization mapping of the predicted value of the target sampling rate adopts the following formula: ; where F is the set of discrete sampling rates supported by the Bluetooth audio device, f i is a specific sampling rate in the set F of discrete sampling rates, f t is the predicted value of the target sampling rate, f q is the finally output quantized sampling rate.

[0105] In some embodiments, the deep learning model includes an input layer, a convolutional layer, a recurrent layer, a fully connected layer, and an output layer; The input layer is used to receive the multi-dimensional input tensor, and this input tensor includes audio signal features and hardware state parameters; The convolutional layer is used to receive the multi-dimensional input tensor, perform local pattern extraction of temporal features using one-dimensional convolution, and output the feature map obtained after the convolution operation; The recurrent layer includes a long short-term memory network or a gated recurrent unit, and is used to receive the feature map output by the convolutional layer and perform global time dependence modeling on the feature map, and output the processed high-dimensional features of the time series; The fully connected layer is used to perform non-linear mapping on the output of the recurrent layer to generate the predicted value of the target sampling rate.

[0106] In some embodiments, the filtering process of selecting the corresponding band-pass filter according to the ratio specifically includes: When the ratio is greater than 1, a low-pass filter is used to remove the high-frequency components of the input signal; When the ratio is less than 1, an anti-aliasing filter is used to limit the bandwidth of the input signal not to exceed the target Nyquist frequency.

[0107] In some embodiments, the low-pass filter and the anti-aliasing filter adopt a preset type of filter, and the preset type of filter includes a FIR filter or an IIR filter; the FIR filter is designed by the window function method, and the window function includes at least one of a Hanning window, a Hamming window, or a Blackman window.

[0108] In some embodiments, determining the interpolation algorithm according to the ratio specifically includes: When the ratio is an integer, a linear interpolation algorithm is selected; when the ratio is a non-integer, a fractional delay interpolation algorithm is selected; the fractional delay interpolation algorithm includes constructing an interpolation kernel function based on a fractional delay filter, and the coefficients of the interpolation kernel function are generated by a preset interpolation polynomial.

[0109] Those skilled in the art can clearly understand that, for the convenience and conciseness of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the present invention. The specific working processes of the units and modules in the above system can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated here.

[0110] Embodiment 2 As Figure 7 shown, this embodiment provides a PCM audio sampling rate up / down control system 100, and the system 100 includes: An audio signal feature extraction module 10, configured to obtain the audio signal features of pulse code modulation (PCM) audio data, where the audio signal features include spectral energy distribution, instantaneous power, signal-to-noise ratio, and peak-to-average ratio; A hardware state real-time monitoring module 20, configured to monitor the hardware state parameters of a Bluetooth audio device in real time, where the hardware state parameters include the smoothed processing result of the processor load rate time series, the Bluetooth link bandwidth occupancy rate, and the current codec mode parameters; A deep learning dynamic prediction module 30, configured to input the audio signal features and the hardware state parameters into a deep learning model, and dynamically output a predicted value of the target sampling rate through the deep learning model; A resampling processing module 40, configured to perform upsampling or downsampling processing on the PCM audio data through a band-pass filter and an interpolation algorithm according to the predicted value of the target sampling rate, and generate PCM audio data matching the target sampling rate; An audio output module 50, configured to input the PCM audio data matching the target sampling rate into the audio output module of the Bluetooth audio device for playback.

[0111] By fusing audio content features and hardware real-time status data in multiple dimensions and combining the dynamic decision-making ability of deep learning models, this system realizes the intelligence and adaptability of the sampling rate adjustment strategy, effectively balancing audio quality fidelity and resource consumption in complex operating environments. At the same time, relying on the collaborative optimization of band-pass filtering and interpolation algorithms, it ensures signal integrity during the resampling process while reducing computational overhead, improving the anti-interference ability and output stability of Bluetooth audio devices in scenarios such as high load and network fluctuations, and ultimately achieving the collaborative optimization of system resource utilization, audio quality, and user experience.

[0112] In some embodiments, the audio signal feature extraction module 10 includes: A multi-channel conversion sub-module for parsing PCM audio data to obtain the number of channels of the PCM audio data, and generating single-channel audio data through weighted averaging or main channel extraction when the number of channels is greater than 1; A time-frequency transformation sub-module for performing time-frequency transformation on the single-channel audio data and extracting the spectral energy distribution; An instantaneous power calculation sub-module for calculating the instantaneous power of the single-channel audio data based on frame-by-frame processing; A signal-to-noise ratio analysis sub-module for calculating the background noise power by detecting the background noise mute section and determining the signal-to-noise ratio in combination with the total signal power; A peak-to-average ratio calculation sub-module for determining the peak-to-average ratio according to the ratio of the audio peak value to the root mean square value.

[0113] In some embodiments, the time-frequency transformation sub-module specifically includes: A frame-by-frame processing unit for framing the single-channel audio data using a sliding window with a preset duration; A windowing processing unit for performing windowing processing on the framed audio data; A frequency-domain analysis unit for calculating the frequency-domain complex spectrum and power spectral density of the windowed data; A frequency band integration unit for obtaining the spectral energy distribution by integrating the power spectral density over different frequency bands.

[0114] In some embodiments, the hardware status real-time monitoring module 20 includes: A load rate monitoring sub-module for periodically obtaining the processor load rate and establishing a time series, and extracting a representative load rate through smoothing processing; A bandwidth monitoring sub-module for real-time obtaining the current Bluetooth link bandwidth occupancy rate through the Bluetooth protocol stack interface; An encoding / decoding parameter parsing sub-module for extracting the current encoding / decoding format, encoding / decoding bit rate, and total transmission delay parameters.

[0115] In some embodiments, the encoding / decoding parameter parsing sub-module includes: An encoding / decoding type acquisition unit, configured to read the currently used audio encoding / decoding type through a Bluetooth protocol stack interface; A delay calculation unit, configured to calculate the total transmission delay according to the inherent delay of the codec, the processing delay of the protocol stack, and the packet transmission delay; A throughput calculation unit, configured to determine the maximum throughput of the Bluetooth link based on the current encoding / decoding bit rate.

[0116] In some embodiments, the deep learning dynamic prediction module 30 includes: A feature tensor construction sub-module, configured to normalize the audio signal features and hardware state parameters and encapsulate them into a multi-dimensional input tensor with n time steps and m-dimensional feature vectors; A model inference sub-module, configured to input the multi-dimensional input tensor into a pre-trained deep learning model to output a predicted value of the target sampling rate; A sampling rate quantization sub-module, configured to map the predicted value to the closest value in the set of discrete sampling rates supported by the Bluetooth device to obtain the final target sampling rate.

[0117] By normalizing and integrating multi-source heterogeneous audio features and hardware parameters, and constructing a time-series multi-dimensional feature tensor to input into the deep learning model, it effectively solves the prediction deviation problem caused by the difference in parameter dimensions and the lack of time-series dynamics in traditional static rule decision-making. At the same time, combined with the discrete sampling rates supported by Bluetooth audio devices for quantization mapping, it not only retains the flexibility of model prediction but also ensures a strict match between the output sampling rate and the actual capabilities of the hardware, thus achieving the unity of the accuracy and engineering feasibility of sampling rate decision-making in a complex and changing operating environment.

[0118] In some embodiments, the resampling processing module 40 includes: A sampling rate ratio calculation sub-module, configured to determine the ratio of the input sampling rate to the target sampling rate, where the input sampling rate is obtained from the format header information of the PCM audio data or the audio decoder interface, and the target sampling rate is predicted by the deep learning model; A floating-point preprocessing sub-module, configured to preprocess the input PCM audio data based on the ratio to obtain preprocessed floating-point format audio data, and the preprocessing includes converting the PCM audio data into a floating-point format and selecting a corresponding band-pass filter according to the ratio to filter the floating-point format audio data; An interpolation resampling sub-module, configured to determine an interpolation algorithm according to the ratio and resample the filtered floating-point format audio data to obtain floating-point format audio data matching the target sampling rate; A data conversion sub-module, configured to convert the floating-point format audio data matching the target sampling rate into target bit-depth format data; A format encapsulation sub-module is used to encapsulate the target bit-depth format data into PCM audio data matching the target sampling rate and provide it to the audio output module for playback.

[0119] Through the collaborative processing of multiple sub-modules, based on dynamically determining the ratio relationship between the input and the target sampling rate, this embodiment combines floating-point precision conversion and band-pass filtering to eliminate signal distortion, and selects an adaptive interpolation algorithm based on the sampling rate ratio, achieving a balance between high-fidelity preservation of the signal frequency-domain characteristics and computational efficiency during the resampling process. At the same time, through target bit-depth conversion and standardized encapsulation, the compatibility of the resampled PCM audio stream with downstream hardware modules is ensured, thereby maintaining the stability and signal integrity of audio output in complex sampling rate switching scenarios.

[0120] In some embodiments, the audio output module 50 includes: A buffer management sub-module is used to store PCM audio data matching the target sampling rate in a dynamic audio buffer and dynamically adjust the buffer size based on the Bluetooth link state, processor load rate, and audio stream rate; A dynamic decoding sub-module is used to read PCM audio data matching the target sampling rate from the dynamic audio buffer and decode it according to the target sampling rate and codec format to obtain the decoded PCM audio data; A gain control sub-module is used to perform dynamic gain control or volume regularization processing on the decoded PCM audio data to obtain floating-point format PCM audio data; A format conversion sub-module is used to convert the floating-point format PCM audio data into integer format PCM audio data; A digital-to-analog conversion sub-module is used to input the integer format PCM audio data into a digital-to-analog conversion unit to generate an analog audio signal; An output sub-module is used to transmit the analog audio signal to a speaker drive unit or a headphone drive unit for playback.

[0121] This embodiment realizes the stability of audio output in complex network environments and hardware resource fluctuation scenarios by dynamically adjusting the buffer capacity to adapt to Bluetooth link fluctuations and changes in the processor load rate, combined with decoding adaptation of the target sampling rate and codec format, optimization of audio signals with dynamic gain, and collaborative processing of digital-to-analog conversion, effectively suppressing problems such as playback stuttering, sudden volume changes, and signal distortion caused by insufficient bandwidth or limited computing power. At the same time, it ensures seamless connection between the resampled audio data and the terminal playback hardware, improving the coherence of the user experience and the controllability of the sound quality.

[0122] In some embodiments, the buffer management sub-module is configured to: When the detected Bluetooth bandwidth occupancy rate exceeds the threshold, automatically reduce the buffer size to reduce latency; When the processor load rate is lower than the preset safety threshold, increase the buffer capacity to improve the anti-jitter ability.

[0123] In some embodiments, the training data of the deep learning model includes: Optimal sampling rate records of various Bluetooth chip sets when processing audio with different spectral characteristics under different load states; Quantitative correlation data of real-time latency, power consumption, and sound quality loss during the dynamic resampling process.

[0124] Embodiment III See Figure 8 , the embodiment of the present application also provides an electronic device 600, which includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the PCM audio sampling rate up and down control method in the foregoing method embodiments.

[0125] The embodiment of the present application also provides a non-transitory computer-readable storage medium, which stores computer instructions for causing the computer to execute the PCM audio sampling rate up and down control method in the foregoing method embodiments.

[0126] Next, refer to Figure 8 , which shows a schematic structural diagram of an electronic device 600 suitable for implementing the embodiment of the present application. The electronic device 600 in the embodiment of the present application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 8 The shown electronic device 600 is only an example and should not impose any limitation on the functions and usage scope of the embodiment of the present application.

[0127] As Figure 8As shown, the electronic device 600 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 601, which may perform various appropriate actions and processes according to a program stored in the read-only memory (ROM) 602 or a program loaded from the storage device 608 into the random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the electronic device 600 are also stored. The processing device 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. The input / output (I / O) interface 605 is also connected to the bus 604.

[0128] Generally, the following devices may be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touchpad, keys, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 608 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 609. The communication device 609 may allow the electronic device 600 to communicate with other devices wirelessly or wiredly to exchange data. Although the figure shows the electronic device 600 having various devices, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices may be implemented or had.

[0129] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart may be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for performing the method shown in the flowchart. In such an embodiment, the computer program may be downloaded and installed from the network through the communication device 609, or installed from the storage device 608, or installed from the ROM 602. When the computer program is executed by the processing device 601, the above functions defined in the method of the embodiments of the present application are executed.

[0130] It should be noted that the computer-readable medium described above in the present disclosure can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. And in the present disclosure, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable signal medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above. The above computer-readable medium can be included in the above electronic device; it can also exist separately without being assembled into the electronic device.

[0131] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order from that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0132] As described above, it is only the specific implementation manner of the present disclosure, but the protection scope of the present disclosure is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present disclosure should be covered within the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure shall be subject to the protection scope of the claims.

Claims

1. A method for controlling the increase or decrease of a PCM audio sampling rate, characterized in that: The method comprises the following steps: S10: Acquire audio signal characteristics of pulse code modulation PCM audio data, where the audio signal characteristics include spectrum energy distribution, instantaneous power, signal-to-noise ratio, and peak-to-average ratio; S20: Real-time monitoring of hardware status parameters of the Bluetooth audio device, wherein the hardware status parameters include a current representative processor load rate, a bandwidth occupancy rate, and a current encoding and decoding mode parameter; S30: Inputting the audio signal feature and the hardware state parameter into a deep learning model, and dynamically outputting a predicted value of a target sampling rate through the deep learning model; S40: resampling the PCM audio data using a filtering method and an interpolation algorithm according to the predicted value of the target sampling rate to obtain PCM audio data matching the target sampling rate; the resampling process includes upsampling or downsampling; S50: Inputting the PCM audio data matching the target sampling rate into the audio output module of the Bluetooth audio device for playing.

2. The method according to claim 1, characterized in that Step S10 specifically includes: S11: parsing the PCM audio data to obtain the number of channels of the PCM audio data, and if the number of channels of the PCM audio data is greater than 1, converting the PCM audio data into single-channel audio data by using a weighted average or main channel extraction method; S12: performing time-frequency transformation processing on the single-channel audio data to extract spectrum energy distribution of the single-channel audio data; S13: Calculate the instantaneous power of the single-channel audio data based on frame processing; S14: Detecting a silent section of background noise and calculating background noise power, calculating a total signal power of the single-channel audio data, and determining a signal-to-noise ratio of the single-channel audio data according to the background noise power and the total signal power; S15: Determine a peak-to-average ratio according to a ratio of a peak value to a root mean square value of the single-channel audio data.

3. The method according to claim 2, characterized in that Step S12 specifically includes: S121: performing frame processing on the single-channel audio data using a sliding window of preset duration to obtain a plurality of framed audio data; S122: performing windowing processing on each frame of audio data to obtain a plurality of windowed frame of audio data; S123: Calculate the frequency domain complex spectrum of each of the windowed framed audio data, and calculate the corresponding power spectral density according to the frequency domain complex spectrum; S124: Obtaining the spectrum energy distribution of the single-channel audio data by integrating the multiple power spectrum densities in different frequency bands.

4. The method according to claim 1, characterized in that: Step S20 specifically includes: S21: periodically acquiring the processor load rate of the Bluetooth audio device, establishing a processor load rate time series, and performing smoothing processing on the processor load rate time series to extract a current representative processor load rate; S22: monitor the bandwidth occupancy rate of the current Bluetooth link in real time through the Bluetooth protocol stack; S23: Acquire current codec mode parameters, where the current codec mode parameters include a codec format, a current codec bit rate, and a current total transmission delay.

5. The method according to claim 4, characterized in that Step S23 includes: S231: Obtain the currently used audio codec type through the Bluetooth protocol stack interface; S232: extracting the codec format, the current codec bit rate, and the inherent delay of the current codec according to the audio codec type, and acquiring the Bluetooth protocol stack processing delay and the data packet transmission delay through the Bluetooth protocol stack interface; S233: Calculate the maximum throughput of the current Bluetooth link according to the current codec bit rate, and calculate the total transmission delay according to the inherent delay of the codec, the Bluetooth protocol stack processing delay and the data packet transmission delay.

6. The method according to claim 1, characterized in that Step S30 specifically includes the following steps: S31: normalizing the audio signal features and the hardware state parameters respectively, encapsulating them into input feature tensors, and serializing them according to time steps to form a multi-dimensional input tensor; S32: Inputting the multidimensional input tensor into a pre-trained deep learning model to obtain a predicted value of a target sampling rate; S33: quantize and map the predicted value of the target sampling rate to the closest value in the discrete sampling rate set supported by the Bluetooth audio device to obtain the final target sampling rate.

7. The method according to claim 6, characterized in that In step S31, the time step serialization includes arranging the input feature vectors of multiple time steps in chronological order to construct a multi-dimensional input tensor with n time steps and m-dimensional feature vectors, where n represents the number of time steps and m represents the feature dimension corresponding to each time step.

8. The method according to claim 1, characterized in that Step S40 specifically includes: S41: Determine a ratio of an input sampling rate to a target sampling rate, wherein the input sampling rate is obtained from format header information of the PCM audio data or an audio decoder interface, and the target sampling rate is predicted by a deep learning model; S42: preprocessing the input PCM audio data based on the ratio to obtain preprocessed floating-point format audio data, wherein the preprocessing includes converting the PCM audio data into a floating-point format, and selecting a corresponding bandpass filter according to the ratio to perform filtering processing on the floating-point format audio data; S43: determining an interpolation algorithm according to the ratio, and resampling the floating-point format audio data after filtering to obtain floating-point format audio data matching the target sampling rate; S44: converting the floating-point format audio data matching the target sampling rate into target bit depth format data; S45: Encapsulate the target bit depth format data into PCM audio data matching the target sampling rate, and provide the data to the audio output module for playback.

9. The method according to claim 1, characterized in that: Step S50 specifically includes: S51: storing PCM audio data matching the target sampling rate in a dynamic audio buffer, and dynamically adjusting the buffer size based on the Bluetooth link status, processor load rate, and audio stream rate; S52: Reading PCM audio data matching the target sampling rate from the dynamic audio buffer, and decoding according to the target sampling rate and encoding and decoding format to obtain decoded PCM audio data; S53: performing dynamic gain control or volume regularization processing on the decoded PCM audio data to obtain floating-point format PCM audio data; S54: converting the floating point format PCM audio data into integer format PCM audio data; S55: inputting the integer format PCM audio data into a digital-to-analog conversion unit to generate an analog audio signal; S56: Transmit the analog audio signal to a speaker driving unit or an earphone driving unit for playback.

10. A PCM audio sampling rate control system, characterized in that: The system comprises: An audio signal feature acquisition module, used to acquire audio signal features of pulse code modulated audio data, wherein the audio signal features include spectrum energy distribution, instantaneous power, signal-to-noise ratio, and peak-to-average ratio; A hardware status monitoring module is used to monitor the hardware status parameters of the Bluetooth audio device in real time, wherein the hardware status parameters include the processor computing load, the Bluetooth bandwidth occupancy status and the current encoding and decoding mode; A deep learning dynamic prediction module, used for inputting the audio signal features and the hardware state parameters into a deep learning model, and dynamically outputting a predicted value of a target sampling rate through the deep learning model; A resampling processing module, used to resample the PCM audio data using a filtering processing method and an interpolation algorithm according to the predicted value of the target sampling rate to obtain PCM audio data matching the target sampling rate, wherein the resampling processing includes upsampling processing or downsampling processing; The audio output module is used to input the PCM audio data matching the target sampling rate into the audio output module of the Bluetooth audio device for playback.

Citation Information

Cited By

  • Multi-mode audio authentic identification system and method based on time-space consistency characteristics

    CN120564759A

  • Spatial audio processing method and system

    CN121397452A

  • Real-time speech recognition method based on Bluetooth audio stream

    CN121545524A

  • Network call method and device, computer equipment, readable storage medium and program product

    CN121771173A

  • Internet calling methods, devices, computer equipment, readable storage media and program products

    CN121771173B