An audio output control method, apparatus, device, and storage medium
Through the combination of MCU control unit and deep neural network, the input voltage and audio signals are detected and analyzed, and the cross attention model is constructed for dynamic audio curve switching and feedback control, which solves the problem of instability of audio output and realizes the intelligence and precision of high-quality audio output.
Patent Information
- Application Number
- CN202510051131.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-13
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-01-13
AI Technical Summary
Traditional audio output control methods cannot be dynamically adjusted according to different power supply conditions and usage scenarios, resulting in unstable audio output quality, especially when voltage changes rapidly under fast charging technology, affecting the audio output quality. The existing solutions lack effective voltage adaptation mechanisms and audio feature extraction methods, resulting in distortion and noise interference in the audio output.
The MCU control unit detects the input voltage and audio signals, generates a mixed feature matrix, uses a deep neural network to perform feature extraction and time-frequency domain analysis, builds a cross-attention model to calculate the audio parameter correlation weight, generates an audio curve switching control sequence, and combines the power amplifier voltage dynamic adjustment and feedback control to perform nonlinear distortion compensation and noise suppression processing.
It realizes the intelligence and precision of audio output control, improves the signal-to-noise ratio and frequency response flatness of audio output, enhances the system's environmental adaptability, and improves the audio output quality.
Smart Images

Figure CN119485118B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of audio processing, and particularly to an audio output control method, apparatus, device, and storage medium. Background Art
[0002] Traditional audio output control methods usually adopt fixed audio processing parameters and cannot be dynamically adjusted according to different power supply conditions and usage scenarios, resulting in unstable audio output quality. Especially in the case of using fast charging technology, the rapid change of voltage will have a significant impact on the audio output quality.
[0003] Currently, the audio processing solutions on the market generally have problems such as insufficient sound quality control accuracy and weak anti-interference ability. Especially in a complex usage environment, due to the lack of effective voltage adaptation mechanisms and audio feature extraction methods, audio output is prone to problems such as distortion and noise interference, seriously affecting the user experience. Although there are also some audio optimization solutions in the prior art, most of these solutions are designed for single scenarios or fixed conditions and lack comprehensive consideration of multi-dimensional factors such as voltage changes and device states. At the same time, these solutions have limited capabilities in dealing with non-linear distortion and noise suppression and cannot meet the requirements of high-quality audio output. Summary of the Invention
[0004] The present invention provides an audio output control method, apparatus, device, and storage medium, effectively solving the problem of unstable audio output under different power supply conditions and realizing the intelligence and precision of audio output control.
[0005] In a first aspect, the present invention provides an audio output control method, which includes:
[0006] Detecting an input voltage signal and collecting an audio input signal through an MCU control unit to obtain a voltage parameter matrix and an audio parameter matrix, and combining the voltage parameter matrix and the audio parameter matrix to form a hybrid feature matrix;
[0007] Inputting the hybrid feature matrix into a deep neural network for feature extraction, performing time-frequency domain analysis on the audio signal to obtain an audio feature space;
[0008] Constructing a cross-attention model according to the audio feature space, calculating the correlation weight values between audio parameters, and generating an audio curve switching control sequence;
[0009] Based on the audio curve switching control sequence, dynamically adjusting the power amplifier voltage through a feedback controller and collecting power amplifier working state data;
[0010] Perform non - linear distortion compensation and noise suppression processing on the power amplifier operating state data to generate audio output quality parameters, where the audio output quality parameters include signal - to - noise ratio parameters, harmonic distortion parameters, and frequency response flatness parameters.
[0011] In a second aspect, the present invention provides an audio output control device, which includes:
[0012] An acquisition module, configured to detect an input voltage signal and acquire an audio input signal through an MCU control unit to obtain a voltage parameter matrix and an audio parameter matrix, and combine the voltage parameter matrix and the audio parameter matrix to form a mixed feature matrix;
[0013] A feature extraction module, configured to input the mixed feature matrix into a deep neural network for feature extraction, perform time - frequency domain analysis on the audio signal to obtain an audio feature space;
[0014] A calculation module, configured to construct a cross - attention model based on the audio feature space, calculate the correlation weight values between audio parameters, and generate an audio curve switching control sequence;
[0015] A dynamic adjustment module, configured to dynamically adjust the power amplifier voltage through a feedback controller based on the audio curve switching control sequence, and acquire the power amplifier operating state data;
[0016] A generation module, configured to perform non - linear distortion compensation and noise suppression processing on the power amplifier operating state data to generate audio output quality parameters, where the audio output quality parameters include signal - to - noise ratio parameters, harmonic distortion parameters, and frequency response flatness parameters.
[0017] In a third aspect, the present invention provides an audio output control device, including: a memory and at least one processor, where instructions are stored in the memory; the at least one processor invokes the instructions in the memory so that the audio output control device executes the above - mentioned audio output control method.
[0018] In a fourth aspect, the present invention provides a computer - readable storage medium, in which instructions are stored, and when the instructions are run on a computer, the computer is made to execute the above - mentioned audio output control method.
[0019] In the technical solution provided by the present invention, the MCU control unit realizes the collaborative detection of the input voltage and audio signal, combines with the deep neural network for feature extraction and analysis, significantly improves the accuracy and adaptive ability of audio output control. The cross-attention model is used to calculate the dynamic weights of audio parameters, realizes the intelligent switching of the audio curve, and effectively solves the problem of unstable audio output under different power supply conditions. Through the dynamic regulation of the power amplifier voltage and the feedback control mechanism, combined with the nonlinear distortion compensation and noise suppression processing, the signal-to-noise ratio and frequency response flatness of the audio output are greatly improved. Innovatively applying the deep learning algorithm to audio feature extraction and parameter optimization realizes the intelligence and precision of audio output control, enabling the system to have stronger environmental adaptability. Adopting the modular design idea, the interfaces between functional units are standardized, improving the scalability and maintainability of the system, and facilitating subsequent function upgrades and optimizations. Through multi-level signal processing and optimization strategies, the overall quality of the audio output is improved, meeting the high-quality audio output requirements in different application scenarios.
[0020] Other features and advantages of the present invention will be described in the following specification, and, in part, will be obvious from the specification, or will be understood by implementing the present invention. The objectives and other advantages of the present invention are achieved and obtained by the structures specifically pointed out in the specification, claims, and drawings.
[0021] To make the above objectives, features, and advantages of the present invention more obvious and understandable, the following specific preferred embodiments are given, and in conjunction with the accompanying drawings, the detailed description is as follows. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 It is a schematic diagram of an embodiment of the audio output control method in the embodiment of the present invention;
[0023] Figure 2 It is a schematic diagram of an embodiment of the audio output control device in the embodiment of the present invention;
[0024] Figure 3 It is a schematic diagram of an embodiment of the audio output control device in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0025] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0026] As used in the embodiments of the present invention, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally further include other unlisted steps or units, or may optionally further include other steps or units inherent to these processes, methods, products or devices.
[0027] For ease of understanding of this embodiment, first, a detailed introduction is given to an audio output control method disclosed in the embodiments of the present invention. As Figure 1 shown, the method includes the following steps:
[0028] 101. Detect the input voltage signal and collect the audio input signal through the MCU control unit to obtain a voltage parameter matrix and an audio parameter matrix, and combine the voltage parameter matrix and the audio parameter matrix to form a mixed feature matrix;
[0029] It can be understood that the execution subject of the present invention can be an audio output control device, or a terminal or a server, and specific limitations are not made here. In the embodiments of the present invention, the server is taken as an example of the execution subject for illustration.
[0030] Specifically, the MCU control unit identifies the fast charging protocol for the input voltage signal to obtain the protocol type identifier, and the protocol type identifier includes USB PD3.0, PD2.0, and BC1.2 boost fast charging protocol identifiers. The identification of the fast charging protocol is achieved through the preliminary analysis of the voltage signal. The MCU performs pattern matching and handshake signal analysis on the input signal according to the different characteristics of the protocol type, so as to determine the protocol standard followed by the current voltage signal and ensure that the system can adapt to different fast charging protocols. After the protocol identification is completed, the input voltage signal is sampled and detected according to the protocol type identifier to generate a voltage sampling sequence. The voltage sampling sequence is implemented by a high-precision ADC module, and its sampling frequency and resolution are dynamically adjusted according to the protocol requirements to capture the fluctuation characteristics and dynamic changes of the voltage signal. The voltage sampling sequence is transmitted to the feature extraction module. In this module, the characteristic parameters of the voltage, including information such as the average value, peak value, and frequency components of the voltage, are extracted through signal processing methods such as Fourier transform. These characteristics are integrated to form a voltage parameter matrix. At the same time, multi-channel sampling is performed on the audio input signal. Through multi-channel sampling technology, the time-domain data of the audio signal is separated into sampling data of different channels, and the data of each channel contains the detailed information of the audio signal on its respective sampling path. The sampled audio data is input into the audio feature extraction module, and the module processes the data through short-time Fourier transform and other frequency-domain analysis methods to extract the frequency response, phase characteristics, and dynamic range data of the audio. These data are integrated to form an audio parameter matrix. Among them, the frequency response data describes the amplitude change of the audio signal at different frequencies and reflects the overall spectral characteristics of the audio signal; the phase characteristic data records the phase distribution information of the audio signal and helps to restore the time-domain form of the audio; the dynamic range data reflects the minimum and maximum amplitude change ranges of the signal and evaluates the performance ability of the signal at different intensities. The voltage parameter matrix is normalized to obtain a normalized voltage matrix, and the audio parameter matrix is normalized to obtain a normalized audio matrix. By using the method of linear transformation, all parameter values are mapped to a unified numerical range (usually 0 to 1). The normalized voltage matrix and the normalized audio matrix are combined into a mixed feature matrix.
[0031] 102. Input the mixed feature matrix into a deep neural network for feature extraction, perform time-frequency domain analysis on the audio signal, and obtain the audio feature space;
[0032] Specifically, the normalized voltage matrix in the hybrid feature matrix is input into the feature extraction module of the deep neural network. This module consists of three convolutional layers. Each convolutional layer uses the ReLU activation function to enhance the non-linear mapping ability. At the same time, a batch normalization layer is introduced after each convolutional layer to accelerate convergence and prevent gradient vanishing. After multi-level processing of the voltage matrix through this structure, a voltage deep feature vector describing the characteristics of the voltage signal is extracted. At the same time, the normalized audio matrix is input into the spectral analysis module in the deep neural network. This module consists of five convolutional neural networks, and each convolutional neural network is followed by a batch normalization layer and a LeakyReLU activation function. LeakyReLU has better performance in dealing with the non-linear relationship of audio signals. It can retain a certain gradient in the negative value interval, thus effectively alleviating the "dead neuron" problem of the neural network. During the convolution process, the audio signal is parsed layer by layer, and its frequency distribution, timing characteristics, etc. are gradually extracted, and finally an audio deep feature vector containing rich spectral information is obtained. The voltage deep feature vector and the audio deep feature vector are feature fused. By concatenating features or using an attention mechanism to combine different dimensions of the two features, a fused feature vector is generated. The fused feature vector represents the joint characteristics of the voltage signal and the audio signal, capturing the potential complex correlations between the two. The fused feature vector is input into a four-layer fully connected neural network in the deep neural network. Using the feature mapping ability of the fully connected layer, the fused features are projected from the low-dimensional space to the high-dimensional space to generate a high-dimensional feature vector. Depth residual calculation is performed on the high-dimensional feature vector. The depth residual network effectively alleviates the gradient vanishing problem in the deep network by introducing skip connections, while retaining the correlation between the low-level information and the high-level information of the features, obtaining a time-frequency feature matrix. Temporal feature extraction is performed on the time-frequency feature matrix through time series analysis to capture the dynamic characteristics of the signal changing with time. Temporal feature extraction uses a recurrent neural network (RNN) or a long short-term memory network (LSTM) to ensure the model's ability to process time correlations. The output of the temporal feature extraction is a dynamic feature vector, which describes the performance of the signal in a dynamic scenario. The dynamic feature vector is input into the decoder network for feature reconstruction. The decoder network contains four transposed convolutional layers and two fully connected layers. The transposed convolutional layers gradually restore the original spatial information of the features, while the fully connected layers adjust the distribution of the output data. Through the processing of the decoder network, an audio feature space is finally generated. The audio feature space describes the frequency response characteristics, phase characteristics, and harmonic characteristics of the audio signal. The frequency response characteristics record the amplitude performance of the signal at different frequencies, reflecting the overall spectral distribution of the audio signal; the phase characteristics describe the phase relationship between different frequency components of the signal; the harmonic characteristics reflect the harmonic components of the audio signal.
[0033] 103. Construct a cross-attention model based on the audio feature space, calculate the correlation weight values between audio parameters, and generate an audio curve switching control sequence;
[0034] Specifically, perform feature projection on the frequency response feature data in the audio feature space. Through linear transformation, map the high-dimensional frequency response features into low-dimensional frequency feature vectors, reducing data redundancy while retaining key information. At the same time, perform feature projection on the phase feature data and harmonic feature data in the audio feature space, extract their main components respectively, and integrate them into an audio representation vector. Use the frequency feature vector as the query vector and the audio representation vector as the key-value vector, and calculate the similarity between the query vector and the key-value vector through the dot product operation in the cross-attention model to generate a similarity matrix. The result of the dot product operation reflects the correlation strength between the frequency features and the phase and harmonic features. Normalize the similarity matrix, and use the softmax function to map the values in the matrix to the range of 0 to 1 to obtain a normalized similarity matrix. Perform weighted summation of the normalized similarity matrix and the audio representation vector to generate an attention feature vector. Based on the attention feature vector, construct an audio feature correlation graph. The nodes of the correlation graph represent different audio parameters, and the weights of the edges represent the correlation strength between the parameters. Through the graph convolutional network, extract features from the correlation graph. The model can extract high-level global information from the local connection structure, generate an audio parameter correlation matrix, and describe the complex relationships of each audio parameter in different dimensions. Perform feature clustering on the audio parameter correlation matrix, use the clustering algorithm to divide similar audio parameters into different categories, and generate an audio curve category matrix. The audio curve category matrix assigns a representative feature to each category of parameters, and these features correspond to different optimization directions in the audio system. Based on the audio curve category matrix, match the curve templates in the preset audio curve library. The curve templates include gain curves, frequency curves, phase curves, etc. The matching process screens out the candidate curve sequences that best match the current audio features by measuring the similarity between the category matrix and the templates in the curve library. Screen and adjust the candidate curve sequences through similarity sorting and optimization. The similarity sorting uses a correlation metric method to sort the priorities of the curves according to their matching degrees with the audio features, and the sequence optimization adjusts the order of the curves through dynamic programming or greedy algorithms to maximize the overall quality of the output audio. The optimized sequence finally generates an audio curve switching control sequence. The control sequence includes gain control parameters, frequency adjustment parameters, and phase compensation parameters, which are used to adjust the gain level, frequency distribution, and phase characteristics of the output audio respectively.
[0035] 104. Based on the audio curve switching control sequence, dynamically adjust the power amplifier voltage through a feedback controller, and collect the power amplifier working state data;
[0036] Specifically, according to the gain control parameters, frequency adjustment parameters, and phase compensation parameters included in the audio curve switching control sequence, the input signal of the power amplifier is preprocessed. These parameters are mapped to the signal adjustment domain, and a preprocessing signal matrix is generated through a specific algorithm. The matrix includes gain compensation data, frequency adjustment data, and phase correction data. Each type of data specifically optimizes a key characteristic of the power amplifier input signal. For example, gain compensation is used to adjust the signal amplitude to match the target output requirements, frequency adjustment data ensures that the spectrum of the audio signal is consistent with the target frequency response, and phase correction data is used to reduce phase distortion to improve sound quality. The preprocessing signal matrix is input into the dynamic gain controller, and the dynamic gain controller calculates the gain adjustment coefficient through a feedback loop. The feedback loop adjusts the gain parameter to minimize the error by real-time monitoring the difference between the output signal and the target signal of the power amplifier, obtaining the power amplifier drive control signal. This signal directly determines the input drive strength of the power amplifier, thereby affecting its output characteristics. A power compensation operation is performed on the power amplifier drive control signal. The power compensation operation is based on the real-time voltage feedback value, and by dynamically adjusting the output power, it ensures that the power amplifier always provides a stable output under different load conditions. The operation result is presented in the form of a power compensation matrix, which includes voltage compensation values and current compensation values, respectively used to optimize the voltage and current supply of the power amplifier. According to the power compensation matrix, a power amplifier drive voltage signal is generated, which directly drives the power amplifier to work. At the same time, the output voltage data is collected through the voltage detection unit to generate a voltage feedback matrix. The voltage feedback matrix records the instantaneous value and the voltage change rate of the output voltage, and these data can reflect the real-time state of the power amplifier output. The voltage feedback matrix is signal-filtered through a loop filter. The filtered voltage characteristic sequence includes voltage stability data and voltage ripple data. The former describes the long-term stability of the output voltage, and the latter records the voltage fluctuations in a short period of time. Based on the voltage characteristic sequence, a power amplifier state monitoring model is established. The model highly accurately monitors the operating state of the power amplifier through multi-point sampling technology, generating a power amplifier state parameter matrix. The state parameter matrix includes temperature data, current data, and efficiency data, respectively describing the heat dissipation performance, current consumption characteristics, and energy conversion efficiency of the power amplifier. Through multi-dimensional analysis of the power amplifier state parameter matrix, the key parameters are evaluated using a threshold judgment circuit. For example, when the temperature data exceeds the safety threshold, the system will determine that the power amplifier has an overheating risk, thereby triggering a protection mechanism; when the efficiency data is lower than the specified threshold, the power compensation strategy is adjusted to optimize the energy utilization rate. The working state data of the power amplifier is generated through state evaluation. The power amplifier working state data includes voltage response characteristic data, temperature change characteristic data, and efficiency characteristic data, and is used to real-time adjust the audio output control parameters, forming an adaptive optimization closed-loop control system.
[0037] 105. Perform nonlinear distortion compensation and noise suppression processing on the power amplifier operating state data to generate audio output quality parameters, which include signal-to-noise ratio parameters, harmonic distortion parameters, and frequency response flatness parameters.
[0038] Specifically, perform a fast Fourier transform on the voltage response characteristic data in the power amplifier operating state data, and obtain a spectral analysis matrix through frequency domain analysis. The spectral analysis matrix describes the fundamental wave component and high-order harmonic component data of the input signal. The fundamental wave component represents the main energy concentration of the signal, while the high-order harmonic component reveals the frequency component shift caused by nonlinear distortion in the signal. Input the spectral analysis matrix into the nonlinear compensation model. The nonlinear compensation model iteratively optimizes the signal characteristics through the backpropagation algorithm in deep learning to calculate the compensation weights. The model determines the optimal distortion compensation coefficient matrix by minimizing the objective function (such as distortion error). The distortion compensation coefficient matrix includes amplitude compensation coefficients and phase compensation coefficients, which are used to correct signal amplitude distortion and phase distortion respectively, so as to make the restoration of the audio signal closer to the real effect. According to this matrix, perform nonlinear correction on the audio signal to generate a compensated signal sequence. The compensated signal sequence has lower harmonic distortion while retaining the key characteristics of the original signal. At the same time, perform wavelet decomposition on the temperature change characteristic data in the power amplifier operating state data to analyze the noise characteristics. Wavelet decomposition is a time-frequency analysis tool that decomposes the signal into sub-signals in different frequency ranges, extracts the noise characteristics, and generates a noise feature vector, which includes background noise data and transient interference data. The background noise data describes the persistent low-frequency noise characteristics in the signal, while the transient interference data records the short-term high-amplitude noise components in the signal. Input the noise feature vector into the Wiener filter, and perform adaptive filtering on the signal through the least mean square error criterion. The Wiener filter calculates the noise reduction coefficient sequence using the statistical characteristics of the signal and noise, and dynamically adjusts the filter parameters to optimize the noise suppression effect. Based on the noise reduction coefficient sequence, perform noise suppression processing on the compensated signal sequence to generate a purified audio matrix. Perform spectral analysis on the purified audio matrix, extract its frequency response characteristic curve, and smooth the curve to reduce spectral fluctuations to generate a frequency response characteristic curve. The flatness of the frequency response characteristic curve is an important indicator to measure the consistency of the audio system, which describes the response uniformity of the system to signals of different frequencies. At the same time, extract the signal-to-noise ratio parameter from the purified audio matrix, which quantifies the clarity of the signal and the noise suppression effect; extract the harmonic distortion parameter from the compensated signal sequence, which reflects the degree of correction of signal nonlinear distortion. Combine the signal-to-noise ratio parameter of the purified audio matrix, the harmonic distortion parameter of the compensated signal sequence, and the flatness parameter of the frequency response characteristic curve to generate audio output quality parameters.
[0039] Classify the fundamental component data and high-order harmonic component data in the spectrum analysis matrix. Through the harmonic order classification method, classify the frequency components of different harmonic orders into independent frequency groups to generate a classified harmonic matrix. The classified harmonic matrix describes each frequency point and its corresponding harmonic components. Perform feature mapping on the classified harmonic matrix to transform the physical attributes of each harmonic component into a feature space suitable for mathematical analysis, generating a harmonic feature mapping matrix. During the feature mapping process, accurately capture the main characteristics of nonlinear distortion by considering the correlation between the amplitude, phase, and frequency of the harmonic components. Based on this mapping matrix, establish a nonlinear compensation model. The nonlinear compensation model analyzes the amplitude and phase characteristics of harmonic distortion through the Taylor series expansion method to obtain the nonlinear feature vector. The Taylor series expansion decomposes complex nonlinear relationships into polynomial approximations, effectively capturing the small nonlinear changes in the signal while providing an accurate description of the high-order distortion components. Calculate the gradient value of the nonlinear feature vector through the backpropagation algorithm, and use the gradient information to optimize and update the network parameters of the model. To ensure the efficiency and stability of the optimization process, use the Adam optimizer to iteratively update the network parameters. The Adam optimizer adaptively adjusts the learning rate, quickly converges to the optimal solution, and generates a weight update matrix. The weight update matrix is an important tool reflecting the magnitude of parameter adjustment during network training, and its result directly determines the accuracy of the nonlinear compensation model. Based on the weight update matrix, construct a Volterra series model to model the dynamic characteristics of the nonlinear system. The Volterra series model is a classic method for nonlinear system modeling, and approximately describes the dynamic response of the signal by superimposing convolution kernel functions of different orders. Through this modeling process, generate a nonlinear dynamic response function. This function can accurately describe the dynamic characteristics of the signal in terms of frequency, amplitude, and phase. After obtaining the nonlinear dynamic response function, optimize it according to the least mean square error criterion and solve the optimal compensation parameters using the conjugate gradient method. The conjugate gradient method accelerates the optimization process using quadratic approximation and finally generates a compensation parameter vector containing amplitude and phase compensation data for each frequency point. To improve the practicality of the compensation parameters, perform parameter smoothing on the compensation parameter vector. The smoothing process reduces the high-frequency fluctuations of the parameters through a low-pass filter or interpolation algorithm, ensuring a smoother transition of the compensation coefficients between different frequency points, thereby enhancing the consistency of the compensation effect. The processed smoothed compensation sequence is used to reconstruct the distortion compensation coefficient matrix. The row vectors of the distortion compensation coefficient matrix represent different frequency points, while the column vectors correspond to the amplitude compensation coefficient and the phase compensation coefficient respectively. This matrix can intuitively describe the distortion characteristics and compensation strategies of the system at each frequency.
[0040] In the embodiments of the present invention, the MCU control unit realizes the collaborative detection of the input voltage and the audio signal, combines with a deep neural network for feature extraction and analysis, and significantly improves the accuracy and adaptability of the audio output control. The cross-attention model is used to calculate the dynamic weights of the audio parameters, realizing the intelligent switching of the audio curve, and effectively solving the problem of unstable audio output under different power supply conditions. Through the dynamic adjustment of the power amplifier voltage and the feedback control mechanism, combined with the non-linear distortion compensation and noise suppression processing, the signal-to-noise ratio and the frequency response flatness of the audio output are greatly improved. Innovatively applying the deep learning algorithm to audio feature extraction and parameter optimization realizes the intelligentization and precision of the audio output control, enabling the system to have stronger environmental adaptability. Adopting the modular design concept, the interfaces between the functional units are standardized, improving the scalability and maintainability of the system, and facilitating subsequent function upgrades and optimizations. Through multi-level signal processing and optimization strategies, the overall quality of the audio output is improved, meeting the high-quality audio output requirements in different application scenarios.
[0041] In a specific embodiment, the process of executing step 101 may specifically include the following steps:
[0042] The MCU control unit identifies the fast charging protocol for the input voltage signal to obtain a protocol type identifier, and the protocol type identifier includes USB PD3.0, PD2.0, and BC1.2 boost fast charging protocol identifiers;
[0043] According to the protocol type identifier, sample and detect the input voltage value to obtain a voltage sampling sequence, and extract voltage features from the voltage sampling sequence to obtain a voltage parameter matrix;
[0044] Perform multi-channel sampling on the audio input signal to obtain audio sampling data, and extract audio features from the audio sampling data to obtain an audio parameter matrix, where the audio parameter matrix includes frequency response data, phase characteristic data, and dynamic range data;
[0045] Perform normalization processing on the voltage parameter matrix to obtain a normalized voltage matrix, perform normalization processing on the audio parameter matrix to obtain a normalized audio matrix, and combine the normalized voltage matrix and the normalized audio matrix to obtain a mixed feature matrix.
[0046] Specifically, the MCU control unit realizes fast charging protocol identification. The fast charging protocol identification is based on the preliminary analysis of the voltage signal. By capturing the handshake signal and the mode flag, the protocol type to which the voltage signal belongs is judged. Assume that the input voltage signal is , and its characteristics include the voltage waveform characteristics under different fast charging protocols. By performing time-domain analysis on the signal, extract the protocol flag bit data , and judge that the protocol type is PD3.0, PD2.0, BC1.2 . According to the protocol type identifier, sample and detect the input voltage signal to generate a voltage sampling sequence , where represents the voltage value at the -th sampling point, is the number of sampling points. The time interval of voltage sampling is determined by the sampling frequency , and the specific sampling value is expressed as: ;
[0047] Through this sequence, extract the features of the voltage signal to generate a voltage parameter matrix , which contains the average value , variance , maximum value and frequency components . For example, the average value and variance are calculated by the following formulas: ;
[0048] The frequency components are calculated by fast Fourier transform: ;
[0049] At the same time, perform multi-channel sampling on the audio input signal . Assume that the audio signal has sampling channels, then the audio sampling data generated by multi-channel sampling is expressed as a matrix , where represents the signal value of the -th channel at the -th sampling point. For the sampling data of each channel, extract the features of frequency response, phase characteristics and dynamic range respectively. For example, the frequency response is calculated by the following formula: ;
[0050] where is the frequency component of the audio output, is the frequency component of the audio input. The phase characteristics represent the phase offset at different frequencies: ;
[0051] The dynamic range represents the ratio of the maximum amplitude to the minimum amplitude of the signal: ;
[0052] After feature extraction, the audio sampling data is integrated into an audio parameter matrix , which contains frequency response data, phase characteristic data and dynamic range data. For the voltage parameter matrix and the audio parameter matrix are normalized to eliminate the differences between different feature dimensions. The normalization method is used as follows:
[0053] ;
[0054] wherein, and are respectively the mean and standard deviation of each column in and are respectively the mean and standard deviation of each column in. The normalized voltage matrix and the normalized audio matrix are combined to generate a hybrid feature matrix , and its form is:
[0055] ;
[0056] The hybrid feature matrix fuses the features of the voltage and audio signals into a unified representation, providing data input for subsequent deep learning or optimization algorithms.
[0057] In a specific embodiment, the process of executing step 102 may specifically include the following steps:
[0058] Input the normalized voltage matrix in the hybrid feature matrix into the feature extraction network including three convolutional layers in the deep neural network for processing. Each convolutional layer uses the ReLU activation function and the batch normalization layer to obtain the voltage deep feature vector;
[0059] Input the normalized audio matrix into the five-layer convolutional neural network in the deep neural network for spectrum analysis. After each layer of the convolutional neural network, there is a BatchNormalization layer and a LeakyReLU activation function to obtain the audio deep feature vector;
[0060] Fuse the voltage deep feature vector and the audio deep feature vector to generate a fused feature vector, and input the fused feature vector into the four-layer fully connected neural network in the deep neural network for mapping transformation to obtain a high-dimensional feature vector;
[0061] Perform deep residual calculation on the high-dimensional feature vector to obtain a time-frequency feature matrix, and perform temporal feature extraction on the time-frequency feature matrix to obtain a dynamic feature vector;
[0062] The dynamic feature vector is input into a decoder network in the deep neural network for feature reconstruction. The decoder network includes four deconvolution layers and two fully connected layers to obtain an audio feature space, which contains frequency response feature data, phase feature data, and harmonic feature data.
[0063] Specifically, the normalized voltage matrix in the mixed feature matrix is input into the feature extraction network of the deep neural network. The feature extraction network consists of three convolutional layers, and each convolutional layer includes a convolution operation, a ReLU activation function, and a batch normalization layer. The convolution operation extracts local features through a two-dimensional convolution kernel The convolution output is expressed as: ;
[0064] where is the value of the output feature map, is the bias term, and are the height and width of the convolution kernel respectively. The ReLU activation function is used to introduce non-linearity, and the activated output is: ;
[0065] Batch Normalization stabilizes the training process by adjusting the mean and variance of the feature map. The normalization formula is: ;
[0066] where and are the mean and variance of the batch data, and are trainable parameters, is a small constant to prevent the denominator from being zero. After being processed by the three-layer convolutional network, a voltage depth feature vector is generated, where is the feature dimension. At the same time, the normalized audio matrix is input into a five-layer convolutional neural network, which is used for spectrum analysis. After each convolution, a Batch Normalization layer and a LeakyReLU activation function are connected. The activation formula of LeakyReLU is: ;
[0067] where is the negative slope parameter, set to 0.01. Through five-layer convolution operations, the spectrum features of the audio signal are extracted, and finally an audio depth feature vector is generated, where is the feature dimension. The voltage depth feature vector and the audio depth feature vector are subjected to feature fusion: ;
[0068] Or perform weighted fusion through an attention mechanism. The fused vector is input into a four-layer fully connected neural network for mapping transformation. Each fully connected operation is represented as: ;
[0069] Where is the output of the th layer, is the weight matrix, is the bias, is the activation function. Finally, a high-dimensional feature vector is generated. The high-dimensional feature vector generates a time-frequency feature matrix through deep residual calculation. Residual calculation retains low-level information by introducing skip connections:
[0070] ;
[0071] Where is a non-linear mapping function. The time-frequency feature matrix is input into a time series model (such as LSTM or GRU) for time series feature extraction to generate a dynamic feature vector . The dynamic feature vector is input into a decoder network for feature reconstruction. The decoder consists of four transposed convolutional layers and two fully connected layers. Transposed convolution is used for upsampling features, and the formula is: ;
[0072] Where is the decoded output, is the value of the input feature map. After processing by the transposed convolutional and fully connected layers, an audio feature space is finally generated, which contains frequency response feature data , phase feature data and harmonic feature data .
[0073] In a specific embodiment, the process of executing step 103 may specifically include the following steps:
[0074] Perform feature projection on the frequency response feature data in the audio feature space to obtain a frequency feature vector, and perform feature projection on the phase feature data and harmonic feature data in the audio feature space to obtain an audio representation vector;
[0075] Use the frequency feature vector as the query vector and the audio representation vector as the key-value vector to obtain a similarity matrix through dot product operation;
[0076] Normalize the similarity matrix and perform weighted summation with the audio representation vector to obtain an attention feature vector, which contains frequency response weights and phase harmonic weights;
[0077] Construct an audio feature correlation graph based on the attention feature vector, extract the correlation features between nodes through a graph convolutional network to obtain an audio parameter correlation matrix, and perform feature clustering on the audio parameter correlation matrix to obtain an audio curve category matrix;
[0078] Match the curve templates in the audio curve library based on the audio curve category matrix to obtain a candidate curve sequence, which includes a gain curve, a frequency curve, and a phase curve;
[0079] Sort the candidate curve sequence according to similarity and perform sequence optimization to generate an audio curve switching control sequence, which includes gain control parameters, frequency adjustment parameters, and phase compensation parameters.
[0080] Specifically, extract frequency response feature data from the audio feature space , which represents the response intensity of the system at different frequencies. Through linear feature projection, map the high-dimensional frequency response features to low-dimensional frequency feature vectors , and the projection formula is: ;
[0081] where is the frequency response feature projection matrix, is the dimension of the original feature, is the dimension after dimensionality reduction. Similarly, perform joint feature projection on the phase feature data and harmonic feature data in the audio feature space to obtain an audio representation vector :
[0082] ;
[0083] where and are the projection matrices of the phase feature and the harmonic feature respectively. Take the frequency feature vector as the query vector and the audio representation vector as the key-value vector, calculate the dot product similarity between the two to obtain a similarity matrix :
[0084] ;
[0085] The dot product similarity measures the correlation between the frequency feature and the audio representation. To ensure the numerical stability and interpretability, normalize the similarity matrix Normalization is performed using the softmax function: ;
[0086] The normalized similarity matrix is used to perform weighted summation on the audio feature vectors to generate attention feature vectors :
[0087] ;
[0088] The attention feature vectors contain frequency response weights and phase harmonic weights and are the core representations describing the relationships of audio features. Based on the attention feature vectors , an audio feature correlation graph is constructed, where the nodes of the graph represent different audio features and the edge weights are determined by the weight values in the attention feature vectors. Feature extraction is performed on the correlation graph through a graph convolutional network, and the formula is: ;
[0089] where is the node feature matrix of the th layer, is the normalized adjacency matrix, is the trainable weight matrix, is the non-linear activation function. After being processed by a multi-layer graph convolutional network, an audio parameter correlation matrix is generated, which describes the complex correlations between various audio parameters. Feature clustering is performed on the audio parameter correlation matrix , and it is divided into categories using the k-means algorithm to generate an audio curve category matrix , where each row represents the feature center of a category. Based on the audio curve category matrix, the closest curve template is matched from the audio curve library, and the matching process is completed by calculating the similarity between the category matrix and the curve library template : ;
[0090] After matching, a candidate curve sequence is obtained, where each curve includes a gain curve, a frequency curve, and a phase curve. The candidate curve sequence is sorted according to the similarity, and sequence optimization is performed through dynamic programming to avoid large jumps during the switching process. Finally, an audio curve switching control sequence is generated, where is the gain control parameter, is the frequency adjustment parameter, is the phase compensation parameter.
[0091] In a specific embodiment, the process of executing step 104 may specifically include the following steps:
[0092] Preprocess the power amplifier input signal according to the gain control parameter, frequency adjustment parameter, and phase compensation parameter in the audio curve switching control sequence to obtain a preprocessed signal matrix, which includes gain compensation data, frequency adjustment data, and phase correction data;
[0093] Input the preprocessed signal matrix into the dynamic gain controller, calculate the gain adjustment coefficient through the feedback loop, and obtain the power amplifier drive control signal;
[0094] Perform power compensation operation on the power amplifier drive control signal, dynamically adjust the output power based on the voltage feedback value, and obtain a power compensation matrix, which includes voltage compensation value and current compensation value;
[0095] Generate a power amplifier drive voltage signal according to the power compensation matrix, and collect output voltage data through the voltage detection unit to obtain a voltage feedback matrix, which includes instantaneous voltage value and voltage change rate;
[0096] Filter the voltage feedback matrix through a loop filter to obtain a filtered voltage characteristic sequence, which includes voltage stability data and voltage ripple data;
[0097] Establish a power amplifier state monitoring model according to the voltage characteristic sequence, obtain the power amplifier working parameters through multi-point sampling, and obtain a power amplifier state parameter matrix, which includes temperature data, current data, and efficiency data;
[0098] Perform multi-dimensional analysis on the power amplifier state parameter matrix, conduct state evaluation based on the threshold judgment circuit, and generate power amplifier working state data, which includes voltage response characteristic data, temperature change characteristic data, and efficiency characteristic data.
[0099] Specifically, based on the gain control parameter 、frequency adjustment parameter and phase compensation parameter in the audio curve switching control sequence, preprocess the power amplifier input signal . Assume the input signal is , where is the signal amplitude, is the frequency, is the initial phase. Gain compensation is achieved by adjusting , and the signal after gain compensation is: ;
[0100] Frequency adjustment is achieved by adding an adjustment amount to , and the signal after frequency adjustment is: ;
[0101] Phase compensation is achieved by adding a compensation amount to . The signal after phase correction is: ; Adding a compensation amount is achieved, and the signal after phase correction is: ;
[0102] After gain, frequency, and phase compensation, a preprocessed signal matrix is finally formed, where the rows represent time sampling points, and the columns represent gain compensation data, frequency adjustment data, and phase correction data respectively. Input into the dynamic gain controller, and the dynamic gain controller adjusts the gain adjustment coefficient through a feedback loop. The feedback loop is based on the error between the target output signal and the actual output signal for gain adjustment: ; its rows represent time sampling points, and the columns represent gain compensation data, frequency adjustment data, and phase correction data respectively. Input into the dynamic gain controller, and the dynamic gain controller adjusts the gain adjustment coefficient through a feedback loop. The feedback loop is based on the error between the target output signal and the actual output signal for gain adjustment: ; into the dynamic gain controller, and the dynamic gain controller adjusts the gain adjustment coefficient through a feedback loop. The feedback loop is based on the error between the target output signal and the actual output signal for gain adjustment: . The feedback loop is based on the error between the target output signal and the actual output signal for gain adjustment: ; ;
[0103] The gain adjustment coefficient is calculated by a proportional-integral-derivative (PID) controller: ; is calculated by a proportional - integral - derivative (PID) controller: ;
[0104] where , , are the proportional, integral, and derivative gains respectively. According to , a power amplifier drive control signal is generated. Input into the power compensation module, and combine it with the voltage feedback value to dynamically adjust the output power, generating a power compensation matrix , which contains a voltage compensation value and a current compensation value : ; a power amplifier drive control signal is generated. Input into the power compensation module, and combine it with the voltage feedback value to dynamically adjust the output power, generating a power compensation matrix , which contains a voltage compensation value and a current compensation value : . Input into the power compensation module, and combine it with the voltage feedback value to dynamically adjust the output power, generating a power compensation matrix , which contains a voltage compensation value and a current compensation value : dynamically adjusts the output power, generating a power compensation matrix , which contains a voltage compensation value and a current compensation value : wherein it contains a voltage compensation value and a current compensation value : ;
[0105] and are the target voltage and current. According to , a power amplifier drive voltage signal is generated, and the output voltage data is collected through a voltage detection unit. By calculating the instantaneous voltage value and the voltage change rate , a voltage feedback matrix is generated. Input into a loop filter for filtering, and the filtered voltage feature sequence contains voltage stability data and voltage ripple data : ; a power amplifier drive voltage signal is generated, and the output voltage data is collected through a voltage detection unit. By calculating the instantaneous voltage value and the voltage change rate , a voltage feedback matrix is generated. Input into a loop filter for filtering, and the filtered voltage feature sequence contains voltage stability data and voltage ripple data : , and the output voltage data is collected through a voltage detection unit. By calculating the instantaneous voltage value and the voltage change rate , a voltage feedback matrix is generated. Input into a loop filter for filtering, and the filtered voltage feature sequence contains voltage stability data and voltage ripple data : . By calculating the instantaneous voltage value and the voltage change rate , a voltage feedback matrix is generated. Input into a loop filter for filtering, and the filtered voltage feature sequence contains voltage stability data and voltage ripple data : into the loop filter for filtering, and the filtered voltage feature sequence contains voltage stability data and voltage ripple data : contains voltage stability data and voltage ripple data : ;
[0106] According to Establish a power amplifier status monitoring model, and record temperature data through multi-point sampling , current data and efficiency data : ;
[0107] Integrate these parameters into a power amplifier status parameter matrix . For , conduct multi-dimensional analysis, and evaluate key parameters through a threshold judgment circuit. If , it is determined that the temperature is too high; if , it is determined that the efficiency is low. After comprehensive evaluation, generate power amplifier working status data, including voltage response characteristics , temperature change characteristics and efficiency characteristics .
[0108] In a specific embodiment, the process of executing step 105 may specifically include the following steps:
[0109] Perform a fast Fourier transform on the voltage response characteristic data in the power amplifier working status data to obtain a spectrum analysis matrix, and the spectrum analysis matrix includes fundamental wave component data and high-order harmonic component data;
[0110] Input the spectrum analysis matrix into a non-linear compensation model, calculate the compensation weight through the backpropagation algorithm, and obtain a distortion compensation coefficient matrix, and the distortion compensation coefficient matrix includes an amplitude compensation coefficient and a phase compensation coefficient;
[0111] Perform non-linear correction on the audio signal according to the distortion compensation coefficient matrix to obtain a compensated signal sequence, and perform noise characteristic analysis on the temperature change characteristic data through wavelet decomposition to obtain a noise characteristic vector, and the noise characteristic vector includes background noise data and transient interference data;
[0112] Input the noise characteristic vector into a Wiener filter, perform adaptive filtering based on the minimum mean square error criterion to obtain a noise reduction coefficient sequence, and perform noise suppression processing on the compensated signal sequence according to the noise reduction coefficient sequence to obtain a purified audio matrix;
[0113] Perform spectrum analysis on the purified audio matrix, obtain a frequency response characteristic curve through smoothing processing, and combine the signal-to-noise ratio parameter of the purified audio matrix, the harmonic distortion parameter of the compensated signal sequence, and the flatness parameter of the frequency response characteristic curve to generate an audio output quality parameter.
[0114] Specifically, extract the voltage response characteristic data from the power amplifier working status data , where is time, represents the change of the power amplifier output voltage over time. In order to analyze the characteristics of the signal in the frequency domain, for Perform a fast Fourier transform to obtain the frequency spectrum , and its calculation formula is:
[0115] ;
[0116] where is the frequency, is the signal length, is the imaginary unit. is a complex number, including the amplitude and the phase . The frequency spectrum analysis matrix consists of the fundamental wave component and the higher harmonic components , where is the fundamental wave frequency, is the harmonic order. Input the frequency spectrum analysis matrix into the non-linear compensation model. The compensation model adopts a deep learning framework and calculates the compensation weights through the backpropagation algorithm. The goal is to minimize the fundamental wave amplitude distortion and the phase distortion , and define the loss function as:
[0117]
[0118] where and are the ideal amplitude and phase, and are the weight coefficients. Obtain the distortion compensation coefficient matrix through gradient descent optimization, which includes the amplitude compensation coefficient and the phase compensation coefficient : ;
[0119] According to perform non-linear correction on the audio signal to obtain the compensated signal sequence : ;
[0120] At the same time, perform wavelet decomposition on the temperature change characteristic data in the power amplifier operating state data. The decomposition formula is: ;
[0121] where is the wavelet basis function, is the wavelet coefficient. Extract the background noise data and the transient interference data , and integrate them into the noise feature vector Input the noise feature vector into the Wiener filter and calculate the noise reduction coefficient sequence based on the minimum mean square error (MMSE) criterion : ;
[0122] where and are the signal power spectrum and the noise power spectrum respectively. According to perform noise suppression on the compensated signal sequence to obtain the purified audio matrix . Perform spectral analysis on the purified audio matrix to extract the frequency response characteristic curve , and perform smoothing processing by the moving average method. The formula is: ;
[0123] where is the smoothing window size. Calculate the signal-to-noise ratio (SNR) of the purified audio matrix. The formula is: ;
[0124] The calculation formula for the total harmonic distortion parameter (THD) of the compensated signal sequence is: ;
[0125] where is the total harmonic order. The flatness parameter of the frequency response characteristic curve is used to measure the uniformity of the response at different frequencies. Its definition is: ;
[0126] Combine the signal-to-noise ratio parameter of the purified audio matrix, the total harmonic distortion parameter of the compensated signal sequence, and the flatness parameter of the frequency response characteristic curve to generate the audio output quality parameter: .
[0127] Among them, before performing noise suppression processing on the compensated signal sequence, it further includes: decomposing the compensated signal sequence into multiple frequency band signals to obtain a multi-scale signal matrix, where the multi-scale signal matrix includes low-frequency signal data, intermediate-frequency signal data, and high-frequency signal data; analyzing the oscillation characteristics of each frequency band signal in the multi-scale signal matrix, extracting the envelope signal through Hilbert transform to obtain an oscillation feature vector, where the oscillation feature vector includes amplitude data, frequency data, and phase data; inputting the oscillation feature vector into a hierarchical prediction controller, establishing a prediction model using the recursive least squares method to obtain a state prediction matrix, where the state prediction matrix includes the predicted states of the next N sampling points; constructing a distributed control network based on the state prediction matrix, allocating control tasks to multiple sub-controllers, where each sub-controller is responsible for the stability control of a specific frequency band to obtain a distributed control instruction sequence; performing collaborative optimization on the distributed control instruction sequence, realizing state synchronization between sub-controllers through a consensus algorithm to obtain a collaborative control matrix, where the collaborative control matrix includes control parameters for each frequency band; inputting the collaborative control matrix into a Lyapunov stability analyzer, calculating the system stability index to obtain a stability evaluation vector, where the stability evaluation vector includes an energy function value and a convergence rate value; dynamically adjusting the control parameters according to the stability evaluation vector, using a sliding mode control method to suppress system oscillation to obtain an oscillation suppression sequence, where the oscillation suppression sequence includes compensation coefficients and stabilization parameters; applying the oscillation suppression sequence to each frequency band signal, and performing signal reconstruction through an adaptive filter to obtain a stabilization signal matrix.
[0128] In a specific embodiment, the process of performing the step of inputting the spectrum analysis matrix into the nonlinear compensation model and calculating the compensation weights through the backpropagation algorithm to obtain the distortion compensation coefficient matrix, where the distortion compensation coefficient matrix includes amplitude compensation coefficients and phase compensation coefficients, may specifically include the following steps:
[0129] Classifying the fundamental wave component data and high-order harmonic component data in the spectrum analysis matrix by harmonic order to obtain a classified harmonic matrix, and performing feature mapping on the classified harmonic matrix to obtain a harmonic feature mapping matrix;
[0130] Establishing a nonlinear compensation model based on the harmonic feature mapping matrix, calculating the amplitude and phase characteristics of harmonic distortion through Taylor series expansion to obtain a nonlinear feature vector;
[0131] Calculating the gradient value of the nonlinear feature vector through the backpropagation algorithm, and iteratively updating the network parameters using an Adam optimizer to obtain a weight update matrix;
[0132] Construct a Volterra series model based on the weight update matrix to model the dynamic characteristics of a nonlinear system, obtain the nonlinear dynamic response function, and optimize the nonlinear dynamic response function with the least mean square error. Solve the optimal compensation parameters through the conjugate gradient method to obtain the compensation parameter vector;
[0133] Perform parameter smoothing on the compensation parameter vector to obtain a smoothed compensation sequence, and reconstruct the smoothed compensation sequence into a distortion compensation coefficient matrix. The row vectors of the distortion compensation coefficient matrix represent different frequency points, and the column vectors represent the corresponding amplitude compensation coefficients and phase compensation coefficients.
[0134] Specifically, extract the fundamental wave component and the high-order harmonic components from the spectral analysis matrix , where is the fundamental wave frequency, and is the harmonic order. The fundamental wave component and the high-order harmonic components respectively describe the main frequency components of the signal and the frequency components introduced by nonlinear distortion. The harmonic order classification is completed by aggregating components with different into the corresponding classification matrix . Assuming the frequency range is , then the th row of the classification matrix represents the amplitude and phase of the th harmonic:
[0135] ;
[0136] Perform feature mapping on the classified harmonic matrix and convert it into a harmonic feature mapping matrix through the mapping matrix : ;
[0137] where is a trainable weight matrix used to capture the complex relationships between harmonic features. The mapped matrix contains the high-level representation of harmonic characteristics and provides the input for the nonlinear compensation model. Based on , establish a nonlinear compensation model. Calculate the amplitude and phase of the harmonic distortion through Taylor series expansion. The Taylor expansion formula is: ;
[0138] where is the ideal signal, is the frequency offset, is the expansion order. By analyzing the nonlinear characteristics of amplitude and phase, a nonlinear feature vector is obtained , which contains the characteristics of high-order harmonics. The nonlinear feature vector is input into the deep neural network, and the gradient value is calculated through the backpropagation algorithm to optimize the network parameters . The loss function is defined as the error between the nonlinear distortion and the ideal characteristic: ;
[0139] The network parameters are iteratively updated through the Adam optimizer, and the update formula is: ;
[0140] where is the learning rate, and are the first-order and second-order momentum estimates of the gradient deviation correction respectively. The optimized weight update matrix is used to construct the Volterra series model. The Volterra model describes the dynamic characteristics through nonlinear convolution, and its expression is:
[0141] ;
[0142] where is order Volterra kernel function, which describes the dynamic characteristics of the system. Through the optimization of the minimum mean square error (MSE), the conjugate gradient method is used to solve the optimal compensation parameter vector : ;
[0143] For parameter smoothing is performed, and a low-pass filter is used to reduce high-frequency fluctuations. The formula is: ;
[0144] The smoothed compensation sequence is reconstructed into a distortion compensation coefficient matrix , whose row vectors represent different frequency points , and the column vectors contain the amplitude compensation coefficient and the phase compensation coefficient : .
[0145] The audio output control method in the embodiments of the present invention has been described above. Next, the audio output control device in the embodiments of the present invention will be described. Please refer to Figure 2 , an embodiment of the audio output control device in the embodiments of the present invention includes:
[0146] The acquisition module 201 is used to detect the input voltage signal and collect the audio input signal through the MCU control unit, obtain the voltage parameter matrix and the audio parameter matrix, and combine the voltage parameter matrix and the audio parameter matrix to form a mixed feature matrix;
[0147] The feature extraction module 202 is used to input the mixed feature matrix into a deep neural network for feature extraction, perform time-frequency domain analysis on the audio signal, and obtain the audio feature space;
[0148] The calculation module 203 is used to construct a cross-attention model based on the audio feature space, calculate the correlation weight values between audio parameters, and generate an audio curve switching control sequence;
[0149] The dynamic adjustment module 204 is used to dynamically adjust the power amplifier voltage based on the audio curve switching control sequence through a feedback controller, and collect the power amplifier working state data;
[0150] The generation module 205 is used to perform nonlinear distortion compensation and noise suppression processing on the power amplifier working state data, generate audio output quality parameters, and the audio output quality parameters include signal-to-noise ratio parameters, harmonic distortion parameters, and frequency response flatness parameters.
[0151] Through the collaborative cooperation of the above-mentioned various components, the MCU control unit realizes the collaborative detection of the input voltage and audio signal, combines the deep neural network for feature extraction and analysis, and significantly improves the accuracy and adaptive ability of the audio output control. The cross-attention model is used to calculate the dynamic weights of audio parameters, realizing the intelligent switching of audio curves, and effectively solving the problem of unstable audio output under different power supply conditions. Through the dynamic adjustment of the power amplifier voltage and the feedback control mechanism, combined with nonlinear distortion compensation and noise suppression processing, the signal-to-noise ratio and frequency response flatness of the audio output are greatly improved. Innovatively applying the deep learning algorithm to audio feature extraction and parameter optimization realizes the intelligentization and precision of audio output control, enabling the system to have stronger environmental adaptability. Adopting the modular design idea, the interfaces between functional units are standardized, improving the scalability and maintainability of the system, and facilitating subsequent function upgrade and optimization. Through multi-level signal processing and optimization strategies, the all-round improvement of audio output quality is realized, meeting the high-quality audio output requirements under different application scenarios.
[0152] Above Figure 2 The audio output control device in the embodiment of the present invention is described in detail from the perspective of modular functional entities. Next, the audio output control device in the embodiment of the present invention is described in detail from the perspective of hardware processing.
[0153] Figure 3FIG. 0 is a schematic structural diagram of an audio output control device provided by an embodiment of the present invention. The audio output control device 300 may vary greatly due to different configurations or performances, and may include one or more processors (central processing units, CPUs) 310 (for example, one or more processors) and a memory 320, and one or more storage media 330 (for example, one or more mass storage device terminals) storing application programs 333 or data 332. Among them, the memory 320 and the storage media 330 may be transient storage or persistent storage. The program stored in the storage media 330 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the audio output control device 300. Further, the processor 310 may be configured to communicate with the storage media 330 and execute a series of instruction operations in the storage media 330 on the audio output control device 300 to implement the steps of the above audio output control method.
[0154] The audio output control device 300 may further include one or more power supplies 340, one or more wired or wireless network interfaces 350, one or more input / output interfaces 360, and / or one or more operating systems 331, such as Windows Serve, Mac OS X, Unix, Linux, FreeBSD, and so on. Those skilled in the art can understand that Figure 3 the shown structural diagram of the audio output control device does not constitute a limitation on the audio output control device provided by the present invention, and may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0155] The present invention also provides a computer-readable storage medium, which may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. Instructions are stored in the computer-readable storage medium. When the instructions are run on a computer, the computer is caused to execute the steps of the audio output control method.
[0156] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described systems, systems, and units can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0157] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0158] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. An audio output control method, characterized in that, The method includes: Detecting an input voltage signal and collecting an audio input signal through an MCU control unit to obtain a voltage parameter matrix and an audio parameter matrix, performing normalization processing on the voltage parameter matrix to obtain a normalized voltage matrix, and performing normalization processing on the audio parameter matrix to obtain a normalized audio matrix, and combining the normalized voltage matrix and the normalized audio matrix in a splicing manner to obtain a mixed feature matrix; Inputting the mixed feature matrix into a deep neural network, respectively extracting a voltage deep feature vector and an audio deep feature vector through the deep neural network, performing feature fusion on the voltage deep feature vector and the audio deep feature vector, and then performing time-frequency domain analysis through the deep neural network to obtain an audio feature space, where the audio feature space includes frequency response feature data, phase feature data, and harmonic feature data; Constructing a cross-attention model according to the audio feature space, calculating the correlation weight values between audio parameters in the audio feature space, and generating an audio curve switching control sequence, where the audio curve switching control sequence includes gain control parameters, frequency adjustment parameters, and phase compensation parameters; Based on the audio curve switching control sequence, dynamically adjusting the power amplifier voltage through a feedback controller, where the feedback controller calculates a gain adjustment coefficient based on the error between the target output signal and the actual output signal through a proportional-integral-derivative controller, and collects power amplifier working state data; Performing nonlinear distortion compensation and noise suppression processing on the power amplifier working state data to generate audio output quality parameters, where the audio output quality parameters include signal-to-noise ratio parameters, harmonic distortion parameters, and frequency response flatness parameters.
2. The audio output control method according to claim 1, wherein The process of detecting an input voltage signal and collecting an audio input signal through an MCU control unit to obtain a voltage parameter matrix and an audio parameter matrix, performing normalization processing on the voltage parameter matrix to obtain a normalized voltage matrix, and performing normalization processing on the audio parameter matrix to obtain a normalized audio matrix, and combining the normalized voltage matrix and the normalized audio matrix in a splicing manner to obtain a mixed feature matrix includes: Identifying a fast charging protocol for the input voltage signal through an MCU control unit to obtain a protocol type identifier, where the protocol type identifier includes USB PD3.0, PD2.0, and BC1.2 boost fast charging protocol identifiers; Sampling and detecting the input voltage value according to the protocol type identifier to obtain a voltage sampling sequence, and extracting voltage features from the voltage sampling sequence to obtain a voltage parameter matrix; Performing multi-channel sampling on the audio input signal to obtain audio sampling data, and extracting audio features from the audio sampling data to obtain an audio parameter matrix, where the audio parameter matrix includes frequency response data, phase characteristic data, and dynamic range data; Normalize the voltage parameter matrix to obtain a normalized voltage matrix, and normalize the audio parameter matrix to obtain a normalized audio matrix. Combine the normalized voltage matrix and the normalized audio matrix to obtain a mixed feature matrix, where the form of the mixed feature matrix is the concatenation of the normalized voltage matrix and the normalized audio matrix.
3. The audio output control method according to claim 2, wherein Input the mixed feature matrix into a deep neural network. Through the deep neural network, extract a voltage deep feature vector and an audio deep feature vector respectively. After fusing the voltage deep feature vector and the audio deep feature vector, perform time-frequency domain analysis through the deep neural network to obtain an audio feature space, where the audio feature space includes frequency response feature data, phase feature data, and harmonic feature data, including: Input the normalized voltage matrix in the mixed feature matrix into a feature extraction network in the deep neural network that contains three convolutional layers. Each convolutional layer uses a ReLU activation function and a batch normalization layer to obtain a voltage deep feature vector; Input the normalized audio matrix into a five-layer convolutional neural network in the deep neural network for spectrum analysis. After each layer of the convolutional neural network, there is a BatchNormalization layer and a LeakyReLU activation function to obtain an audio deep feature vector; Fuse the voltage deep feature vector and the audio deep feature vector. The feature fusion is achieved by concatenating the voltage deep feature vector and the audio deep feature vector or by weighted fusion through an attention mechanism to generate a fused feature vector. Then input the fused feature vector into a four-layer fully connected neural network in the deep neural network for mapping transformation to obtain a high-dimensional feature vector; Perform a deep residual calculation on the high-dimensional feature vector to obtain a time-frequency feature matrix, and extract temporal features from the time-frequency feature matrix to obtain a dynamic feature vector; Input the dynamic feature vector into a decoder network in the deep neural network. The decoder network contains four transposed convolutional layers and two fully connected layers to obtain an audio feature space, where the audio feature space includes frequency response feature data, phase feature data, and harmonic feature data.
4. The audio output control method according to claim 3, wherein Construct a cross-attention model based on the audio feature space, calculate the correlation weight values between audio parameters in the audio feature space, and generate an audio curve switching control sequence. The audio curve switching control sequence includes gain control parameters, frequency adjustment parameters, and phase compensation parameters, including: Perform feature projection on the frequency response feature data in the audio feature space to obtain a frequency feature vector through linear transformation, and perform feature projection on the phase feature data and harmonic feature data in the audio feature space to obtain an audio representation vector; Use the frequency feature vector as the query vector and the audio representation vector as the key-value vector to obtain a similarity matrix through dot product operation; Normalize the similarity matrix and perform weighted summation with the audio feature vector to obtain an attention feature vector, which includes a frequency response weight and a phase harmonic weight; Construct an audio feature correlation graph based on the attention feature vector, extract the correlation features between nodes through a graph convolutional network to obtain an audio parameter correlation matrix, and perform feature clustering on the audio parameter correlation matrix to obtain an audio curve category matrix; Match the curve templates in the audio curve library based on the audio curve category matrix. The matching process is completed by calculating the similarity between the audio curve category matrix and the templates in the curve library to obtain a candidate curve sequence, which includes a gain curve, a frequency curve, and a phase curve; Sort the candidate curve sequence according to the similarity and perform sequence optimization to generate an audio curve switching control sequence, which includes a gain control parameter, a frequency adjustment parameter, and a phase compensation parameter.
5. The audio output control method according to claim 4, wherein Based on the audio curve switching control sequence, dynamically adjust the power amplifier voltage through a feedback controller. The feedback controller calculates the gain adjustment coefficient based on the error between the target output signal and the actual output signal through a proportional-integral-derivative controller, and collects the power amplifier operating state data, including: According to the gain control parameter, the frequency adjustment parameter, and the phase compensation parameter in the audio curve switching control sequence, preprocess the power amplifier input signal to obtain a preprocessing signal matrix, which includes gain compensation data, frequency adjustment data, and phase correction data; Input the preprocessing signal matrix into a dynamic gain controller, calculate the gain adjustment coefficient based on the error between the target output signal and the actual output signal through a feedback loop. The gain adjustment coefficient is calculated through a proportional-integral-derivative controller to obtain a power amplifier drive control signal; Perform power compensation operation on the power amplifier drive control signal, calculate the voltage compensation value and the current compensation value by comparing the difference between the target voltage / current and the actual voltage / current, and dynamically adjust the output power based on the voltage feedback value to obtain a power compensation matrix, which includes the voltage compensation value and the current compensation value; Generate a power amplifier drive voltage signal according to the power compensation matrix, and collect the output voltage data through a voltage detection unit to obtain a voltage feedback matrix, which includes the instantaneous voltage value and the voltage change rate; Filter the voltage feedback matrix through a loop filter to obtain a filtered voltage feature sequence, which includes voltage stability data and voltage ripple data; Establish a power amplifier state monitoring model based on the voltage feature sequence, obtain the power amplifier operating parameters through multi-point sampling to obtain a power amplifier state parameter matrix, which includes temperature data, current data, and efficiency data; Perform multi-dimensional analysis on the power amplifier state parameter matrix, conduct state assessment based on the threshold judgment circuit, trigger corresponding state judgments when the temperature exceeds the preset maximum value or the efficiency is lower than the preset minimum value, and generate power amplifier working state data, where the power amplifier working state data includes voltage response characteristic data, temperature change characteristic data, and efficiency characteristic data.
6. The audio output control method according to claim 5, wherein Perform non-linear distortion compensation and noise suppression processing on the power amplifier working state data to generate audio output quality parameters, where the audio output quality parameters include signal-to-noise ratio parameters, harmonic distortion parameters, and frequency response flatness parameters, including: Perform fast Fourier transform on the voltage response characteristic data in the power amplifier working state data to obtain a spectrum analysis matrix, where the spectrum analysis matrix includes fundamental wave component data and high-order harmonic component data; Input the spectrum analysis matrix into the non-linear compensation model, calculate the compensation weights through the backpropagation algorithm to obtain a distortion compensation coefficient matrix, where the distortion compensation coefficient matrix includes amplitude compensation coefficients and phase compensation coefficients; Perform non-linear correction on the audio signal according to the distortion compensation coefficient matrix to obtain a compensated signal sequence, and conduct noise feature analysis on the temperature change characteristic data through wavelet decomposition to obtain a noise feature vector, where the noise feature vector includes background noise data and transient interference data; Input the noise feature vector into a Wiener filter, perform adaptive filtering based on the minimum mean square error criterion to obtain a noise reduction coefficient sequence, and perform noise suppression processing on the compensated signal sequence according to the noise reduction coefficient sequence to obtain a purified audio matrix; Conduct spectrum analysis on the purified audio matrix, obtain a frequency response characteristic curve through smoothing processing, and combine the signal-to-noise ratio parameter of the purified audio matrix, the harmonic distortion parameter of the compensated signal sequence, and the flatness parameter of the frequency response characteristic curve to generate audio output quality parameters.
7. The audio output control method according to claim 6, characterized in that Input the spectrum analysis matrix into the non-linear compensation model, calculate the compensation weights through the backpropagation algorithm to obtain a distortion compensation coefficient matrix, where the distortion compensation coefficient matrix includes amplitude compensation coefficients and phase compensation coefficients, including: Classify the fundamental wave component data and high-order harmonic component data in the spectrum analysis matrix by harmonic order to obtain a classified harmonic matrix, and perform feature mapping on the classified harmonic matrix to obtain a harmonic feature mapping matrix; Establish a non-linear compensation model according to the harmonic feature mapping matrix, calculate the amplitude and phase characteristics of harmonic distortion through Taylor series expansion to obtain a non-linear feature vector; Calculate the gradient value of the non-linear feature vector through the backpropagation algorithm, and use the Adam optimizer to iteratively update the network parameters to obtain a weight update matrix; Construct a Volterra series model based on the weight update matrix, model the dynamic characteristics of the non-linear system to obtain a non-linear dynamic response function, and perform minimum mean square error optimization on the non-linear dynamic response function, and solve the optimal compensation parameters through the conjugate gradient method to obtain a compensation parameter vector; Perform parameter smoothing on the compensation parameter vector to obtain a smoothed compensation sequence, and reconstruct the smoothed compensation sequence into a distortion compensation coefficient matrix. The row vectors of the distortion compensation coefficient matrix represent different frequency points, and the column vectors represent the corresponding amplitude compensation coefficients and phase compensation coefficients.
8. An audio output control device, characterized in that, For implementing the audio output control method according to any one of claims 1-7, the apparatus includes: An acquisition module, configured to detect an input voltage signal and acquire an audio input signal through an MCU control unit, obtain a voltage parameter matrix and an audio parameter matrix, perform normalization processing on the voltage parameter matrix to obtain a normalized voltage matrix, perform normalization processing on the audio parameter matrix to obtain a normalized audio matrix, and combine the normalized voltage matrix and the normalized audio matrix in a splicing manner to obtain a mixed feature matrix; A feature extraction module, configured to input the mixed feature matrix into a deep neural network, respectively extract a voltage deep feature vector and an audio deep feature vector through the deep neural network, perform feature fusion on the voltage deep feature vector and the audio deep feature vector, and perform time-frequency domain analysis through the deep neural network to obtain an audio feature space, where the audio feature space includes frequency response feature data, phase feature data, and harmonic feature data; A calculation module, configured to construct a cross-attention model according to the audio feature space, calculate the correlation weight values between audio parameters in the audio feature space, and generate an audio curve switching control sequence, where the audio curve switching control sequence includes a gain control parameter, a frequency adjustment parameter, and a phase compensation parameter; A dynamic adjustment module, configured to dynamically adjust the power amplifier voltage based on the audio curve switching control sequence through a feedback controller. The feedback controller calculates a gain adjustment coefficient based on the error between the target output signal and the actual output signal through a proportional-integral-derivative controller, and acquires power amplifier working state data; A generation module, configured to perform nonlinear distortion compensation and noise suppression processing on the power amplifier working state data to generate an audio output quality parameter, where the audio output quality parameter includes a signal-to-noise ratio parameter, a harmonic distortion parameter, and a frequency response flatness parameter.
9. An audio output control device, characterized in that, The audio output control device includes: a memory and at least one processor, and instructions are stored in the memory; The at least one processor calls the instructions in the memory to cause the audio output control device to execute the audio output control method according to any one of claims 1-7.
10. A computer-readable storage medium, on which instructions are stored, characterized in that, When the instructions are executed by the processor, the audio output control method according to any one of claims 1-7 is implemented.
Citation Information
Patent Citations
Non-linear control of loudspeakers
CN106664481A
Integrated MEMS loudspeaker and microphone mixed signal processing method and system
CN118660262A