Frequency Response Analysis Method, Device and Equipment of Intelligent Audio Testing System

Through the frequency response analysis method of the intelligent audio test system, using technical means such as multi-level feature extraction, deep neural network and environmental compensation, the shortcomings of traditional audio test systems in terms of test accuracy, noise resistance and automation degree are solved, and high-precision and high-automation frequency response analysis are achieved.

CN119446188BActive Publication Date: 2025-05-27SHENZHEN DEWEI AUDIO CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411605171.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-12
Publication Date
2025-05-27
Estimated Expiration
2044-11-12

AI Technical Summary

Technical Problem

Traditional audio testing systems have shortcomings in testing accuracy, noise immunity, automation degree and adaptability, especially in complex environments, it is difficult to effectively deal with noise interference and environmental changes.

Method used

The frequency response analysis method of the intelligent audio testing system is adopted to realize high-precision frequency response analysis of audio equipment through audio parameter initialization, dual-channel acquisition and time-frequency feature extraction, dual-branch data processing, segmented recursive calculation and feature aggregation, environmental compensation and nonlinear calibration, as well as feature recognition and pattern classification of deep neural networks.

Benefits of technology

It significantly improves the test accuracy, anti-interference ability, automation degree and adaptability, ensuring the accuracy of full-band tests and the reliability of test results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119446188B_ABST
    Figure CN119446188B_ABST
Patent Text Reader

Abstract

The present application relates to the field of audio testing technologies, and discloses a frequency response analysis method, device and equipment for an intelligent audio testing system. The method includes: initializing audio parameters of the intelligent audio testing system to obtain a test parameter matrix and an excitation signal; performing dual-channel acquisition and time-frequency feature extraction to obtain a time-frequency feature data set; performing dual-branch processing of time-domain noise reduction and frequency-domain enhancement to obtain a purified feature data set and a feature weighting matrix; performing segmented recursive calculation and feature aggregation to obtain a frequency response feature vector; decomposing the frequency response feature vector according to test frequency bands, and performing environmental compensation and non-linear calibration to obtain a standard frequency response data set; inputting the standard frequency response data set into a deep neural network for feature recognition and pattern classification to obtain a frequency response curve and a phase characteristic curve, thereby improving the test accuracy, anti-interference ability, automation degree and adaptability of the intelligent audio testing system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of audio testing, and particularly to a method, device, and equipment for frequency response analysis of an intelligent audio testing system. Background Art

[0002] Traditional audio testing systems mainly rely on manual operation and simple digital signal processing, and have obvious deficiencies in terms of testing accuracy, anti-noise ability, and automation level. Especially in the actual testing environment, external factors such as environmental noise, temperature, and humidity changes will significantly affect the accuracy of the test results, and traditional testing methods are difficult to effectively cope with these interferences.

[0003] Current audio frequency response testing methods generally have problems such as unstable signal quality, insufficient feature extraction, and low data processing efficiency. Especially in complex environments, the test results are easily affected by noise interference and environmental changes, resulting in a decrease in measurement accuracy. At the same time, traditional testing systems lack intelligent data processing capabilities, cannot automatically identify and compensate for various errors during the testing process, and require testers to have rich professional experience to obtain reliable test results. Existing frequency response analysis methods often use a single signal processing algorithm, making it difficult to comprehensively capture the characteristic performance of audio devices in different frequency bands. Especially when dealing with complex characteristics such as nonlinear distortion and phase response, the performance of traditional methods is not ideal. In addition, the lack of effective adaptive mechanisms and intelligent analysis means makes it difficult for the testing system to adapt to the testing requirements of different types of audio devices, restricting the further improvement of testing efficiency and accuracy. Summary of the Invention

[0004] This application provides a method, device, and equipment for frequency response analysis of an intelligent audio testing system, thereby improving the testing accuracy, anti-interference ability, automation level, and adaptability of the intelligent audio testing system.

[0005] In the first aspect of this application, a method for frequency response analysis of an intelligent audio testing system is provided. The method for frequency response analysis of the intelligent audio testing system includes:

[0006] Initialize the audio parameters of the intelligent audio testing system to obtain a test parameter matrix and an excitation signal;

[0007] Based on the test parameter matrix, perform dual-channel acquisition and time-frequency feature extraction on the excitation signal and the response signal to obtain a time-frequency feature data set;

[0008] Perform dual-branch processing of time-domain noise reduction and frequency-domain enhancement on the time-frequency feature data set to obtain a purified feature data set and a feature weighting matrix;

[0009] Perform segmented recursive calculation and feature aggregation on the purified feature dataset according to the feature weighting matrix to obtain a frequency response feature vector;

[0010] Decompose the frequency response feature vector according to the test frequency band, and perform environmental compensation and non-linear calibration to obtain a standard frequency response dataset;

[0011] Input the standard frequency response dataset into a deep neural network for feature recognition and pattern classification to obtain a frequency response curve and a phase characteristic curve.

[0012] The second aspect of the present application provides a frequency response analysis device for an intelligent audio test system. The frequency response analysis device for the intelligent audio test system includes:

[0013] An initialization module for initializing audio parameters of the intelligent audio test system to obtain a test parameter matrix and an excitation signal;

[0014] An acquisition module for performing two-channel acquisition and time-frequency feature extraction on the excitation signal and the response signal based on the test parameter matrix to obtain a time-frequency feature dataset;

[0015] A processing module for performing dual-branch processing of time-domain noise reduction and frequency-domain enhancement on the time-frequency feature dataset to obtain a purified feature dataset and a feature weighting matrix;

[0016] A calculation module for performing segmented recursive calculation and feature aggregation on the purified feature dataset according to the feature weighting matrix to obtain a frequency response feature vector;

[0017] A calibration module for decomposing the frequency response feature vector according to the test frequency band, and performing environmental compensation and non-linear calibration to obtain a standard frequency response dataset;

[0018] A classification module for inputting the standard frequency response dataset into a deep neural network for feature recognition and pattern classification to obtain a frequency response curve and a phase characteristic curve.

[0019] The third aspect of the present application provides an electronic device, including: a memory and at least one processor, wherein instructions are stored in the memory; the at least one processor calls the instructions in the memory so that the electronic device executes the above-mentioned frequency response analysis method for the intelligent audio test system.

[0020] The fourth aspect of the present application provides a computer-readable storage medium, wherein instructions are stored in the computer-readable storage medium, and when it runs on a computer, it causes the computer to execute the above-mentioned frequency response analysis method for the intelligent audio test system.

[0021] Compared with the prior art, the present application has the following beneficial effects: By introducing a dual-branch data processing architecture, the collaborative enhancement of time-domain and frequency-domain signals is achieved, effectively improving the anti-noise ability of the system; adopting a multi-level feature extraction and recursive analysis strategy significantly improves the extraction accuracy of frequency response features; innovatively designing an environmental compensation and non-linear calibration mechanism effectively eliminates the interference effects of environmental factors; introducing a deep neural network for feature recognition and pattern classification realizes the intelligence and automation of the testing process. This method ensures the accuracy of full-band testing through segmented recursive calculation and feature aggregation technology; adopts a two-stream attention network and a residual learning mechanism to enhance the system's ability to capture complex frequency response features; ensures the reliability of test results through a multi-level calibration and compensation mechanism; and finally realizes the high-precision reconstruction of the frequency response curve through the intelligent analysis of the deep learning model. The present invention has achieved remarkable improvements in terms of testing accuracy, anti-interference ability, automation level, and adaptability. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0023] The structures, ratios, sizes, etc. shown in the drawings of this specification are only used to cooperate with the content disclosed in the specification for those familiar with this technology to understand and read, and are not used to limit the limited conditions under which the present invention can be implemented. Therefore, they do not have substantial technical significance. Any modification of the structure, change in the proportional relationship, or adjustment of the size, without affecting the effects that the present invention can produce and the purposes that can be achieved, should still fall within the scope that can be covered by the technical content disclosed in the present invention.

[0024] Figure 1 is a schematic flowchart of the frequency response analysis method of the intelligent audio testing system provided by the embodiment of the present invention;

[0025] Figure 2 is a schematic block diagram of the structure of the frequency response analysis device of the intelligent audio testing system provided by the embodiment of the present invention;

[0026] Figure 3 is a schematic block diagram of the structure of the electronic device provided by the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0027] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0028] The flowcharts shown in the accompanying drawings are only illustrative examples, and do not necessarily include all the contents and operations / steps, nor do they necessarily need to be executed in the described order. For example, some operations / steps can also be decomposed, combined, or partially merged. Therefore, the actual execution order may change according to the actual situation.

[0029] It should also be understood that the terms used in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application. As used in the specification of this application and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms.

[0030] It should be further understood that the term "and / or" used in the specification of this application and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations. Please refer to Figure 1 , an embodiment of the frequency response analysis method of the intelligent audio test system in the embodiments of this application includes:

[0031] Step 100: Initialize the audio parameters of the intelligent audio test system to obtain a test parameter matrix and an excitation signal;

[0032] It can be understood that the execution subject of this application can be a frequency response analysis device of an intelligent audio test system, or a terminal or a server. Specifically, it is not limited here. In the embodiments of this application, the server is used as an example of the execution subject for illustration.

[0033] Specifically, the digital signal processing (DSP) algorithm is configured for the processor of the intelligent audio test system to optimize and enhance the audio processing ability of the system. During this process, an all-pass filter and a dynamic equalizer (DEQ) are loaded into the intelligent audio test system. The all-pass filter is used to keep the amplitude of the signal unchanged while adjusting its phase response, and the dynamic equalizer is used to adjust the frequency response in real time to make the system's response to different frequency bands more balanced and accurate. By loading these audio processing modules, DSP initialization parameters are obtained. Based on the DSP initialization parameters, the system sampling parameters are set to obtain the basic sampling parameters. The basic sampling parameters include parameters related to audio signal acquisition such as the sampling rate and bit depth. The reasonable setting of these parameters can effectively ensure the accuracy of signal acquisition and the stability of the system. According to the basic sampling parameters, the signal ranges of 14 digital gain channels of the system are set to obtain a channel parameter set, ensuring that the signal gain range of each channel is reasonable, so that each channel can operate within its predetermined dynamic range without overloading or underperforming. The gain step of the channel parameter set is set to precisely control the amplitude adjustment of each gain channel, and in this way, a gain control matrix is generated. The setting of the gain step enables the system to be tested under different gain conditions, covering a wider dynamic range to ensure that the test results are representative and comprehensive. Based on the gain control matrix, the dynamic range of the system is verified. By performing the measurement of the noise floor index, the obtained dynamic range verification data reflects the noise performance and dynamic response ability of the system under different gain conditions. According to the dynamic range verification data, the signal-to-noise ratio index is measured and calculated to obtain the performance parameters of the system. By calculating the signal-to-noise ratio, the anti-interference ability of the system in the face of different noise levels is evaluated. The system performance parameters and the gain control matrix are combined and calculated to obtain a test parameter matrix. The test parameter matrix contains all the key parameters in the audio test process and is used to guide and standardize the entire test process. Based on the generated test parameter matrix, a test signal is generated to obtain an excitation signal. The excitation signal is the core part of the audio test and is used to excite the audio system to analyze the frequency response characteristics and dynamic performance of the system. By reasonably configuring the test parameter matrix, the excitation signal can cover all the key frequency points in the test frequency band to ensure the comprehensiveness and accuracy of the test results. The generation of the excitation signal also needs to consider the non-linear characteristics and dynamic response of the system so that the test can truly reflect the actual performance of the system.

[0034] Step 200: Based on the test parameter matrix, two-channel acquisition and time-frequency feature extraction are performed on the excitation signal and the response signal to obtain a time-frequency feature data set;

[0035] Specifically, perform 24-bit analog-to-digital (A / D) conversion on the excitation signal and the response signal, set the number of sampling points to 2048 points / frame, sample the signal 2048 times in each frame to obtain the original sampling data. Based on the test parameter matrix, perform signal allocation for the original sampling data in 14 channels to obtain channel sampling data. Each channel corresponds to a specific frequency band or signal feature, ensuring that signals in each frequency band can be independently analyzed and processed, improving the resolution and accuracy of different frequency components. Perform 1dB step gain adjustment on the channel sampling data, and finely control the amplitude of the signal through the processor to obtain gain control data. The 1dB step amplitude can finely adjust the signal under different gain states to ensure that the amplitude of the signal is always within a reasonable range during the test, avoiding distortion or information loss caused by overly strong or weak signals. The gain adjustment step ensures the stability and effectiveness of subsequent signal processing steps by controlling the amplitude of the signal. Input the gain control data into an all-pass filter for signal preprocessing. The all-pass filter adjusts the phase response without changing the signal amplitude to optimize the time-domain characteristics of the signal. By performing dynamic equalization algorithm operations, adjust the frequency response characteristics to eliminate the imbalance phenomenon in the spectrum and obtain filtered processing data. Perform real-time Fourier transform on the filtered signal to convert the time-domain signal into a frequency-domain signal, obtain frequency-domain characteristic data, and obtain the amplitude and phase information of the signal at different frequencies. Perform phase characteristic extraction on the frequency-domain characteristic data, calculate the group delay and phase response to obtain phase characteristic data. The group delay reflects the time delay of different frequency components passing through the system, while the phase response reflects the phase change of the signal. Perform information fusion operations on the frequency-domain characteristic data and the phase characteristic data, and perform feature alignment processing on them to obtain fusion characteristic data. Perform normalization processing on the fusion characteristic data to obtain a time-frequency characteristic data set. Adjust the eigenvalues to the same dimension range to avoid deviations caused by different feature scales. The normalized feature data set has better comparability and stability.

[0036] Step 300: Perform dual-branch processing of time-domain noise reduction and frequency-domain enhancement on the time-frequency characteristic data set to obtain a purified feature data set and a feature weighting matrix;

[0037] It should be noted that the time-frequency feature dataset is branched and divided, and the signal is divided into a time-domain processing branch and a frequency-domain processing branch to form a dual-branch dataset. Perform 4-level wavelet decomposition operation on the time-domain branch in the dual-branch dataset. Through multi-level decomposition, the signal is decomposed into details and approximation components of different scales. Set the threshold to 3 times the standard deviation of the noise to distinguish the signal from the noise. In this way, the noise components in the signal are extracted at different scales, so as to effectively suppress it and obtain time-domain denoised data. At the same time, for the frequency-domain branch in the dual-branch dataset, perform spectral subtraction operation to enhance the frequency-domain signal characteristics. In the spectral subtraction operation, set the frequency-related suppression factor to 0.8, which means that 80% of the noise is suppressed in the signal spectrum, while the effective components of the signal are retained. The use of spectral subtraction effectively reduces the influence of background noise and enhances the spectral characteristics of the signal, making the frequency-domain characteristics of the signal more prominent and clear. The obtained frequency-domain enhanced data has a high spectral signal-to-noise ratio and frequency-domain smoothness. Calculate the signal-to-noise ratio index based on the time-domain denoised data, and set the signal-to-noise ratio threshold to 40 dB to ensure that the characteristics of the signal are more accurately extracted under the condition of low noise level. Through this calculation, the time-domain weight coefficient is obtained, which reflects the contribution degree of time-domain processing to the signal. At the same time, calculate the spectral flatness parameter based on the frequency-domain enhanced data, and set the flatness deviation range to ±3 dB to ensure that the spectral characteristics have high consistency and smoothness. Through this calculation, the frequency-domain weight coefficient is obtained, and the frequency-domain weight coefficient reflects the effect of frequency-domain enhancement processing and the smoothness of the signal. Perform weighted combination on the time-domain weight coefficient and the frequency-domain weight coefficient, and perform weight normalization operation to obtain the feature weighted matrix. Standardize the weights of different features so that the feature contributions of the time domain and the frequency domain can be balancedly considered in subsequent feature fusion, and avoid a certain feature occupying too large a weight. Based on the feature weighted matrix, perform feature fusion on the time-domain denoised data and the frequency-domain enhanced data to obtain a fused feature set. Based on indicators such as frequency response flatness, phase linearity, and group delay fluctuation, perform feature screening operation on the fused feature set to obtain a purified feature dataset. Frequency response flatness is used to evaluate the consistency performance of the system at different frequencies, phase linearity is used to evaluate whether the phase response of the system changes linearly, and group delay fluctuation reflects whether the delay characteristics of the system are stable at different frequencies. Through the evaluation of these indicators, the features that make significant contributions to the frequency response characteristics of the system are screened out from the fused feature set, and the insignificant or unhelpful parts for analysis are removed to obtain a purified feature dataset.

[0038] Step 400: Perform segmented recursive calculation and feature aggregation on the purified feature dataset according to the feature weighted matrix to obtain a frequency response feature vector;

[0039] Specifically, the purified feature dataset is divided into frequency bands. The frequency range from 20 Hz to 20 kHz is divided into 128 frequency bands, and each frequency band represents a specific frequency range for more refined signal analysis. Band-pass filtering operations are performed within each divided frequency band. The role of band-pass filtering is to retain the signal components within a specific frequency range while suppressing the noise and irrelevant signals in other frequency bands, obtaining frequency band division data. The frequency band division data is encoded and transformed according to the feature weighting matrix to construct a 16-dimensional frequency encoding vector. In this encoding vector, the amplitude information occupies 8 dimensions, the phase information occupies 4 dimensions, and the time information occupies 4 dimensions. This encoding method can fully express various characteristics of the frequency, including the strength change of the signal (amplitude), the change characteristics of the phase (phase), and the dynamic characteristics in the time domain (time). The finally obtained frequency encoding data has the ability to describe multi-dimensional information. Recursive sampling processing is performed on the frequency encoding data. In the normal region, the basic sampling density is set to 4 sampling points per octave to ensure the sampling uniformity and coverage rate of the signal across the entire frequency band. In the region where the frequency response changes abruptly, in order to capture the subtle changes in the signal, the sampling density is increased to 16 sampling points per octave. The method of dynamically adjusting the sampling density can effectively handle the regions with drastic changes in the frequency response, making the sampling data more representative and obtaining multi-density sampling data. The multi-density sampling data is input into a three-level feature extraction unit for processing. The first-level feature extraction unit performs time-domain convolution operations to extract the transient features of the signal, and the transient features reflect the rapid changes of the signal in time. The second-level feature extraction unit performs frequency-domain transformation to extract the harmonic features of the signal, and the harmonic features reflect the structural properties of the signal in the frequency domain. The third-level feature extraction unit extracts the correlation features of the signal through correlation analysis, analyzing the internal connections between different parts of the signal. The obtained three-dimensional feature data comprehensively describes the time domain, frequency domain, and their internal correlation characteristics of the signal. A two-layer prediction unit is constructed for the three-dimensional feature data. The first layer processes the temporal relationship through a long short-term memory network (LSTM). The advantage of LSTM is that it can remember the long-term dependence relationship of the signal. Through the processing of this layer, a temporal prediction value is obtained to describe the evolution trend of the signal in time. The second layer fuses the frequency-domain information through a feed-forward neural network to obtain a frequency response prediction value, and the frequency response prediction value can describe the response characteristics of the signal in frequency. Through the joint processing of these two layers, prediction feature data is obtained. A three-level residual network is constructed based on the prediction feature data to eliminate the biases and errors in the features. The first-level residual network calculates the amplitude residual to eliminate the errors in the signal amplitude and ensure the accuracy of the signal amplitude characteristics. The second-level residual network calculates the phase residual to ensure the consistency and accuracy of the signal in phase through phase correction. The third-level residual network calculates the time delay residual, and the time delay residual reflects the delay characteristics of the signal during propagation. Through this step, the errors caused by time delay are eliminated, obtaining residual feature data.Input the residual feature data into the feature fusion network, and perform feature dimensionality reduction and optimization processing through the fully connected layer. In the fully connected layer, the number of output layer nodes is set to 32 to ensure that the features retain their main information components while reducing the dimensionality, and obtain the optimized feature data. Perform normalization processing on the optimized feature data to eliminate the influence caused by different dimensions between different features. Compress the three-dimensional features into a one-dimensional vector through weighted summation operation, and finally obtain the frequency response feature vector.

[0040] Step 500: Decompose the frequency response feature vector according to the test frequency band, and perform environmental compensation and nonlinear calibration to obtain the standard frequency response data set;

[0041] Specifically, a three-segment frequency division is performed on the frequency response eigenvector. The frequency range from 20 Hz to 200 Hz is defined as the low-frequency band, the range from 200 Hz to 2 kHz is defined as the mid-frequency band, and the range from 2 kHz to 20 kHz is divided into the high-frequency band, obtaining segmented frequency response data. Targeted analysis and processing are carried out according to the sensitivity of the human ear to different frequencies and the performance characteristics of audio devices in different frequency bands. An adaptive interval model is constructed based on the segmented frequency response data. By setting the confidence level to 95%, the calculation of the maximum deviation is performed on the frequency response data to obtain the interval boundary data for each frequency band. The interval boundary data is used to describe the variation range of the frequency response characteristics within different frequency bands. By calculating the maximum deviation, the upper and lower limits of these variations are determined, thereby reflecting whether the frequency response of the system is within the expected range. To describe the uncertainty and volatility of the frequency response characteristics, a mixed probability modeling is performed on the interval boundary data, where the weight of the Gaussian distribution is set to 0.7, and the weight of the uniform distribution is set to 0.3. The purpose of the mixed probability modeling is to combine the central tendency (Gaussian distribution) and uniform variation (uniform distribution) of the signal to better describe the distribution of the frequency response characteristics in different frequency bands, obtaining probability distribution data. According to the influence of the actual test environment, an environmental compensation matrix is constructed. Considering that temperature and humidity are important factors affecting audio test results, the environmental compensation matrix is constructed based on the temperature coefficient of ±0.01 dB / °C and the humidity coefficient of ±0.02 dB / 10%RH. A compensation operation is performed on the probability distribution data to obtain environmental compensation data. The changes in temperature and humidity will affect the response characteristics of audio devices. By adjusting the signal through the compensation matrix, the error introduced by the change in environmental conditions is effectively reduced, making the frequency response data more accurate and reliable. The environmental compensation data is compared with the 1 kHz / -20 dBFS reference signal for calculation to obtain the error parameters of the system, obtaining calibration reference data. The 1 kHz / -20 dBFS reference signal is a commonly used reference signal in audio testing. By comparing with the compensation data, the error of the system under different conditions is evaluated. According to these error parameters, a 12th-order finite impulse response (FIR) filter is constructed to perform the compensation operation for nonlinear errors, obtaining compensated calibration data. The FIR filter is a filter used for signal processing. The higher its order, the stronger its ability to adjust the signal. The nonlinear error of the system is effectively compensated by the 12th-order FIR filter, making the calibrated data more conform to the standard response. Based on the compensated calibration data, the segmented frequency response data is calibrated and corrected. The correction coefficients for the low-frequency band, mid-frequency band, and high-frequency band are calculated respectively to obtain frequency band calibration data. The correction coefficients for different frequency bands reflect the amount of compensation required within different frequency ranges. Through precise correction, the frequency response deviation of the device within a specific frequency range can be eliminated, making the response of the system more tend to the ideal state. After correcting the calibration data for each frequency band, data integration processing is performed on the calibration data for the entire frequency band to ensure smooth transition and no mutation between frequency bands, obtaining a standard frequency response data set.

[0042] Step 600: Input the standard frequency response data set into a deep neural network for feature recognition and pattern classification to obtain a frequency response curve and a phase characteristic curve.

[0043] Specifically, the standard frequency response data set is input into the feature extraction network, which consists of an 8-layer convolutional neural network. Each convolutional layer uses the ReLU activation function. The activation function can introduce non-linearity and improve the network's ability to express complex patterns. A batch normalization layer is added to each convolutional layer to stabilize the training process, accelerate convergence, and mitigate the overfitting problem of the network. Through these convolutional operations, a multi-scale feature map containing rich frequency response features is extracted, capturing the detailed information of the signal in different time and frequency ranges. The feature pyramid operation is performed on the multi-scale feature map. Through this operation, 4 feature layers with different resolutions are constructed, namely the original resolution layer and three feature layers with downsampling ratios of 1 / 2, 1 / 4, and 1 / 8. These feature layers reflect the characteristics of the signal at different scales. The use of the feature pyramid ensures a multi-level analysis of the frequency response data, enabling the network to simultaneously focus on the global structure and local details. Through the multi-level feature representation, a more refined description of the frequency response features is obtained. The multi-level feature data is input into the dual-stream attention network for processing. The dual-stream attention network consists of a frequency channel attention module and a temporal-spatial attention module. The frequency channel attention module uses a 1×1 convolutional layer to extract the weights of the channels. These weights dynamically adjust the contributions of different frequency channels, making the network pay more attention to the frequency bands that are important for the frequency response. The temporal-spatial attention module uses a 5×5 convolutional layer to extract the position weights in space, capturing the important characteristics of the signal in time and space, and enhancing the network's ability to focus on key features. Through the processing of the dual-stream attention network, attention feature data is obtained, containing the weight information of different channels and different positions. Feature aggregation is performed on the attention feature data. Through the cross-layer feature fusion module, the feature maps with different resolutions are weighted and fused. This module learns the fusion weights between the feature maps of different layers through a 1×1 convolutional layer, so that the features of each layer can reasonably allocate weights during the fusion process, avoiding feature loss or overemphasis caused by different scales. Through feature fusion, fused feature data is obtained. The fused feature data is input into the classification branch network, which contains three fully connected layers with the number of nodes being 512, 256, and 128 respectively. A Dropout layer with a dropout rate of 0.5 is connected after each fully connected layer, which is used to randomly discard half of the nodes to prevent overfitting and thus improve the generalization ability of the network. Through the processing of the fully connected layers, a frequency classification result is obtained, which is used to determine which frequency response pattern the input frequency response data belongs to. At the same time, regression branch processing is performed on the fused feature data. Deep features are extracted through three residual blocks, and each residual block contains two 3×3 convolutional layers and a shortcut connection. The introduction of the residual block aims to alleviate the vanishing gradient problem in the training of deep networks. Through the shortcut connection, information is transmitted more smoothly in the deep network, improving the learning ability of deep features. Through the processing of these three residual blocks, regression feature data describing the change trend of the frequency response characteristics is obtained.The regression feature data is input into the dual-head output layer. The first output head contains two fully connected layers for predicting the frequency response value. Through this output head, the amplitude characteristics of the signal at different frequencies are obtained. The second output head contains two fully connected layers for predicting the phase response value and obtaining the phase change characteristics of the signal. Through the processing of the dual-head output layer, the complete prediction data of the frequency response and phase response are obtained, reflecting the frequency response characteristics and phase characteristics of the signal. The frequency classification results and response prediction data are processed continuously by the spline interpolation algorithm. The spline interpolation can construct a smooth curve between discrete prediction points, so that the frequency response curve and the phase characteristic curve are continuous, which meets the frequency response measurement requirements in practical applications, and finally the frequency response curve and the phase characteristic curve are obtained.

[0044] In the embodiment of the present application, by introducing a dual-branch data processing architecture, the coordinated enhancement of time domain and frequency domain signals is achieved, and the anti-noise ability of the system is effectively improved; the multi-level feature extraction and recursive analysis strategy is adopted to significantly improve the extraction accuracy of frequency response features; the environmental compensation and nonlinear calibration mechanism are innovatively designed to effectively eliminate the interference of environmental factors; the deep neural network is introduced for feature recognition and pattern classification to realize the intelligence and automation of the test process. This method ensures the accuracy of full-band testing through segmented recursive calculation and feature aggregation technology; adopts a dual-stream attention network and residual learning mechanism to enhance the system's ability to capture complex frequency response features; through a multi-level calibration and compensation mechanism, the reliability of the test results is ensured; and finally, through the intelligent analysis of the deep learning model, high-precision frequency response curve reconstruction is achieved. The present invention has achieved significant improvements in test accuracy, anti-interference ability, degree of automation and adaptability.

[0045] In a specific embodiment, the process of executing step 100 may specifically include the following steps:

[0046] Perform DSP algorithm configuration on the processor of the intelligent audio test system, load the all-pass filter and dynamic equalizer DEQ into the intelligent audio test system, and obtain DSP initialization parameters;

[0047] The system sampling parameters are set based on the DSP initialization parameters to obtain the basic sampling parameters, and the signal ranges of the 14 digital gain channels are set according to the basic sampling parameters to obtain the channel parameter set;

[0048] Gain step settings are performed on the channel parameter set to obtain a gain control matrix, and the system dynamic range is verified based on the gain control matrix. The noise floor index is measured to obtain dynamic range verification data.

[0049] Verify the data according to the dynamic range, measure and calculate the signal-to-noise ratio index to obtain the system performance parameters, and perform a combined operation on the system performance parameters and the gain control matrix to obtain the test parameter matrix. Generate a test signal based on the test parameter matrix to obtain the excitation signal.

[0050] Specifically, configure the DSP (Digital Signal Processing) algorithm for the processor of the intelligent audio test system, and load the all-pass filter and dynamic equalizer modules into the intelligent audio test system. The all-pass filter is a filter that adjusts the signal phase while keeping the signal amplitude unchanged, and is expressed in the following transfer function form:

[0051] ;

[0052] Among them, is the transfer function of the filter, is the complex frequency variable, is the filter parameter used to adjust the phase response. The purpose of loading the all-pass filter is to optimize the phase response without affecting the amplitude characteristics. The dynamic equalizer is a filter that dynamically adjusts the frequency response and is used to control the gain of specific frequency bands so that the response of the signal at each frequency is more balanced. The function of the dynamic equalizer is expressed by the following formula:

[0053] ;

[0054] Among them, is the gain at a certain frequency, is the basic gain, is the frequency at which the gain adjustment amount, and the gain adjustment of the th frequency band consists of the equalization control of the th frequency band. After loading these modules into the DSP system, the DSP initialization parameters are obtained. Based on the DSP initialization parameters, set the sampling parameters of the system to obtain the basic sampling parameters. The sampling parameters include key attributes such as the sampling frequency and quantization bit depth. The sampling frequency is usually set to kHz, that is, sampling 48000 times per second, to ensure sufficient time-domain resolution for the audio signal. The quantization bit depth is set to bits, and each sampling point represents different levels, thus ensuring the high fidelity of the signal. According to these basic sampling parameters, set the signal range for 14 digital gain channels in the system to obtain the channel parameter set. The gain range setting for each channel is designed to ensure that the signal will not be overloaded or underloaded under different gain conditions. Let the gain range be , and the channel signal range is expressed by the following formula:

[0055] ;

[0056] Among them, is the output signal voltage, is the input signal voltage, is the gain, and its value range is between and . By setting a reasonable gain range, stable output characteristics are maintained under various input conditions. The gain step setting is performed on the channel parameter set to obtain the gain control matrix. The signals at different gain levels are adjusted and tested step by step to find the optimal gain configuration. The step size of the gain step is set to 1 dB, and the adjustment from the minimum gain to the maximum gain is expressed as:

[0057] ;

[0058] Among them, is the step size, usually set to 1 dB, is the number of steps. The gain control matrix generated in this way contains the gain configurations of each channel in different gain states. The dynamic range of the system is verified according to the gain control matrix, and the background noise index measurement is performed to obtain the dynamic range verification data. The dynamic range is expressed as the ratio of the maximum output voltage of the system to the minimum detectable output voltage, expressed as:

[0059] ;

[0060] Among them, is the dynamic range, is the maximum output voltage, is the background noise voltage. By measuring the background noise voltage of the system, the dynamic range of the system is calculated to evaluate the overall performance of the system. Based on the dynamic range verification data, the signal-to-noise ratio index is measured and calculated to obtain the performance parameters of the system. The calculation formula of the signal-to-noise ratio is:

[0061] ;

[0062] Among them, is the voltage of the effective signal, is the background noise voltage. The signal-to-noise ratio is an important index to evaluate the quality of the audio system, and the larger its value, the better the signal quality of the system. After obtaining the system performance parameters, the system performance parameters and the gain control matrix are combined and calculated to form a test parameter matrix containing the performance of the system in different states. The test parameter matrix is used to describe the performance of the system under different gains and different input conditions, and its form is expressed as a matrix , where is the number of gain steps, is the number of each performance parameter (such as dynamic range, signal-to-noise ratio, etc.). Each row of the matrix represents the system performance under a specific gain, and each column represents a specific performance parameter. Based on the test parameter matrix, a test signal is generated to obtain an excitation signal. The type of the excitation signal is a swept-frequency signal, whose frequency changes linearly or logarithmically with time, covering the entire frequency range of audio testing. The mathematical expression of the swept-frequency signal is:

[0063] ;

[0064] where is the value of the excitation signal at time , is the signal amplitude, is the starting frequency, is the frequency change rate (i.e., the swept-frequency rate). The generated excitation signal is used as an input to drive the audio device under test, thereby measuring its frequency response characteristics.

[0065] In a specific embodiment, the process of executing step 200 may specifically include the following steps:

[0066] Perform 24-bit analog-to-digital conversion on the excitation signal and the response signal, set the number of sampling points to 2048 points / frame to obtain the original sampling data, and perform 14-channel signal allocation on the original sampling data based on the test parameter matrix to obtain channel sampling data;

[0067] Perform 1dB step gain adjustment on the channel sampling data, control the signal amplitude through the processor to obtain gain control data;

[0068] Input the gain control data into an all-pass filter for signal preprocessing, perform dynamic equalization DEQ algorithm operations to obtain filtered processing data, and perform real-time Fourier transform on the filtered processing data to obtain frequency-domain characteristic data;

[0069] Perform phase characteristic extraction on the frequency-domain characteristic data, calculate the group delay and phase response to obtain phase characteristic data, and perform information fusion operations on the frequency-domain characteristic data and the phase characteristic data, perform feature alignment processing to obtain fusion characteristic data;

[0070] Perform normalization processing on the fusion characteristic data to obtain a time-frequency characteristic data set.

[0071] Specifically, perform 24-bit analog-to-digital conversion (A / D conversion) on the excitation signal and the response signal. The 24-bit quantization depth divides the signal into A number of different level values. The number of sampling points is set to 2048 points per frame, and the number of points for signal sampling within each frame is 2048. The obtained original sampling data is the discrete representation of the signal in the time domain. Based on the test parameter matrix, the original sampling data is distributed into 14 channels of signals. The test parameter matrix defines specific parameters for different channels, such as gain range, frequency band division, etc. Assuming there are 14 channels, each channel is used for the analysis of a specific frequency range, and the distribution is represented mathematically as:

[0072] ;

[0073] where, represents the output signal of the th channel, is the signal processing function related to the channel, is the original input signal. In this process, the signal is distributed into different channels for analysis and processing within their respective frequency ranges, and the obtained channel sampling data is the specific manifestation of the signal in each channel. A 1dB step gain adjustment is performed on the channel sampling data. Through step-by-step gain adjustment, the characteristics of the signal under different gain conditions and the non-linear behavior of the system are effectively tested. The gain adjustment is described by the following formula:

[0074] ;

[0075] where, is the signal after gain adjustment, is the gain factor, is the gain value, with the unit of decibel (dB). The step size is 1dB, which means that in each round of adjustment, the gain value increases or decreases by 1dB. By controlling the amplitude of the signal through the processor, the signal performance under different gains is obtained, thereby obtaining the gain control data. The gain control data is input into an all-pass filter for signal preprocessing. The all-pass filter is used to adjust the phase response of the signal while keeping the amplitude unchanged. The transfer function of the all-pass filter is expressed as:

[0076] ;

[0077] where, is the transfer function of the filter, is the complex frequency variable, is the filter parameter used to adjust the phase response of the filter. After the all-pass filter adjusts the phase of the signal, the operation of the dynamic equalization algorithm is performed. The gain is dynamically adjusted according to the spectral characteristics of the signal to enhance specific frequency components and equalize the frequency response of the signal. The equalized signal is expressed by the following formula:

[0078] ;

[0079] Among them, is the signal after equalization processing, is the dynamic gain factor, is the impulse response of the equalizer, is the number of frequency bands. Through equalization processing, the filtered data is obtained. The filtered data is subjected to real-time Fourier transform to convert the signal from the time domain to the frequency domain, and the frequency domain characteristic data is obtained. The formula for Fourier transform is:

[0080] ;

[0081] Among them, is the th spectrum value of the frequency component, is the number of points of the Fourier transform, is the time-domain signal, is the complex exponential basis function. Through Fourier transform operation, the amplitude and phase characteristics of the signal in the frequency domain are obtained. The phase characteristic extraction is performed on the frequency domain characteristic data to calculate the group delay and phase response of the signal. The phase response describes the phase change of different frequency components of the signal, and the group delay is defined as the derivative of the phase response with respect to frequency, expressed as:

[0082] ;

[0083] Among them, is the group delay, is the phase response, is the frequency. By calculating the group delay, the delay characteristics of the signal at different frequencies are judged, and the phase characteristic data is obtained. The frequency domain characteristic data and the phase characteristic data are subjected to information fusion operation, and feature alignment processing is performed to effectively combine the two. The purpose of feature alignment is to ensure the correspondence relationship of different types of features at the same time and frequency. The information fusion is expressed by the following formula:

[0084]

[0085] Among them, is the fused feature data, is the frequency domain characteristic data, is the phase characteristic data, and are the fusion weights, which are used to control the relative importance of the two features. The fused feature data is normalized to obtain the time-frequency feature dataset.

[0086] In a specific embodiment, the process of executing step 300 may specifically include the following steps:

[0087] The time-frequency feature dataset is branched and divided, and the signal is divided into a time-domain processing branch and a frequency-domain processing branch to obtain a two-branch dataset;

[0088] Perform a 4-level wavelet decomposition operation on the time-domain branch in the two-branch dataset, and set the threshold to 3 times the standard deviation of the noise to obtain time-domain noise-reduced data;

[0089] Perform spectral subtraction on the frequency-domain branch in the two-branch dataset, and set the frequency-related suppression factor to 0.8 to obtain frequency-domain enhanced data;

[0090] Calculate the signal-to-noise ratio index based on the time-domain noise-reduced data, set the signal-to-noise ratio threshold to 40 dB to obtain the time-domain weight coefficient, and calculate the spectral flatness parameter based on the frequency-domain enhanced data, and set the flatness deviation range to ±3 dB to obtain the frequency-domain weight coefficient;

[0091] Perform weighted combination on the time-domain weight coefficient and the frequency-domain weight coefficient, perform weight normalization operation to obtain a feature weighting matrix, and perform feature fusion on the time-domain noise-reduced data and the frequency-domain enhanced data according to the feature weighting matrix to obtain a fused feature set;

[0092] Based on the frequency response flatness, phase linearity, and group delay fluctuation indexes, perform feature screening operation on the fused feature set to obtain a purified feature dataset.

[0093] Specifically, the time-frequency feature dataset is branched and divided, and the signal is divided into a time-domain processing branch and a frequency-domain processing branch to obtain a two-branch dataset. The time-domain and frequency-domain characteristics of the signal are processed separately to fully extract the important information of the signal in the time and frequency dimensions. Time-domain processing is used to analyze and reduce the instantaneous noise in the signal, while frequency-domain processing is used to enhance the frequency characteristics of the signal and improve the flatness of the spectrum. In the time-domain processing branch, a 4-level wavelet decomposition operation is performed on the time-domain data to separate the effective information and noise components in the signal. Wavelet decomposition is a time-frequency localization method that decomposes the signal into details and approximations of different scales by using wavelet functions. The wavelet decomposition is expressed by the following formula:

[0094] ;

[0095] where, represents the original input signal, is the scaling function, used to represent the low-frequency components, is the wavelet function, representing the detail components of different frequencies, and are the approximation and detail coefficients respectively. Among them, represents the decomposition level, which is equal to 4 levels. After the decomposition is completed, in order to remove the noise in the signal, the threshold is set to 3 times the standard deviation of the noise:

[0096] ;

[0097] Wherein, represents the threshold for noise suppression, represents the standard deviation of the noise. Perform soft threshold processing on the wavelet coefficients, retain the coefficients exceeding the threshold, and suppress the coefficients less than the threshold to obtain the data after time-domain noise reduction. The soft threshold function is expressed as:

[0098] ;

[0099] Wherein, is the processed detail coefficient, represents the sign of the coefficient. Through this processing, the noise interference is effectively reduced to obtain a clean time-domain signal. At the same time, in the frequency-domain processing branch, perform spectral subtraction on the frequency-domain data to enhance the frequency characteristics of the signal. Spectral subtraction is an effective noise reduction method, which estimates the spectrum of the noise and subtracts it from the spectrum of the signal to achieve the effect of enhancing the signal. Let represent the spectrum of the input signal, represent the noise spectrum, then the spectral subtraction operation is expressed as:

[0100] ;

[0101] Wherein, is the signal enhanced in the frequency domain, is the frequency-dependent suppression factor, which is set to 0.8, meaning that 80% of the noise is suppressed. Through this operation, the proportion of noise in the spectrum is reduced, the clarity of the signal is improved, and the data enhanced in the frequency domain is obtained. Calculate the signal-to-noise ratio (SNR) based on the time-domain noise-reduced data, and set the SNR threshold to 40 dB to obtain the time-domain weight coefficient. The calculation formula for the signal-to-noise ratio is:

[0102] ;

[0103] Wherein, SNR represents the signal-to-noise ratio, represents the power of the signal, represents the power of the noise. If the calculated signal-to-noise ratio is higher than 40 dB, it is considered that the signal quality is good, corresponding to a higher time-domain weight coefficient. The value of the weight coefficient is linearly mapped according to the size of the signal-to-noise ratio, for example:

[0104] ;

[0105] Wherein, is the time-domain weight coefficient, is the maximum possible value of the signal-to-noise ratio. At the same time, the spectral flatness parameter is calculated based on the frequency-domain enhanced data to evaluate the uniformity of the spectrum. Spectral flatness is an index that measures the degree of uniformity of the signal distribution in the frequency domain, and its calculation formula is:

[0106] ;

[0107] where, represents the spectral flatness, is the number of spectral components, represents the th spectral component amplitude. If the flatness deviation range is within ±3 dB, the spectrum is considered to be evenly distributed, and the corresponding frequency-domain weight coefficient is relatively high. The frequency-domain weight coefficient is calculated as follows:

[0108] ;

[0109] where, is the frequency-domain weight coefficient, is the reference spectral flatness value, is the flatness deviation range, set to 3 dB. The time-domain weight coefficient and the frequency-domain weight coefficient are weighted and combined, and a weight normalization operation is performed to obtain the feature weighted matrix. The weighted combination is expressed as:

[0110] ;

[0111] where, is the feature weighted matrix, is a small numerical constant to prevent the denominator from being zero. According to the feature weighted matrix, the time-domain noise-reduced data and the frequency-domain enhanced data are feature-fused, and the fusion process is described by the following formula:

[0112] ;

[0113] where, is the time-domain noise-reduced data, is the frequency-domain enhanced data. Through feature fusion, the time-domain and frequency-domain features are combined to obtain a more comprehensive and accurate feature description. Feature screening operations are performed on the fused feature data to obtain a purified feature data set. The screening criteria are based on the frequency response flatness, phase linearity, and group delay fluctuation indicators. The frequency response flatness is used to measure whether the response of the signal in different frequency bands is balanced, the phase linearity reflects whether the phase response changes linearly, and the group delay fluctuation describes whether the delays of different frequency components passing through the system are consistent. The calculation formula of the group delay is:

[0114] ;

[0115] where, is the group delay, is the frequency The phase response at. When performing feature screening on the fused feature set, if some features do not meet the above indicators, these features are eliminated or adjusted to obtain the final purified feature data set.

[0116] In a specific embodiment, the process of performing step 400 may specifically include the following steps:

[0117] Perform frequency band division on the purified feature data set, divide the range of 20 Hz - 20 kHz into 128 frequency bands, and perform band-pass filtering operations on each frequency band to obtain frequency band division data;

[0118] Perform coding conversion on the frequency band division data according to the feature weighting matrix to construct a 16-dimensional frequency coding vector, where the amplitude occupies 8 dimensions, the phase occupies 4 dimensions, and the time occupies 4 dimensions, to obtain frequency coding data;

[0119] Perform recursive sampling processing on the frequency coding data, set the basic sampling density to 4 points / octave, and set the sampling density to 16 points / octave in the frequency response mutation region to obtain multi-density sampling data;

[0120] Input the multi-density sampling data into a three-level feature extraction unit. The first layer performs time-domain convolution operations to extract transient features, the second layer performs frequency-domain transformation to extract harmonic features, and the third layer performs correlation analysis to extract correlation features to obtain three-dimensional feature data;

[0121] Construct a two-layer prediction unit for the three-dimensional feature data. The first layer processes the time series relationship through an LSTM network to obtain a time series prediction value, and the second layer fuses the frequency-domain information through a feed-forward network to obtain a frequency response prediction value to obtain prediction feature data;

[0122] Construct a three-level residual network based on the prediction feature data. The first level calculates the amplitude residual, the second level calculates the phase residual, and the third level calculates the time delay residual to obtain residual feature data;

[0123] Input the residual feature data into a feature fusion network, perform feature dimensionality reduction and optimization through a fully connected layer, set the number of output layer nodes to 32 to obtain optimized feature data, perform normalization processing on the optimized feature data, and compress the three-dimensional features into a one-dimensional vector through weighted summation operations to obtain a frequency response feature vector.

[0124] Specifically, perform frequency band division on the purified feature data set, and divide the frequency range from 20 Hz to 20 kHz into 128 frequency bands. Each divided frequency band range is represented by the following formula:

[0125] ;

[0126] where is the The center frequency of a frequency band, is the minimum frequency (20 Hz), is the width of the frequency band, and its value is obtained by dividing the total frequency range (20 kHz - 20 Hz) by 128. A band-pass filtering operation is performed on each frequency band to retain the signal components of that frequency band and suppress the signal interference of other frequency bands, which is achieved through the transfer function of the band-pass filter:

[0127] ;

[0128] Among them, is the th transfer function of the band-pass filter, is the center frequency of this frequency band, is the bandwidth of the frequency band. After band-pass filtering, the data after frequency band division is obtained, and the signal components within each frequency band can be analyzed independently. According to the feature weighting matrix, the frequency band division data is encoded and transformed to construct a 16-dimensional frequency encoding vector. In this encoding vector, the amplitude feature occupies 8 dimensions and is used to describe the energy distribution of the signal in each frequency band. The amplitude feature is expressed as:

[0129] ;

[0130] Among them, is the th dimension of the amplitude feature, is the complex spectrum value of the frequency band division data. The phase feature occupies 4 dimensions and is used to describe the phase change of the signal in different frequency bands, which is expressed by the following formula:

[0131] ;

[0132] Among them, is the phase feature, represents the th phase value of the frequency band. The time feature occupies 4 dimensions and is used to describe the transient change of the signal on the time axis. After encoding, a 16-dimensional frequency encoding vector is obtained. Recursive sampling processing is performed on the frequency encoding data. The basic sampling density is set to 4 points / octave, and the sampling density is increased to 16 points / octave in the region of sudden frequency response change. Higher sampling accuracy is obtained in the region where the frequency response changes violently in order to capture finer feature changes. The formula for recursive sampling is expressed as:

[0133]

[0134] Among them, represents the th sampling density of the frequency band, is the threshold for the sudden change in frequency response. When the amplitude of the frequency band change is greater than this threshold, the sampling density increases to 16 points / octave to ensure that sufficient information is obtained in these regions. Through the sampling strategy, multi-density sampling data is obtained, providing rich features in different frequency ranges. The multi-density sampling data is input into a three-level feature extraction unit to extract features at different levels. In the first layer, a time-domain convolution operation is performed to extract transient features, describing the sudden changes and short-term variations of the signal in time. The calculation of the time-domain convolution is expressed as:

[0135] ;

[0136] where, is the output after convolution, is the input signal, is the convolution kernel, is the length of the convolution kernel. The transient features help analyze the noise bursts and instantaneous changes in the signal. In the second layer, a frequency-domain transformation is performed to extract the harmonic features of the signal. The harmonic features are realized through the fast Fourier transform, which is used to identify the harmonic components in the signal, and these components reflect the frequency structure of the signal. In the third layer, a correlation analysis is carried out to extract the correlation features between different frequency bands. The result of the correlation analysis is represented by a correlation coefficient matrix:

[0137] ;

[0138] where, is the correlation coefficient between the and the frequency bands, and are the means of the and the frequency bands respectively. The correlation analysis helps identify the dependence relationships between different frequency components in the signal, obtaining three-dimensional feature data. A two-layer prediction unit is constructed for the three-dimensional feature data. In the first layer, the time-series relationship is processed through an LSTM network to obtain the time-series prediction value. LSTM (Long Short-Term Memory network) is a recurrent neural network suitable for processing time-series data, which can effectively capture the long-term dependence relationships of the signal. The state update formula of the LSTM network is:

[0139] ;

[0140] where, is the hidden state at the current moment, is the output gate, is the cell state. In this way, the LSTM can process the time-series data and extract the time-series features of the signal. In the second layer, the frequency-domain information is fused through a feedforward network to obtain the frequency response prediction value. The calculation of the feedforward network is expressed as:

[0141] ;

[0142] wherein, is the output of the network, is the weight matrix, is the input data, is the bias, is the activation function. Through the processing of these two layers, the predicted feature data is obtained. Based on the predicted feature data, a three-level residual network is constructed to calculate the amplitude residual for eliminating the error in amplitude prediction. The amplitude residual is expressed as:

[0143] ;

[0144] wherein, is the amplitude residual, is the true amplitude value, is the predicted amplitude value. Calculate the phase residual:

[0145] ;

[0146] wherein, is the phase residual, is the true phase value, is the predicted phase value. Calculate the delay residual:

[0147] ;

[0148] wherein, is the delay residual, and are the true and predicted group delays respectively. Through the three-level residual network, the error in prediction is effectively corrected to obtain the residual feature data. The residual feature data is input into the feature fusion network and undergoes feature dimensionality reduction and optimization processing through the fully connected layer. The fully connected layer is used to reduce the high-dimensional feature data to a lower dimension for further analysis. The number of output layer nodes is set to 32, and the most important information is extracted through dimensionality reduction while reducing the data complexity. The optimized feature data after dimensionality reduction undergoes normalization processing to eliminate the scale difference between different features. Through the weighted summation operation, the three-dimensional features are compressed into a one-dimensional vector to obtain the frequency response feature vector, and the formula is:

[0149] ;

[0150] wherein, is the final frequency response feature vector, is the weighting coefficient, is the i-th dimension in the three-dimensional feature data. Through the weighted summation operation, different features are effectively fused to obtain the frequency response feature vector describing the frequency response characteristics of the entire system.

[0151] In a specific embodiment, the process of executing step 500 may specifically include the following steps:

[0152] Perform a three - segment frequency division on the frequency response feature vector, divide 20Hz - 200Hz into the low - frequency band, 200Hz - 2kHz into the mid - frequency band, and 2kHz - 20kHz into the high - frequency band to obtain segmented frequency response data;

[0153] Based on the segmented frequency response data, construct an adaptive interval model, set the confidence level to 95%, perform the maximum deviation calculation to obtain interval boundary data, and perform a mixed probability modeling on the interval boundary data, set the Gaussian distribution weight to 0.7 and the uniform distribution weight to 0.3 to obtain probability distribution data;

[0154] Construct an environmental compensation matrix according to the temperature coefficient ±0.01dB / °C and the humidity coefficient ±0.02dB / 10%RH, perform a compensation operation on the probability distribution data to obtain environmental compensation data;

[0155] Compare the environmental compensation data with the 1kHz / -20dBFS reference signal, calculate the system error parameter to obtain calibration reference data, and construct a 12 - order FIR filter according to the calibration reference data to perform a non - linear error compensation operation to obtain compensated calibration data;

[0156] Based on the compensated calibration data, calibrate and correct the segmented frequency response data, calculate the correction coefficients for the low - frequency band, mid - frequency band, and high - frequency band respectively to obtain band - calibrated data, and perform data integration processing on the band - calibrated data to obtain a standard frequency response data set.

[0157] Specifically, divide the frequency range from 20Hz to 20kHz into three different frequency bands, where 20Hz to 200Hz is defined as the low - frequency band, 200Hz to 2kHz is defined as the mid - frequency band, and 2kHz to 20kHz is defined as the high - frequency band. Through the division, analyze the frequency response characteristics of the low - frequency, mid - frequency, and high - frequency bands respectively to obtain segmented frequency response data. Based on the segmented frequency response data, construct an adaptive interval model, set the confidence level to 95%, and perform the maximum deviation calculation to determine the response range within each frequency band. The construction of the adaptive interval model is based on the statistical characteristics of the signal. Calculate the mean and standard deviation of each frequency band for constructing the confidence interval. Taking the low - frequency band as an example, let its mean be and the standard deviation be , then at the 95% confidence level, its confidence interval is expressed as:

[0158] ;

[0159] Among them, Indicates the confidence interval for the low frequency band, and 1.96 is the coefficient corresponding to the 95% confidence level. Using the same method, the confidence intervals for the mid-frequency and high-frequency bands are calculated to obtain the interval boundary data for all frequency bands. To describe the characteristics of these interval boundaries, a mixture probability model is performed, combining the Gaussian distribution and the uniform distribution. Among them, the weight of the Gaussian distribution is set to 0.7, and the weight of the uniform distribution is set to 0.3 to reflect the concentration and randomness in the signal characteristics. The mixture probability density function is expressed as:

[0160] ;

[0161] where, represents a Gaussian distribution with a mean of and a standard deviation of , represents a uniform distribution on the interval . Through this mixture model, more accurate probability distribution data is obtained. Environmental compensation operations are performed on the probability distribution data according to the temperature coefficient and humidity coefficient. Environmental conditions, such as temperature and humidity, will affect the characteristics of the frequency response. An environmental compensation matrix is constructed according to the temperature coefficient of ±0.01 dB / °C and the humidity coefficient of ±0.02 dB / 10%RH to correct the signal. The compensation operation is expressed as:

[0162] ;

[0163] where, is the frequency response data after compensation, and are the change amounts of temperature and humidity respectively, and are the compensation coefficients of temperature and humidity respectively. By performing compensation operations on the signal, the obtained environmental compensation data can better reflect the system performance in the real environment. The environmental compensation data is compared with the reference signal of 1 kHz / -20 dBFS to calculate the error parameters of the system and obtain the calibration reference data. The reference signal of 1 kHz / -20 dBFS is a commonly used reference signal in audio testing. By comparing with the compensation data, the deviation of the system under different conditions is calculated. The system error is expressed as:

[0164] ;

[0165] where, is the error value, is the frequency response characteristic of the reference signal. Based on the error parameters, a 12th-order finite impulse response filter (FIR) is constructed to perform non-linear error compensation. The design goal of the FIR filter is to make the frequency response characteristic of the system as close as possible to the ideal response. The impulse response of the filter is expressed as:

[0166] ;

[0167] Among them, is the impulse response of the filter, is the filter coefficient, is the unit impulse function. The error signal is compensated by this filter to obtain compensated calibration data, making the response of the system closer to the target characteristics. Based on the compensated calibration data, the segmented frequency response data is calibrated and corrected, and the correction coefficients for the low-frequency band, middle-frequency band, and high-frequency band are calculated respectively. The correction coefficients for the low-frequency band, middle-frequency band, and high-frequency band are expressed by the following formulas:

[0168] ;

[0169] Among them, is the correction coefficient for the frequency band , is the error data for the frequency band , is the original response data for the frequency band . By calculating the correction coefficient, the frequency response characteristics within the frequency band are effectively adjusted, making the response of each frequency band meet the expected standard. The frequency band calibration data is subjected to data integration processing to obtain a standard frequency response data set.

[0170] In a specific embodiment, the process of executing step 600 may specifically include the following steps:

[0171] Input the standard frequency response data set into the feature extraction network. The feature extraction network includes an 8-layer convolutional neural network, and each layer uses a ReLU activation function and a batch normalization layer to obtain a multi-scale feature map;

[0172] Perform feature pyramid operations on the multi-scale feature map to construct 4 feature layers with different resolutions. The downsampling ratios of each feature layer are 1, 1 / 2, 1 / 4, and 1 / 8 respectively to obtain multi-level feature data;

[0173] Input the multi-level feature data into the dual-stream attention network. The dual-stream attention network includes a frequency channel attention module and a temporal-spatial attention module. Among them, the frequency channel attention module uses a 1×1 convolutional layer to extract channel weights, and the temporal-spatial attention module uses a 5×5 convolutional layer to extract position weights to obtain attention feature data;

[0174] Perform feature aggregation on the attention feature data. Through the cross-layer feature fusion module, the feature maps with different resolutions are weighted and fused, and the weight coefficients are learned through a 1×1 convolutional layer to obtain fused feature data;

[0175] Input the fused feature data into the classification branch network. The classification branch network contains 3 fully connected layers with the number of nodes being 512, 256, and 128 respectively. A Dropout layer is connected after each layer, and the dropout rate is set to 0.5 to obtain the frequency classification result;

[0176] Perform regression branch processing on the fused feature data, extract deep features through 3 residual blocks. Each residual block contains two 3×3 convolutional layers and a shortcut connection to obtain the regression feature data;

[0177] Input the regression feature data into the dual-head output layer. The first output head predicts the frequency response value through two fully connected layers, and the second output head predicts the phase response value through two fully connected layers to obtain the response prediction data;

[0178] Perform prediction point continuity processing on the frequency classification result and the response prediction data through the spline interpolation algorithm to obtain the frequency response curve and the phase characteristic curve.

[0179] Specifically, input the standard frequency response dataset into the feature extraction network and design a feature extraction module containing an 8-layer convolutional neural network. These convolutional layers are used to gradually extract the high-level features of the signal and help the network learn more complex representations layer by layer. Each convolutional layer applies the ReLU activation function to introduce non-linearity and enable the network to process complex patterns. The formula of the ReLU activation function is:

[0180] ;

[0181] Among them, is the activated output, is the input of the convolutional layer. Through this operation, the problem of gradient disappearance is effectively prevented and the expression ability of the network is improved. Each convolutional layer includes a batch normalization layer, which is used to normalize the data of each mini-batch to make the training of the network more stable. After 8 layers of convolution and batch normalization processing, a multi-scale feature map is obtained, which contains frequency response features at different levels. Perform feature pyramid operation on the multi-scale feature map. The feature pyramid is used to construct feature layers with different resolutions to capture the characteristics of the signal at different scales. Construct 4 feature layers with different resolutions, and the downsampling ratios of each feature layer are 1, 1 / 2, 1 / 4, and 1 / 8 respectively. Downsampling is achieved through pooling operations or strided convolutions, gradually reducing the resolution while retaining the global information to facilitate the fusion and processing of features at different scales. Let the input feature map be , the -th layer feature layer of the feature pyramid is , then the downsampling ratio is:

[0182] ;

[0183] Among them, Indicates a pooling operation Indicates the downsampling ratio, which are 1, 1 / 2, 1 / 4, and 1 / 8 respectively. After the feature pyramid operation, multi-level feature data is obtained, which respectively reflect the frequency response characteristics at different scales. The multi-level feature data is input into the dual-stream attention network, which consists of a frequency channel attention module and a temporal-spatial attention module. Among them, the frequency channel attention module uses a 1×1 convolutional layer to extract the weights of the channels. The role of the 1×1 convolution is to emphasize important frequency band features by adjusting the weights of each channel. Let the input feature be , and the output of the channel attention be , then:

[0184] ;

[0185] Among them, Indicates Convolution operation is the Sigmoid activation function, which is used to map the output to the range (0, 1) to adjust the weights of each channel. The temporal-spatial attention module uses a 5×5 convolutional layer to extract the weights of the positions, capturing the important characteristics of the signal in time and space. Let its output be , then:

[0186] ;

[0187] The output of the dual-stream attention network is the attention feature data. Through this mechanism, the attention to specific frequency bands and time points is enhanced, making the feature extraction more effective. Feature aggregation is performed on the attention feature data, and the feature maps with different resolutions are weighted and fused through a cross-layer feature fusion module. During the fusion process, the weight coefficients are learned through Convolutional layer to ensure the best effect of the fusion between features of different scales. Let the feature map be , then the fused feature data Is expressed as:

[0188] ;

[0189] Among them, Is the weight coefficient learned through the 1×1 convolutional layer, Is the Layer feature maps. Through weighted fusion, the feature information of multiple layers is combined to obtain more representative fused feature data. The fused feature data is input into the classification branch network. The classification branch network contains 3 fully connected layers with the number of nodes being 512, 256, and 128 respectively. A Dropout layer is connected after each fully connected layer, and the dropout rate is set to 0.5. The purpose of Dropout is to randomly discard half of the nodes during the training process to prevent overfitting of the network. The formula is:

[0190] ;

[0191] Among them, is the output after Dropout, is the output of the fully connected layer, is the dropout rate, set to 0.5. After being processed by the classification branch network, the result of frequency classification is obtained, which is used to identify which frequency response pattern the input signal belongs to. At the same time, the fused feature data is subjected to regression branch processing to extract deep features through 3 residual blocks. Each residual block contains two convolutional layers and a shortcut connection. The introduction of the residual block helps to solve the problem of gradient disappearance in the training of deep networks and enables information to be transmitted more smoothly in the deep network. The output of the residual block is expressed as:

[0192] ;

[0193] Among them, represents the transformation through two 3×3 convolutional layers, is the input signal, is the output of the residual block. Through 3 residual blocks, the deep features of the signal are extracted to obtain regression feature data. The regression feature data is input into a dual-head output layer. The first output head predicts the frequency response value through two fully connected layers, and the second output head predicts the phase response value through two fully connected layers. Let the input feature be , then the frequency response value and the phase response value are respectively expressed as:

[0194] ;

[0195] ;

[0196] Among them, and are weight matrices, and are biases, and Activation functions for predicting frequency response values and phase response values respectively. Through the dual-head output layer, prediction data for frequency response and phase response are obtained. The spline interpolation algorithm is used to continuousize the prediction points for the frequency classification results and response prediction data to obtain the frequency response curve and phase characteristic curve. Spline interpolation is a smoothing method used to connect discrete prediction points into a continuous curve. Let the discrete points be , then the goal of spline interpolation is to construct a smooth function that satisfies:

[0197] ;

[0198] Through spline interpolation, a smooth curve is generated to describe the frequency response and phase characteristics, making the final output more in line with the actual frequency response characteristics.

[0199] The above described the frequency response analysis method of the intelligent audio test system in the embodiment of the present application. Next, the frequency response analysis device 10 of the intelligent audio test system in the embodiment of the present application will be described. Please refer to Figure 2 , an embodiment of the frequency response analysis device 10 of the intelligent audio test system in the embodiment of the present application includes:

[0200] Initialization module 11, used to initialize the audio parameters of the intelligent audio test system to obtain a test parameter matrix and an excitation signal;

[0201] Acquisition module 12, used to perform dual-channel acquisition and time-frequency feature extraction on the excitation signal and response signal based on the test parameter matrix to obtain a time-frequency feature dataset;

[0202] Processing module 13, used to perform dual-branch processing of time-domain noise reduction and frequency-domain enhancement on the time-frequency feature dataset to obtain a purified feature dataset and a feature weighting matrix;

[0203] Calculation module 14, used to perform segmented recursive calculation and feature aggregation on the purified feature dataset according to the feature weighting matrix to obtain a frequency response feature vector;

[0204] Calibration module 15, used to decompose the frequency response feature vector according to the test frequency band and perform environmental compensation and nonlinear calibration to obtain a standard frequency response dataset;

[0205] Classification module 16, used to input the standard frequency response dataset into a deep neural network for feature recognition and pattern classification to obtain a frequency response curve and a phase characteristic curve.

[0206] Through the collaborative cooperation of the above-mentioned various components, by introducing a dual-branch data processing architecture, the collaborative enhancement of time-domain and frequency-domain signals is achieved, effectively improving the anti-noise ability of the system; adopting a multi-level feature extraction and recursive analysis strategy significantly improves the extraction accuracy of frequency response features; innovatively designing an environmental compensation and non-linear calibration mechanism effectively eliminates the interference effects of environmental factors; introducing a deep neural network for feature recognition and pattern classification realizes the intelligence and automation of the testing process. This method ensures the accuracy of full-band testing through segmented recursive calculation and feature aggregation technology; adopts a dual-stream attention network and a residual learning mechanism to enhance the system's ability to capture complex frequency response features; ensures the reliability of test results through a multi-level calibration and compensation mechanism; and finally realizes the reconstruction of high-precision frequency response curves through the intelligent analysis of a deep learning model. The present invention has achieved remarkable improvements in terms of testing accuracy, anti-interference ability, automation level, and adaptability.

[0207] Please refer to Figure 3 , Figure 3 which is a schematic block diagram of the structure of the electronic device 300 provided by an embodiment of the present application. The electronic device 300 includes a processor 301 and a memory 302. The processor 301 and the memory 302 are connected through a device bus 303. Among them, the memory 302 may include a non-volatile storage medium and an internal memory.

[0208] The non-volatile storage medium can store a computer program. The computer program includes program instructions. When the program instructions are executed by the processor 301, the processor 301 can be enabled to execute any of the above-mentioned frequency response analysis methods of the intelligent audio testing system.

[0209] The processor 301 is used to provide computing and control capabilities to support the operation of the entire electronic device 300.

[0210] The internal memory provides an environment for the operation of the computer program in the non-volatile storage medium. When the computer program is executed by the processor 301, the processor 301 can be enabled to execute any of the above-mentioned frequency response analysis methods of the intelligent audio testing system.

[0211] Those skilled in the art can understand that Figure 3 the structure shown in

[0212] It should be understood that the processor 301 may be a central processing unit (CPU), and the processor 301 may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0213] It should be noted that those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working process of the above-described electronic device 300 can refer to the corresponding process of the foregoing intelligent audio test system frequency response analysis method, which will not be elaborated herein.

[0214] The embodiment of the present application further provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by one or more processors, the one or more processors are caused to implement the intelligent audio test system frequency response analysis method provided by the embodiment of the present application.

[0215] Among them, the computer-readable storage medium may be an internal storage unit of the foregoing embodiment of the electronic device 300, such as the hard disk or memory of the electronic device 300. The computer-readable storage medium may also be an external storage device of the electronic device 300, such as a plug-in hard disk equipped with the electronic device 300, a smart media card (SMC), a secure digital (SD) card, a flash card, etc.

[0216] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described system, device, and unit can refer to the corresponding processes in the foregoing method embodiments, which will not be elaborated herein.

[0217] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing an electronic device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.

[0218] As described above, the above embodiments are only used to illustrate the technical solutions of this application and are not intended to limit them; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of various embodiments of this application.

Claims

1. A frequency response analysis method for an intelligent audio test system, characterized in that: The method comprises: Initialize audio parameters of the intelligent audio test system to obtain a test parameter matrix and an excitation signal; Based on the test parameter matrix, dual-channel acquisition and time-frequency feature extraction are performed on the excitation signal and the response signal to obtain a time-frequency feature data set; Performing dual-branch processing of time domain noise reduction and frequency domain enhancement on the time-frequency feature data set to obtain a purified feature data set and a feature weighting matrix; Performing segmented recursive calculation and feature aggregation on the purified feature data set according to the feature weighting matrix to obtain a frequency response feature vector; Decomposing the frequency response characteristic vector according to the test frequency band, and performing environmental compensation and nonlinear calibration to obtain a standard frequency response data set; The standard frequency response data set is input into a deep neural network for feature recognition and pattern classification to obtain a frequency response curve and a phase characteristic curve.

2. The frequency response analysis method of the intelligent audio test system according to claim 1, characterized in that: The method of initializing audio parameters of the intelligent audio test system to obtain a test parameter matrix and an excitation signal includes: Perform DSP algorithm configuration on the processor of the intelligent audio test system, load the all-pass filter and dynamic equalizer DEQ into the intelligent audio test system, and obtain DSP initialization parameters; The system sampling parameters are set based on the DSP initialization parameters to obtain basic sampling parameters, and the signal ranges of 14 digital gain channels are set according to the basic sampling parameters to obtain a channel parameter set; Performing gain step setting on the channel parameter set to obtain a gain control matrix, verifying the system dynamic range according to the gain control matrix, performing background noise index measurement, and obtaining dynamic range verification data; According to the dynamic range verification data, the signal-to-noise ratio index is measured and calculated to obtain system performance parameters, and the system performance parameters and the gain control matrix are combined to obtain a test parameter matrix. Based on the test parameter matrix, a test signal is generated to obtain an excitation signal.

3. The frequency response analysis method of the intelligent audio test system according to claim 2, characterized in that: The step of performing dual-channel acquisition and time-frequency feature extraction on the excitation signal and the response signal based on the test parameter matrix to obtain a time-frequency feature data set includes: Performing 24-bit analog-to-digital conversion on the excitation signal and the response signal, setting the number of sampling points to 2048 points / frame to obtain original sampling data, and performing 14-channel signal allocation on the original sampling data based on the test parameter matrix to obtain channel sampling data; Performing 1 dB step gain adjustment on the channel sampling data, performing signal amplitude control through a processor, and obtaining gain control data; Input the gain control data into an all-pass filter for signal preprocessing, execute dynamic equalization DEQ algorithm operation to obtain filtered processing data, and perform real-time Fourier transform on the filtered processing data to obtain frequency domain feature data; Performing phase characteristic extraction on the frequency domain feature data, calculating group delay and phase response to obtain phase feature data, and performing information fusion operation on the frequency domain feature data and the phase feature data, performing feature alignment processing to obtain fused feature data; The fused feature data is standardized to obtain a time-frequency feature data set.

4. The frequency response analysis method of the intelligent audio test system according to claim 3, characterized in that: The step of performing dual-branch processing of time domain noise reduction and frequency domain enhancement on the time-frequency feature data set to obtain a purified feature data set and a feature weighting matrix includes: Performing branch division on the time-frequency feature data set, dividing the signal into a time domain processing branch and a frequency domain processing branch, to obtain a dual-branch data set; Performing a 4-level wavelet decomposition operation on the time domain branch in the dual-branch data set, setting the threshold to 3 times the noise standard deviation, to obtain time domain denoised data; Performing a spectral subtraction operation on the frequency domain branch in the dual-branch data set, setting the frequency-related suppression factor to 0.8, and obtaining frequency domain enhanced data; Calculate the signal-to-noise ratio index according to the time domain noise reduction data, set the signal-to-noise ratio threshold to 40 dB, obtain the time domain weight coefficient, and calculate the spectrum flatness parameter based on the frequency domain enhancement data, set the flatness deviation range to ±3 dB, and obtain the frequency domain weight coefficient; Performing a weighted combination on the time domain weight coefficient and the frequency domain weight coefficient, performing a weight normalization operation to obtain a feature weight matrix, and performing feature fusion on the time domain denoised data and the frequency domain enhanced data according to the feature weight matrix to obtain a fused feature set; Based on frequency response flatness, phase linearity and group delay fluctuation indicators, a feature screening operation is performed on the fused feature set to obtain a purified feature data set.

5. The frequency response analysis method of the intelligent audio test system according to claim 4, characterized in that: The step of performing segmented recursive calculation and feature aggregation on the purified feature data set according to the feature weighting matrix to obtain a frequency response feature vector includes: Performing frequency band division on the purified feature data set, dividing the range of 20 Hz-20 kHz into 128 frequency bands, performing a bandpass filtering operation on each frequency band, and obtaining frequency band division data; The frequency band division data is encoded and converted according to the feature weighting matrix to construct a 16-dimensional frequency encoding vector, in which the amplitude occupies 8 dimensions, the phase occupies 4 dimensions, and the time occupies 4 dimensions, to obtain frequency encoding data; Performing recursive sampling processing on the frequency-encoded data, setting the basic sampling density to 4 points / octave, and setting the sampling density to 16 points / octave in the frequency response mutation region, to obtain multi-density sampling data; The multi-density sampling data is input into a three-level feature extraction unit, the first layer performs a time domain convolution operation to extract transient features, the second layer performs a frequency domain transformation to extract harmonic features, and the third layer performs a correlation analysis to extract associated features, thereby obtaining three-dimensional feature data; A two-layer prediction unit is constructed for the three-dimensional feature data, wherein the first layer processes the time series relationship through an LSTM network to obtain a time series prediction value, and the second layer obtains a frequency response prediction value by fusing frequency domain information through a feedforward network to obtain prediction feature data; Based on the predicted feature data, a three-level residual network is constructed, wherein the first level calculates the amplitude residual, the second level calculates the phase residual, and the third level calculates the delay residual to obtain residual feature data; The residual feature data is input into the feature fusion network, and feature dimension reduction and optimization are performed through the fully connected layer. The number of output layer nodes is set to 32 to obtain optimized feature data, and normalization is performed on the optimized feature data. The three-dimensional features are compressed into a one-dimensional vector through a weighted sum operation to obtain a frequency response feature vector.

6. The frequency response analysis method of the intelligent audio test system according to claim 5, characterized in that: Decomposing the frequency response characteristic vector according to the test frequency band, and performing environmental compensation and nonlinear calibration to obtain a standard frequency response data set includes: Performing three-segment frequency division on the frequency response characteristic vector, dividing 20 Hz-200 Hz into a low frequency segment, 200 Hz-2 kHz into a medium frequency segment, and 2 kHz-20 kHz into a high frequency segment, to obtain segmented frequency response data; An adaptive interval model is constructed based on the segmented frequency response data, the confidence level is set to 95%, a maximum deviation calculation is performed to obtain interval boundary data, and mixed probability modeling is performed on the interval boundary data, the Gaussian distribution weight is set to 0.7, and the uniform distribution weight is set to 0.3 to obtain probability distribution data; An environmental compensation matrix is ​​constructed according to a temperature coefficient of ±0.01 dB / °C and a humidity coefficient of ±0.02 dB / 10% RH, and a compensation operation is performed on the probability distribution data to obtain environmental compensation data; Compare the environmental compensation data with a 1kHz / -20dBFS reference signal, calculate system error parameters, obtain calibration reference data, and construct a 12th-order FIR filter based on the calibration reference data, perform nonlinear error compensation operations, and obtain compensation calibration data; The segmented frequency response data is calibrated and corrected based on the compensation calibration data, correction coefficients of the low frequency band, the middle frequency band and the high frequency band are calculated respectively to obtain frequency band calibration data, and data integration processing is performed on the frequency band calibration data to obtain a standard frequency response data set.

7. The frequency response analysis method of the intelligent audio test system according to claim 6, characterized in that: The step of inputting the standard frequency response data set into a deep neural network for feature recognition and pattern classification to obtain a frequency response curve and a phase characteristic curve includes: Inputting the standard frequency response data set into a feature extraction network, wherein the feature extraction network comprises an 8-layer convolutional neural network, each layer uses a ReLU activation function and a batch normalization layer to obtain a multi-scale feature map; Performing feature pyramid operation on the multi-scale feature map to construct four feature layers with different resolutions, with downsampling ratios of each feature layer being 1, 1 / 2, 1 / 4 and 1 / 8 respectively, to obtain multi-level feature data; Input the multi-level feature data into a dual-stream attention network, wherein the dual-stream attention network includes a frequency channel attention module and a temporal space attention module, wherein the frequency channel attention module uses a 1×1 convolution layer to extract channel weights, and the temporal space attention module uses a 5×5 convolution layer to extract position weights, to obtain attention feature data; Performing feature aggregation on the attention feature data, weighted fusion of feature maps of different resolutions through a cross-layer feature fusion module, wherein the weight coefficient is learned through a 1×1 convolutional layer to obtain fused feature data; The fused feature data is input into a classification branch network, wherein the classification branch network comprises three fully connected layers, the number of nodes of which are 512, 256 and 128 respectively, each layer is followed by a Dropout layer, and the dropout rate is set to 0.5, to obtain a frequency classification result; Performing regression branch processing on the fused feature data, extracting deep features through three layers of residual blocks, each residual block including two 3×3 convolutional layers and a short-circuit connection, to obtain regression feature data; The regression feature data is input into a dual-head output layer, the first output head predicts a frequency response value through two fully connected layers, and the second output head predicts a phase response value through two fully connected layers to obtain response prediction data; The frequency classification result and the response prediction data are processed into prediction point continuity by using a spline interpolation algorithm to obtain a frequency response curve and a phase characteristic curve.

8. An intelligent audio test system frequency response analysis device, characterized in that: Used to execute the intelligent audio test system frequency response analysis method according to any one of claims 1 to 7, the intelligent audio test system frequency response analysis device comprises: An initialization module is used to initialize the audio parameters of the intelligent audio test system to obtain a test parameter matrix and an excitation signal; An acquisition module, used for performing dual-channel acquisition and time-frequency feature extraction on the excitation signal and the response signal based on the test parameter matrix to obtain a time-frequency feature data set; A processing module, used for performing a dual-branch process of time domain noise reduction and frequency domain enhancement on the time-frequency feature data set to obtain a purified feature data set and a feature weighting matrix; A calculation module, used for performing segmented recursive calculation and feature aggregation on the purified feature data set according to the feature weighting matrix to obtain a frequency response feature vector; A calibration module, used for decomposing the frequency response characteristic vector according to the test frequency band, and performing environmental compensation and nonlinear calibration to obtain a standard frequency response data set; The classification module is used to input the standard frequency response data set into a deep neural network for feature recognition and pattern classification to obtain a frequency response curve and a phase characteristic curve.

9. An electronic device, characterized in that: The electronic device comprises: a memory and at least one processor, wherein instructions are stored in the memory; The at least one processor calls the instruction in the memory to enable the electronic device to execute the frequency response analysis method of the intelligent audio test system according to any one of claims 1 to 7.

10. A computer-readable storage medium having instructions stored thereon, characterized in that: When the instructions are executed by the processor, the frequency response analysis method of the intelligent audio test system according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Frequency response analysis method and device, equipment and storage medium

    CN116564332A

  • Sound event detection method and system based on grouping feature calibration

    CN116778919A