Cockpit voice noise reduction function automation test method, device and equipment
The cockpit voice noise reduction function was tested using a fully automated testing method, which solved the problem of poor accuracy caused by manual testing and achieved high consistency and efficiency in the test results.
Patent Information
- Application Number
- CN202611142069.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-30
- Publication Date
- 2026-08-25
AI Technical Summary
Existing technologies rely on manual or semi-automated testing methods for cockpit voice noise reduction functions, resulting in poor accuracy of test results and susceptibility to human error and transmission distortion caused by interruptions.
A fully automated testing method is adopted. By collecting mixed audio data streams from the target vehicle's cabin, analyzing and denoising them, a comprehensive score for frequency band denoising and a defect distribution map are generated. A closed-loop process of audio acquisition, denoising, testing and analysis and result output is constructed to eliminate human error and avoid audio data distortion.
It improved the consistency and accuracy of testing, shortened the testing cycle, increased testing efficiency, and ensured the reliability of test results.
Smart Images

Figure CN122637814A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of noise reduction testing technology, and in particular to automated testing methods, devices and equipment for cockpit voice noise reduction functions. Background Technology
[0002] With the rapid development of intelligent cockpit technology, in-vehicle voice interaction systems have become the core entry point for human-machine interaction. To improve voice recognition rates in noisy driving environments, various noise reduction algorithms based on digital signal processing (such as spectral subtraction and Wiener filtering) or deep learning models are widely used. Currently, the testing and verification of cockpit voice noise reduction functions generally employs manual or semi-automated testing methods. These methods are susceptible to the subjective operation of testers and transmission distortion due to interruptions, resulting in poor accuracy of test results.
[0003] The above content is only used to help understand the technical solution of the present invention and does not represent an admission that the above content is prior art. Summary of the Invention
[0004] The main objective of this application is to provide an automated testing method, apparatus, and equipment for cockpit voice noise reduction function, aiming to solve the technical problem of poor accuracy of test results in traditional cockpit voice noise reduction function testing methods in the prior art.
[0005] Firstly, this application provides an automated testing method for cockpit voice noise reduction function, the method comprising: The system collects a mixed audio data stream from the target vehicle's cockpit, parses the mixed audio data stream, and obtains the noise audio data. The noisy audio data is denoised to obtain denoised audio data, and the denoised audio data is then converted into test audio data in the target playback format. Based on the differences between the test audio data and the reference audio data in different target frequency bands, the comprehensive score of frequency band noise reduction and the defect distribution map of the test audio data are determined. Based on the comprehensive score of frequency band noise reduction and the defect distribution map of the test audio data, the test results of the speech noise reduction function are generated.
[0006] Compared with related technologies, this automated testing method for cockpit voice noise reduction has at least the following advantages: it integrates discrete test steps into a continuous automated flow, constructs a fully automated test closed loop of audio acquisition, noise reduction processing, test analysis, and result output, utilizes standardized automated processes to minimize differences caused by human operation, and largely avoids distortion of audio data during transmission or conversion. This improves test consistency and evaluation accuracy, ensures the accuracy of test results, and shortens the test cycle, thereby increasing test efficiency.
[0007] In some possible implementations of the first aspect, the steps of denoising the noisy audio data to obtain denoised audio data include: Extracting time-frequency features from noisy audio data; Based on time-frequency characteristics, the noise type of the noisy audio data is determined; Based on the noise type, determine the matching noise reduction strategy for the noisy audio data; Based on the matching noise reduction strategy, the noisy audio data is processed to obtain the corresponding noise-reduced audio data.
[0008] Among some possible implementations of the first aspect, a fully automated test closed loop is constructed, which includes audio acquisition, noise reduction processing, test analysis, and result output. During noise reduction processing, the optimal noise reduction strategy is adaptively matched according to the noise type, thereby enabling targeted noise reduction processing of audio data, ensuring the effectiveness of noise reduction processing, and improving the accuracy of evaluation during test analysis to obtain accurate test results.
[0009] In some possible implementations of the first aspect, the step of determining the noise type of the noise audio data based on its time-frequency characteristics includes: When the time-frequency characteristics meet the steady-state conditions, the noise type of the noise audio data is determined to be steady-state noise; When the time-frequency characteristics do not meet the steady-state conditions and the energy proportion of the noise audio data in the preset frequency band is less than or equal to the preset proportion, the noise type of the noise audio data is determined to be non-steady-state noise. When the time-frequency characteristics do not meet the steady-state conditions and the energy proportion of the noise audio data in the preset frequency band is greater than the preset proportion, the noise type of the noise audio data is determined to be mixed noise.
[0010] In some possible implementations of the first aspect, time-frequency features are extracted from the noisy audio data, and the noise type of the current noisy audio data is accurately determined using the time-frequency features. This allows for adaptive matching of the optimal noise reduction strategy according to the noise type during noise reduction processing, enabling targeted noise reduction of the audio data to ensure the effectiveness of the noise reduction process.
[0011] Among some possible implementations of the first aspect, the time-frequency features include energy variance and inter-frame change rate, and the methods also include: When the energy variance of the noisy audio data is less than or equal to the energy stationary threshold and the inter-frame change rate is less than or equal to the feature mutation threshold, the time-frequency characteristics of the noisy audio data are determined to meet the steady-state condition. When the energy variance of the noisy audio data is greater than the energy stationary threshold or the inter-frame change rate is greater than the feature mutation threshold, the determination of the time-frequency features does not meet the steady-state condition.
[0012] In some possible implementations of the first aspect, energy variance and inter-frame change rate are used as time-frequency features. By using time-frequency features, the noise type of the current noisy audio data can be accurately determined. Thus, during the noise reduction process, the optimal noise reduction strategy can be adaptively matched according to the noise type to perform targeted noise reduction on the audio data, thereby ensuring the effectiveness of the noise reduction process.
[0013] In some possible implementations of the first aspect, the steps for determining the overall frequency band noise reduction score and defect distribution map of the test audio data based on the differences between the test audio data and the reference audio data in different target frequency bands include: Based on the frequency range of the test audio data and the reference audio data, multiple consecutive target frequency bands are determined in the frequency domain, and the difference between the test audio data and the reference audio data in each target frequency band is calculated. Based on the difference between the test audio data and the reference audio data in each target frequency band, a difference distribution map is generated, and defect points are identified in the difference distribution map. Based on the differential distribution map and the time period evaluation results of the defect points in the differential distribution map, a defect distribution map of the test audio data is generated. Based on the difference values corresponding to each target frequency band and the evaluation weight of each target frequency band, the comprehensive score of frequency band noise reduction for the test audio data is determined.
[0014] Among some possible implementations of the first aspect, a fully automated test closed loop is constructed, which includes audio acquisition, noise reduction processing, test analysis, and result output. During test analysis, the differences between the test audio data and the reference audio data are evaluated by frequency band, and the defective sites are identified. Then, a refined evaluation is performed by time period to obtain a specific defect distribution map. Through adaptive weighting, a comprehensive score of the noise reduction effect is calculated to improve the evaluation accuracy and obtain accurate test results.
[0015] In some possible implementations of the first aspect, the step of generating a defect distribution map of the test audio data based on the difference distribution map and the time-period evaluation results of defect points in the difference distribution map further includes: Based on the time period type corresponding to the defect point, determine the differentiated evaluation index for the defect point. Defects are evaluated based on differentiated evaluation indicators to obtain time-period evaluation results.
[0016] In some possible implementations of the first aspect, the difference between the test audio data and the reference audio data is evaluated by frequency band, and obvious defective sites are identified. Then, a more refined evaluation is performed in different time periods to improve the accuracy of the evaluation and obtain accurate test results.
[0017] In some possible implementations of the first aspect, the steps for determining the differentiated evaluation indicators for defect points based on the time period type corresponding to the defect points include: When the time period corresponding to the defect point is a silent period, the noise suppression ratio is determined based on the original average noise energy and residual noise energy of the test audio data, and the noise suppression ratio is used as the differential evaluation index of the defect point. When the time period corresponding to the defect point is a speech time period, the speech distortion is determined based on the instantaneous error energy of the test audio data, and the speech distortion is used as the differential evaluation index of the defect point. When the time period corresponding to the defect point is a transition period, the envelope delay difference and the relative attenuation of the envelope slope of the test audio data are used as the differential evaluation indicators of the defect point.
[0018] In some possible implementations of the first aspect, the difference between the test audio data and the reference audio data is evaluated by frequency band, and obvious defective sites are identified. Then, a refined evaluation is carried out in different time periods to comprehensively determine the defect distribution map and noise reduction effect, so as to ensure the accuracy of the evaluation and improve the accuracy of the test results.
[0019] In some possible implementations of the first aspect, the step of determining the overall score of frequency band noise reduction for the test audio data based on the difference values corresponding to each target frequency band and the evaluation weight of each target frequency band further includes: Calculate the background noise energy of each target frequency band and the total background noise energy in the frequency domain; Obtain the correspondence between the background noise energy of the target frequency band, the total background noise energy in the frequency domain, the frequency band perception correction coefficient, and the evaluation weight of the target frequency band; Based on the correspondence, the total background noise energy in the frequency domain, and the background noise energy and frequency band perception correction coefficient of each target frequency band, the evaluation weight of each target frequency band is obtained.
[0020] In some possible implementations of the first aspect, the evaluation weight depends on the relative energy proportion of background noise in each frequency band. Using the evaluation weight, the difference values are aggregated in a time-frequency two-dimensional weighted manner, and the score of the noise reduction effect is calculated in a comprehensive manner to ensure the accuracy of the evaluation and improve the accuracy of the test results.
[0021] In some possible implementations of the first aspect, the steps for calculating the differences between the test audio data and the reference audio data in each target frequency band include: Based on the power spectral density of the test audio data, the logarithmic power spectral density of the test audio data is determined, and based on the power spectral density of the reference audio data, the logarithmic power spectral density of the reference audio data is determined. Based on the log power spectral density of the test audio data and the log power spectral density of the reference audio data, calculate the log power spectral distance between the test audio data and the reference audio data in each target frequency band. The logarithmic power spectrum distance corresponding to each target frequency band is used as the difference value corresponding to each target frequency band.
[0022] In some possible implementations of the first aspect, during test analysis, the difference between the test audio data and the reference audio data is evaluated by frequency band, and the logarithmic power spectrum distance is used as the difference value to find the defective sites. Then, a refined evaluation is carried out in different time periods to comprehensively determine the defect distribution map and noise reduction effect, so as to improve the evaluation accuracy and obtain accurate test results.
[0023] In some possible implementations of the first aspect, the step of determining the overall frequency band noise reduction score and defect distribution map of the test audio data based on the difference between the test audio data and the reference audio data in different target frequency bands also includes: Extract the feature vectors from the test audio data and the reference audio data respectively; Based on the distance between the feature vectors of the test audio data and the feature vectors of the reference audio data, a nonlinear mapping relationship between the test audio data and the reference audio data is determined. The time delay difference of the test audio data is determined based on the nonlinear mapping relationship; Based on the time delay difference of the test audio data, the test audio data is time-aligned with the reference audio data.
[0024] In some possible implementations of the first aspect, dynamic time warping is used to calculate the delay, and the audio under test is aligned with the reference signal at the sub-millisecond level. This ensures that the difference between the test audio data and the reference audio data can be accurately evaluated during subsequent test analysis, thereby ensuring the accuracy of the evaluation and improving the accuracy of the test results.
[0025] In some possible implementations of the first aspect, the steps of parsing the mixed audio data stream to obtain noisy audio data include: Based on the header signature of each audio data stream in the mixed audio data stream, the encapsulation format of each audio data stream in the mixed audio data stream is determined; Based on the encapsulation format of each audio data stream in the mixed audio data stream, the mixed audio data stream is decoded into the original noisy audio data; The original noise audio data is standardized to obtain the noise audio data.
[0026] Among some possible implementations of the first aspect, a fully automated test closed loop is constructed, encompassing audio acquisition, noise reduction processing, test analysis, and result output. During audio acquisition, the encapsulation format is automatically determined by reading the header feature code of the data stream, and decapsulation and decoding are automatically performed. Standardized processing is then carried out to provide a standardized data foundation for subsequent noise reduction processing and test analysis, thereby improving the consistency and accuracy of the test.
[0027] Secondly, this application provides an automated testing device for cockpit voice noise reduction function, the automated testing device for cockpit voice noise reduction function includes: The acquisition and parsing module is used to acquire the mixed audio data stream from the target vehicle's cockpit, parse the mixed audio data stream, and obtain the noise audio data. The noise reduction and synthesis module is used to process noisy audio data to obtain noise-reduced frequency data, and convert the noise-reduced frequency data into test audio data in the target playback format; The functional testing module is used to determine the overall frequency band noise reduction score and defect distribution map of the test audio data based on the difference between the test audio data and the reference audio data in different target frequency bands. The functional testing module is also used to generate test results for the speech noise reduction function based on the comprehensive score of frequency band noise reduction and the defect distribution map of the test audio data.
[0028] Compared with related technologies, this automated testing device for cockpit voice noise reduction has at least the following advantages: it integrates discrete testing steps into a continuous automated flow, constructs a fully automated testing closed loop of audio acquisition, noise reduction processing, test analysis, and result output, utilizes standardized automated processes to minimize differences caused by human operation, and largely avoids distortion of audio data during transmission or conversion, thereby improving the consistency and accuracy of testing, while also shortening the testing cycle and increasing testing efficiency.
[0029] Thirdly, this application provides an automated testing device for cockpit voice noise reduction function, which includes: a memory, a processor, and a computer program stored in the memory and executable on the processor. The computer program is configured to implement the steps of the automated testing method for cockpit voice noise reduction function as described above.
[0030] Fourthly, this application provides a storage medium, which is a computer-readable storage medium, and stores a computer program on the storage medium. When the computer program is executed by a processor, it implements the steps of the automated testing method for cockpit voice noise reduction function as described above.
[0031] Fifthly, this application provides a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the automated testing method for cockpit voice noise reduction function as described above. Attached Figure Description
[0032] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0033] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0034] Figure 1 This is a flowchart illustrating an embodiment of the automated testing method for cockpit voice noise reduction function of this application. Figure 2 A schematic diagram illustrating the acquisition and parsing process of an automated testing method for cockpit voice noise reduction function provided in an embodiment of this application; Figure 3 A schematic diagram of the noise reduction processing flow of an automated testing method for cockpit voice noise reduction function provided in an embodiment of this application; Figure 4 A schematic diagram of a new audio synthesis process for an automated testing method for cockpit voice noise reduction function provided in an embodiment of this application; Figure 5 A schematic diagram of the test evaluation process for an automated test method for cockpit voice noise reduction function provided in an embodiment of this application; Figure 6 This is a flowchart illustrating another embodiment of the automated testing method for cockpit voice noise reduction function of this application; Figure 7 A simplified flowchart illustrating an embodiment of the automated testing method for cockpit voice noise reduction function provided in this application; Figure 8 This is a schematic diagram of the module structure of the automated testing device for cockpit voice noise reduction function according to an embodiment of this application; Figure 9 This is a schematic diagram of the equipment structure of the hardware operating environment involved in the automated testing method for cockpit voice noise reduction function in the embodiments of this application.
[0035] The realization of the purpose, functional features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0036] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.
[0037] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.
[0038] Currently, the testing and verification of cockpit voice noise reduction functions generally employs manual or semi-automated testing methods. For example, testers typically play pre-recorded standard noisy speech through speakers in the cockpit while simultaneously controlling the vehicle's infotainment system manually or via scripts to record audio data including ambient noise. The collected audio is then exported and compared using offline analysis software or through manual listening to determine indicators such as signal-to-noise ratio and speech clarity before and after noise reduction. This method is susceptible to subjective operation by the tester and transmission distortion due to interruptions, resulting in relatively poor accuracy of the test results.
[0039] This application provides a feasible solution that involves acquiring a mixed audio data stream from the target vehicle's cockpit, parsing the mixed audio data stream to obtain noisy audio data, performing noise reduction processing on the noisy audio data to obtain noise-reduced frequency data, and converting the noise-reduced frequency data into test audio data in the target playback format. Based on the differences between the test audio data and reference audio data in different target frequency bands, the comprehensive noise reduction score and defect distribution map of the test audio data are determined. Based on the comprehensive noise reduction score and defect distribution map of the test audio data, test results for the speech noise reduction function are generated. This integrates discrete test steps into a continuous automated flow, constructing a fully automated test closed loop encompassing audio acquisition, noise reduction processing, test analysis, and result output. Utilizing standardized automated processes, it minimizes differences caused by human operation and largely avoids distortion of audio data during transmission or conversion, improving test consistency and accuracy, while also shortening the test cycle and increasing test efficiency.
[0040] The executing entity in this embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, or mobile phone; or an electronic device capable of performing the above functions, such as an automated testing device for cockpit voice noise reduction. This embodiment does not specifically limit the specific implementation. The following description uses an automated testing device for cockpit voice noise reduction as an example to illustrate this embodiment and the subsequent embodiments.
[0041] This application provides an automated testing method for cockpit voice noise reduction function, referring to... Figure 1 , Figure 1 This is a flowchart illustrating a feasible embodiment of the automated testing method for cockpit voice noise reduction function of this application.
[0042] In this embodiment, the automated testing method for cockpit voice noise reduction function includes steps S10~S40: Step S10: Collect the mixed audio data stream of the target vehicle's cabin, parse the mixed audio data stream, and obtain the noise audio data; The target vehicle is the vehicle for which the cockpit speech noise reduction function test is currently being conducted. The mixed audio data stream refers to audio data streams of different formats collected from different target data sources. Target data sources are data sources capable of generating noisy audio, such as: Android system data sources, Bluetooth data sources, DSP (Digital Signal Processing) data sources, and bus data sources. Specifically, the Android system data source can be an Android multimedia channel, the Bluetooth data source can be a Bluetooth call channel or a Bluetooth music transmission channel, the DSP data source can be a DSP noise reduction output channel, and the bus data source can be an in-vehicle Ethernet forwarding channel. This embodiment does not specifically limit these.
[0043] In some feasible implementations, the step of acquiring a mixed audio data stream from the target vehicle cabin includes: acquiring the target data stream transmitted from the target data source in the vehicle cabin through an acquisition probe of the target data source, and generating a mixed audio data stream.
[0044] The target data stream is the audio data stream with noise generated in the target data source. The format of the target data stream is usually different in different target data sources.
[0045] In this embodiment, a corresponding acquisition probe is deployed for each target data source. The acquisition probes are used to acquire the target data stream of each target data source. The types of acquisition probes deployed for different target data sources are different. The data acquired by all acquisition probes are encapsulated into a unified data frame object, which includes the original binary payload and timestamp.
[0046] For example, for Android system data sources, an ADB (Android Debug Bridge) probe is deployed to directly capture the mixing data from the AudioFlinger layer or the raw PCM (Pulse Code Modulation) stream from the HAL (Hardware Abstraction Layer). For Bluetooth data sources, an HCI (Bluetooth Host Controller Interface) probe is deployed to capture the A2DP (Advanced Audio Distribution Profile) or HFP (Hands-Free Profile) transport streams by listening to HCI data packets. For DSP and bus data sources, hardware probes are deployed to capture I2S (Inter-IC Sound) / TDM (Time Division Multiplexing) frame data via Ethernet or CAN (Controller Area Network) buses.
[0047] In practice, collection commands can be sent to the collection probes deployed in the target data source through the interfaces applicable to each target data source. Different interfaces use different transmission protocols.
[0048] Audio data streams from different target data sources converge on the cockpit audio bus to form a mixed audio data stream. Real-time capture of this mixed audio data stream in a non-intrusive manner reduces the overhead and latency interference caused by calling upper-level operating system APIs (Application Programming Interfaces). The mixed audio data stream can encompass raw data in PCM / I2S format, AVB (Audio Video Bridging) data streams encapsulated in Ethernet, or compressed audio streams. After parsing, the data source can be automatically traced and tagged based on the data source identifier at the time of capture.
[0049] This embodiment adopts an adaptive blind detection mechanism, which does not require manual preset configuration. It can directly determine the encapsulation format by reading the header feature code of the audio data stream, automatically call the corresponding decapsulation logic, decode and restore the audio data stream, and realize automated parsing.
[0050] In some feasible implementations, the step of parsing the mixed audio data stream to obtain noisy audio data includes steps A11-A13: Step A11: Determine the encapsulation format of each audio data stream in the mixed audio data stream based on the header feature code of each audio data stream in the mixed audio data stream; Based on the header signatures of each audio data stream in the mixed audio data stream, the corresponding encapsulation format can be automatically determined.
[0051] For example, if a timing synchronization feature, such as the Frame Synchronization Signal (LRCK), is detected in the header signature, the audio data stream's encapsulation format is determined to be the standard PCM / I2S format. If an Ethernet frame header with an AVTP protocol type field is detected in the header signature, the audio data stream's encapsulation format is determined to be AVB format. If ID3 (the metadata tag identifier for MP3 files) or a specific frame header (such as FF or FB) is detected in the header signature, the audio data stream's encapsulation format is determined to be MP3 / AAC (Advanced Audio Coding) compression format.
[0052] Step A12: Based on the encapsulation format of each audio data stream in the mixed audio data stream, decode the mixed audio data stream into the original noise audio data; The raw noise audio data is the complete noise audio obtained through decoding. For standard PCM / I2S format audio data streams, no decapsulation is required; the raw noise audio data can be obtained directly. For AVB format audio data streams, the MAC header, IP header, UDP header, and AVTP protocol header are automatically stripped, and the valid audio payload is extracted as the raw noise audio data. For MP3 / AAC compressed format audio data streams, the built-in decoder is used to losslessly decode the compressed data into linear PCM data, ensuring the integrity of the audio content, thus obtaining the raw noise audio data.
[0053] Step A13: Standardize the original noise audio data to obtain noise audio data.
[0054] In this embodiment, standardization processing refers to converting the bit depth of raw noise audio data from different data sources into a unified high-precision 32-bit floating-point format according to an internal standard, in order to eliminate as much as possible the calculation error caused by the difference in quantization precision of different hardware.
[0055] For the raw noise audio data decoded from TDM or interleaved data streams, it can be precisely decomposed into independent single-channel audio data according to timing logic to ensure the synchronization of noise audio data on the time axis.
[0056] In the specific implementation, refer to Figure 2When noisy audio data is desired, a collection command is sent, and the relevant systems in the cockpit respond. The system can capture the mixed audio data stream in real time via microphone / bus. The collected audio data is temporarily cached in memory. The system performs protocol parsing on the collected audio data, identifies the sampling rate and format, removes the transmission protocol header to decapsulate, and obtains complete audio data. The obtained complete audio data is then cleaned and standardized, and the standardized audio data is output for audio noise reduction processing.
[0057] During audio acquisition, the encapsulation format is automatically determined by reading the header feature code of the data stream, and decapsulation and decoding are automatically performed. Standardization processing is also carried out to provide a standardized data foundation for subsequent noise reduction processing and test analysis, thereby improving the consistency and accuracy of testing.
[0058] Step S20: Perform noise reduction processing on the noisy audio data to obtain noise-reduced frequency data, and convert the noise-reduced frequency data into test audio data in the target playback format; Noise-reduced audio data refers to the audio data obtained after noise reduction processing. The target playback format is the set format that can be played, such as WAV (Waveform Audio File Format) or MP3. This embodiment does not specifically limit this. The noise-reduced audio data needs to be re-encoded according to the target playback format and encapsulated, with a header containing information such as sampling rate and number of channels added. The encapsulated audio file is the test audio data, or reference signal, used for subsequent noise reduction function testing.
[0059] In some feasible implementations, the step of denoising the noisy audio data to obtain denoised audio data may include steps B11-B14: Step B11: Extract time-frequency features from the noisy audio data; Time-frequency characteristics include energy variance With inter-frame change rate Energy variance refers to short-time energy variance. It is calculated by first determining the short-time energy of each frame of the noisy audio data, and then calculating the variance of these short-time energy values. Inter-frame variation rate refers to the inter-frame variation rate of MFCC (Mel-frequency cepstral coefficients), i.e., the first-order difference MFCC (Delta-MFCC), which describes the rate of change of the MFCC coefficients of the current frame relative to adjacent frames.
[0060] Step B12: Based on the time-frequency characteristics, determine the noise type of the noisy audio data; In this embodiment, the noise type includes at least steady-state noise, non-steady-state noise, and mixed noise. For example, steady-state noise can be engine noise or continuous wind noise, non-steady-state noise can be impact noise or sudden abnormal noise, and mixed noise can be engine roaring on a bumpy road surface. No specific limitations are imposed on these types. The time-frequency characteristics of the noise audio data can be used to determine the noise type of the noise audio data.
[0061] In some feasible implementations, step B12 may include: determining the noise type of the noise audio data as steady-state noise when the time-frequency characteristics meet the steady-state conditions; determining the noise type of the noise audio data as non-steady-state noise when the time-frequency characteristics do not meet the steady-state conditions and the energy proportion of the noise audio data in the preset frequency band is less than or equal to the preset proportion; and determining the noise type of the noise audio data as mixed noise when the time-frequency characteristics do not meet the steady-state conditions and the energy proportion of the noise audio data in the preset frequency band is greater than the preset proportion.
[0062] The preset frequency band is the specified low frequency band, such as 20Hz-500Hz. The preset ratio is the threshold value of the energy proportion set in advance. The specific value can be flexibly adjusted according to the actual situation.
[0063] If the time-frequency characteristics of the noise audio data meet the steady-state conditions, the noise type of the noise audio data is considered to be steady-state noise. If the time-frequency characteristics of the noise audio data do not meet the steady-state conditions and the proportion of low-frequency energy is less than or equal to a preset proportion, the noise type of the noise audio data is considered to be non-steady-state noise. If the time-frequency characteristics of the noise audio data do not meet the steady-state conditions and the proportion of low-frequency energy is greater than a preset proportion, the noise type of the noise audio data is considered to be mixed noise.
[0064] The steady-state conditions are that the energy variance is less than or equal to the energy stability threshold and the inter-frame change rate is less than or equal to the characteristic mutation threshold. The energy stability threshold is a set threshold for the energy variance used to measure whether the short-term energy is stable; exceeding this threshold indicates that the short-term energy is not stable. The characteristic mutation threshold is a set threshold for the inter-frame change rate used to measure whether the MFCC coefficients undergo abrupt changes; exceeding this threshold indicates that the MFCC coefficients have undergone abrupt changes.
[0065] In some feasible implementations, when the energy variance of the noise audio data is less than or equal to the energy stability threshold and the inter-frame change rate is less than or equal to the feature mutation threshold, the time-frequency characteristics of the noise audio data are determined to meet the steady-state condition; when the energy variance of the noise audio data is greater than the energy stability threshold or the inter-frame change rate is greater than the feature mutation threshold, the time-frequency characteristics are determined to not meet the steady-state condition.
[0066] If the energy variance of the noise audio data is less than or equal to the energy stationarity threshold and the inter-frame change rate is less than or equal to the characteristic mutation threshold, it is considered to meet the steady-state condition. If the energy variance of the noise audio data is greater than the energy stationarity threshold and the inter-frame change rate is greater than the characteristic mutation threshold, it is considered not to meet the steady-state condition. Similarly, if the energy variance of the noise audio data is greater than the energy stationarity threshold and the inter-frame change rate is less than or equal to the characteristic mutation threshold, or if the energy variance of the noise audio data is less than or equal to the energy stationarity threshold and the inter-frame change rate is greater than the characteristic mutation threshold, it is also considered not to meet the steady-state condition.
[0067] Steady-state noise can also be divided into low-frequency steady-state noise and broadband steady-state noise. In this case, another time-domain characteristic needs to be introduced, namely spectral flatness. Spectral flatness measures the uniformity of the energy distribution in a signal's spectrum. Under steady-state conditions, if the spectral flatness of the noise audio data is less than or equal to the flatness threshold, the noise type of the noise audio data is considered to be low-frequency steady-state noise (such as engine noise). If the spectral flatness of the noise audio data is greater than the flatness threshold, the noise type of the noise audio data is considered to be broadband steady-state noise (such as continuous wind noise).
[0068] Step B13: Determine the matching noise reduction strategy for the noise audio data based on the noise type; Based on the noise type of the audio data, the optimal noise reduction algorithm is selected, i.e., a matching noise reduction strategy. This allows for targeted noise reduction processing of the audio data, ensuring a high degree of matching between the noise reduction algorithm used and the acoustic scene. For example, for steady-state noise, a DSP algorithm can be used as a matching noise reduction strategy, while for non-steady-state noise, a deep learning noise reduction model can be used.
[0069] Step B14: Based on the matching noise reduction strategy, the noisy audio data is denoised to obtain the corresponding denoised audio data.
[0070] Since the noisy audio data has been standardized, it can be considered a standardized digital signal and can be directly processed for noise reduction.
[0071] In the specific implementation, refer to Figure 3 After acquiring the standardized digital input signal, a suitable noise reduction algorithm type is selected. If Digital Signal Processing (DSP) is selected, the DSP algorithm engine is loaded, and the DSP algorithm performs frequency domain analysis, spectral subtraction, or filtering on the noisy audio data to remove noise. If a deep learning model is selected, a neural network model is loaded, and speech enhancement is achieved through feature extraction and model inference, thereby outputting a clean, noise-reduced signal as the noise-reduced signal. (Reference) Figure 4The system encodes the denoised digital signal according to the target playback format (such as WAV, MP3, etc.), encapsulates it, adds a file header containing information such as sampling rate and number of channels, re-encapsulates it into a complete audio file structure, synthesizes a new audio file (test audio data), and saves the newly generated audio file in the specified storage path for playback or analysis.
[0072] By extracting time-frequency features from noisy audio data and accurately determining the noise type of the current noisy audio data, the optimal noise reduction strategy can be adaptively matched according to the noise type during noise reduction processing. This allows for targeted noise reduction of the audio data, ensuring the effectiveness of the noise reduction process. Consequently, it can improve the accuracy of evaluation during test analysis and obtain accurate test results.
[0073] Step S30: Based on the difference between the test audio data and the reference audio data in different target frequency bands, determine the comprehensive score of frequency band noise reduction and the defect distribution map of the test audio data; This embodiment defines multiple consecutive target frequency bands in the frequency domain, with different target frequency bands having different frequency ranges. In specific implementation, the target frequency bands can be set according to the frequency ranges of the test audio data and the reference audio data. The reference audio data is the clean audio used as a reference and can be considered as a reference signal.
[0074] The defect distribution map is a visual map that shows the defects present in the test audio data. The frequency band noise reduction score is the noise reduction score of the test audio data across the entire frequency band after integrating the noise reduction results of all target frequency bands. The difference between the test audio data and the reference audio data in each target frequency band is calculated frame by frame. Based on this, a visual multidimensional defect distribution map is generated, and the frequency band noise reduction score is calculated.
[0075] Before calculating the difference, the test audio data is usually time-aligned with the reference audio data. For example, Dynamic Time Warping (DTW) is used to calculate the time delay difference between the test audio data and the reference audio data, thereby achieving sub-millisecond time alignment between the test audio data and the reference audio data.
[0076] In some feasible implementations, step S30 may include: extracting feature vectors of test audio data and reference audio data respectively; determining a nonlinear mapping relationship between test audio data and reference audio data based on the distance between the feature vectors of test audio data and reference audio data; determining the time delay difference of test audio data based on the nonlinear mapping relationship; and aligning test audio data and reference audio data in time based on the time delay difference of test audio data.
[0077] First, the test audio data is preprocessed by de-encapsulating and decoding to extract the PCM audio payload. Then, the sampling rate and bit depth of the test audio data are adaptively aligned according to the reference audio data to eliminate pseudo-errors introduced by format conversion. Subsequently, the test audio data and the reference audio data are processed by framing and windowing.
[0078] To improve robustness and reduce computational load, this embodiment does not directly use the original waveform amplitude. Instead, it extracts the feature vectors of the test audio data and the reference audio data separately. The feature vectors can be Mel-frequency cepstral coefficients (MFCCs) or power spectral density; there is no specific limitation. Assume the reference audio data is... The feature vector of the frame is Reference audio data total A frame, then the feature vector sequence of the reference audio data can be denoted as... Test audio data The feature vector of the frame is The test audio data totaled If the frame is a sequence of frames, then the feature vector sequence of the test audio data can be denoted as: .
[0079] The distance (similarity) between the feature vectors of the test audio data and the feature vectors of the reference audio data is calculated using the following formula:
[0080] In the formula, For the frame index of the reference audio data, , To test the frame index of the audio data, , For reference audio data, the first The feature vector of the frame and the test audio data The distance between the feature vectors of the frames For reference audio data, the first The feature vector of a frame, For testing audio data, the first The feature vector of a frame, is the dimension of the feature vector.
[0081] By using the distance between the feature vectors of the test audio data and the feature vectors of the reference audio data, the nonlinear mapping relationship between the test audio data and the reference audio data can be determined.
[0082] In some feasible implementations, a local cost matrix is constructed based on the distance between the feature vectors of the test audio data and the feature vectors of the reference audio data; a cumulative cost matrix is constructed based on the local cost matrix; a target regularization path is determined based on the cumulative cost matrix; and a nonlinear mapping relationship between the test audio data and the reference audio data is determined based on the target regularization path.
[0083] Construct a [database] based on the distance between the feature vectors of the test audio data and the feature vectors of the reference audio data. Local cost matrix Local cost matrix elements in Indicates the reference audio data number The feature vector of the frame and the test audio data The distance between feature vectors of frames. Based on the local cost matrix. , build a Cumulative cost matrix The cumulative cost matrix Used to record from the starting point to any point The minimum cumulative distance. Product cost matrix. elements The recursion can be performed in the following way:
[0084] In the formula, Indicates arrival point The minimum cumulative cost, This represents the local cost of the preceding point. This indicates that the frame was moved from the diagonal direction of the previous frame (corresponding to normal speed matching). This indicates that the data was transferred from the previous frame of the test audio data (corresponding to the test audio data being compressed / the reference audio data being stretched). This indicates that the data was transferred from the previous frame of the reference audio data (corresponding to the test audio data being stretched / the reference audio data being compressed).
[0085] When the cumulative cost matrix After the calculation is completed, start from the endpoint Start by backtracking backwards along the direction with the least cumulative cost to the starting point. The optimal regularization path, i.e. the target regularization path, is obtained. , of which Path points The target path is regularized. A non-linear mapping relationship was directly established between the test audio data and the reference audio data. Due to the dynamic processing delay and local waveform distortion introduced during the noise reduction process, the time axes of the test audio data and the reference audio data no longer exhibit a simple linear shift (i.e., Instead, it manifests as nonlinear stretching or compression. Target regularized path coordinate pairs in It accurately recorded the alignment state under this deformation, that is, for the reference audio data, the first... Frames, through this mapping relationship, can determine the unique frame index that is actually aligned in the test audio signal. By utilizing nonlinear mapping relationships, dynamic time delay differences are extracted through quantization, thereby converting the nonlinear offset of the frame index into a physical time scale, thus obtaining the test audio data. Time delay difference between frames and reference audio data The calculation formula is:
[0086] In the formula, Indicates the test audio data number The time delay difference between the frame and the reference audio data. This indicates that on the target normalization path, the reference audio data is at the [missing information]. The corresponding test audio data frame index. Indicates the reference audio data number The duration of a frame (in seconds), for example: if the frame shift is 10ms, then... It takes 0.01 seconds.
[0087] Based on the calculated time delay difference, the test audio data is interpolated or shifted along the time axis to align it with the reference audio data in the time dimension, thereby eliminating the impact of noise reduction processing delay on subsequent comparative evaluation.
[0088] By using dynamic time warping to calculate the time delay, the audio under test and the reference signal are aligned in sub-millisecond time, ensuring that the difference between the test audio data and the reference audio data can be accurately evaluated during subsequent test analysis, thus ensuring the accuracy of the evaluation and improving the accuracy of the test results.
[0089] Step S40: Based on the comprehensive score of frequency band noise reduction and the defect distribution map of the test audio data, generate the test results of the speech noise reduction function.
[0090] The lower the overall score of the frequency band noise reduction, the higher the noise reduction fidelity (the ability to retain the true characteristics of the original audio while eliminating background noise), and the better the noise reduction effect.
[0091] Multiple score ranges can be set, with different score ranges corresponding to different noise reduction effect levels. For example, if the overall score for frequency band noise reduction is calculated in the range of 0 to 10, then three score ranges can be set: [0,2], (2,6], and (6,10]. The noise reduction effect level corresponding to the score range (6,10] is unqualified, the noise reduction effect level corresponding to the score range (2,6] is qualified, and the noise reduction effect level corresponding to the score range [0,2] is good.
[0092] Within a set number of score intervals, a matching score interval for the overall frequency band noise reduction score is determined. The noise reduction effect level corresponding to this matching score interval is used as the noise reduction effect level of the test audio data. Simultaneously, by combining this with the defect distribution map to show the defects present in the test audio data, the final test result is obtained, and a corresponding test report is generated.
[0093] In the specific implementation, refer to Figure 5 The test audio data is compared and calculated with the reference audio data. Various indicators in the test audio data are analyzed to generate test results for the noise reduction function. The test results are verified by outputting the test audio data.
[0094] This embodiment acquires a mixed audio data stream from the target vehicle's cockpit, parses the mixed audio data stream to obtain noisy audio data, performs noise reduction processing on the noisy audio data to obtain noise-reduced frequency data, and converts the noise-reduced frequency data into test audio data in the target playback format. Based on the differences between the test audio data and reference audio data in different target frequency bands, the comprehensive noise reduction score and defect distribution map of the test audio data are determined. Based on the defect distribution map and comprehensive noise reduction score of the test audio data, the test results of the speech noise reduction function are generated. This embodiment integrates discrete test steps into a continuous automated flow, constructing a fully automated test closed loop of audio acquisition, noise reduction processing, test analysis, and result output. By utilizing standardized automated processes, it minimizes differences caused by human operation and largely avoids distortion of audio data during transmission or conversion, thereby improving test consistency and accuracy, shortening the test cycle, and increasing test efficiency.
[0095] In one feasible embodiment of this application, content that is the same as or similar to the above embodiments can be referred to the above description, and will not be repeated hereafter. Based on this, please refer to... Figure 6 Step S30 may include steps S301 to S304: Step S301: Based on the frequency range of the test audio data and the reference audio data, determine multiple consecutive target frequency bands in the frequency domain, and calculate the difference value between the test audio data and the reference audio data in each target frequency band; This embodiment determines multiple consecutive target frequency bands in the frequency domain based on the frequency range of the test audio data and the reference audio data. Different target frequency bands have different frequency ranges. For example, assuming the frequency range of the test audio data and the reference audio data is 20Hz-8kHz, three target frequency bands can be divided: low frequency band (BLow), mid frequency band (BMid), and high frequency band (BHigh). The frequency range corresponding to the low frequency band (BLow) is 20Hz-500Hz, the frequency range corresponding to the mid frequency band (BMid) is 500Hz-2kHz, and the frequency range corresponding to the high frequency band (BHigh) is 2kHz-8kHz.
[0096] The low-frequency band typically contains engine roar and mechanical vibration noise. Low-frequency energy is concentrated in this band; insufficient noise reduction in the low-frequency portion of the audio data will produce a muffled, booming sound. The mid-frequency band typically contains tire noise, road noise, and the fundamental tone of the voice. The mid-frequency band is the core area for speech intelligibility and can be used to assess the balance between speech fidelity and noise residue. The high-frequency band typically contains wind noise and air conditioning vent noise. The high-frequency band significantly affects the brightness of the voice; excessive noise reduction in the high-frequency portion of the audio data will result in a muffled sound.
[0097] In some feasible implementations, the step of calculating the difference between the test audio data and the reference audio data in each target frequency band may include: determining the logarithmic power spectral density of the test audio data based on the power spectral density of the test audio data, and determining the logarithmic power spectral density of the reference audio data based on the power spectral density of the reference audio data; calculating the logarithmic power spectral distance between the test audio data and the reference audio data in each target frequency band based on the logarithmic power spectral density of the test audio data and the logarithmic power spectral density of the reference audio data; and using the logarithmic power spectral distance corresponding to each target frequency band as the difference value corresponding to each target frequency band.
[0098] Logarithmic transformation is performed on the power spectral density of the test audio data to obtain its logarithmic power spectral density. Similarly, a logarithmic transformation is performed on the power spectral density of the reference audio data to obtain its logarithmic power spectral density. The mean square error of the logarithmic power spectral density between the test audio data and the reference audio data in each target frequency band is calculated frame by frame to obtain the logarithmic power spectral distance between the test audio data and the reference audio data in each target frequency band. This logarithmic power spectral distance between the test audio data and the reference audio data in each target frequency band is used as the difference value between the test audio data and the reference audio data in each target frequency band. At this point, for the... The first frame The difference between the target frequency bands can be denoted as: .
[0099] Step S302: Based on the difference between the test audio data and the reference audio data in each target frequency band, generate a difference distribution map and determine the defect points in the difference distribution map; By timeline The x-axis represents the index value of the target frequency band. Using color blocks of different shades as the vertical axis, the difference values are represented. The darker the color of the color block, the greater the difference value; the lighter the color of the color block, the smaller the difference value. The resulting heat map is the difference distribution map.
[0100] Each target frequency band has a corresponding difference threshold set. If the test audio data is of the [missing value], the difference threshold is [missing value]. The first frame corresponding to The difference value of the target frequency band is greater than that of the first target frequency band. If the difference threshold of a target frequency band is calculated, the location of that difference value in the difference distribution map is marked as a defect point. In other words, the difference value of a defect point in the difference distribution map is greater than the difference threshold of the corresponding target frequency band. Defect points can be marked in a more conspicuous way, such as as red warning points; this embodiment does not specifically limit this.
[0101] Step S303: Based on the differential distribution map and the time period evaluation results of the defect points in the differential distribution map, generate a defect distribution map of the test audio data; This embodiment divides the test audio data into different types of time periods: silent periods, speech periods, and transition periods. A silent period refers to the time when the VAD (Voice Activity Detection) detection result is non-speech and the energy is below a preset energy threshold. The preset energy threshold is the maximum energy of the silent period set in advance, and the specific value can be set according to actual conditions. A speech period refers to the time when the VAD detection result is speech and the duration exceeds a preset duration, which is a preset duration threshold, for example, 100ms. A transition period refers to the time period before the start of the speech and the time period after the end of the speech. The duration of the transition period can be set according to actual conditions, for example, 50ms before the start and 100ms after the end.
[0102] According to the timeline This allows us to determine the time period in which the defect point is located. If the defect point is within a silent period, the time period type corresponding to the defect point is a silent period. If the defect point is within a speech period, the time period type corresponding to the defect point is a speech period. If the defect point is within a transition period, the time period type corresponding to the defect point is a transition period.
[0103] Different evaluation metrics are used to assess noise reduction defects in different time periods. Quiet periods focus more on noise residue, speech periods on speech fidelity, and transition periods on the transient response of the noise reduction algorithm. In practice, the most suitable evaluation metric, i.e., the differentiated evaluation metric, is found based on the time period type corresponding to the defect. This differentiated evaluation metric is then used to evaluate the noise reduction effect of the defect across different time periods, and the resulting evaluation is the time period evaluation result.
[0104] In some feasible implementations, steps C11-C12 are included before step S303: Step C11: Determine the differentiated evaluation indicators for the defect points based on the time period type corresponding to the defect points; When the time period corresponding to the defect point is a silent period, the noise suppression ratio is determined based on the original average noise energy and residual noise energy of the test audio data, and the noise suppression ratio is used as the differential evaluation index of the defect point.
[0105] The original average noise energy is the average noise energy of the test audio data during the silent period when the defect point is located, before noise reduction. In other words, it represents the average noise energy of the noise audio data during the silent period when the defect point is located, and can characterize the actual environmental noise energy level inside the cockpit before noise reduction. Assuming the time window of the silent period when the defect point is located is [t1, t2], the noise audio data at this time... Integrating the squares of the amplitudes and summing them, we obtain the original average noise energy. ,Right now .
[0106] Residual noise energy refers to the residual energy of the test audio data after noise reduction processing within the silent period where the defect point is located. It characterizes the residual background noise that could not be filtered out after noise reduction processing, as well as the potential algorithmic background noise (such as music noise) energy. Assuming the time window of the silent period where the defect point is located is [t1, t2], the test audio data at this time... Integrating the squares of the amplitudes and summing them, we obtain the residual noise energy. ,Right now .
[0107] The noise suppression ratio (RSR) is the degree to which the noise reduction algorithm attenuates background noise. A higher RRS indicates that the background noise is eliminated more effectively during the corresponding quiet period. If the time period corresponding to the defect is a quiet period, the original average noise energy and residual noise energy of the test audio data are calculated. Based on these values, the RRS is calculated and used as a differentiated evaluation index for the defect. The formula for calculating the RRS is shown below:
[0108] In the formula, Noise suppression ratio, This represents the original average noise energy. This is residual noise energy.
[0109] When the time period corresponding to the defect point is a speech period, the speech distortion degree is determined based on the instantaneous error energy of the test audio data, and the speech distortion degree is used as the differential evaluation index of the defect point.
[0110] Instantaneous error energy refers to the energy corresponding to the error between the test audio data and the reference audio data in each frame. Speech distortion refers to the degree of distortion in the speech segment where the defect point is located. The smaller the speech distortion, the closer the waveform of the noise-reduced test audio data is to that of the reference audio data, and the higher the fidelity.
[0111] If the time period corresponding to the defect is a speech time period, then the instantaneous error energy of the test audio data is calculated. Based on the instantaneous error energy, the speech distortion is calculated and used as the differential evaluation index for this defect. The calculation formula for speech distortion is as follows:
[0112] In the formula, For speech distortion, This represents the total number of frames in the speech segment where the defect occurs. For reference audio data in the first The time-domain waveform amplitude sequence of the frame. To test audio data at the first The time-domain waveform amplitude sequence of the frame. For the first Instantaneous error energy of a frame.
[0113] When the time period corresponding to the defect point is a transition period, the envelope delay difference and the relative attenuation of the envelope slope of the test audio data are used as the differential evaluation indicators of the defect point.
[0114] If the time period corresponding to the defect point is a transition period, the microscopic energy envelopes of the test audio data (test signal) and the reference audio data (reference signal) during the transition period are extracted and compared. First, the onset delay of the envelope is compared, and the onset point of the energy envelopes of the two signals (i.e., the moment when the energy first crosses the set low threshold) is detected to obtain the onset point of the reference signal. The starting point of the test signal Calculate the envelope delay difference , This can be used to evaluate the signal hysteresis introduced by the noise reduction algorithm. Secondly, by comparing the envelope slope, the slope of the energy envelope of the reference signal during the transition period at the defect point is calculated to obtain the reference slope. , , For the reference signal during the transition period at the defect point Energy envelope difference within the range, The starting time point selected for calculating the reference slope. The selected end time point for calculating the reference slope. The test slope is obtained by setting a time step (the specific value can be adjusted flexibly according to the actual situation, and there is no specific limitation on it) and calculating the slope of the energy envelope of the test signal during the transition period at the defect point. , , For the test signal during the transition period at the defect point Energy envelope difference within the range, The selected starting time point for calculating the test slope. The energy envelope of the reference signal at the selected end time point for calculating the test slope is... The variation trend within the range and the energy envelope of the test signal are in The trend of change within the range must be consistent, typically showing a continuous upward or downward trend, to ensure that the signs of the reference slope and the test slope are consistent. The corresponding energy envelope value can be equal to The corresponding energy envelope value. Calculate the relative attenuation of the slope. , This can be used to assess the losses caused by transient impact forces. The envelope delay difference and the relative attenuation of the envelope slope are used as the differentiation evaluation indicators for this defect point.
[0115] Step C12: Evaluate the defect points based on the differentiated evaluation indicators of the defect points to obtain the time period evaluation results of the defect points.
[0116] If the noise suppression ratio is lower than the set suppression ratio threshold, the time period evaluation result is determined to be that the noise reduction algorithm failed or had insufficient gain in the time period of the defect point, and a "residual noise leakage" label can be added at the defect point. If the speech distortion is greater than the set distortion threshold, the time period evaluation result is determined to be that there is speech clipping, missing words, or timbre variation, and a "musical noise" label can be added at the defect point. If the envelope delay difference is greater than the set delay threshold, or the relative attenuation of the envelope slope is greater than the set attenuation threshold, the time period evaluation result is determined to be that the noise reduction algorithm has insufficient transient response, and a "transient response defect" label can be added at the defect point.
[0117] Step S304: Based on the difference values corresponding to each target frequency band and the evaluation weight of each target frequency band, determine the comprehensive score of frequency band noise reduction for the test audio data.
[0118] The difference values corresponding to each target frequency band are aggregated in two dimensions, time and frequency domains, to calculate the total score across the entire frequency band, i.e., the comprehensive score for frequency band noise reduction. In this embodiment, the weighting strategy is as follows: the mid-frequency band has a higher weight than the low-frequency and high-frequency bands (consistent with the characteristics of human hearing sensitivity), and the speech period has a higher weight than the silence period.
[0119] In some feasible implementations, before step S304, the method further includes: calculating the background noise energy of each target frequency band and the total background noise energy in the frequency domain; obtaining the correspondence between the background noise energy of the target frequency band, the total background noise energy in the frequency domain, the frequency band perception correction coefficient, and the evaluation weight of the target frequency band; and obtaining the evaluation weight of each target frequency band based on the correspondence, the total background noise energy in the frequency domain, the background noise energy of each target frequency band, and the frequency band perception correction coefficient.
[0120] Background noise energy refers to the root mean square energy value of the test signal during all silent periods in the target frequency band. Total background noise energy in the frequency domain refers to the total background noise energy of the test signal across the entire auditory range. The frequency band perception correction coefficient is a preset constant, for example, 1.5 for mid-frequency speech periods and 1.0 for low-frequency periods. This is used to manually compensate for the weighting of different frequency bands based on the equal loudness curve characteristics of human hearing.
[0121] The evaluation weight, which represents the proportion of the difference value of each target frequency band in the overall score of frequency band noise reduction, is calculated using the background noise energy of the target frequency band, the total background noise energy in the frequency domain, and the frequency band perception correction coefficient. The correspondence between the background noise energy of the target frequency band, the total background noise energy in the frequency domain, the frequency band perception correction coefficient, and the evaluation weight of the target frequency band is the calculation formula for the evaluation weight, as shown below:
[0122] In the formula, No. The target frequency band is used as the evaluation weight. For the first Background noise energy of each target frequency band This represents the total background noise energy in the frequency domain. For the first Frequency band sensing correction coefficients for each target frequency band.
[0123] Based on the differences between the test audio data and the reference audio data calculated frame by frame in each target frequency band, the average difference value of all frames within each target frequency band is determined. Then, according to the evaluation weight of the target frequency band, the average difference values corresponding to each target frequency band are weighted and summed to obtain the final comprehensive noise reduction score, as shown below:
[0124] In the formula, The overall score is for frequency band noise reduction. No. The target frequency band is used as the evaluation weight. For the first The average of the difference values corresponding to each target frequency band The number of target frequency bands.
[0125] This embodiment divides the frequency domain into multiple target frequency bands and calculates the difference between the test audio data and the reference audio data in each target frequency band. Based on the difference between the test audio data and the reference audio data in each target frequency band, a difference distribution map is generated, and defect points are identified in the difference distribution map. Based on the difference distribution map and the time-period evaluation results of the defect points in the difference distribution map, a defect distribution map of the test audio data is generated. Based on the difference values corresponding to each target frequency band and the evaluation weight of each target frequency band, the comprehensive score of frequency band noise reduction for the test audio data is determined. This embodiment constructs a fully automated test closed loop of audio acquisition, noise reduction processing, test analysis, and result output. During test analysis, the difference between the test audio data and the reference audio data is evaluated by frequency band, and defective sites are identified. Then, a refined evaluation is performed by time period to obtain a specific defect distribution map. Through adaptive weighting, a comprehensive score of noise reduction effect is calculated to improve the evaluation accuracy and obtain accurate test results.
[0126] For example, to help understand the implementation process of the automated testing method for cockpit voice noise reduction function obtained in this embodiment in conjunction with the above embodiments, please refer to... Figure 7 A simplified flowchart, specifically: Audio data acquisition and analysis. The system acquires noisy frequency streams within the cockpit in real time and simultaneously obtains a clean speech reference signal. The acquired audio data is parsed to extract sampling rate, bit depth, and channel information, and standardized and packaged. Simultaneously, short-time energy and spectral characteristics of the audio are extracted to provide data support for subsequent noise reduction strategy selection.
[0127] Intelligent noise reduction. It determines the type of noise (steady-state noise, non-steady-state noise, or mixed noise), adaptively matches the optimal noise reduction strategy, calls digital signal processing (DSP) algorithms for steady-state noise, and calls deep learning noise reduction models for non-steady-state noise, and performs targeted noise reduction processing on audio data to ensure a high degree of matching between the processing strategy and the acoustic scene.
[0128] The synthesis and output of new audio. The noise-reduced audio data stream is re-encoded and synthesized into a test audio file according to a preset standard format, which will be used as the output for subsequent quality evaluation.
[0129] Closed-loop verification and feedback. First, the dynamic time warping (DTW) algorithm is used to calculate the time delay, and the audio under test is aligned with the reference signal at the sub-millisecond level. Then, frequency band adaptive weighted evaluation and time-segmented fine comparison are performed.
[0130] Generate test reports. Summarize quantitative indicators for each frequency band and time period, automatically generate diagnostic reports, and intuitively show the shortcomings in noise reduction performance.
[0131] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the automated testing method for cockpit voice noise reduction function of this application. Any simple modifications based on this technical concept are within the protection scope of this application.
[0132] This application also provides an automated testing device for cockpit voice noise reduction function; please refer to... Figure 8 The automated testing device for cockpit voice noise reduction function includes: The acquisition and parsing module 10 is used to acquire the mixed audio data stream of the target vehicle's cockpit, parse the mixed audio data stream, and obtain the noise audio data. The noise reduction and synthesis module 20 is used to perform noise reduction processing on the noisy audio data to obtain noise-reduced frequency data, and convert the noise-reduced frequency data into test audio data in the target playback format; Functional test module 30 is used to determine the frequency band noise reduction comprehensive score and defect distribution map of the test audio data based on the difference between the test audio data and the reference audio data in different target frequency bands; The functional test module 30 is also used to generate test results for the speech noise reduction function based on the comprehensive score of frequency band noise reduction and the defect distribution map of the test audio data.
[0133] In one feasible implementation, the noise reduction and synthesis module 20 is also used to extract time-frequency features from the noisy audio data; Based on time-frequency characteristics, the noise type of the noisy audio data is determined; Based on the noise type, determine the matching noise reduction strategy for the noisy audio data; Based on the matching noise reduction strategy, the noisy audio data is processed to obtain the corresponding noise-reduced audio data.
[0134] In one feasible implementation, the noise reduction and synthesis module 20 is further configured to determine that the noise type of the noise audio data is steady-state noise when the time-frequency characteristics meet the steady-state conditions; When the time-frequency characteristics do not meet the steady-state conditions and the energy proportion of the noise audio data in the preset frequency band is less than or equal to the preset proportion, the noise type of the noise audio data is determined to be non-steady-state noise. When the time-frequency characteristics do not meet the steady-state conditions and the energy proportion of the noise audio data in the preset frequency band is greater than the preset proportion, the noise type of the noise audio data is determined to be mixed noise.
[0135] In one feasible implementation, the noise reduction and synthesis module 20 is further configured to determine that the time-frequency characteristics of the noise audio data meet the steady-state condition when the energy variance of the noise audio data is less than or equal to the energy stability threshold and the inter-frame change rate is less than or equal to the feature mutation threshold. When the energy variance of the noisy audio data is greater than the energy stationary threshold or the inter-frame change rate is greater than the feature mutation threshold, the determination of the time-frequency features does not meet the steady-state condition.
[0136] In one feasible implementation, the functional test module 30 is further configured to determine multiple consecutive target frequency bands in the frequency domain based on the frequency range of the test audio data and the reference audio data, and calculate the difference value between the test audio data and the reference audio data in each target frequency band. Based on the difference between the test audio data and the reference audio data in each target frequency band, a difference distribution map is generated, and defect points are identified in the difference distribution map. Based on the differential distribution map and the time period evaluation results of the defect points in the differential distribution map, a defect distribution map of the test audio data is generated. Based on the difference values corresponding to each target frequency band and the evaluation weight of each target frequency band, the comprehensive score of frequency band noise reduction for the test audio data is determined.
[0137] In one feasible implementation, the functional testing module 30 is also used to determine the differentiated evaluation index of the defect point based on the time period type corresponding to the defect point. Defects are evaluated based on differentiated evaluation indicators to obtain time-period evaluation results.
[0138] In one feasible implementation, the functional test module 30 is also used to determine the noise suppression ratio based on the original average noise energy and residual noise energy of the test audio data when the time period type corresponding to the defect point is a silent time period, and to use the noise suppression ratio as a differential evaluation index of the defect point. When the time period corresponding to the defect point is a speech time period, the speech distortion is determined based on the instantaneous error energy of the test audio data, and the speech distortion is used as the differential evaluation index of the defect point. When the time period corresponding to the defect point is a transition period, the envelope delay difference and the relative attenuation of the envelope slope of the test audio data are used as the differential evaluation indicators of the defect point.
[0139] In one feasible implementation, the functional test module 30 is also used to calculate the background noise energy of each target frequency band and the total background noise energy in the frequency domain. Obtain the correspondence between the background noise energy of the target frequency band, the total background noise energy in the frequency domain, the frequency band perception correction coefficient, and the evaluation weight of the target frequency band; Based on the correspondence, the total background noise energy in the frequency domain, and the background noise energy and frequency band perception correction coefficient of each target frequency band, the evaluation weight of each target frequency band is obtained.
[0140] In one feasible implementation, the functional test module 30 is further configured to determine the logarithmic power spectral density of the test audio data based on the power spectral density of the test audio data, and to determine the logarithmic power spectral density of the reference audio data based on the power spectral density of the reference audio data. The logarithmic power spectrum distance corresponding to each target frequency band is used as the difference value corresponding to each target frequency band; The logarithmic power spectrum distance between the test audio data and the reference audio data in each target frequency band is used as the difference value between the test audio data and the reference audio data in each target frequency band.
[0141] In one feasible implementation, the functional test module 30 is also used to extract the feature vectors of the test audio data and the feature vectors of the reference audio data, respectively. Based on the distance between the feature vectors of the test audio data and the feature vectors of the reference audio data, a nonlinear mapping relationship between the test audio data and the reference audio data is determined. The time delay difference of the test audio data is determined based on the nonlinear mapping relationship; Based on the time delay difference of the test audio data, the test audio data is time-aligned with the reference audio data.
[0142] In one feasible implementation, the acquisition and parsing module 10 is further used to determine the encapsulation format of each audio data stream in the mixed audio data stream based on the header feature code of each audio data stream in the mixed audio data stream. Based on the encapsulation format of each audio data stream in the mixed audio data stream, the mixed data stream is decoded into the original noisy audio data; The original noise audio data is standardized to obtain the noise audio data.
[0143] The automated testing device for cockpit voice noise reduction function provided in this application, employing the automated testing method for cockpit voice noise reduction function in the above embodiments, can solve the technical problem of poor accuracy of test results in traditional cockpit voice noise reduction function testing methods. Compared with the prior art, the beneficial effects of the automated testing device for cockpit voice noise reduction function provided in this application are the same as those of the automated testing method for cockpit voice noise reduction function provided in the above embodiments, and other technical features in the automated testing device for cockpit voice noise reduction function are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.
[0144] This application provides an automated testing device for cockpit voice noise reduction function. The automated testing device for cockpit voice noise reduction function includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the automated testing method for cockpit voice noise reduction function in the above embodiment 1.
[0145] The following is for reference. Figure 9 The diagram illustrates a structural schematic of an automated testing device suitable for implementing the cockpit voice noise reduction function in the embodiments of this application. The automated testing device for cockpit voice noise reduction function in the embodiments of this application may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital radio receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Description), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 9 The automated testing equipment for cockpit voice noise reduction shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.
[0146] like Figure 9As shown, the automated testing equipment for cockpit voice noise reduction function may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to the program stored in ROM (Read Only Memory) 1002 or the program loaded from storage device 1003 into RAM (Random Access Memory) 1004. RAM 1004 also stores various programs and data required for the operation of the automated testing equipment for cockpit voice noise reduction function. The processing unit 1001, ROM 1002, and RAM 1004 are interconnected via bus 1005. Input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to I / O interface 1006: input devices 1007 including, for example, touchscreens, touchpads, keyboards, mice, image sensors, microphones, accelerometers, gyroscopes, etc.; output devices 1008 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 1003 including, for example, magnetic tapes, hard disks, etc.; and communication devices 1009. Communication device 1009 allows the automated cockpit voice noise reduction function test equipment to communicate wirelessly or wiredly with other devices to exchange data. Although the figure shows an automated cockpit voice noise reduction function test equipment with various systems, it should be understood that it is not required to implement or possess all the systems shown. More or fewer systems may be implemented alternatively.
[0147] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from ROM 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.
[0148] The automated testing equipment for cockpit voice noise reduction function provided in this application, employing the automated testing method for cockpit voice noise reduction function in the above embodiments, can solve the technical problem of poor accuracy of test results in traditional cockpit voice noise reduction function testing methods. Compared with the prior art, the beneficial effects of the automated testing equipment for cockpit voice noise reduction function provided in this application are the same as the beneficial effects of the automated testing method for cockpit voice noise reduction function provided in the above embodiments, and other technical features in this automated testing equipment for cockpit voice noise reduction function are the same as those disclosed in the method of the previous embodiment, and will not be repeated here.
[0149] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.
[0150] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0151] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, the computer-readable program instructions being used to execute the automated testing method for cockpit voice noise reduction function in the above embodiments.
[0152] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.
[0153] The aforementioned computer-readable storage medium may be included in the automated testing equipment for cockpit voice noise reduction function; or it may exist independently and not be installed in the automated testing equipment for cockpit voice noise reduction function.
[0154] The aforementioned computer-readable storage medium carries one or more programs. When these programs are executed by the automated testing equipment for cockpit voice noise reduction, the automated testing equipment for cockpit voice noise reduction performs the following actions: parses the mixed audio data stream to obtain noisy audio data; performs noise reduction processing on the noisy audio data to obtain noise-reduced frequency data, and converts the noise-reduced frequency data into test audio data in a target playback format; determines the comprehensive noise reduction score and defect distribution map of the test audio data based on the difference between the test audio data and the reference audio data in different target frequency bands; and generates the test results for the voice noise reduction function based on the defect distribution map and the comprehensive noise reduction score of the test audio data.
[0155] Computer program code for performing the operations of this application can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0156] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0157] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.
[0158] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the above-described automated testing method for cockpit voice noise reduction function. This solves the technical problem of poor accuracy in traditional cockpit voice noise reduction function testing methods. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the automated testing method for cockpit voice noise reduction function provided in the above embodiments, and will not be repeated here.
[0159] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the above-described automated testing method for cockpit voice noise reduction function.
[0160] The computer program product provided in this application can solve the technical problem of poor accuracy in test results of traditional cockpit voice noise reduction function testing methods. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as those of the automated cockpit voice noise reduction function testing method provided in the above embodiments, and will not be repeated here.
[0161] The above are only some embodiments of this application and do not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.
Claims
1. An automated testing method for cockpit voice noise reduction function, characterized in that, The method includes: Collect a mixed audio data stream from the target vehicle's cockpit, and parse the mixed audio data stream to obtain noise audio data; The noise audio data is denoised to obtain denoised audio data, and the denoised audio data is converted into test audio data in the target playback format. Based on the difference between the test audio data and the reference audio data in different target frequency bands, the frequency band noise reduction comprehensive score and defect distribution map of the test audio data are determined. Based on the comprehensive score of frequency band noise reduction and the defect distribution map of the test audio data, the test results of the speech noise reduction function are generated.
2. The method as described in claim 1, characterized in that, The step of performing noise reduction processing on the noise audio data to obtain noise-reduced audio data includes: Extract time-frequency features from the noise audio data; Based on the time-frequency characteristics, the noise type of the noise audio data is determined; Based on the noise type, determine the matching noise reduction strategy for the noise audio data; Based on the matching noise reduction strategy, the noise audio data is processed to obtain the corresponding noise-reduced audio data.
3. The method as described in claim 2, characterized in that, The step of determining the noise type of the noise audio data based on the time-frequency characteristics includes: When the time-frequency characteristics meet the steady-state conditions, the noise type of the noise audio data is determined to be steady-state noise; When the time-frequency characteristics do not meet the steady-state conditions and the energy proportion of the noise audio data in the preset frequency band is less than or equal to the preset proportion, the noise type of the noise audio data is determined to be non-steady-state noise. When the time-frequency characteristics do not meet the steady-state conditions and the energy proportion of the noise audio data in the preset frequency band is greater than the preset proportion, the noise type of the noise audio data is determined to be mixed noise.
4. The method as described in claim 3, characterized in that, The time-frequency features include energy variance and inter-frame change rate, and the method further includes: When the energy variance of the noise audio data is less than or equal to the energy stationary threshold and the inter-frame change rate is less than or equal to the feature mutation threshold, it is determined that the time-frequency characteristics of the noise audio data meet the steady-state condition. When the energy variance of the noise audio data is greater than the energy stability threshold or the inter-frame change rate is greater than the feature mutation threshold, it is determined that the time-frequency feature does not meet the steady-state condition.
5. The method as described in claim 1, characterized in that, The step of determining the comprehensive frequency band noise reduction score and defect distribution map of the test audio data based on the difference between the test audio data and the reference audio data in different target frequency bands includes: Based on the frequency range of the test audio data and the reference audio data, multiple consecutive target frequency bands are determined in the frequency domain, and the difference between the test audio data and the reference audio data in each target frequency band is calculated. Based on the difference between the test audio data and the reference audio data in each target frequency band, a difference distribution map is generated, and defect points are identified in the difference distribution map. Based on the difference distribution map and the time period evaluation results of the defect points in the difference distribution map, a defect distribution map of the test audio data is generated; Based on the difference values corresponding to each target frequency band and the evaluation weight of each target frequency band, the comprehensive score of frequency band noise reduction for the test audio data is determined.
6. The method as described in claim 5, characterized in that, Before the step of generating the defect distribution map of the test audio data based on the difference distribution map and the time-period evaluation results of the defect points in the difference distribution map, the method further includes: Based on the time period type corresponding to the defect point, determine the differentiated evaluation index of the defect point; Based on the differentiated evaluation indicators of the defect points, the defect points are evaluated to obtain the time period evaluation results of the defect points.
7. The method as described in claim 6, characterized in that, The step of determining the differentiated evaluation index of the defect point based on the time period type corresponding to the defect point includes: When the time period corresponding to the defect point is a silent time period, the noise suppression ratio is determined based on the original average noise energy and residual noise energy of the test audio data, and the noise suppression ratio is used as the differential evaluation index of the defect point. When the time period type corresponding to the defect point is a speech time period, the speech distortion degree is determined based on the instantaneous error energy of the test audio data, and the speech distortion degree is used as the differential evaluation index of the defect point. When the time period corresponding to the defect point is a transition period, the envelope delay difference and the relative attenuation of the envelope slope of the test audio data are used as the differential evaluation indicators of the defect point.
8. The method as described in claim 5, characterized in that, Before the step of determining the comprehensive score of frequency band noise reduction for the test audio data based on the difference values corresponding to each target frequency band and the evaluation weight of each target frequency band, the method further includes: Calculate the background noise energy of each target frequency band and the total background noise energy in the frequency domain; Obtain the correspondence between the background noise energy of the target frequency band, the total background noise energy in the frequency domain, the frequency band perception correction coefficient, and the evaluation weight of the target frequency band; Based on the correspondence, the total background noise energy in the frequency domain, and the background noise energy and frequency band perception correction coefficient of each target frequency band, the evaluation weight of each target frequency band is obtained.
9. The method as described in claim 5, characterized in that, The step of calculating the difference between the test audio data and the reference audio data in each target frequency band includes: Based on the power spectral density of the test audio data, the logarithmic power spectral density of the test audio data is determined, and based on the power spectral density of the reference audio data, the logarithmic power spectral density of the reference audio data is determined. Based on the log power spectral density of the test audio data and the log power spectral density of the reference audio data, calculate the log power spectral distance between the test audio data and the reference audio data in each target frequency band; The logarithmic power spectrum distance corresponding to each target frequency band is used as the difference value corresponding to each target frequency band.
10. The method as described in claim 1, characterized in that, Before the step of determining the comprehensive frequency band noise reduction score and defect distribution map of the test audio data based on the difference between the test audio data and the reference audio data in different target frequency bands, the following steps are also included: Extract the feature vectors of the test audio data and the reference audio data respectively; Based on the distance between the feature vector of the test audio data and the feature vector of the reference audio data, a nonlinear mapping relationship between the test audio data and the reference audio data is determined; Based on the nonlinear mapping relationship, the time delay difference of the test audio data is determined; Based on the time delay difference of the test audio data, the test audio data and the reference audio data are time-aligned.
11. The method according to any one of claims 1 to 10, characterized in that, The step of parsing the mixed audio data stream to obtain noisy audio data includes: Based on the header feature codes of each audio data stream in the mixed audio data stream, the encapsulation format of each audio data stream in the mixed audio data stream is determined; Based on the encapsulation format of each audio data stream in the mixed audio data stream, the mixed audio data stream is decoded into raw noise audio data; The original noise audio data is standardized to obtain noise audio data.
12. An automated testing device for cockpit voice noise reduction function, characterized in that, The device includes: The acquisition and parsing module is used to acquire the mixed audio data stream of the target vehicle's cockpit, and parse the mixed audio data stream to obtain noise audio data; The noise reduction and synthesis module is used to perform noise reduction processing on the noise audio data to obtain noise-reduced frequency data, and convert the noise-reduced frequency data into test audio data in the target playback format; The functional testing module is used to determine the frequency band noise reduction comprehensive score and defect distribution map of the test audio data based on the difference between the test audio data and the reference audio data in different target frequency bands. The functional testing module is also used to generate test results for the speech noise reduction function based on the comprehensive score of frequency band noise reduction and the defect distribution map of the test audio data.
13. An automated testing device for cockpit voice noise reduction function, characterized in that, The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the automated testing method for cockpit voice noise reduction function as described in any one of claims 1 to 11.
14. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the automated testing method for cockpit voice noise reduction function as described in any one of claims 1 to 11.