New energy automobile cabin abnormal noise detection method and device based on sound image fusion
By combining multi-channel microphone arrays and cameras in the cockpit of new energy vehicles, using audio-visual fusion technology and neural network models, the abnormal noise of new energy vehicles is identified and positioned, and the problems of low single-mode detection efficiency and accuracy in the existing technology are solved, and more efficient and accurate noise detection is achieved.
Patent Information
- Application Number
- CN202510322256.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-06-27
AI Technical Summary
In the prior art, the abnormal noise detection of cockpits of new energy vehicles mainly relies on single-modal data, such as acoustics or images, resulting in low detection efficiency and accuracy, and the inability to intuitively locate the noise source position.
Using a detection method based on audio-visual fusion, audio and video data are synchronized by installing a multi-channel microphone array and camera on the top of the cockpit of a new energy vehicle. Anomaly noise is identified and positioned, and the noise source is displayed in the image as a sound field thermal map using an abnormality monitoring model and an improved random search fast method (SRC).
It realizes more efficient and intuitive noise abnormality detection of new energy vehicle cockpits, significantly improving the accuracy and detection efficiency of noise source positioning, and can be widely used in noise monitoring and fault diagnosis of new energy vehicle cockpits.
Smart Images

Figure CN120213209A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method and device for detecting abnormal noise in a new energy vehicle cockpit based on audio-visual fusion, belonging to the technical field of noise detection. Background Technique
[0002] With the rapid development of new energy vehicles, especially pure electric vehicles and hybrid electric vehicles, their market share has been continuously expanding. Compared with traditional fuel vehicles, new energy vehicles are favored by users for their advantages such as high efficiency, low emissions, and quietness. However, at the same time, the noise problem in the new energy vehicle cockpit has increasingly become a key technical challenge affecting user experience and comfort.
[0003] Compared with traditional fuel vehicles, the noise characteristics of new energy vehicles are numerous types, noise-sensitive, and complex noise. At present, the detection and analysis of abnormal cockpit noise are mainly based on acoustic detection or image detection. However, single-modal methods such as acoustic or image are difficult to comprehensively reflect the spatial and temporal characteristics of abnormal noise. Existing methods only rely on single-modal data, and the detection accuracy in a complex noise environment is limited. The acoustic detection method cannot intuitively locate the position of the internal noise source in the cockpit, and the image detection method lacks the ability to identify acoustic features. Summary of the Invention
[0004] The purpose of the present invention is to provide a method and device for detecting abnormal noise in a new energy vehicle cockpit based on audio-visual fusion to solve the problems of low efficiency and accuracy of abnormal noise detection existing in the prior art that only rely on single-modal data such as acoustic or image.
[0005] The technical solution of the present invention is as follows:
[0006] A method for detecting abnormal noise in a new energy vehicle cockpit based on audio-visual fusion includes the following steps:
[0007] S1. Install a multi-channel microphone array and a camera on the top of the new energy vehicle cockpit, synchronously collect multi-channel audio data and video data through the multi-channel microphone array and the camera respectively, and after preprocessing the obtained multi-channel audio data, obtain the preprocessed audio data;
[0008] S2. Identify and detect whether there is an abnormality in the preprocessed audio data through an abnormal monitoring model based on a neural network. If an abnormality exists, determine the type of abnormal noise and extract the abnormal noise signal, and enter the next step S3; otherwise, repeat step S2 until the detection of the preprocessed audio data is completed;
[0009] S3. Locate the abnormal noise through an improved spatial shrinkage fast method based on random search, that is, the improved SRC method, to obtain the abnormal noise location;
[0010] S4. Calibrate the abnormal noise localization position extracted and the images in the video data, and display the noise source in the form of a sound field heat map in the image to obtain the sound field map after sound-image fusion.
[0011] Further, in step S1, the multi-channel microphone array uses multiple microphones arranged in multiple concentric circular arrays from the inside out.
[0012] Further, in step S2, the abnormal types include at least one of the following: abnormal vibration, suspected abnormal vibration, partial discharge, suspected partial discharge, and abnormal sound.
[0013] Further, in step S2, the abnormal monitoring model based on the neural network includes a time-frequency conversion module, a convolutional neural network CNN, a spatio-temporal conversion module, a transformer model, and a classification head.
[0014] Time-frequency conversion module: Perform time-frequency conversion on the preprocessed audio data input to obtain a Mel spectrogram.
[0015] Convolutional neural network CNN: Extract spectral features from the input Mel spectrogram and output the feature map to the spatio-temporal conversion module.
[0016] Spatio-temporal conversion module: Reshape the features of the feature map into a time series and then output it to the transformer model.
[0017] Transformer model: Perform position encoding, dynamic feature association, and multi-scale feature fusion on the input time series, capture the time series dependence relationship, and output a feature vector containing local features and global context information to the classification head.
[0018] Classification head: Output the classification result according to the input feature vector containing local features and global context information.
[0019] Further, step S3 is specifically as follows.
[0020] S31. For the preprocessed audio data obtained in step S1, use spectral subtraction and Wiener filtering for signal enhancement and denoising.
[0021] S32. Perform short-time Fourier transform on each frame of the signal to obtain the frequency-domain signal X m (f), where m is the microphone index and f is the frequency.
[0022] S33. Calculate the phase transformation weighted cross-correlation R ij τ:
[0023]
[0024] where τ represents the time delay difference between the ith and jth microphones in a pair of microphones in the microphone array when the sound wave arrives, and X i (f) represents the spectrum of the ith microphone signal, represents the complex conjugate of the spectrum of the jth microphone signal, and e is the natural constant;
[0025] S34. Divide the initial search space into several sound source surfaces at intervals of L along the normal direction of the multi-channel microphone array from near to far, discretize each sound source surface into grid points in a mesh pattern, and determine the three-dimensional spatial position coordinates of each grid point. Each grid point forms a grid point set G;
[0026] S35. Calculate the theoretical time delay τ ij (p) from each grid point p in the grid point set G to the ith and jth microphones in each pair of microphones;
[0027] S36. Construct a positioning model based on the compressed sensing technology:
[0028] z = Ψθ
[0029] where the estimated value z of the distance difference between the sound source and the two microphones in each microphone pair calculated based on the received signals is a C×1 vector, and C is the total number of microphone pairs in the multi-channel microphone array; Ψ is a C×Q dictionary matrix, and each column element of the dictionary matrix is the theoretical value of the assumed distance difference of the sound source at the grid point position of each microphone pair; the spatial position estimation parameter θ of Q grid points in the search space is a Q×1 vector;
[0030] S37. Randomly select J grid points in the grid point set G, and calculate the superimposed value SRP of the cross-power spectra of the received signals of all microphone pairs for each grid point: Select the N k grid points with the largest superimposed value SRP of the cross-power spectra. For the N k grid points, construct a spatial dictionary Ψ k according to the time delay relationship, and then solve the position estimation parameter θ in the positioning model constructed based on the compressed sensing technology through the orthogonal matching pursuit algorithm OMP. Obtain the grid points with the position estimation parameter θ greater than the set threshold as new grid points and get the number N k ` of the new grid points, update N k = the number N k ` of the new grid points, shrink the search area to a cuboid area containing these N k points, form a new grid point set G` with the grid points of the cuboid area, and update the grid point set G = G`; Iterate step S37 until it shrinks to the smallest search space, that is, the search space with only one grid point, and enter the next step S38;
[0031] S38. After the iteration is completed, the grid point with the largest superimposed value of the cross-power spectrum is used as the abnormal noise localization position.
[0032] Further, step S4 is specifically as follows:
[0033] S41. Extract the characteristic data of the abnormal noise signal from the audio data collected by the microphone array, including energy data, power spectral density, and frequency bandwidth;
[0034] S42. Perform normalization processing on the characteristic data, pair the normalized characteristic data with the abnormal noise localization position, and construct the corresponding mapping relationship;
[0035] S43. Use the checkerboard calibration method to determine the internal parameters of the camera, including the focal coordinates and the optical center pixel coordinates, and the external parameters, including the rotation matrix and the translation vector, establish the mapping relationship between the image pixel coordinates and the real three-dimensional space, and calibrate the abnormal noise localization position with the camera image pixel coordinates;
[0036] S44. According to the mapping relationship, convert the characteristic data into a heat map through a color mapping scheme, where the color mapping scheme uses colors with different brightness levels to represent different characteristic data.
[0037] Further, in step S43, calibrating the abnormal noise localization position with the camera image pixel coordinates is specifically to convert the abnormal noise localization position P s =(X, Y, Z) into the camera image pixel coordinates (u, v):
[0038]
[0039] where R is the rotation matrix, t is the translation vector, and the internal parameter matrix of the camera where f x , f y is the focal coordinate, c x , c y is the optical center pixel coordinate.
[0040] An abnormal noise detection device for a new energy vehicle cockpit based on acoustic and image fusion for implementing the method described above, comprising a multi-channel microphone array, an image acquisition module, a data processing module, a display module, and a power supply module for power supply.
[0041] Multi-channel microphone array: Installed on the top of the new energy vehicle cockpit to collect multi-channel audio data;
[0042] Image acquisition module: The video data is collected by a camera installed on the top of the new energy vehicle cockpit;
[0043] Data processing module: After preprocessing the multi-channel audio data, the preprocessed audio data is obtained; the preprocessed audio data is identified and detected for abnormalities through an anomaly detection model based on a neural network. If an anomaly exists, the type of abnormal noise is determined and the abnormal noise signal is extracted until the detection of the preprocessed audio data is completed; through an improved spatial shrinkage fast method based on random search, i.e., the improved SRC method, the abnormal noise is located to obtain the location of the abnormal noise; the extracted location of the abnormal noise is calibrated with the image in the video data, and the noise source is displayed in the image in the form of a sound field heat map to obtain the sound-image fusion sound field map.
[0044] Display module: Used to display the sound-image fusion sound field map.
[0045] The beneficial effects of the present invention are: This method and device for detecting abnormal noise in a new energy vehicle cockpit based on sound-image fusion can achieve more efficient and intuitive detection of abnormal noise in the new energy vehicle cockpit, can significantly improve the accuracy of noise source localization, can greatly improve the efficiency of detecting abnormal noise in the new energy vehicle cockpit, and can be widely applied to the noise monitoring and fault diagnosis of the new energy vehicle cockpit. By combining acoustic signals and image data, through a multi-modal fusion method, the deficiencies of single-modal detection are made up for. The microphone array is used to achieve the spatial localization of cockpit noise, and the camera is combined to obtain the dynamic image inside the cockpit to achieve the intuitive capture of the abnormal noise source. By jointly analyzing the acoustic features and image information, the accuracy and real-time performance of abnormal noise detection can be significantly improved, thereby enhancing the user experience and vehicle safety. Brief Description of the Drawings
[0046] Figure 1 is a schematic flowchart of the method for detecting abnormal noise in a new energy vehicle cockpit based on sound-image fusion according to an embodiment of the present invention;
[0047] Figure 2 is a schematic illustration of a multi-channel microphone array installed on the top of a new energy vehicle cockpit in an embodiment;
[0048] Figure 3 is a schematic illustration of a multi-channel microphone array and a camera in an embodiment;
[0049] Figure 4 is a schematic illustration of the device for detecting abnormal noise in a new energy vehicle cockpit based on sound-image fusion according to an embodiment;
[0050] Figure 5 is a schematic illustration of the sound source position before calibration and the actual sound position in an embodiment;
[0051] Figure 6 is a schematic illustration of the sound source position after calibration and the actual sound position in an embodiment;
[0052] Among them, 1 is a multi-channel microphone array, and 2 is a camera. Specific implementation manner
[0053] The preferred embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0054] The embodiment provides a method for detecting abnormal noise in a new energy vehicle cockpit based on audio-visual fusion, as Figure 1 , including the following steps,
[0055] S1. Install a multi-channel microphone array 1 and a camera 2 on the top of the new energy vehicle cockpit, as Figure 2 and Figure 3 . Synchronously collect multi-channel audio data and video data through the multi-channel microphone array 1 and the camera 2 respectively, and after preprocessing the obtained multi-channel audio data, obtain the preprocessed audio data.
[0056] In step S1, the multi-channel microphone array 1 adopts a plurality of microphones arranged in a plurality of concentric circular arrays from the inside out, as Figure 3 . The preprocessing includes data cleaning, outlier removal, and signal smoothing. Among them, the standard score Z-Score test is used to identify and eliminate abnormal data, and a calibration value Z0 and a threshold T are set, where x represents the collected audio data, μ is the mean value, and ρ is the standard deviation; when the calibration value |Z0| > the threshold T, it is determined as abnormal data, and conversely, when |Z| ≤ the threshold T, it is determined as normal data. The signal smoothing process smooths the original audio signal by applying methods such as moving average or median filtering to reduce random fluctuations.
[0057] S2. Identify and detect whether there is an abnormality in the preprocessed audio data through an abnormal monitoring model based on a neural network. If an abnormality exists, determine the type of abnormal noise and extract the abnormal noise signal, and enter the next step S3; otherwise, repeat step S2 until the detection of the preprocessed audio data is completed.
[0058] In step S2, the abnormal types include at least one of the following: abnormal vibration, suspected abnormal vibration, partial discharge, suspected partial discharge, and abnormal sound. Among them, abnormal vibration is a mechanical vibration with a vibration amplitude exceeding the set threshold; suspected abnormal vibration is a potential fault vibration with abnormal vibration spectrum characteristics, that is, the audio signal contains the spectrum of the vibration signal, but the vibration amplitude is lower than the set threshold; partial discharge is an initial discharge signal with a partial discharge amount higher than the detection threshold; suspected partial discharge is an initial discharge signal with a partial discharge amount lower than the detection threshold.
[0059] In step S2, the anomaly monitoring model based on a neural network includes a time-frequency conversion module, a convolutional neural network (CNN), a spatio-temporal conversion module, a Transformer model, and a classification head.
[0060] Time-frequency conversion module: It performs time-frequency conversion on the preprocessed audio data input to obtain a Mel spectrogram.
[0061] Convolutional neural network (CNN): It extracts spectral features from the input Mel spectrogram and outputs a feature map to the spatio-temporal conversion module.
[0062] Spatio-temporal conversion module: It reshapes the features of the feature map into a time series and then outputs it to the Transformer model.
[0063] Transformer model: It performs positional encoding, dynamic feature association, and multi-scale feature fusion on the input time series, captures the temporal dependence relationship, and outputs a feature vector containing local features and global context information to the classification head.
[0064] Classification head: It outputs a classification result based on the input feature vector containing local features and global context information.
[0065] The multi-task loss function is:
[0066] L = αL class - βL logic
[0067] where α and β are weight coefficients, and the categorical cross-entropy loss L logic : where y i is the true label, is the predicted probability, N is the number of classes, and the localization mean squared error loss L loc : where p j is the true position, is the predicted position, and M is the number of localization points.
[0068] S3. Use the improved spatial shrinkage rapid method based on random search, that is, the improved SRC method, to locate the abnormal noise and obtain the abnormal noise localization position.
[0069] S31. For the preprocessed audio data obtained in step S1, use spectral subtraction and Wiener filtering for signal enhancement and denoising.
[0070] In step S31, spectral subtraction and Wiener filtering are used as signal processing techniques to remove the background noise of each microphone channel and improve the signal-to-noise ratio of the target signal. Among them, spectral subtraction subtracts the estimated noise spectrum N(f) from the audio signal spectrum Y(f) to obtain the enhanced signal spectrum: where γ is the over - reduction factor. In the Prony domain, the estimate of Wiener filtering is: where and are the power spectral densities of the abnormal noise and the audio signal respectively.
[0071] S32. Perform short - time Fourier transform on each frame of the signal to obtain the frequency - domain signal X m (f), where m is the microphone index and f is the frequency;
[0072] S33. Calculate the phase - transform weighted cross - correlation R ij τ of each pair of microphone signals:
[0073]
[0074] where τ represents the time - delay difference between the i - th and j - th microphones in a pair of microphones of the microphone array when the sound wave arrives, X i (f) represents the spectrum of the i - th microphone signal, represents the complex conjugate of the spectrum of the j - th microphone signal, and e is the natural constant;
[0075] S34. Divide the initial search space into several sound - source surfaces along the normal direction of the multi - channel microphone array from near to far at intervals of L, discretize each sound - source surface into grid points in a mesh - like manner, and determine the three - dimensional spatial position coordinates of each grid point. Each grid point forms a grid - point set G;
[0076] S35. Calculate the theoretical time - delay τ ij (p) from each grid point p in the grid - point set G to the i - th and j - th microphones in each pair of microphones;
[0077] S36. Construct a positioning model based on the compressed sensing technology:
[0078] z = Ψθ
[0079] where the estimated value z of the distance difference between the sound source and the two microphones in the microphone pair calculated according to the received signals by each microphone pair is a C×1 vector, C is the total number of microphone pairs of the multi - channel microphone array 1; Ψ is a C×Q dictionary matrix, and each column element of the dictionary matrix is the theoretical value of the assumed distance difference of the sound source at the grid - point position of each microphone pair; the spatial position estimation parameter θ of Q grid points in the search space is a Q×1 vector;
[0080] S361. The signals received by the microphone array come from one or more point sound sources in space, and there are M microphones at known positions: where m ix ,m iy ,m izDenote the x-axis, y-axis, and z-axis coordinates of the i-th microphone. The number of microphone pairs is C = (M - 1)M / 2, and the set of microphone pairs is: P = p j | j = 1…C}, where the j-th pair of microphones p j contains two three-dimensional spatial position vectors: and
[0081] S362. Divide the initial search space into Q grid points, then the set of grid points where the k-th grid point where, x kx , x ky , x kz denote the x-axis, y-axis, and z-axis coordinates of the k-th grid point ;
[0082] S363. For the k-th grid point there is a virtual sound source. The distance difference between the k-th grid point and the position p j is
[0083]
[0084] S364. Assume the position where, s x , s y , s z denote the x-axis, y-axis, and z-axis coordinates of this position. There is a sound source. Calculate the TDOA estimate value j of the microphone pair p and the estimated value of the distance difference between the position of p j by the generalized cross-correlation algorithm based on the received signals from the microphones:
[0085]
[0086] where, denotes the theoretical TDOA value of the microphone pair p j receiving the virtual sound source at . is the point in the spatial set G that satisfies the equality of the theoretical value and the estimated value of TDOA, that is there is:
[0087]
[0088] S365. For all microphone pairs p j , should satisfy the above formula to the greatest extent. Assume the vector Denote all microphone pairs as p j and the theoretical value of the assumed sound source distance difference at
[0089] S366. Form of the positioning model of the improved algorithm:
[0090] z = Ψθ
[0091] where z represents the sound source position calculated from the received signals by all microphone pairs and p j the estimated value of the distance difference, z is a C×1 vector θ represents the spatial position estimation parameter of Q points in the set G Ψ represents a dictionary matrix of C×Q Each column element of the dictionary matrix Ψ consists of For all constitute the elements in the dictionary Ψ
[0092] The vector z can be regarded as the input corresponding to different parameters θ in the proposed model, and under the function of the dictionary Ψ, the output is generated
[0093] S37. Randomly select J lattice points in the lattice point set G, and calculate the superimposed value SRP of the cross-power spectra of the received signals of all microphone pairs for each lattice point: Select the N lattice points with the largest superimposed value SRP of the cross-power spectra from them k For the N k lattice points, construct a spatial dictionary Ψ according to the time delay relationship k , and then solve the position estimation parameter θ in the positioning model constructed based on the compressed sensing technology through the orthogonal matching pursuit algorithm OMP. Obtain the lattice points with the position estimation parameter θ greater than the set threshold as new lattice points and get the number of new lattice points N' k , update N k = the number of new lattice points N' k , shrink the search area to a cuboid area containing these N k points, form a new lattice point set G' with the lattice points in the cuboid area, and update the lattice point set G = G'; Iterate step S37 until it shrinks to the smallest search space, that is, the search space with only one lattice point, and enter the next step S38;
[0094] S38. After the iteration is completed, use the lattice point with the largest superimposed value of the cross-power spectrum as the abnormal noise positioning position
[0095] In step S3, the abnormal noise localization in the new energy vehicle cockpit can be regarded as a process of matching and searching in the space dictionary of the new energy vehicle cockpit. All the spaces in the reasoning refer to the space of the new energy vehicle cockpit. Compared with the original SRC method, in the improved SRC method, in each iteration, the principle of region contraction is not only based on the superimposed value SRP of the cross-power spectra of the microphones at the position points within the region, but also based on the spatial position parameter θ of these spatial points. The improved SRC method retains the advantages of the original SRC method: in each iteration, instead of examining the SRP of all the divided position points within the region, only a subset of them is calculated, and the iteration shrinks towards the sub-region with a greater possibility of containing the sound source. In the process of region contraction of the improved SRC method, a matching search method is combined, which combines the randomness of the SRC method and the directivity of the orthogonal matching pursuit algorithm OMP, and can improve the efficiency of randomly distributing points in the search region in each iteration.
[0096] S4. Calibrate the abnormal noise localization position extracted and the image in the video data, and display the noise source in the image in the form of a sound field heat map to obtain the sound field map after sound-image fusion.
[0097] S41. Extract the characteristic data of the abnormal noise signal from the audio data collected by the microphone array, including energy data, power spectral density, and frequency bandwidth.
[0098] S42. Perform normalization processing on the characteristic data, pair the normalized characteristic data with the abnormal noise localization position, and construct the corresponding mapping relationship.
[0099] S43. Use the checkerboard calibration method to determine the internal parameters of camera 2, including the focal point coordinates and the optical center pixel coordinates, and the external parameters, including the rotation matrix and the translation vector, establish the mapping relationship between the image pixel coordinates and the real three-dimensional space, and calibrate the abnormal noise localization position with the image pixel coordinates of camera 2.
[0100] In step S43, calibrating the abnormal noise localization position with the image pixel coordinates of camera 2 specifically means converting the abnormal noise localization position P s =(X, Y, Z) into the image pixel coordinates (u, v) of camera 2:
[0101]
[0102] where R is the rotation matrix and t is the translation vector, and the internal parameter matrix of camera 2 where f x and f y are the focal point coordinates, and c x and c y are the optical center pixel coordinates. The calibration of the external parameters including the rotation matrix and the translation vector determines the position of camera 2 in the search space.
[0103] S44. According to the mapping relationship, convert the feature data into a heat map through a color mapping scheme, where the color mapping scheme uses colors with different brightness levels to represent different feature data respectively.
[0104] Before the calibration in step S4, for the sound source position obtained by positioning in step S3, the sound source position before calibration is obtained by three-dimensional coordinate conversion through the following formula: The illustration of the sound source position before calibration and the actual sound source position is as Figure 5 ; The illustration of the sound source position after calibration in step S4 and the actual sound source position is as Figure 6 , from Figure 5 and Figure 6 it can be seen that after the calibration of the embodiment, high-accuracy sound source localization can be achieved.
[0105] This method for detecting abnormal noise in a new energy vehicle cockpit based on sound-image fusion can achieve more efficient and intuitive detection of abnormal noise in the new energy vehicle cockpit, can significantly improve the accuracy of sound source localization, can greatly improve the efficiency of detecting abnormal noise in the new energy vehicle cockpit, and can be widely applied to noise monitoring and fault diagnosis of new energy vehicle cockpits. By combining acoustic signals and image data, through a multi-modal fusion method, it makes up for the deficiencies of single-modal detection, uses a microphone array to achieve spatial localization of cockpit noise, combines with a camera 2 to obtain dynamic images inside the cockpit, realizes intuitive capture of abnormal noise sources, and jointly analyzes acoustic features and image information, which can significantly improve the accuracy and real-time performance of abnormal noise detection, thereby enhancing the user experience and vehicle safety.
[0106] This method for detecting abnormal noise in a new energy vehicle cockpit based on sound-image fusion first collects audio data of each channel of the microphone array and performs preprocessing; uses an abnormal monitoring algorithm to analyze the data to determine whether there is abnormal noise; uses an improved SRC algorithm to locate the abnormal noise source; at the same time, obtains real-time video data of the camera 2; finally, calibrates the position of the abnormal noise with the image information to achieve visual localization of the abnormal noise. The present invention can be widely applied to noise monitoring and fault diagnosis of new energy vehicle cockpits, and significantly improves the accuracy and working efficiency of sound source localization.
[0107] As Figure 4 , the embodiment also provides a device for detecting abnormal noise in a new energy vehicle cockpit based on sound-image fusion for implementing the method of any one of the above, which includes a multi-channel microphone array 1, an image acquisition module, a data processing module, a display module, and a power supply module for power supply,
[0108] Multi-channel microphone array 1: Installed on the top of the new energy vehicle cockpit to collect multi-channel audio data;
[0109] Image acquisition module: The camera 2 installed on the top of the cockpit of the new energy vehicle collects video data;
[0110] Data processing module: After preprocessing the multi-channel audio data, the preprocessed audio data is obtained; The preprocessed audio data is identified and detected for abnormalities through an anomaly monitoring model based on a neural network. If there are abnormalities, the type of abnormal noise is determined and the abnormal noise signal is extracted until the detection of the preprocessed audio data is completed; Through an improved spatial shrinkage fast method based on random search, namely the improved SRC method, the abnormal noise is located to obtain the location of the abnormal noise; The location of the extracted abnormal noise is calibrated with the image in the video data, and the noise source is displayed in the image in the form of a sound field heat map to obtain the sound-image fused sound field map;
[0111] Display module: Used to display the sound-image fused sound field map.
[0112] Multiple Monte Carlo experiments are carried out on the SRC method and the improved SRC method. The positioning test results of the four-element microphone are shown in Table 1, and the positioning test results of the eight-element microphone are shown in Table 2:
[0113] Table 1 Positioning test results of four-element microphone
[0114]
[0115] Table 2 Positioning test results of eight-element microphone
[0116]
[0117] It can be seen from the data in Table 1 and Table 2 that in the case of the same number of microphones, the improved SRC method reduces the regional shrinkage time, the number of iterations, and the number of calculations of the SRP value by nearly half compared with the original SRC method, and obtains a positioning accuracy not lower than that of the original SRC algorithm.
[0118] This abnormal noise detection device for the cockpit of a new energy vehicle based on sound-image fusion is used for monitoring abnormal noise in the cabin of a new energy vehicle. It simultaneously collects audio and video data, detects whether there are any abnormalities through an anomaly monitoring algorithm. If not, it maintains the monitoring state. Otherwise, it locates the sound source through a sound source localization algorithm and finally performs visual localization.
[0119] The above describes the main technical features, basic principles and related advantages of the present invention. For those skilled in the art, it is obvious that the present invention is not limited to the details of the above-mentioned exemplary specific embodiments, and can be implemented in other specific forms without departing from the concept or basic features of the present invention.
Claims
1. A method for detecting abnormal noise in the cockpit of a new energy vehicle based on sound and image fusion, characterized in that: The following steps are included: S1. Install a multi-channel microphone array and a camera on the top of the new energy vehicle cockpit, synchronously collect multi-channel audio data and video data through the multi-channel microphone array and the camera, and pre-process the acquired multi-channel audio data to obtain pre-processed audio data; S2, the pre-processed audio data is identified and detected for abnormality through an abnormality monitoring model based on a neural network. If an abnormality exists, the abnormal noise type is determined and the abnormal noise signal is extracted, and the next step S3 is entered; Otherwise, repeat step S2 until the detection of the preprocessed audio data is completed; S3, locating the abnormal noise by using an improved random search-based space shrinkage fast method, namely, an improved SRC method, to obtain the abnormal noise location; S4, calibrating the extracted abnormal noise location and the image in the video data, displaying the noise source in the image in the form of a sound field heat map, and obtaining a sound field map after sound and image fusion.
2. The abnormal noise detection method for a new energy vehicle cabin based on sound and image fusion as claimed in claim 1, characterized in that: In step S1, a multi-channel microphone array uses a plurality of microphones arranged in a plurality of circular arrays from inside to outside.
3. The abnormal noise detection method for a new energy vehicle cabin based on sound and image fusion as claimed in claim 1, characterized in that: In step S2, the abnormality type includes at least one of the following: abnormal vibration, suspected abnormal vibration, partial discharge, suspected partial discharge and abnormal noise.
4. The abnormal noise detection method for a new energy vehicle cabin based on sound and image fusion as claimed in claim 1, characterized in that: In step S2, the neural network-based anomaly monitoring model includes a time-frequency conversion module, a convolutional neural network (CNN), a space-time conversion module, a transformer model, and a classification head. Time-frequency conversion module: performs time-frequency conversion on the input pre-processed audio data to obtain a Mel-spectrogram; Convolutional neural network (CNN): extracts spectral features from the input Mel-spectrogram and outputs the feature map to the spatiotemporal conversion module; Spatiotemporal conversion module: reshapes the feature map into a time series and outputs it to the transformer model; Transformer model: It performs position encoding, dynamic feature association, and multi-scale feature fusion on the input time series to capture the temporal dependency and output a feature vector containing local features and global context information to the classification head. Classification head: Outputs the classification result based on the input feature vector containing local features and global context information.
5. The abnormal noise detection method for a new energy vehicle cabin based on sound and image fusion according to any one of claims 1 to 4, characterized in that: Step S3, specifically, S31, performing signal enhancement and denoising on the preprocessed audio data obtained in step S1 by using spectral subtraction and Wiener filtering; S32, perform short-time Fourier transform on each frame signal to obtain a frequency domain signal X m (f), where m is the microphone index and f is the frequency; S33, calculate the phase transformation weighted cross-correlation R of each pair of microphone signals ij (τ): Where τ represents the time delay difference between the sound wave reaching the i-th and j-th microphones in a pair of microphones in the microphone array, and X i (f) represents the spectrum of the i-th microphone signal, represents the complex conjugate of the j-th microphone signal spectrum, and e is a natural constant; S34, dividing the initial search space into a number of sound source planes from near to far along the normal direction of the multi-channel microphone array at a spacing L, and discretizing each sound source plane into grid points in a mesh shape and determining the three-dimensional spatial position coordinates of each grid point, and each grid point forms a grid point set G; S35, calculate the theoretical delay τ from each grid point p in the grid point set G to the i-th and j-th microphones in each pair of microphones ij (p); S36. Build a positioning model based on compressed sensing technology: z=Ψθ Among them, the estimated value z of the distance difference between the sound source and the two microphones in the microphone pair calculated according to the received signal of each microphone pair is a C×1 vector, and C is the total number of microphone pairs in the multi-channel microphone array; Ψ is a C×Q dictionary matrix, and each column element of the dictionary matrix is the theoretical value of the assumed sound source distance difference of each microphone pair at the grid point position in the search space; the spatial position estimation parameter θ of the Q grid points in the search space is a Q×1 vector; S37. Randomly select J grid points in the grid point set G, and calculate the superposition value SRP of the cross power spectra of all microphone pairs receiving signals at each grid point: Select N with the largest superposition value SRP of the cross power spectrum from them k grid points, for N k The grid points construct a spatial dictionary Ψ according to the time delay relationship k Then, the orthogonal matching pursuit algorithm OMP is used to solve the position estimation parameter θ in the positioning model based on compressed sensing technology, and the grid points whose position estimation parameter θ is greater than the set threshold are obtained as new grid points and the number of new grid points N' is obtained. k , update N k =Number of new grid points N' k , shrink the search area to include these N k A rectangular region with 100 points is formed, and a new grid point set G' is formed with the grid points of the rectangular region, and the grid point set G=G' is updated; step S37 is iterated until the search space is shrunk to the minimum search space, that is, the search space with only one grid point, and the next step S38 is entered; S38. After the iteration is completed, the grid point with the largest superposition value of the cross-power spectrum is used as the abnormal noise location position.
6. The abnormal noise detection method for a new energy vehicle cabin based on sound and image fusion according to any one of claims 1 to 4, characterized in that: Step S4 is specifically, S41, extracting characteristic data of abnormal noise signals including energy data, power spectrum density and frequency bandwidth; S42, normalizing the feature data, pairing the normalized feature data with the abnormal noise location, and constructing a corresponding mapping relationship; S43, using a checkerboard calibration method to determine the camera's internal parameters including focal coordinates and optical center pixel coordinates and external parameters including a rotation matrix and a translation vector, establish a mapping relationship between the image pixel coordinates and the real three-dimensional space, and calibrate the abnormal noise location with the camera image pixel coordinates; S44. According to the mapping relationship, the feature data is converted into a heat map through a color mapping scheme, wherein the color mapping scheme uses colors of different brightness to represent different feature data.
7. The abnormal noise detection method for a new energy vehicle cabin based on sound and image fusion as claimed in claim 6, characterized in that: In step S43, the abnormal noise location position is calibrated with the camera image pixel coordinates. Specifically, the abnormal noise location position P s =(X,Y,Z) converted to camera image pixel coordinates (u,v): Among them, R is the rotation matrix, t is the translation vector, and the camera's intrinsic parameter matrix Among them, f x , f y is the focal coordinate, c x , c y is the pixel coordinate of the optical center.
8. A device for detecting abnormal noise in the cabin of a new energy vehicle based on sound and image fusion, which implements the method of any one of claims 1 to 7, characterized in that: It includes a multi-channel microphone array, an image acquisition module, a data processing module, a display module and a power supply module. Multi-channel microphone array: installed on the top of the new energy vehicle cabin to collect multi-channel audio data; Image acquisition module: The camera installed on the top of the new energy vehicle cabin collects video data; Data processing module: after preprocessing the multi-channel audio data, obtaining the preprocessed audio data; The preprocessed audio data is identified and detected for abnormalities through an abnormality monitoring model based on a neural network. If an abnormality exists, the abnormal noise type is determined and the abnormal noise signal is extracted until the detection of the preprocessed audio data is completed; the abnormal noise is located through an improved random search-based spatial shrinkage fast method, namely an improved SRC method, to obtain the abnormal noise location; the extracted abnormal noise location and the image in the video data are calibrated, and the noise source is displayed in the image in the form of a sound field heat map to obtain a sound field map after sound and image fusion; Display module: used to display the sound field image after the fusion of sound and image.
Citation Information
Patent Citations
Sound source positioning system based on distributed microphone array
CN107102296A
Abnormal sound event identification method based on MFCC+MP fusion characteristic
CN109785857A
GIS equipment mechanical fault detection system and method based on acoustic imaging
CN113267330A
Acoustic imaging positioning system and method for intelligent monitoring of transformer substation domain faults
CN114414963A
Track abnormal condition detection method based on noise and image information fusion
CN114841966A
Cited By
Wind turbine generator fault diagnosis method based on audio and video identification
CN121382552A