A method, device, electronic device and storage medium for identifying abnormal noises in a passenger car

Through the combination of data expansion and parallel deep learning network, the MFCC characteristics in vehicle abnormal noise are extracted, which solves the problems of lack of data and high computing costs in vehicle abnormal noise diagnosis, and achieves efficient abnormal noise recognition.

CN114065809BActive Publication Date: 2025-06-27ZHEJIANG GEELY HLDG GRP CO LTD +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202111293794.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-03
Publication Date
2025-06-27
Estimated Expiration
2041-11-03

AI Technical Summary

Technical Problem

The prior art has problems such as lack of data, high computing costs and complexity of classification models in vehicle noise diagnosis, resulting in low diagnostic efficiency and uncertain results.

Method used

The limited data set is augmented through data augmentation technology, and the spatial and time-series information of MFCC features is extracted using the parallel deep learning network mechanism, and combined with the convolutional neural network and the Transformer encoder stack for training and classification.

Benefits of technology

It effectively solves the problem of data scarcity, reduces calculation costs, improves diagnostic efficiency, and realizes accurate identification of vehicle abnormal noises.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114065809B_ABST
    Figure CN114065809B_ABST
Patent Text Reader

Abstract

The present invention relates to a method, device, electronic device and storage medium for identifying abnormal noises in passenger cars. The method includes: Step S1) For a limited data set, use data augmentation technology to complete the augmentation of the data set; Step S2) Use a parallel deep learning network mechanism to complete the extraction of spatial and temporal information of MFCC features, and complete training and classification; Step S3) Use the trained model to identify abnormal noise data of passenger cars. Compared with the prior art, the present invention has the advantages of solving the problem of data scarcity in the field of vehicle abnormal noise identification and greatly reducing the computational load of the network, etc.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to vehicle abnormal noise diagnosis technology, and in particular to a method and device for identifying abnormal noise of a passenger vehicle, an electronic device, and a storage medium. Background Art

[0002] There are a large number of rotating mechanical components and interior components in the vehicle system. Due to processing errors or physical factors during the working process, these components will cause phenomena such as wear and displacement of the components. If the machine components are damaged due to the above reasons during operation, it will directly relate to property and even safety losses. At present, relying on a large number of effective data sets, fault diagnosis technologies based on various signals have been widely applied in factory assembly lines and nuclear factories. However, for the abnormal noise problems existing in the vehicle system, the lack of relevant data sets restricts the development of diagnostic technologies in this field. Most vehicle enterprises and after-sales service institutions still use the method of combining artificial auscultation in various specific experimental environments. This method not only has low efficiency but also heavily relies on the professional knowledge and diagnostic experience of the detection personnel, and the final diagnostic results may also vary from person to person. Therefore, it is also necessary to develop a vehicle abnormal noise diagnosis method based on easily obtainable signals.

[0003] After retrieval, Chinese Patent Publication No. CN105841797A proposes a method for detecting abnormal noise of a window motor based on MFCC and SVM. It extracts the MFCC features of the collected abnormal noise audio, uses SVM as a classifier model to complete training and classification. In order to improve the classification accuracy, during the windowing process of extracting MFCC, the original Hanning window is replaced with a self-convolution Hanning window, and the artificial bee colony algorithm is applied to optimize the model, making this method have the advantages of strong practicability and high reliability.

[0004] CN109949823A provides a method for identifying abnormal noise in a vehicle based on DWPT-MFCC and GMM. It also extracts MFCC features, and introduces wavelet transform during the extraction process to optimize the MFCC features to obtain DWPT-MFCC, making the features more illustrative. Subsequently, the Gaussian mixture model (GMM) is combined to complete the training and classification tasks.

[0005] CN112149498A proposes an online intelligent identification system and method for abnormal noise of complex automotive components. Its core lies in using the FBank (WT-FBank) coefficient map added with wavelet transform as the input of the convolutional neural network (for obtaining deep abnormal noise features), and then performing dimensionality reduction processing and combining with an SVM classifier to complete the classification task. This method combines the map features with the excellent spatial information extraction ability of CNN and the classification performance of SVM, making this method more practical.

[0006] CN 112735468A is a method for detecting abnormal noise of automotive seat motors based on MFCC. The specific details are as follows: Extract the MFCC features of the abnormal noise of the seat (without optimization), and use the BP neural network to complete the training and classification tasks. It saves computational costs in the feature extraction part and utilizes the BP neural network to complete the recognition task faster in the actual application process.

[0007] Among them, in the process of extracting MFCC features in CN105841797A, CN109949823A, and CN112149498A, additional signal processing knowledge is respectively added, such as wavelet transform, second-order self-convolution Hanning window, etc. This undoubtedly requires more prior knowledge and will increase the computational cost of the computer. As the amount of data increases, the above methods will be difficult to handle. In addition, the SVM classifiers adopted by CN105841797A and CN112149498A have many limitations: (1) The input features need to be one-dimensional data, and (2) It is a typical binary classification mathematical model. If you want to complete the multi-classification task, you need to establish an SVM model between any two samples of all classifications, increasing the computational load.

[0008] The GMM model adopted by CN109949823A does not utilize the context information of the time series signal and cannot learn deep non-linear feature transformations; the BP neural network algorithm adopted by CN112735468A has a risk of tending to local minimization during the training process. Moreover, it is trained in the form of the gradient descent method, resulting in low training efficiency and unable to fine-tune the network purposefully, resulting in low development efficiency. In addition, it depends on the quality and scale of the training samples. For example, it is difficult to train a classification model with strong generalization ability with small training samples. However, there is no scientific and large dataset in this field. Summary of the Invention

[0009] The purpose of the present invention is to provide a method, device, electronic device, and storage medium for identifying abnormal noises in passenger cars to overcome the defects of the above-mentioned existing technologies.

[0010] The purpose of the present invention can be achieved through the following technical solutions:

[0011] According to the first aspect of the present invention, a method for identifying abnormal noises in passenger cars is provided. The method includes:

[0012] Step S1) For the limited dataset, use data augmentation technology to complete the augmentation of the dataset;

[0013] Step S2) Use the parallel deep learning network mechanism to complete the extraction of the spatial and temporal information of the MFCC features, and complete the training and classification;

[0014] Step S3) Use the trained model to identify abnormal noise data of passenger cars.

[0015] As a preferred technical solution, in the step S1), the limited data set is the abnormal noises collected from various parts, including gear whine, reducer knocking, gear impact, valve system abnormal noise, armrest box abnormal noise, glove box abnormal noise, and seat abnormal noise.

[0016] As a preferred technical solution, the data augmentation in the step S1) includes audio cropping and data enhancement.

[0017] As a preferred technical solution, the specific audio cropping is as follows: Cut an audio of a certain duration into several small time blocks to increase the quantity.

[0018] As a preferred technical solution, the data enhancement includes:

[0019] Time stretching, on the premise of unchanged pitch, change the speed of the audio signal, and change the speed of the original audio by setting the stretching parameter ν. When v ∈ (1, +∞) or v ∈ (0, 1), theoretically it means accelerating or decelerating the audio rate to v times that of the original audio;

[0020] Time shifting, keep the pitch unchanged, and shift a set distance in the time domain range. The shifting parameter σ can be set to a positive or negative value, representing forward or backward shifting of the audio data respectively;

[0021] Adding noise, adding background noise to the original audio data;

[0022] Pitch correction, on the premise of unchanged sound speed, change the pitch of the original audio, and move the pitch up or down by several steps by setting the correction parameter ρ.

[0023] As a preferred technical solution, in the parallel deep learning network mechanism of the step S2), set two parallel CNN lines CNN1 and CNN2 for extracting spatial information, and an encoder stack line Transformer for extracting temporal information.

[0024] As a preferred technical solution, for the input 2D features, three convolutional layers are set in CNN1, using 3×3 micro convolutional kernels; three convolutional layers are also set in CNN2, and the 3×3 convolutional kernels are replaced by 3×1 and 1×3 asymmetric convolutional kernels;

[0025] In the Transformer, first perform pooling on the input feature map, and then use an encoder stack composed of several encoder units in series to capture temporal information; then combine and concatenate the spatial and temporal information extracted by the three parallel lines and then perform a linear transformation to the fully connected layer, and finally output the probabilities of each noise type through a softmax classifier.

[0026] According to the second aspect of the present invention, there is provided a device for identifying abnormal noises in a passenger car, the device comprising:

[0027] A dataset augmentation module, configured to augment a limited dataset by using data augmentation techniques for the limited dataset;

[0028] An identification model construction module, configured to extract spatial and temporal information of MFCC features by using a parallel deep learning network mechanism, and complete training and classification;

[0029] A data identification module, configured to identify abnormal noise data of a passenger car by using the trained model.

[0030] According to the third aspect of the present invention, there is provided an electronic device, comprising a memory and a processor, wherein a computer program is stored on the memory, and the method as described above is implemented when the processor executes the program.

[0031] According to the fourth aspect of the present invention, there is provided a computer-readable storage medium, on which a computer program is stored, and the method as described above is implemented when the program is executed by a processor.

[0032] Compared with the prior art, the present invention has the following advantages:

[0033] 1) The data augmentation method of the present invention well solves the problem of data scarcity in the field of vehicle abnormal noise identification;

[0034] 2) For the classification and identification task of time series signals, the parallel mechanism adopted by the present invention can simultaneously take into account the spatial and context information of the signals, and alleviate the computational cost brought by deeper networks. Description of the Drawings

[0035] Figure 1 is a working flowchart of the method of the present invention;

[0036] Figure 2 is a schematic structural diagram of the device of the present invention;

[0037] Figure 3 is a flowchart of MFCC feature extraction of the present invention;

[0038] Figure 4 is a schematic structural diagram of a simplified CNN structure containing one typical layer and two fully connected layers;

[0039] Figure 5 is a schematic diagram of an encoder stack;

[0040] Figure 6 is a schematic diagram of a parallel deep learning network architecture;

[0041] Figure 7It is the overall flowchart for abnormal sound recognition;

[0042] Figure 8 It is the loss curve graph during the training process;

[0043] Figure 9 It is the schematic diagram of the confusion matrix for noise types. Specific implementation manners

[0044] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0045] As Figure 1 shown, a method for abnormal sound recognition of a passenger car, the method includes:

[0046] Step S1) For the limited data set, use data augmentation technology to complete the augmentation of the data set;

[0047] Step S2) Use the parallel deep learning network mechanism to complete the extraction of the spatial and temporal information of the MFCC features, and complete the training and classification;

[0048] Step S3) Use the trained model to identify the abnormal sound data of the passenger car.

[0049] The specific process of the present invention is as follows:

[0050] (i) First, apply data augmentation methods to the abnormal noises collected from various parts (including gear howling, reducer knocking, gear impact, valve system abnormal sound, armrest box abnormal sound, glove box abnormal sound, seat abnormal sound):

[0051] Audio cropping: Cut an audio of a certain duration into several small time blocks to achieve the purpose of increasing the quantity. We analyzed the waveform diagrams of each abnormal noise signal, determined the time interval between abnormal frequencies, and finally determined the small time block of t = 3s. To avoid information loss between adjacent time blocks during the cropping process, we set the cropping step size of stride = 2s, that is, there will be 1s of information overlap between two adjacent time blocks. If the last small time block is less than 3s, it is still retained.

[0052] Time stretching: Without changing the pitch, change the speed of the audio signal. By setting the stretching parameter v to change the speed of the original audio, when v ∈ (1, +∞) or v ∈ (0, 1), it theoretically means accelerating or decelerating the audio rate to v times that of the original audio. To prevent audio distortion, a set of stretching parameters v = {0.8, 2} is set for each audio data in this article;

[0053] Time shift: Keeping the pitch unchanged, shift a certain distance in the time domain. The shift parameter σ can be set as a positive or negative value, representing the forward or backward shift of the audio data respectively (positive for forward, negative for backward). In this paper, the shift parameter σ = fs / 2 is set for each audio data, where fs represents the sampling frequency (48KHz). In this paper, σ = {-fs / 2, fs / 2} is taken.

[0054] Adding noise: Adding noise is a commonly used enhancement technique in natural language processing and image recognition. Adding noise in the field of sound recognition means adding background noise, such as Gaussian noise, ambient sound, etc., to the original audio data. In this paper, Gaussian white noise with a mean of 0 and a standard deviation of 1 is selected to be added.

[0055] Pitch correction: On the premise of constant sound speed, change the pitch of the original audio. In fact, the change in pitch does not affect the label of the fault feature. Therefore, pitch correction will be a very beneficial enhancement for this research. By setting the correction parameter ρ, the pitch is moved up or down by several steps (in units of half - tone, ρ being positive represents upward movement, and vice versa for downward movement). In this paper, ρ = {-6, 3, 6} is taken.

[0056] (ii) Extract MFCC features for the abnormal sound obtained in step (i), and the extraction flow chart is as Figure 3 shown;

[0057] The specific description of each step is as follows:

[0058] (1) Pre - emphasis:

[0059] The noise data undergoes pre - emphasis to balance the spectrum and improve the signal - to - noise ratio. For the time - domain signal X(n), the output after pre - emphasis is:

[0060] X ′ (n) = X(n) - αX(n - 1)(1)

[0061] where α is the filter coefficient.

[0062] (2) Frame division and windowing

[0063] The signal is divided into several short time periods using a Hamming window. Each short time period is called an analysis frame. In this way, it can be considered that within the analysis frame, the frequency of the signal is stable. To maintain the continuity of the signal and avoid signal distortion, usually adjacent analysis frames overlap. The overlapping part is called the frame shift. The frame - division operation is represented by the following equation:

[0064] X i ′ (n) = ω(n)·X(n)(2)

[0065]

[0066] where X i ′ (n) represents the i-th frame signal after frame division, ω(n) represents the Hamming window function, and N is the window length.

[0067] (3) Fourier transform and power spectrum calculation

[0068] Perform Fourier transform on each analysis frame to convert the time-domain signal into a frequency-domain power distribution, and then calculate the power spectrum through the following formula:

[0069]

[0070] where P i (k) represents the power spectrum corresponding to the i-th frame, and K represents the length of the Fourier transform.

[0071] (4) Mel filter bank

[0072] The power spectrum passes through a Mel-scale triangular filter bank to obtain a frequency range that conforms to the human ear's perception. The relationship between the actual frequency of the audio and the Mel-scale frequency is as follows:

[0073]

[0074] (5) Logarithmic energy

[0075] Take the logarithm of the output of each filter to obtain the logarithmic energy. We name the logarithmic energy output by the filter bank as Logfbank, and the logarithmic energy output is:

[0076]

[0077] where H m (k) is the definition function of the triangular filter bank, M is the number of filters, and f(m) represents the center frequency.

[0078] (6) Discrete cosine transform

[0079] Remove the high correlation between FBank features to obtain more abstract features (MFCC)

[0080]

[0081] where n represents the order of MFCC.

[0082] (iii) Construction of a parallel deep learning network architecture:

[0083] (1) Convolutional neural network

[0084] Convolutional neural networks are still the mainstream idea in computer vision. The reason is that they can share weight parameters and can establish sparse connections with relatively few weight parameters. The above characteristics make the network easier to optimize and also reduce the risk of overfitting.

[0085] A convolutional network consists of several typical layers. A typical layer generally includes a convolutional layer and a pooling layer. Among them, the convolutional stage sweeps local information by performing a convolutional operation on the input tensor using a small convolutional kernel.

[0086] It is also necessary to use a non-linear activation function (usually the ReLU function) to accelerate the feature learning ability. The pooling layer is used to extract important local information and improve the computational efficiency. Generally, max pooling and average pooling are used. Finally, the classification function is achieved through the fully connected layer. Figure 4 Shows a simplified CNN structure containing one typical layer and two fully connected layers.

[0087] (2) Transformer Encoder Stack

[0088] Transformer has gradually replaced the traditional Seq2Seq model represented by RNN. It uses an encoder stack composed of several encoders in series to replace the encoder with RNN as the core. As Figure 5 shown, each encoder in the encoder stack is composed of a multi-head attention (MHA) unit and a feed-forward neural network unit in series, and each unit is attached with a residual connection. The reason for adding the residual connection is that the distribution of parameters may change continuously during training. The residual connection can normalize the feature parameters of the network, so that more effective gradients can be learned and the model is easier to learn.

[0089] (3) Parallel Architecture Construction

[0090] For complex input features, the convolutional neural network makes the output approximate a non-linear function that can match the signal features as much as possible through forward and backward propagation, so as to obtain the information of the input features at the spatial scale. The Transformer encoder captures the hidden relationship between each time series of the continuous signal through the multi-head attention mechanism plus the residual connection, so as to obtain the time series information of the continuous features of the input. In order to improve the diagnostic ability of the diagnostic model and intend to obtain the spatial information and time series relationship information of the signal at the same time, this paper sets up an architecture in which the deep convolutional network and the Transformer encoder stack work simultaneously to improve the diagnostic performance.

[0091] Figure 6The proposed parallel architecture is shown. In this architecture, the input is the 2D MFCC extracted in step (ii), which is a feature matrix of 40×282 (MFCC feature order × time dimension). The feature is provided with two parallel CNN lines (CNN1, CNN2) for extracting spatial information and an encoder stack line (Transformer) for extracting temporal information. For the input 2D features, we set three convolutional layers in CNN1, using a 3×3 micro convolutional kernel. In CNN2, we also set three convolutional layers. Different from CNN1, the 3×3 convolutional kernel is replaced by 3×1 and 1×3 asymmetric convolutional kernels, which not only greatly reduces the calculation parameters but also can obtain more additional spatial information. In addition, a pooling operation is set at the end of each convolutional layer to reduce the number of parameters and thus speed up the training. In the Transformer, first, pooling (downsampling) is performed on the input feature map, and then an encoder stack composed of several encoder units in series is used to capture temporal information. Finally, the spatial and temporal information extracted by the three parallel lines is fused and then linearly transformed to the fully connected layer, and finally the softmax function is used to output the probabilities of each noise type. The parallel lines can enable the CNN and the Transformer to work together, avoiding the computational load brought by the deep network. The network also adds a batch normalization (BN) layer, which has an obvious gain for the network training efficiency and optimizing the gradient problem. For the gradient problem in backpropagation, the stochastic gradient descent (SGD) optimization technique is selected, and the optimization parameters in SGD are set as follows: learning rate = 0.01, weight decay coefficient = 0.001, momentum = 0.8. In addition, in the convolutional layer, in order to avoid losing the edge information in the feature map, we uniformly use zero-padding convolution. And after the pooling stage of each convolutional layer

[0092] we all adopt the Dropout technique, which avoids the problem of poor generalization ability of the model due to overfitting by randomly discarding parameters. In addition, the cross-entropy loss function is used to calculate the network cost.

[0093] The diagnostic process of the present invention: The overall process of abnormal sound recognition is as Figure 7 shown. For the accuracy of the network's later training, the data augmentation method is set after dividing the data set (to avoid the augmented data from the same original audio being distributed to the training set, validation set, and test set at the same time).

[0094] It should be noted that in order to verify the classification performance of the model, it is necessary to visualize the loss curve of the model training and evaluate the performance of the test set in the trained model, which is reflected by the confusion matrix, as Figure 8 and 9 shown: Figure 8For the loss curve during the training process, it can be observed that the model loss can converge rapidly at the beginning, but then starts to converge slowly. After 200 iterations, both the training loss and the validation loss basically stabilize. To further illustrate the diagnostic performance of the architecture for seven types of noise, we present the confusion matrix of the noise types and perform normalization, as Figure 9 shown. The model makes misjudgments between abnormal noises of valve system rattling and whistling, and between abnormal vibrations of car seats and abnormal vibrations of glove boxes, but the probability is about 4%. The specific evaluation indicators are calculated as follows: Accuracy = 0.9831, Precision = 0.9760, Recall = 0.9824, F1score = 0.9787.

[0095] The above is the introduction of the method embodiment. Next, through the device embodiment, the solution of the present invention will be further described.

[0096] As Figure 2 shown, the device of the present invention includes:

[0097] The dataset augmentation module 1 is used to complete the augmentation of the dataset for a limited dataset by using data augmentation techniques;

[0098] The recognition model construction module 2 is used to complete the extraction of the spatial and temporal information of MFCC features by using a parallel deep learning network mechanism, and complete training and classification;

[0099] The data recognition module 3 is used to identify abnormal noise data of passenger cars by using the trained model.

[0100] Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the described modules can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated here.

[0101] The electronic device of the present invention includes a central processing unit (CPU), which can execute various appropriate actions and processes according to computer program instructions stored in a read-only memory (ROM) or computer program instructions loaded from a storage unit into a random access memory (RAM). In the RAM, various programs and data required for device operation can also be stored. The CPU, ROM, and RAM are connected to each other through a bus. The input / output (I / O) interface is also connected to the bus.

[0102] Multiple components in the device are connected to the I / O interface, including: an input unit, such as a keyboard, a mouse, etc.; an output unit, such as various types of displays, speakers, etc.; a storage unit, such as a disk, an optical disc, etc.; and a communication unit, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit allows the device to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0103] The processing unit executes the various methods and processes described above, such as methods S1 - S3. For example, in some embodiments, methods S1 - S3 can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device via the ROM and / or the communication unit. When the computer program is loaded into the RAM and executed by the CPU, one or more steps of methods S1 - S3 described above can be executed. Alternatively, in other embodiments, the CPU can be configured to execute methods S1 - S3 by any other suitable means (e.g., by means of firmware).

[0104] The functions described above herein can be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: Field Programmable Gate Arrays (FPGAs), Application Specific Integrated Circuits (ASICs), Application Specific Standard Products (ASSPs), Systems on Chip (SOCs), Complex Programmable Logic Devices (CPLDs), and so on.

[0105] The program code for implementing the method of the present invention can be written in any combination of one or more programming languages. These program codes can be provided to the processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0106] In the context of the present invention, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0107] As described above, the foregoing are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily conceive of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.

Claims

1. A method for identifying abnormal noises in a passenger car, characterized in that, The method includes: Step S1) For the limited dataset of abnormal noises in various parts of passenger cars, use data augmentation technology to complete the augmentation of the dataset; Step S2) Extract the MFCC features of the augmented dataset; Step S3) In the parallel deep learning network mechanism, set two parallel CNN lines CNN1 and CNN2 for extracting spatial information, and an encoder stack line Transformer for extracting temporal information, so as to use the parallel deep learning network mechanism to complete the extraction of spatial and temporal information of the MFCC features of the augmented dataset, and complete training and classification; Step S4) Use the trained model to identify abnormal noise data of passenger cars.

2. The method for identifying abnormal noises in a passenger car according to claim 1, characterized in that, The limited dataset in the step S1) is the abnormal noises collected in various parts, including gear whine, reducer knocking, gear impact, valve system abnormal noise, armrest box abnormal noise, glove box abnormal noise, and seat abnormal noise.

3. A method for identifying abnormal noises in a passenger car according to claim 1 or 2, characterized in that The data augmentation in the step S1) includes audio cropping and data enhancement.

4. The method for identifying abnormal noises of a passenger car according to claim 3, wherein The specific audio cropping is: cutting an audio of a certain duration into several small time blocks to increase the quantity.

5. A method for identifying abnormal noises in a passenger car according to claim 3, characterized in that The data enhancement includes: Time stretching, without changing the pitch, change the speed of the audio signal, and change the speed of the original audio by setting the stretching parameter ν. When v ∈ (1, +∞) or v ∈ (0, 1), it theoretically means accelerating or decelerating the audio rate to v times that of the original audio; Time shifting, keeping the pitch unchanged, shift a set distance in the time domain range. The shifting parameter σ can be set to a positive or negative value, representing forward or backward shifting of the audio data respectively; Adding noise, adding background noise to the original audio data; Pitch correction, without changing the sound speed, change the pitch of the original audio, and move the pitch up or down by several steps by setting the correction parameter ρ.

6. The method for identifying abnormal noises in a passenger car according to claim 1, wherein For the input 2D features, three convolutional layers are set in CNN1, using 3×3 micro convolutional kernels; three convolutional layers are also set in CNN2, and the 3×3 convolutional kernels are replaced by 3×1 and 1×3 asymmetric convolutional kernels; In the Transformer, first perform pooling on the input feature map, and then use an encoder stack composed of several encoder units in series to capture temporal information; then combine and concatenate the spatial and temporal information extracted by the three parallel lines, and then perform a linear transformation to the fully connected layer, and finally output the probabilities of each noise type through a softmax classifier.

7. An abnormal noise recognition device for a passenger car, characterized in that, The device includes: A dataset augmentation module, which is used to complete the augmentation of the dataset for the limited dataset of abnormal noises in various parts of passenger cars by using data augmentation technology; A feature extraction module, which is used to extract the MFCC features of the augmented dataset; An identification model construction module, which is used to set two parallel CNN lines CNN1 and CNN2 for extracting spatial information, and an encoder stack line Transformer for extracting temporal information in the parallel deep learning network mechanism, so as to use the parallel deep learning network mechanism to complete the extraction of spatial and temporal information of the MFCC features of the augmented dataset, and complete training and classification; A data recognition module, configured to recognize abnormal noise data of a passenger car by using a trained model.

8. An electronic device, comprising a memory and a processor, wherein a computer program is stored on the memory, characterized in that, When the processor executes the program, the method described in any one of claims 1 to 6 is implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, the method described in any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Window motor abnormal noise detection method and apparatus based on MFCC and SVM

    CN105841797A

  • Online intelligent identification system and method for abnormal sounds of complex parts of automobile

    CN112149498A

  • Automobile seat motor abnormal noise detection method based on MFCC

    CN112735468A

  • Method for identifying abnormal sounds in vehicle on basis of DWPT-MFCC and GMM

    CN109949823A