Vehicle microphone fault identification method and device, electronic equipment and storage medium

By extracting and normalizing the sample noise data of vehicle microphones, and combining it with the scene recognition model of the attention module, the problem of low accuracy in vehicle microphone fault identification was solved, and a more efficient fault identification effect was achieved.

CN119785825BActive Publication Date: 2026-02-06HEBEI CHUGUANG AUTO PARTS CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202411835434.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-13
Publication Date
2026-02-06
Estimated Expiration
2044-12-13

AI Technical Summary

Technical Problem

Existing technologies for vehicle microphone fault identification have low accuracy, especially in complex and variable outdoor environments, resulting in poor background noise filtering.

Method used

By acquiring sample noise data from vehicle microphones and preprocessing it, training and testing sets are divided. A first scene recognition model, consisting of a feature extraction module, a normalization module, and an attention module, is used for training and testing to identify whether the vehicle microphone is malfunctioning. By combining feature extraction, normalization, and feature channel enhancement, a preset optional sound database is matched to determine the target noise scene.

Benefits of technology

It improves the accuracy of vehicle microphone fault identification, reduces the probability of identification errors caused by environmental complexity, and achieves more accurate fault identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119785825B_ABST
    Figure CN119785825B_ABST
Patent Text Reader

Abstract

The application provides a vehicle microphone fault recognition method and device, electronic equipment and a storage medium, the method comprising: obtaining sample noise data and preprocessing and dividing; inputting the divided training set and test set into a first scene recognition model for training and testing; obtaining first to-be-tested noise data through a vehicle microphone and preprocessing; inputting the first to-be-tested noise data into the first scene recognition model, sequentially passing through a feature extraction module, a normalization module and an attention module to obtain enhanced features; determining a target noise scene from the selectable scene classification according to the enhanced features; and in response to a vehicle microphone fault recognition instruction, performing fault recognition processing on the vehicle microphone to obtain a first matching result of the vehicle microphone, wherein the first matching result is used to represent whether the vehicle microphone has failed, thereby improving the accuracy of vehicle microphone fault recognition.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of vehicle microphone fault identification, and in particular relates to a vehicle microphone fault identification method and device, an electronic device and a storage medium. BACKGROUND

[0002] External noise monitoring is very important for vehicles. In order to improve the comfort of users, background noise needs to be filtered according to external noise monitoring, and external noise monitoring needs to be based on noise collection of a microphone. Therefore, vehicle microphone fault identification is increasingly valued by people. However, microphone fault detection is still at the stage of using a radio to detect, but because the outdoor environment is very variable, and microphone fault conditions are very diverse, sometimes the microphone can collect sound, but the collected sound changes because of internal faults, which will cause a large error when filtering background noise subsequently. The accuracy of using a radio to identify vehicle microphone faults is very low. SUMMARY

[0003] The present application aims to at least solve one of the technical problems in the prior art. To this end, the present application provides a vehicle microphone fault identification method and device, an electronic device and a storage medium, which can improve the accuracy of vehicle microphone fault identification.

[0004] In a first aspect, an embodiment of the present application provides a vehicle microphone fault identification method applied to a vehicle microphone fault identification device, the vehicle microphone fault identification device being connected to a vehicle microphone, the vehicle microphone being used to collect background noise data outside the vehicle, and the method comprising:

[0005] obtaining sample noise data of the vehicle microphone, and preprocessing the sample noise data, wherein the sample noise data is used to indicate historical off-vehicle background noise data collected by the vehicle microphone;

[0006] dividing the preprocessed sample noise data to obtain a training set and a test set;

[0007] inputting the training set into a preset first scene identification model for training to obtain a trained first scene identification model, wherein the first scene identification model comprises a feature extraction module, a normalization module and an attention module;

[0008] inputting the test set into the trained first scene identification model to obtain the first scene identification model after completion of testing;

[0009] In response to the vehicle microphone fault identification instruction, a fault identification process is performed on the vehicle microphone to obtain a first matching result of the vehicle microphone, wherein the first matching result is used to indicate whether the vehicle microphone has failed, and the fault identification process comprises:

[0010] The first to-be-tested noise data is obtained through the vehicle microphone, and the first to-be-tested noise data is preprocessed, wherein the first to-be-tested noise data is used to indicate real-time off-vehicle background noise data collected by the vehicle microphone;

[0011] The preprocessed first to-be-tested noise data is input into the first scene identification model that has been trained;

[0012] The first to-be-tested noise data is extracted through the feature extraction module to obtain sound features;

[0013] The extracted sound features are normalized through the normalization module, and the normalized sound features are enhanced in the feature channel through the attention module;

[0014] According to the enhanced sound features, a target noise scene is determined from a preset selectable scene classification;

[0015] According to the target noise scene, a target database is determined from a preset first selectable sound database;

[0016] The first to-be-tested noise data and the sound data in the target database are matched and calculated to obtain the first matching result of the vehicle microphone.

[0017] In some embodiments of the present application, the sample noise data acquisition step comprises:

[0018] When the vehicle is started, initial noise is collected through the vehicle microphone;

[0019] The initial noise is separated through different filters in sequence to obtain wind noise, traffic road noise and engine noise, respectively, wherein the sample noise data comprises the wind noise, the traffic road noise and the engine noise.

[0020] In some embodiments of the present application, the first to-be-tested noise data is extracted through the feature extraction module to obtain the mel cepstrum coefficient and the linear cepstrum coefficient of the first to-be-tested noise data;

[0021] The mel cepstrum coefficient and the linear cepstrum coefficient are arranged and combined to obtain a feature vector of the first to-be-tested noise data as the sound features.

[0022] In some embodiments of the present application, the feature extraction module extracts features from the first noise data to be tested to obtain the mel-frequency cepstral coefficients and linear frequency cepstral coefficients of the first noise data to be tested, including:

[0023] The feature extraction module pre-emphasizes the first noise data to be tested;

[0024] The first noise data to be tested is framed and windowed;

[0025] The first noise data to be tested after framing and windowing is subjected to fast Fourier transform processing;

[0026] The first noise data to be tested after fast Fourier transform processing is converted by a mel filter to obtain a mel frequency domain;

[0027] The mel frequency domain is subjected to logarithmic processing and discrete cosine transform to obtain the mel-frequency cepstral coefficients;

[0028] The feature extraction module pre-emphasizes the first noise data to be tested;

[0029] The first noise data to be tested is framed and windowed;

[0030] The first noise data to be tested after framing and windowing is subjected to linear prediction analysis;

[0031] The first noise data to be tested after linear prediction analysis is subjected to cepstrum transform to obtain linear frequency cepstral coefficients.

[0032] In some embodiments of the present application, the first optional sound database acquisition step includes:

[0033] The historical noise data of the vehicle microphone corresponding to each scenario is obtained;

[0034] The historical noise data of each category is subjected to feature extraction to obtain historical waveform features, and the historical waveform features of each category are used as the first optional sound database.

[0035] In some embodiments of the present application, the first noise data to be tested is matched with the sound data in the target database to obtain the first matching result of the vehicle microphone, including:

[0036] The first noise data to be tested is subjected to feature extraction to obtain target waveform features;

[0037] The target waveform features are sequentially compared with each historical waveform feature in the target database to obtain a plurality of first similarities;

[0038] linearly weighting the first similarity to obtain a second similarity;

[0039] When the second similarity is less than a preset value, it is determined that the first matching result of the vehicle microphone is that the microphone has a fault.

[0040] In some embodiments of the present application, a vibration sensor, a counter and a timer are arranged in the vehicle microphone, and the generation of the vehicle microphone fault identification instruction further comprises:

[0041] initializing the counter and the timer;

[0042] When the vehicle starts to work, the timer starts to record;

[0043] Whenever the value of the vibration sensor is greater than a preset vibration threshold, the value of the counter is incremented by 1;

[0044] Obtaining the value of the timer and the value of the counter, when the value of the counter is greater than a preset number threshold, or the value of the timer is greater than a preset time threshold, or a user start detection instruction input by a vehicle PC is received, the vehicle microphone fault identification instruction is generated.

[0045] In a second aspect, the embodiments of the present application provide a vehicle microphone fault identification method and device, the vehicle microphone fault identification device is connected with a vehicle microphone, the vehicle microphone is used to collect background noise data outside the vehicle, and the vehicle microphone fault identification device is used to execute the vehicle microphone fault identification method of the first aspect.

[0046] In a third aspect, the embodiments of the present application provide an electronic device comprising the vehicle microphone fault identification method device of the second aspect.

[0047] In a fourth aspect, the embodiments of the present application provide a computer readable storage medium storing computer executable instructions, and the computer executable instructions are used to execute the vehicle microphone fault identification method of the first aspect.

[0048] The vehicle microphone fault identification method according to the embodiments of the present application has at least the following beneficial effects: applied to a vehicle microphone fault identification device, the vehicle microphone fault identification device is connected with a vehicle microphone, the vehicle microphone is used to collect background noise data outside the vehicle, and the method comprises:

[0049] Obtaining sample noise data of the vehicle microphone, and preprocessing the sample noise data, wherein the sample noise data is used to indicate historical vehicle exterior background noise data collected by the vehicle microphone;

[0050] Divide the preprocessed sample noise data to obtain a training set and a test set;

[0051] Input the training set into a preset first scene recognition model for training to obtain a trained first scene recognition model, wherein the first scene recognition model comprises a feature extraction module, a normalization module, and an attention module;

[0052] Input the test set into the trained first scene recognition model to obtain the first scene recognition model that has completed testing;

[0053] In response to a vehicle microphone fault identification instruction, perform fault identification processing on the vehicle microphone to obtain a first matching result of the vehicle microphone, wherein the first matching result is used to represent whether the vehicle microphone has failed, and the fault identification processing step comprises:

[0054] Obtain first test noise data through the vehicle microphone, and pre-process the first test noise data, wherein the first test noise data is used to indicate real-time off-vehicle background noise data collected by the vehicle microphone;

[0055] Input the preprocessed first test noise data into the first scene recognition model that has completed training;

[0056] Extract features of the first test noise data through the feature extraction module to obtain sound features;

[0057] Perform normalization processing on the extracted sound features through the normalization module, and enhance the sound features through the attention module;

[0058] According to the enhanced sound features, determine a target noise scene from a preset selectable scene classification;

[0059] According to the target noise scene, determine a target database from a preset first selectable sound database;

[0060] Match and calculate the first test noise data and sound data in the target database to obtain the first matching result of the vehicle microphone.

[0061] According to the sample noise data collected by the vehicle microphone, the model is trained, and then the first to-be-tested noise data is input into the trained first scene recognition model, wherein the features in the sound are first extracted by the feature extraction module for feature learning, then the scale is standardized by the normalization module, and finally the sound features are enhanced by the attention module, so as to better understand the sound scene collected by the microphone, and different databases are called according to different scenes to match the first to-be-tested noise data, accurately determine the scene of the vehicle, and reduce the probability of fault recognition failure caused by complex environment. Then, the first to-be-tested noise data is compared with the target database to accurately identify whether the first to-be-tested noise data is abnormal data, and when it is abnormal data, it means that the vehicle microphone has failed. Therefore, the accuracy is improved. BRIEF DESCRIPTION OF DRAWINGS

[0062] Figure 1 is a flow chart of a vehicle microphone fault recognition method provided by an embodiment of the application;

[0063] Figure 2 is a structural diagram of a vehicle microphone fault recognition device provided by another embodiment of the application. DETAILED DESCRIPTION

[0064] The embodiments of the application are described in detail below, and examples of the embodiments are shown in the drawings, wherein the same or similar reference signs represent the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the drawings are exemplary and are only used to explain the application, and cannot be understood as a limitation on the application.

[0065] In the description of the application, it should be understood that the orientation description, such as the orientation or position relationship indicated by up, down, front, back, left, right, etc. is based on the orientation or position relationship shown in the drawings, and is only for the convenience of describing the application and simplifying the description, and does not indicate or imply that the indicated device or element must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as a limitation on the application.

[0066] In the description of the application, the meaning of several is one or more, and the meaning of multiple is more than two, greater than, less than, more than, etc. are understood as not including the number, and above, below, etc. are understood as including the number. If the first, second, etc. are described, they are only used for the purpose of distinguishing technical features, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of indicated technical features or the sequence of indicated technical features.

[0067] In the description of the present application, the words such as arrangement, installation, connection, etc. should be understood broadly unless otherwise explicitly limited, and the person skilled in the art can reasonably determine the specific meaning of the above words in the present application in combination with the specific content of the technical solution.

[0068] The control method of the embodiment of the present application is further described below based on the drawings.

[0069] Referring to Figure 1 , Figure 1 A flowchart of a vehicle microphone fault identification method provided by the embodiment of the present application, the vehicle microphone fault identification method is applied to a vehicle microphone fault identification device, the vehicle microphone fault identification device is connected with a vehicle microphone, the vehicle microphone is used to collect background noise data outside the vehicle, and the vehicle microphone fault identification method includes but is not limited to the following steps:

[0070] In step S110, sample noise data of the vehicle microphone is obtained, and the sample noise data is preprocessed, wherein the sample noise data is used to indicate historical off-vehicle background noise data collected by the vehicle microphone;

[0071] In step S120, the preprocessed sample noise data is divided to obtain a training set and a test set;

[0072] In step S130, the training set is input into a preset first scene identification model for training to obtain a trained first scene identification model, wherein the first scene identification model includes a feature extraction module, a normalization module and an attention module;

[0073] In step S140, the test set is input into the trained first scene identification model to obtain a first scene identification model after completion of the test;

[0074] In step S150, in response to a vehicle microphone fault identification instruction, a fault identification process is performed on the vehicle microphone to obtain a first matching result of the vehicle microphone, wherein the first matching result is used to represent whether the vehicle microphone has a fault, and the fault identification process includes:

[0075] First to-be-tested noise data is obtained by the vehicle microphone, and the first to-be-tested noise data is preprocessed, wherein the first to-be-tested noise data is used to indicate real-time off-vehicle background noise data collected by the vehicle microphone;

[0076] The preprocessed first to-be-tested noise data is input into the first scene identification model after completion of the training;

[0077] The feature extraction module is used to extract features of the first to-be-tested noise data to obtain sound features;

[0078] The normalized sound features are processed by a normalization module, and the normalized sound features are enhanced in feature channels by an attention module.

[0079] A target noise scene is determined from the enhanced sound features in a preset selectable scene classification.

[0080] A target database is determined from the target noise scene in a preset first selectable sound database.

[0081] The first noise data to be tested is matched with the sound data in the target database to obtain a first matching result of the vehicle microphone.

[0082] It should be noted that the vehicle microphone is used to collect the sound outside the vehicle. It can be understood that the purpose of the first scene recognition model is to determine the current scene of the vehicle, such as driving in different scenes such as high speed, urban area, and mountainous area. The noise generated by the surrounding environment is different and diverse, so the importance of a specific feature channel is enhanced by using the attention module, so that accurate classification can be performed even if there are many features. The first scene recognition model can dynamically identify different parts of the feature sequence to determine the scene of the vehicle, which can reduce the decrease in accuracy of microphone fault recognition caused by the complexity and large changes of the vehicle driving environment. Then, the target database is determined from the target noise scene in the preset first selectable sound database. Further, the first selectable sound database is obtained by processing sample noise data. First, the sample noise data is classified according to different scenes, and then the normal sound in each scene is preprocessed, such as data cleaning. Then, the preprocessed sample noise data is extracted to obtain features such as time domain features and frequency domain features. The features of the normal sound are stored. It should be noted that the first noise data to be tested also needs to be preprocessed and extracted before being matched with the sound data in the target database. Then, the features of the first noise data to be tested are calculated with the features of the normal sound stored in the database. When the difference is large, it is considered that the microphone has a fault.

[0083] In another embodiment, the step of obtaining sample noise data includes:

[0084] When the vehicle is started, the initial noise is collected by the vehicle microphone.

[0085] The initial noise is separated by different filters in sequence to obtain wind noise, traffic road noise, and engine noise, respectively, wherein the sample noise data includes wind noise, traffic road noise, and engine noise.

[0086] It should be noted that there are many noises during the driving of the vehicle, and the noises mainly used by us are wind noise, traffic road noise and engine noise. It can be understood that the traffic road noise mainly refers to the sound of people. In this way, the first scene recognition model can judge different scenes according to the characteristics in the subsequent scene recognition model. The first scene recognition model further extracts the characteristics of wind noise, traffic road noise and engine noise in the feature extraction module, for example, the wind noise can extract the intensity, intensity change range and other characteristics. After normalization and feature enhancement, the classifier is used to classify and identify the characteristics of the scene. Therefore, the pre-designed band-pass filter is used to separate the noise in a specific frequency band. For example, filter the human voice frequency band (about 70~350Hz) to extract the human voice signal; filter the engine noise frequency band to extract the engine noise signal.

[0087] In another embodiment, the feature extraction module extracts human voice information from the traffic road noise, enhances the sound characteristics, and determines the target noise scene from the pre-set optional scene classification through the classifier, including:

[0088] Determine the enhanced sound characteristics through the classifier;

[0089] When the sound intensity feature of the wind noise is greater than the pre-set wind noise threshold, and the intensity change range feature of the wind noise is greater than the pre-set first range value, the sound intensity feature of the engine noise is greater than the first engine noise threshold, and the sound intensity change range feature of the human voice information is less than the pre-set second range value, it is determined that the scene where the first to-be-tested noise data is located is a high-speed scene;

[0090] When the intensity change range of the wind noise is greater than the first range value, the sound intensity feature of the engine noise is greater than the second engine noise threshold, and the sound intensity change range of the human voice information is less than the pre-set second range value, it is determined that the scene where the first to-be-tested noise data is located is a mountainous area scene, wherein the second engine noise threshold is greater than the first engine noise threshold;

[0091] The sound intensity feature of the wind noise is less than the pre-set wind noise threshold, and the intensity change range feature of the wind noise is less than the pre-set first range value of the engine noise, the sound intensity feature of the engine noise is greater than the second engine noise threshold, and the sound intensity change range feature of the human voice information is greater than the pre-set second range value. It is determined that the scene where the first to-be-tested noise data is located is an urban area scene.

[0092] It should be noted that in different scenarios, such as at high speed, because the increase in vehicle speed causes the relative movement between air and vehicle body to intensify, so the intensity of the wind sound is larger, but the wind direction is stable, so the intensity change range is smaller, at the same time, the engine needs higher speed to provide power at high speed, so the engine noise also increases, and pedestrians and the like are very few at high speed, and the fewer the people, the smaller the sound intensity change range, and the more people, the larger the sound intensity change range. Therefore, when the sound intensity of the wind sound is greater than the preset wind noise threshold, and the intensity change range of the wind sound is greater than the preset first range value, at the same time, the engine noise is greater than the first engine noise threshold, and the sound intensity change range of the human voice information is smaller than the preset second range value, it is determined that the sample noise data corresponds to the label of the high-speed scene; similarly, in the mountainous area, because of the terrain, first of all, there are fewer people, and when the wind blows through the microphone from different directions, its intensity may be different. By monitoring the intensity change of the wind noise, the source direction of the wind can be inferred. In addition, in the mountainous area, because more power is needed to climb the slope, the engine noise will increase due to the high load operation of the engine, which is much larger than the engine noise at high speed, so when the intensity change range of the wind sound is greater than the first range value, the engine noise is greater than the second engine noise threshold, and the sound intensity change range of the human voice information is smaller than the preset second range value, it is determined that the sample noise data corresponds to the label of the mountainous area scene, wherein the second engine noise threshold is greater than the first engine noise threshold; similarly, in the urban area, because the vehicle speed is relatively low, the wind sound is also relatively low, and the wind change is also less, but in the urban area, because of the frequent start-stop and low-speed driving, the engine noise will be relatively obvious, and the biggest difference between the urban area and other environments is that there are very many pedestrians, so when the sound intensity of the wind sound is less than the preset wind noise threshold, and the intensity change range of the wind sound is less than the preset first range value, the engine noise is greater than the second engine noise threshold, and the sound intensity change range of the human voice information is greater than the preset second range value, it is determined that the sample noise data corresponds to the label of the urban area scene. It can be understood that the above is only an example of some scenarios.

[0093] In another embodiment, the feature extraction module is used to extract features from the first to-be-tested noise data to obtain sound features, including:

[0094] The feature extraction module is used to extract features from the first to-be-tested noise data to obtain the mel cepstrum coefficient and the linear cepstrum coefficient of the first to-be-tested noise data;

[0095] The mel cepstrum coefficient and the linear cepstrum coefficient are arranged and combined to obtain a feature vector of the first to-be-tested noise data as the sound features.

[0096] It should be noted that the mel-spectral coefficient and the linear predictive cepstral coefficient obtain different features in the first to-be-tested noise data respectively, the mel-spectral coefficient is used to obtain features related to human auditory characteristics, and the linear predictive cepstral coefficient is used to obtain a spectral envelope of a speech signal, and the combination of the two can more comprehensively obtain features of the speech signal. The process of arranging and combining the two is mainly to tile the mel-spectral coefficient and the linear predictive cepstral coefficient respectively, and then connect them in a preset order to become a new feature, that is, the final sound feature. He is a kind of composite feature vector, which can more accurately describe the first to-be-tested noise data and improve the accuracy of classification.

[0097] In another embodiment, the step of obtaining the mel-spectral coefficient of the first to-be-tested noise data comprises:

[0098] The first to-be-tested noise data is pre-emphasized by the feature extraction module;

[0099] The first to-be-tested noise data is framed and windowed;

[0100] The first to-be-tested noise data after framing and windowing is subjected to fast Fourier transform processing;

[0101] The first to-be-tested noise data after fast Fourier transform processing is subjected to signal conversion by a mel filter to obtain a mel frequency domain;

[0102] The mel frequency domain is subjected to logarithmic processing and discrete cosine transform to obtain the mel-spectral coefficient.

[0103] It should be noted that the mel-spectral coefficient can obtain features related to human auditory characteristics, and the specific obtaining steps are as follows: pre-emphasis processing mainly performs high-pass filtering on the first to-be-tested noise data, that is, highlights the high-frequency part and reduces the interference of the low-frequency part. Then frame and window processing is performed to eliminate edge effects, and a Hamming window function is used for windowing. The data after windowing processing of each frame is subjected to fast Fourier transform conversion to the frequency domain and input to the mel filter for conversion to the mel frequency domain. The mel frequency domain is taken logarithm to obtain the logarithmic mel spectrum, and the logarithmic mel spectrum is subjected to discrete cosine transform to obtain the mel-spectral coefficient.

[0104] In another embodiment, the step of obtaining the linear predictive cepstral coefficient of the first to-be-tested noise data comprises:

[0105] The first to-be-tested noise data is pre-emphasized by the feature extraction module;

[0106] The first to-be-tested noise data is framed and windowed;

[0107] The first to-be-tested noise data after framing and windowing is subjected to linear predictive analysis;

[0108] The first to-be-tested noise data is subjected to a cepstrum transform to obtain linear cepstrum coefficients.

[0109] It should be noted that the linear prediction cepstrum coefficients can obtain the spectral envelope of the speech signal, and the specific steps are as follows: the pre-emphasis processing is mainly high-pass filtering of the first to-be-tested noise data, that is, highlighting the high-frequency part and reducing the interference of the low-frequency part. Then, the frame division and windowing processing are performed to eliminate the edge effect, and the Hamming window function is used for windowing. Then, the linear prediction analysis is performed on each frame of the windowed speech signal, and the linear prediction coefficients are determined by the least mean square error criterion. The linear prediction coefficients are converted into the cepstrum coefficients of each frame of the speech signal by the cepstrum transform.

[0110] In another embodiment, the obtaining step of the first optional sound database comprises:

[0111] The historical noise data of the vehicle microphone corresponding to each category is obtained in each scene;

[0112] The feature extraction is performed on the historical noise data of each category to obtain historical waveform features, and the historical waveform features of each category are taken as the first optional sound database.

[0113] It should be noted that, because the vehicle will arrive at different scenes, such as mountainous areas, urban areas, highways, etc., the noise of different scenes is also different, and therefore the corresponding database corresponding to each scene needs to be constructed for subsequent comparison of whether the microphone is abnormal. Direct comparison of sound segments with sound segments is easy to cause inaccuracy, and therefore the features of the historical noise data corresponding to each type of scene need to be obtained first to obtain the historical waveform features as the database for subsequent comparison. The waveform quantizes the sound, and is more accurate for subsequent comparison.

[0114] In another embodiment, the first to-be-tested noise data is matched with the sound data in the target database to obtain a first matching result of the vehicle microphone, including:

[0115] The feature extraction is performed on the first to-be-tested noise data to obtain target waveform features;

[0116] The target waveform features are sequentially compared with each historical waveform feature in the target database to obtain a plurality of first similarities;

[0117] The linear weighting calculation is performed on the first similarities to obtain second similarities;

[0118] When the second similarity is less than a preset value, it is determined that the first matching result of the vehicle microphone is that the microphone has a fault.

[0119] It should be noted that the similarity between the sounds is determined by calculating the similarity between the waveforms of different sound data, so as to determine whether the sound collected by the microphone is a normal sound. Therefore, in order to improve the accuracy, the waveforms corresponding to the historical sound collected multiple times are stored in the target database, the target waveform feature is sequentially calculated with each historical waveform feature in the target database to obtain a first similarity, and then the first similarity is linearly weighted to obtain a second similarity. The weight can be preset according to the actual situation. When the second similarity is closer to 1, it means complete positive correlation, close to -1, it means complete negative correlation, and it is irrelevant. Therefore, a preset value is set, and when the second similarity is less than the preset value, it means that the waveform of the first to-be-tested noise data is very different from the waveform in the target database, and it is considered that the microphone is malfunctioning, so the collected sound also has a problem.

[0120] In another embodiment, a vibration sensor, a counter and a timer are arranged in the vehicle microphone, and the generation step of the vehicle microphone fault identification instruction further comprises:

[0121] Initializing the counter and the timer;

[0122] When the vehicle starts to work, the timer starts to record;

[0123] The value of the counter is increased by 1 every time the value of the vibration sensor is greater than the preset vibration threshold;

[0124] Obtaining the value of the timer and the value of the counter, and when the value of the counter is greater than the preset number threshold, or the value of the timer is greater than the preset time threshold, or the user starts the detection instruction input by the vehicle PC end is received, the vehicle microphone fault identification instruction is generated.

[0125] It should be noted that the timing of the microphone fault identification needs to be adjusted. When the vehicle suddenly brakes, or is hit, or drives on a bumpy road, the microphone is most likely to be affected and malfunction. Therefore, a vibration sensor is arranged. Because the vibration sensor is in the microphone, and the vehicle microphone is in the vehicle, when the vehicle vibration causes the microphone to vibrate, the vibration sensor will feel it and record the number of times greater than the preset vibration threshold. When the number of times reaches the preset number threshold, that is, when the vehicle is subjected to excessive vibration, the microphone is most likely to be damaged, and the vehicle microphone fault identification instruction is generated to start the automatic fault identification of the vehicle microphone. In addition, the microphone is also prone to malfunction when used for a long time, so when the value of the timer is greater than the preset time threshold, the vehicle microphone fault identification instruction is also generated. In addition, if the user actively inputs the user start detection instruction on the vehicle PC end, the vehicle microphone fault identification instruction can also be generated.

[0126] AsFigure 2 As shown, Figure 2 is a structural diagram of a vehicle microphone fault identification device provided by an embodiment of the present application. The present application also provides a vehicle microphone fault identification device, comprising:

[0127] The processor 201 can be implemented in a general-purpose central processing unit (CPU), a microprocessor, an application specific integrated circuit (ASIC), or one or more integrated circuits, and is configured to execute related programs to implement the technical solutions provided by the embodiments of the present application.

[0128] The memory 202 can be implemented in a read only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 202 can store an operating system and other application programs. When the technical solutions provided by the embodiments of the present application are implemented by software or firmware, the related program codes are stored in the memory 202 and are called and executed by the processor 201 to implement the vehicle microphone fault identification method of the embodiments of the present application.

[0129] The input / output interface 203 is configured to realize information input and output.

[0130] The communication interface 204 is configured to realize the communication interaction between the device and other devices. The communication can be realized by a wired manner (for example, a USB, a network cable, etc.) or a wireless manner (for example, a mobile network, WIFI, Bluetooth, etc.).

[0131] The bus 205 is configured to transmit information between various components (for example, the processor 201, the memory 202, the input / output interface 203, and the communication interface 204) of the device.

[0132] The processor 201, the memory 202, the input / output interface 203, and the communication interface 204 are connected to each other through the bus 205 to realize the communication connection between them in the device.

[0133] The embodiments of the present application also provide an electronic device, which comprises the vehicle microphone fault identification device as described above.

[0134] The embodiments of the present application also provide a storage medium, which is a computer readable storage medium. The storage medium stores a computer program. When the computer program is executed by a processor, the vehicle microphone fault identification method described above is realized.

[0135] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory can include a high-speed random access memory, and can also include a non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state memory device. In some embodiments, the memory can optionally include a memory disposed remotely with respect to the processor, which can be connected to the processor through a network. Examples of the above network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof. The above-described device embodiments are only illustrative, and units described as separate components can or can not be physically separated, implemented in one place, or distributed to multiple network units. Part or all of the modules can be selected according to actual needs to achieve the purpose of the present embodiment.

[0136] Those of ordinary skill in the art can understand that all or some steps in the above disclosed method and system can be implemented as software, firmware, hardware, and appropriate combinations thereof. Some or all physical components can be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, which can include computer storage media (or non-transitory media) and communication media (or transitory media). As known to those of ordinary skill in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storage of information such as computer readable instructions, data structures, program modules or other data. Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tapes, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and can be accessed by a computer. In addition, as known to those of ordinary skill in the art, communication media generally includes computer readable instructions, data structures, program modules or other data in modulated data signals such as carrier waves or other transport mechanisms, and can include any information delivery medium.

[0137] The above is a specific description of the preferred embodiment of the present application, but the present application is not limited to the above-described embodiments, and those skilled in the art can make various equivalent modifications or replacements without departing from the spirit of the present application. These equivalent modifications or replacements are all included in the scope defined by the claims of the present application.

Claims

1. A vehicle microphone fault identification method, characterized by, The application is applied to a vehicle microphone fault recognition device connected with a vehicle microphone used for collecting background noise data outside the vehicle, and the method comprises: Obtaining sample noise data of the vehicle microphone, and preprocessing the sample noise data, wherein the sample noise data is used to indicate historical off-vehicle background noise data collected by the vehicle microphone, and specifically, the sample noise data obtaining step comprises: collecting initial noise through the vehicle microphone when the vehicle starts; separating the initial noise through different filters in turn to obtain wind noise, traffic road noise and engine noise, wherein the sample noise data comprises the wind noise, the traffic road noise and the engine noise; Dividing the preprocessed sample noise data to obtain a training set and a test set; Inputting the training set into a preset first scene recognition model for training to obtain a trained first scene recognition model, wherein the first scene recognition model comprises a feature extraction module, a normalization module and an attention module; Inputting the test set into the trained first scene recognition model to obtain the first scene recognition model after completing the test; In response to a vehicle microphone fault recognition instruction, performing fault recognition processing on the vehicle microphone to obtain a first matching result of the vehicle microphone, wherein the first matching result is used to represent whether the vehicle microphone has a fault, and the fault recognition processing step comprises: Obtaining first test noise data through the vehicle microphone, and preprocessing the first test noise data, wherein the first test noise data is used to indicate real-time off-vehicle background noise data collected by the vehicle microphone; Inputting the preprocessed first test noise data into the first scene recognition model after completing the training; Extracting features of the first test noise data through the feature extraction module to obtain sound features; Performing normalization processing on the extracted sound features through the normalization module, and enhancing the sound features after the normalization processing through the attention module; According to the enhanced sound features, determining a target noise scene from a preset optional scene classification through a classifier; According to the target noise scene, determining a target database from a preset first optional sound database; Matching and calculating the first test noise data and sound data in the target database to obtain the first matching result of the vehicle microphone; Wherein, the feature extraction module extracts human voice information from the traffic road noise, and according to the enhanced sound features, the target noise scene is determined from the preset optional scene classification through the classifier, which comprises: Judging the enhanced sound features through the classifier; determining that the first to-be-tested noise data is in a high-speed scene when the sound intensity feature of the wind sound is greater than a preset wind noise threshold, the intensity change range feature of the wind sound is greater than a preset first range value, the sound intensity feature of the engine noise is greater than a first engine noise threshold, and the sound intensity change range feature of the human voice information is less than a preset second range value; determining that the first to-be-tested noise data is in a mountainous area scene when the intensity change range of the wind sound is greater than the first range value, the sound intensity feature of the engine noise is greater than a second engine noise threshold, and the sound intensity change range of the human voice information is less than a preset second range value, wherein the second engine noise threshold is greater than the first engine noise threshold; determining that the first to-be-tested noise data is in an urban area scene when the sound intensity feature of the wind sound is less than a preset wind noise threshold, the intensity change range feature of the wind sound is less than a preset first range value, the sound intensity feature of the engine noise is greater than a second engine noise threshold, and the sound intensity change range feature of the human voice information is greater than a preset second range value; wherein the matching calculation of the first to-be-tested noise data and the sound data in the target database to obtain the first matching result of the vehicle microphone includes: obtaining historical noise data of the vehicle microphone in each scene of a corresponding category; extracting features from each category of the historical noise data to obtain historical waveform features, and taking each category of the historical waveform features as the first selectable sound database; extracting features from the first to-be-tested noise data to obtain target waveform features; sequentially performing similarity calculation on the target waveform features and each historical waveform feature in the target database to obtain a plurality of first similarities; performing linear weighting calculation on the first similarities to obtain a second similarity; when the second similarity is less than a preset value, determining that the first matching result of the vehicle microphone is that the microphone has a fault.

2. The vehicle microphone fault identification method of claim 1, wherein, the feature extraction of the first to-be-tested noise data by the feature extraction module to obtain sound features includes: extracting, by the feature extraction module, the first to-be-tested noise data to obtain mel-frequency cepstrum coefficients and linear cepstrum coefficients of the first to-be-tested noise data; arranging and combining the mel-frequency cepstrum coefficients and the linear cepstrum coefficients to obtain a feature vector of the first to-be-tested noise data as the sound features.

3. The vehicle microphone fault identification method of claim 2, wherein, the feature extraction of the first to-be-tested noise data by the feature extraction module to obtain mel-frequency cepstrum coefficients and linear cepstrum coefficients of the first to-be-tested noise data includes: performing pre-emphasis processing on the first to-be-tested noise data by the feature extraction module; performing framing and windowing processing on the first to-be-tested noise data; performing fast Fourier transform processing on the first to-be-tested noise data after framing and windowing processing; performing signal conversion on the first to-be-tested noise data after fast Fourier transform processing by a mel filter to obtain a mel frequency domain; Logarithm processing and discrete cosine transform are performed on the mel frequency domain to obtain the mel cepstrum coefficients; The first noise data to be measured is pre-emphasized by the feature extraction module; The first noise data to be measured is framed and windowed; Linear prediction analysis is performed on the first noise data to be measured after framing and windowing; Linear prediction analysis is performed on the first noise data to be measured after framing and windowing.

4. The vehicle microphone fault identification method of claim 1, wherein, The vehicle microphone is provided with a vibration sensor, a counter and a timer, and the generation of the vehicle microphone fault recognition instruction further comprises: Initializing the counter and the timer; When the vehicle starts to work, the timer starts to record; Every time the value of the vibration sensor is greater than the preset vibration threshold, the value of the counter is increased by 1; Obtain the value of the timer and the value of the counter, and when the value of the counter is greater than the preset number threshold, or the value of the timer is greater than the preset time threshold, or a user starts the detection instruction input by the vehicle PC, the vehicle microphone fault recognition instruction is generated.

5. A vehicle microphone malfunction recognition apparatus characterized by comprising: The vehicle microphone fault recognition device is connected with the vehicle microphone, and the vehicle microphone is used to collect background noise data outside the vehicle. The vehicle microphone fault recognition device is used to: Obtain the sample noise data of the vehicle microphone, and pre-process the sample noise data, wherein the sample noise data is used to indicate the historical off-road background noise data collected by the vehicle microphone, and specifically, the sample noise data acquisition step comprises: collecting initial noise through the vehicle microphone when the vehicle starts; separating the initial noise through different filters in turn to obtain wind noise, traffic road noise and engine noise, wherein the sample noise data includes the wind noise, the traffic road noise and the engine noise; The pre-processed sample noise data is divided to obtain a training set and a test set; The training set is input into a preset first scene recognition model for training to obtain a trained first scene recognition model, wherein the first scene recognition model comprises a feature extraction module, a normalization module and an attention module; The test set is input into the trained first scene recognition model to obtain a completed test first scene recognition model; In response to the vehicle microphone fault recognition instruction, the vehicle microphone is subjected to fault recognition processing to obtain a first matching result of the vehicle microphone, wherein the first matching result is used to represent whether the vehicle microphone has failed, and the fault recognition processing step comprises: Obtain the sample noise data of the vehicle microphone, and pre-process the sample noise data, wherein the sample noise data is used to indicate the historical off-road background noise data collected by the vehicle microphone, and specifically, the sample noise data acquisition step comprises: collecting initial noise through the vehicle microphone when the vehicle starts; separating the initial noise through different filters in turn to obtain wind noise, traffic road noise and engine noise, wherein the sample noise data includes the wind noise, the traffic road noise and the engine noise; The pre-processed sample noise data is divided to obtain a training set and a test set; The training set is input into a preset first scene recognition model for training to obtain a trained first scene recognition model, wherein the first scene recognition model comprises a feature extraction module, a normalization module and an attention module; The test set is input into the trained first scene recognition model to obtain a completed test first scene recognition model; In response to the vehicle microphone fault recognition instruction, the vehicle microphone is subjected to fault recognition processing to obtain a first matching result of the vehicle microphone, wherein the first matching result is used to represent whether the vehicle microphone has failed, and the fault recognition processing step comprises: The normalized module is used for normalizing the extracted sound features, and the attention module is used for enhancing the sound features in the channel; According to the enhanced sound features, a target noise scene is determined from a preset optional scene classification; According to the target noise scene, a target database is determined from a preset first optional sound database; The first test noise data is matched with the sound data in the target database to obtain the first matching result of the vehicle microphone; The feature extraction module extracts human voice information from the traffic road noise, and according to the enhanced sound features, a target noise scene is determined from a preset optional scene classification by using a classifier, which includes: The classifier is used for judging the enhanced sound features; When the sound intensity feature of the wind sound is greater than a preset wind noise threshold, the intensity change range feature of the wind sound is greater than a preset first range value, the sound intensity feature of the engine noise is greater than a first engine noise threshold, and the sound intensity change range feature of the human voice information is less than a preset second range value, it is determined that the first test noise data is in a high-speed scene; When the intensity change range of the wind sound is greater than the first range value, the sound intensity feature of the engine noise is greater than a second engine noise threshold, and the sound intensity change range of the human voice information is less than a preset second range value, it is determined that the first test noise data is in a mountainous area scene, wherein the second engine noise threshold is greater than the first engine noise threshold; When the sound intensity feature of the wind sound is less than a preset wind noise threshold, the intensity change range feature of the wind sound is less than a preset first range value, the sound intensity feature of the engine noise is greater than a second engine noise threshold, and the sound intensity change range feature of the human voice information is greater than a preset second range value, it is determined that the first test noise data is in an urban area scene; The first test noise data is matched with the sound data in the target database to obtain the first matching result of the vehicle microphone, which includes: The historical noise data of the vehicle microphone in each scene is obtained; The historical noise data of each category is subjected to feature extraction to obtain historical waveform features, and the historical waveform features of each category are used as the first optional sound database; The first test noise data is subjected to feature extraction to obtain target waveform features; The target waveform features are sequentially subjected to similarity calculation with each historical waveform feature in the target database to obtain a plurality of first similarities; The first similarities are subjected to linear weighting calculation to obtain a second similarity; When the second similarity is less than a preset value, it is determined that the first matching result of the vehicle microphone is a microphone fault.

6. An electronic device, comprising: The vehicle microphone fault identification device of claim 5 is included.

7. A computer readable storage medium characterized in that, The computer readable storage medium stores computer executable instructions for causing a computer to perform the vehicle microphone fault identification method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Acoustic event detecting method under hospital noise environment

    CN108648748A

  • Noise generation model training method and device, equipment and medium

    CN115035911A

  • Speech recognition method and device, computer readable storage medium and computer equipment

    CN116935852A

  • External microphone fault identification device, fault identification method and vehicle

    CN117641220A