Road surface type recognition method and apparatus, vehicle, and cloud server
By collecting tire noise frequency data from vehicle tires, using a deep learning model to extract and fuse features from the time-frequency map, and combining this with arbitration of multi-tire recognition results, the problem of road surface type recognition being affected by the external environment has been solved, thus improving the accuracy and reliability of recognition.
Patent Information
- Application Number
- PCT/CN2025/094884
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-27
- Filing Date
- 2025-05-14
- Publication Date
- 2025-12-04
AI Technical Summary
In existing technologies, road surface type recognition is easily affected by external environmental factors such as lighting and weather, resulting in low recognition accuracy.
By collecting tire noise frequency data from vehicle tires, a deep learning model is used to extract and fuse features from the time-frequency map. The results of identification from multiple tires of the vehicle are then used for arbitration to determine the final road surface type.
It improves the accuracy of road surface type identification, reduces the impact of external environmental factors, and enhances the reliability and precision of the identification results.
Smart Images

Figure CN2025094884_04122025_PF_FP_ABST
Abstract
Description
Road surface type identification method and device, vehicle and cloud server
[0001] This application claims priority to Chinese Patent Application No. 202410662884.0, filed on May 27, 2024, the entire contents of which are incorporated herein by reference. TECHNICAL FIELD
[0002] The present disclosure relates to the technical field of intelligent vehicles, and in particular to a road surface type identification method, device, vehicle and cloud server. BACKGROUND
[0003] With the progress of intelligent driving technology, accurate identification of the road surface type of the driving road becomes increasingly important, and the identification result is directly related to the safety of driving and the comfort experience of passengers. SUMMARY
[0004] The present disclosure provides a road surface type identification method, device, vehicle and cloud server, which can solve the problem of low accuracy of road surface type identification.
[0005] In a first aspect, some embodiments of the present disclosure provide a road surface type identification method, comprising:
[0006] obtaining first audio data, the first audio data being tire noise audio data of a first tire of a vehicle driving on a road surface to be identified;
[0007] obtaining a first road surface type of the road surface to be identified based on the first audio data.
[0008] In some embodiments, feature extraction is performed on the first audio data to obtain time-frequency graph features corresponding to the first audio data.
[0009] Performing road surface type identification based on the time-frequency graph features corresponding to the first audio data to obtain the first road surface type of the road surface to be identified.
[0010] In some embodiments, the performing road surface type identification based on the time-frequency graph features corresponding to the first audio data to obtain the first road surface type of the road surface to be identified comprises:
[0011] inputting the time-frequency graph features corresponding to the first audio data into a deep learning model to extract features to obtain deep learning features corresponding to the time-frequency graph features;
[0012] obtaining the first road surface type of the road surface to be identified based on the deep learning features.
[0013] In some embodiments, the time-frequency graph features corresponding to the first audio data comprise at least two types of time-frequency graph features.
[0014] The inputting of the time-frequency graph features corresponding to the first audio data into a deep learning model for feature extraction to obtain deep learning features corresponding to the time-frequency graph features comprises:
[0015] The inputting of the at least two types of time-frequency graph features into deep learning models respectively for feature extraction to obtain deep learning features corresponding to each type of time-frequency graph feature in the at least two types of time-frequency graph features comprises:
[0016] The fusion processing of the obtained deep learning features corresponding to each type of time-frequency graph feature in the at least two types of time-frequency graph features to obtain fusion features comprises:
[0017] The obtaining of the first road surface type of the to-be-identified road surface based on the deep learning features comprises:
[0018] The obtaining of the first road surface type of the to-be-identified road surface based on the fusion features comprises:
[0019] In some embodiments, the at least two types of time-frequency graph features comprise: a mel time-frequency graph feature and a wavelet transform time-frequency graph feature.
[0020] The inputting of the at least two types of time-frequency graph features into deep learning models respectively for feature extraction to obtain deep learning features corresponding to each type of time-frequency graph feature in the at least two types of time-frequency graph features comprises:
[0021] The inputting of the mel time-frequency graph feature into a deep convolutional neural network to obtain deep learning features corresponding to the mel time-frequency graph feature comprises:
[0022] The inputting of the wavelet transform time-frequency graph feature into a deep residual network to obtain deep learning features of the wavelet transform time-frequency graph comprises.
[0023] In some embodiments, the feature extraction of the first audio data to obtain time-frequency graph features corresponding to the first audio data comprises:
[0024] The short-time Fourier transform of the first audio data to obtain a frequency spectrum of the first audio data, and the filtering processing of the frequency spectrum by a mel filter bank to obtain the mel time-frequency graph feature corresponding to the first audio data.
[0025] The wavelet transform of the first audio data to obtain the wavelet transform time-frequency graph feature corresponding to the first audio data.
[0026] In some embodiments, the identification method further comprises:
[0027] The statistical feature extraction of the first audio data to obtain statistical features corresponding to the first audio data.
[0028] inputting the statistical features into a deep learning model for feature extraction to obtain deep learning features corresponding to the statistical features;
[0029] The deep learning features corresponding to each of the at least two types of time-frequency graph features obtained are fused to obtain a fusion feature, including:
[0030] The deep learning features corresponding to each of the at least two types of time-frequency graph features obtained and the deep learning features corresponding to the statistical features are fused to obtain a fusion feature.
[0031] In some embodiments, the statistical features are input into a deep learning model for feature extraction to obtain deep learning features corresponding to the statistical features, including:
[0032] The statistical features are input into a deep perception machine to obtain deep learning features corresponding to the statistical features.
[0033] In some embodiments, the feature extraction of the first audio data to obtain time-frequency graph features corresponding to the first audio data includes:
[0034] The first audio data is subjected to denoising processing of environmental interference noise to obtain denoised first audio data;
[0035] The denoised first audio data is subjected to feature extraction to obtain time-frequency graph features corresponding to the first audio data.
[0036] In some embodiments, the denoising processing of environmental interference noise on the first audio data to obtain denoised first audio data includes:
[0037] A low-pass filter is used to perform denoising processing of environmental interference noise on the first audio data to obtain the denoised first audio data.
[0038] In some embodiments, the frequency threshold of the low-pass filter is determined according to the driving speed of the vehicle.
[0039] In some embodiments, the identification method further includes:
[0040] Obtaining second audio data; the second audio data is the tire noise audio data of the second tire of the vehicle driving on the road to be identified;
[0041] Performing road type identification based on the second audio data to obtain a second road type of the road to be identified;
[0042] obtain a target road surface type of the to-be-identified road surface based on the first road surface type and the second road surface type.
[0043] In some embodiments, the road surface type identification based on the second audio data comprises:
[0044] perform feature extraction on the second audio data to obtain time-frequency graph features corresponding to the second audio data;
[0045] perform road surface type identification based on the time-frequency graph features corresponding to the second audio data to obtain a second road surface type of the to-be-identified road surface.
[0046] In some embodiments, the obtaining of the target road surface type of the to-be-identified road surface based on the first road surface type and the second road surface type comprises:
[0047] if the first road surface type is the same as the second road surface type, determining the first road surface type or the second road surface type as the target road surface type;
[0048] if the first road surface type is different from the second road surface type, determining a target road surface type of the to-be-identified road surface according to confidences corresponding to the first road surface type and the second road surface type respectively.
[0049] In some embodiments, the determining of the target road surface type of the to-be-identified road surface according to the confidences corresponding to the first road surface type and the second road surface type respectively comprises:
[0050] if one of the confidences corresponding to the first road surface type and the second road surface type respectively is greater than a preset threshold value and the other is less than or equal to the preset threshold value, determining the road surface type with the confidence greater than the preset threshold value as the target road surface type;
[0051] if the confidences corresponding to the first road surface type and the second road surface type respectively are all greater than the preset threshold value or are all less than the preset threshold value, determining the road surface type with a higher priority as the target road surface type according to priorities of the first road surface type and the second road surface type.
[0052] In some embodiments, the determining of the target road surface type of the to-be-identified road surface according to the confidences corresponding to the first road surface type and the second road surface type respectively comprises:
[0053] determining the road surface type with a higher confidence as the target road surface type of the to-be-identified road surface.
[0054] In some embodiments, the road surface type identification based on the time-frequency graph feature corresponding to the first audio data obtains a first road surface type of the to-be-identified road surface, and the method comprises the following steps:
[0055] The time-frequency graph feature corresponding to the first audio data is input into a target road surface type identification model to obtain the first road surface type of the to-be-identified road surface.
[0056] In some embodiments, the target road surface type identification model is obtained by training a candidate road surface type identification model based on training samples, and the training samples comprise: an audio data sample; the audio data sample comprises: audio data content and a road surface type label corresponding to the audio data.
[0057] In some embodiments, the identification method further comprises: obtaining a plurality of audio data samples by performing data enhancement processing on the audio data sample;
[0058] The data enhancement processing comprises at least one of the following: pitch enhancement processing, pitch height enhancement processing, time translation processing, or time mask processing.
[0059] In a second aspect, some embodiments of the present disclosure provide a road surface type identification device, comprising an acquisition module and an identification module.
[0060] The acquisition module is configured to acquire first audio data, wherein the first audio data is tire noise audio data of a first tire of a vehicle driving on a to-be-identified road surface;
[0061] The identification module is configured to obtain a first road surface type of the to-be-identified road surface based on the first audio data.
[0062] In some embodiments, the identification module is configured to extract features from the first audio data to obtain a time-frequency graph feature corresponding to the first audio data, and to identify a road surface type based on the time-frequency graph feature corresponding to the first audio data to obtain a first road surface type of the to-be-identified road surface.
[0063] In some embodiments, the identification module is configured to input the time-frequency graph feature corresponding to the first audio data into a deep learning model to extract features to obtain a deep learning feature corresponding to the time-frequency graph feature.
[0064] The first road surface type of the to-be-identified road surface is obtained based on the deep learning feature.
[0065] In some embodiments, the time-frequency graph feature corresponding to the first audio data comprises at least two types of time-frequency graph features.
[0066] The identification module is configured to input the at least two types of time-frequency map features into a deep learning model respectively for feature extraction to obtain deep learning features corresponding to each type of time-frequency map feature in the at least two types of time-frequency map features; perform fusion processing on the obtained deep learning features corresponding to each type of time-frequency map feature in the at least two types of time-frequency map features to obtain fusion features; and obtain the first pavement type of the pavement to be identified based on the fusion features.
[0067] In some embodiments, the at least two types of time-frequency map features include: a mel time-frequency map feature and a wavelet transform time-frequency map feature; the identification module is configured to input the mel time-frequency map feature into a deep convolutional neural network to obtain deep learning features corresponding to the mel time-frequency map feature; and input the wavelet transform time-frequency map feature into a deep residual network to obtain deep learning features of the wavelet transform time-frequency map.
[0068] In some embodiments, the pavement type identification apparatus further includes a feature extraction module configured to perform short-time Fourier transform on the first audio data to obtain a frequency spectrum of the first audio data, perform filter processing on the frequency spectrum through a mel filter bank to obtain the mel time-frequency map feature corresponding to the first audio data, and perform wavelet transform on the first audio data to obtain the wavelet transform time-frequency map feature corresponding to the first audio data.
[0069] In some embodiments, the feature extraction module is further configured to perform statistical feature extraction on the first audio data to obtain statistical features corresponding to the first audio data.
[0070] The identification module is configured to input the statistical features into a deep learning model for feature extraction to obtain deep learning features corresponding to the statistical features; and perform fusion processing on the obtained deep learning features corresponding to each type of time-frequency map feature in the at least two types of time-frequency map features and the deep learning features corresponding to the statistical features to obtain fusion features.
[0071] In some embodiments, the identification module is configured to input the statistical features into a deep perception machine to obtain deep learning features corresponding to the statistical features.
[0072] In some embodiments, the feature extraction module is configured to perform denoising processing on the first audio data to obtain denoised first audio data, and perform feature extraction on the denoised first audio data to obtain time-frequency map features corresponding to the first audio data.
[0073] In some embodiments, the feature extraction module is configured to perform denoising processing on the first audio data to obtain denoised first audio data by using a low-pass filter.
[0074] In some embodiments, a frequency threshold of the low-pass filter is determined according to a driving speed of the vehicle.
[0075] In some embodiments, the acquisition module is further configured to acquire second audio data, the second audio data being tire noise audio data of a second tire of the vehicle driving on the to-be-identified road surface;
[0076] The identification module is further configured to identify a road surface type based on the second audio data, to obtain a second road surface type of the to-be-identified road surface.
[0077] The road surface type identification apparatus further comprises an arbitration module configured to obtain a target road surface type of the to-be-identified road surface based on the first road surface type and the second road surface type.
[0078] In some embodiments, the identification module is further configured to perform feature extraction on the second audio data to obtain a time-frequency map feature corresponding to the second audio data, and perform road surface type identification based on the time-frequency map feature corresponding to the second audio data to obtain a second road surface type of the to-be-identified road surface.
[0079] In some embodiments, the arbitration module is configured to: if the first road surface type is the same as the second road surface type, determine the first road surface type or the second road surface type as the target road surface type; and if the first road surface type is different from the second road surface type, determine a target road surface type of the to-be-identified road surface according to confidences corresponding to the first road surface type and the second road surface type respectively.
[0080] In some embodiments, the arbitration module is configured to: if one of the confidences corresponding to the first road surface type and the second road surface type respectively is greater than a preset threshold, and the other is less than or equal to the preset threshold, determine the road surface type with the confidence greater than the preset threshold as the target road surface type.
[0081] If the confidences corresponding to the first road surface type and the second road surface type respectively are all greater than the preset threshold or all less than the preset threshold, determine the road surface type with a higher priority as the target road surface type according to priorities of the first road surface type and the second road surface type.
[0082] In some embodiments, the arbitration module is configured to determine the road surface type with a higher confidence in the first road surface type and the second road surface type as the target road surface type of the to-be-identified road surface.
[0083] In some embodiments, the identification module is configured to input the time-frequency graph feature corresponding to the first audio data into a target road surface type identification model to obtain a first road surface type of the road surface to be identified.
[0084] In some embodiments, the target road surface type identification model is obtained by training a candidate road surface type identification model based on training samples, and the training samples include: an audio data sample, the audio data sample including: audio data content and a road surface type label corresponding to the audio data.
[0085] In some embodiments, the road surface type identification apparatus further includes a preprocessing module configured to obtain a plurality of audio data samples by performing data enhancement processing on the audio data sample.
[0086] The data enhancement processing includes at least one of: pitch enhancement processing, pitch height enhancement processing, time translation processing, or time mask processing.
[0087] In a third aspect, some embodiments of the present disclosure provide a controller, including: a processor and a memory, the memory storing programs or instructions executable on the processor, and the programs or instructions, when executed by the processor, implement the steps of the road surface type identification method according to any one of the first aspect.
[0088] In a fourth aspect, some embodiments of the present disclosure provide a vehicle, including: a processor and a memory, the memory storing programs or instructions executable on the processor, and the programs or instructions, when executed by the processor, implement the steps of the road surface type identification method according to any one of the first aspect.
[0089] In a fifth aspect, some embodiments of the present disclosure provide a cloud server, including: a processor and a memory, the memory storing programs or instructions executable on the processor, and the programs or instructions, when executed by the processor, implement the steps of the road surface type identification method according to any one of the first aspect.
[0090] In a sixth aspect, some embodiments of the present disclosure provide a computer readable storage medium, the readable storage medium storing programs or instructions, and the programs or instructions, when executed by a processor, implement the steps of the road surface type identification method according to any one of the first aspect.
[0091] In a seventh aspect, some embodiments of the present disclosure provide a computer program product, the program product, when executed by a processor of a vehicle or a cloud server, implements the steps of the road surface type identification method according to any one of the first aspect.
[0092] Some embodiments of the present disclosure obtain first audio data, obtain a first road surface type of a to-be-identified road surface based on the first audio data, and reflect the road surface state through the first audio data. Since the first audio data is not affected by factors such as light and weather, the accuracy of road surface identification can be improved. In addition, the road surface type is identified based on the time-frequency graph feature corresponding to the first audio data, and the first road surface type of the to-be-identified road surface is obtained. Since the time-frequency graph feature is a two-dimensional feature, that is, the time-frequency graph feature includes time information and frequency information, it has more information reflecting the road surface state, and therefore the accuracy of road surface type identification can be further improved. BRIEF DESCRIPTION OF DRAWINGS
[0093] FIG. 1 is a flowchart of a road surface type identification method according to some embodiments;
[0094] FIG. 2 is a flowchart of a road surface type identification method according to some embodiments;
[0095] FIG. 3A is a schematic diagram of the installation position of an acoustic sensor according to some embodiments;
[0096] FIG. 3B is a schematic diagram of the installation position of another acoustic sensor according to some embodiments;
[0097] FIG. 4 is a flowchart of another road surface type identification method according to some embodiments;
[0098] FIG. 5 is a schematic diagram of the structure of a deep convolutional neural network according to some embodiments;
[0099] FIG. 6 is a schematic diagram of the structure of a deep residual network according to some embodiments;
[0100] FIG. 7 is a flowchart of another road surface type identification method according to some embodiments;
[0101] FIG. 8 is a flowchart of another road surface type identification method according to some embodiments;
[0102] FIG. 9 is a schematic diagram of the structure of a road surface type identification device according to some embodiments;
[0103] FIG. 10 is a flowchart of another road surface type identification method according to some embodiments;
[0104] FIG. 11 is a schematic diagram of the structure of another road surface type identification device according to some embodiments;
[0105] FIG. 12 is a flowchart of another road surface type identification method according to some embodiments;
[0106] FIG. 13 is a flowchart of a road surface type identification method according to some embodiments;
[0107] FIG. 14 is a flowchart of yet another road surface type identification method according to some embodiments;
[0108] FIG. 15 is a flowchart of an arbitration method according to some embodiments;
[0109] FIG. 16 is a flowchart of a model training according to some embodiments;
[0110] FIG. 17 is an architecture diagram of a road surface type identification system according to some embodiments;
[0111] FIG. 18 is a flowchart of a data acquisition and preprocessing method according to some embodiments;
[0112] FIG. 19A is a schematic diagram of a training process of a candidate road surface type identification model according to some embodiments;
[0113] FIG. 19B is a schematic diagram of a training process of another candidate road surface type identification model according to some embodiments;
[0114] FIG. 20 is a confusion matrix diagram of a validation set according to some embodiments;
[0115] FIG. 21 is a block diagram of a road surface type identification device according to some embodiments. DETAILED DESCRIPTION
[0116] The technical solutions in the embodiments of the present disclosure will be clearly described below with reference to the drawings of the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, rather than all the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by a person of ordinary skill in the art belong to the scope of protection of the present disclosure.
[0117] In order to improve the safety and comfort of vehicle driving, intelligent vehicles are equipped with more and more driving modes matched with road surface types, so that accurate identification of road surface types becomes a key technology for intelligent vehicles. In related technologies, road surface types are usually distinguished by collecting image information of the road surface on which the vehicle is driving. For example, the road surface image is captured by using the vehicle-mounted camera system, and then the system can identify different road surface types by analyzing the image data, and adjust the driving mode of the vehicle accordingly to adapt to the road conditions. However, this way is easily disturbed by external environmental factors, such as weather conditions and light, which may affect the collection of image data that can fully reflect the road surface state, resulting in low accuracy of road surface type identification.
[0118] In order to overcome the above problems, some embodiments of the present disclosure utilize the characteristics of auditory signals that are not affected by factors such as light and weather, and reflect the road surface state through tire noise audio data to improve the accuracy of road surface type identification.
[0119] To further improve the accuracy of the road surface type identification, some embodiments of the present disclosure further perform feature extraction on the tire noise audio data, and identify the road surface type based on the time-frequency graph features corresponding to the tire noise audio data. Since the time-frequency graph features are two-dimensional features, they have more abundant information reflecting the road surface state, and thus can improve the accuracy of the road surface type identification.
[0120] Some embodiments of the present disclosure further extract at least two types of time-frequency graph features based on the tire noise audio data, perform feature fusion on the at least two types of time-frequency graph features, and identify the road surface type based on the fused features. Since the at least two types of time-frequency graph features can more comprehensively and abundantly reflect the road surface state information from different angles, the accuracy of the road surface type identification is further improved.
[0121] Some embodiments of the present disclosure further identify the road surface type for at least two tires of the vehicle respectively, and arbitrate the identification results of the road surface type identification for the at least two tires respectively, and further improve the accuracy of the road surface type identification.
[0122] In some embodiments, in the arbitration process described above, the priority of the road surface type is combined to determine the final road surface type identification result, so that the determined road surface type is more reasonable.
[0123] Some embodiments of the present disclosure are described below.
[0124] As shown in FIG. 1, some embodiments of the present disclosure provide a road surface type identification method, which can be executed by a cloud server in communication connection with a vehicle, can be executed by the vehicle, or can be executed by the cloud server and the vehicle in cooperation. The execution mode can be determined according to the actual application scenario, and the present disclosure does not limit this. The method of some embodiments of the present disclosure includes S10 to S11.
[0125] S10: Obtain first audio data.
[0126] Here, the first audio data is tire noise audio data of a first tire of a vehicle driving on a road surface to be identified.
[0127] In some embodiments, the first audio data can be collected by an acoustic sensor, which can be a microphone. The acoustic sensor is arranged on a suspension close to the tire to be collected. The collected tire can be any one of the tires of the vehicle, which can be determined according to the driving condition of the vehicle. For example, the rear tire of the vehicle, such as the left rear tire or the right rear tire, can be collected. When the left rear tire or the right rear tire is collected, the installation position of the microphone 21 is shown in FIGS. 3A and 3B.
[0128] In some embodiments, the first audio data can be acquired in a real-time manner. The first audio data can also be acquired in a condition-triggered manner, for example, the acquisition of the first audio data can be triggered in combination with vehicle positioning data, so as to reduce the energy consumption requirement of the vehicle without real-time acquisition all the time. For example, a highway is generally a bituminous pavement, and thus the positioning data can be collected, and the acquisition of the first audio data can be triggered to stop after it is detected that the vehicle enters the highway, and the acquisition of the first audio data can be triggered to start after it is detected that the vehicle exits the highway.
[0129] S11: obtaining a first road surface type of the to-be-identified road surface based on the first audio data.
[0130] In some embodiments, as shown in FIG. 2, S11 can include S111 to S112.
[0131] S111: performing feature extraction on the first audio data to obtain a time-frequency map feature corresponding to the first audio data.
[0132] In some embodiments, the first audio data can be subjected to feature extraction by a mel-frequency spectrum or wavelet transform to obtain a time-frequency map feature corresponding to the first audio data.
[0133] S112: performing road surface type identification based on the time-frequency map feature corresponding to the first audio data to obtain the first road surface type of the to-be-identified road surface.
[0134] In some embodiments, a corresponding relationship between the time-frequency map feature and the road surface type can be established, and the first road surface type of the to-be-identified road surface can be obtained based on the corresponding relationship.
[0135] In some embodiments, the time-frequency map feature corresponding to the first audio data can be input into a target road surface type identification model to obtain the first road surface type of the to-be-identified road surface. Here, the target road surface type identification model is obtained by training based on training samples, and the training samples can include audio data samples. The audio data samples include audio data content and a road surface type label corresponding to the audio data.
[0136] The first road surface type can be dry bituminous, wet bituminous, dry grassland, wet grassland, dry mud, wet mud, dry sand, wet sand, hard snow, soft snow, etc.
[0137] In some embodiments of the present disclosure, a first road surface type of the to-be-identified road surface is obtained by acquiring first audio data and based on the first audio data. The road surface state is reflected by the first audio data, and since the first audio data is not affected by factors such as light and weather, the accuracy of road surface identification can be improved. Moreover, the first road surface type of the to-be-identified road surface is obtained by identifying the road surface type based on the time-frequency graph feature corresponding to the first audio data, and since the time-frequency graph feature is a two-dimensional feature, that is, the time-frequency graph feature includes time information and frequency information, the time-frequency graph feature has more information reflecting the road surface state, and thus the accuracy of road surface type identification can be further improved.
[0138] In some embodiments, on the basis of the embodiment shown in FIG. 2, as shown in FIG. 4, S112 can further include S1121-S1122.
[0139] S1121: inputting the time-frequency graph feature corresponding to the first audio data into a deep learning model to perform feature extraction, to obtain a deep learning feature corresponding to the time-frequency graph feature.
[0140] In some embodiments, the deep learning feature extraction on the time-frequency graph feature can be performed by a deep convolutional neural network or a deep residual network, to obtain the deep learning feature corresponding to the time-frequency graph feature.
[0141] In some embodiments, the structure of the deep convolutional neural network is shown in FIG. 5, which is an 83-layer neural network model containing 13 identical convolution modules. Each convolution module in the 13 identical convolution modules is composed of a convolution layer, a batch normalization layer, an activation function layer, a group convolution layer, a batch normalization layer and an activation function layer in sequence. Here, the convolution kernel size of the convolution layer is 3x3. The group convolution layer is used to reduce the model parameters, and the size of the convolution kernel thereof is 3x3, and the number of convolution groups is equal to the number of channels of the input feature map. After the last convolution module in the 13 identical convolution modules, a convolution layer, a batch normalization layer, an activation function layer, a global average pooling layer and a fully connected layer are sequentially followed, and the size of the output of the fully connected layer is the number of deep learning feature extraction, which is 96 in some embodiments.
[0142] In some embodiments, the structure of the deep residual network is shown in FIG. 6, which contains 8 residual modules with attention mechanism. The wavelet transform time-frequency graph can be used as the input thereof, and then the wavelet transform time-frequency graph is first subjected to a convolution layer with a convolution kernel size of 7x7 and a step size of 2, and then subjected to a batch normalization layer, an activation function layer and a maximum pooling layer with a pooling kernel size of 3x3 and a step size of 2 in sequence, and then transmitted to the residual module with attention mechanism for feature extraction. The residual module with attention mechanism is composed of a main path and a shortcut path.
[0143] Here, the main path comprises a convolution module and an attention module, the convolution module comprises, in sequence, a convolution layer, a batch normalization layer, an activation function layer, a convolution layer, and a batch normalization layer, and is configured to extract features; the attention module comprises, in sequence, a global average pooling layer, a fully connected layer, an activation function layer, a fully connected layer, and an activation function layer, the last activation function layer is a Sigmoid function, configured to fix the output value in the range of [0, 1] as a weight; finally, the product of the feature map extracted by the convolution module and the weight information extracted by the attention module is taken as the output of the main path. The shortcut path comprises a convolution layer and a batch normalization layer, and is configured to extract features. After the features are extracted by the main path and the shortcut path, the two feature maps are added, and then the output of the residual module with the attention mechanism is obtained after the activation function layer. After the last residual module with the attention mechanism in the eight residual modules with the attention mechanism, a global average pooling layer and a fully connected layer are sequentially followed, and the size of the output of the fully connected layer is the number of deep learning features, which is 96 in some embodiments.
[0144] S1122: Obtain the first road surface type of the road surface to be identified based on the deep learning feature.
[0145] In some embodiments, the deep learning feature is input into a classifier to obtain the first road surface type of the road surface to be identified.
[0146] In some embodiments of the present disclosure, the deep learning feature corresponding to the time-frequency graph feature of the first audio data is obtained by deep learning feature extraction on the time-frequency graph feature, and the first road surface type of the road surface to be identified is obtained based on the deep learning feature. The first road surface type is extracted by deep learning feature extraction, which can further extract more rich feature information, and therefore can more accurately reflect the road surface condition and further improve the accuracy of road surface type identification.
[0147] On the basis of the embodiment shown in FIG. 4, in the case where the time-frequency graph feature corresponding to the first audio data comprises at least two types of time-frequency graph features, as shown in FIG. 7, S1121 further comprises a process of fusion processing of deep learning features corresponding to the at least two types of time-frequency graph features, which comprises S11211 to S11212.
[0148] S11211: Input the at least two types of time-frequency graph features into a deep learning model for feature extraction to obtain deep learning features corresponding to each type of time-frequency graph feature in the at least two types of time-frequency graph features.
[0149] In some embodiments, the at least two types of time-frequency map features include a mel-frequency time-frequency map feature and a wavelet transform time-frequency map feature. At this time, the mel-frequency time-frequency map feature is input into a deep convolutional neural network, where the structure of the deep convolutional neural network can be as shown in FIG. 5, to obtain a deep learning feature corresponding to the mel-frequency time-frequency map feature; and the wavelet transform time-frequency map feature is input into a deep residual network, where the structure of the deep residual network can be as shown in FIG. 6, to obtain a deep learning feature of the wavelet transform time-frequency map.
[0150] S11212: performing fusion processing on the deep learning feature corresponding to each type of time-frequency map feature in the obtained at least two types of time-frequency map features, to obtain a fusion feature.
[0151] In some embodiments, the fusion processing can be performed by feature map splicing, that is, two 96-dimensional feature vectors are spliced into one 192-dimensional feature vector.
[0152] In some embodiments, the implementation of S1122 is as shown in S1122'.
[0153] S1122': obtaining the first road surface type of the to-be-identified road surface based on the fusion feature.
[0154] In some embodiments, the fusion feature can be input into a classifier to obtain the first road surface type of the to-be-identified road surface.
[0155] In some embodiments of the present disclosure, the deep learning feature corresponding to each type of time-frequency map feature in the at least two types of time-frequency map features is obtained by performing deep learning feature extraction on the at least two types of time-frequency map features respectively, the fusion feature is obtained by performing fusion processing on the deep learning feature corresponding to each type of time-frequency map feature in the obtained at least two types of time-frequency map features, and the first road surface type of the to-be-identified road surface is obtained based on the fusion feature. Since the at least two types of time-frequency map features are extracted by the corresponding deep neural network respectively, more comprehensive and rich road surface state information is obtained from different angles, and therefore the accuracy of road surface type identification can be further improved.
[0156] In some embodiments, FIG. 8 is described based on the embodiment shown in FIG. 7, taking the mel-frequency time-frequency map feature and the wavelet transform time-frequency map feature as examples.
[0157] In some embodiments, as shown in FIG. 9, some embodiments of the present disclosure further provide a road surface type identification device. The road surface type identification device includes an input module and a target road surface type identification model.
[0158] The target road surface type recognition model comprises a deep learning feature extraction module, a feature fusion module and a classifier. The input module is configured to perform the steps of receiving the first audio data and feature extraction, the feature extraction step (such as S111) comprising S1111 and S1112 (shown in FIG. 8).
[0159] The deep learning feature extraction module is a multi-input neural network model, which inputs are mel-frequency-time-frequency feature and wavelet transform time-frequency feature; and a deep learning feature extraction network corresponding to each input is designed, which are deep convolutional neural network and deep residual network respectively.
[0160] The deep learning feature extraction module is configured to perform S112111 and S112112 (shown in FIG. 8), then the deep learning features of the two inputs are fused by the feature fusion module, the feature fusion module is configured to perform S11212' (shown in FIG. 8), and finally the fused features are recognized by the classifier, the classifier is configured to perform S1122' (shown in FIG. 8).
[0161] S1111: performing short-time Fourier transform on the first audio data to obtain the frequency spectrum of the first audio data, filtering the frequency spectrum by a mel filter bank to obtain the mel-frequency-time-frequency feature corresponding to the first audio data.
[0162] In some embodiments, the first audio data is subjected to short-time Fourier transform to obtain the frequency spectrum of the first audio data. Taking the sampling frequency of the microphone as 44100Hz for example, the first audio data needs to be subjected to framing and windowing processing before short-time Fourier transform. Here, the frame length can be set to 2205, the overlapping part between frames is set to 1764, the Hamming window function is used for windowing processing and the window length is set to 2205, and then Fourier transform is performed on each frame of the first audio data with the Fourier length set to 4096. Then, the frequency spectrum is filtered by a mel filter bank with the number of mel filter banks set to 64, and after processing, logarithmic operation is performed to finally obtain the mel-frequency-time-frequency feature corresponding to the first audio data.
[0163] S1112: performing wavelet transform on the first audio data to obtain the wavelet transform time-frequency feature corresponding to the first audio data.
[0164] In some embodiments, the wavelet transform is performed on the first audio data, where the wavelet basis function is Morse potential function, the wavelet transform time-frequency feature corresponding to the first audio data is generated, and it is converted into an RGB image with a size of 224x224.
[0165] S112111: input the mel-frequency spectrogram feature into a deep convolutional neural network to obtain a deep learning feature corresponding to the mel-frequency spectrogram feature.
[0166] Here, the description of the deep convolutional neural network is detailed in the foregoing detailed description of FIG. 5, which will not be repeated here.
[0167] S112112: input the wavelet transform spectrogram feature into a deep residual network to obtain a deep learning feature of the wavelet transform spectrogram.
[0168] Here, the description of the deep residual neural network is detailed in the foregoing detailed description of FIG. 6, which will not be repeated here.
[0169] S11212’: fuse the deep learning feature corresponding to the mel-frequency spectrogram feature and the deep learning feature of the wavelet transform spectrogram feature to obtain a fusion feature.
[0170] S1122’: obtain the first road surface type of the road surface to be identified based on the fusion feature.
[0171] In some embodiments, the fusion feature can be input into a classifier to obtain the first road surface type of the road surface to be identified.
[0172] In some embodiments, the classifier network is composed of a fully connected layer, an activation function layer, a dropout layer, a fully connected layer, and a normalization exponential (Softmax) layer. The first fully connected layer has an input size of 192 and an output size of 256 or 512; the activation function layer is a rectified linear unit (ReLU) activation function; the dropout rate of the dropout layer is 0.2, in order to increase the robustness of the model; the last fully connected layer has an input size of 256 or 512 and an output size of the number of road surface types; and finally, the Softmax layer is used to obtain the prediction probability of each road surface type, i.e., the confidence.
[0173] In some embodiments of the present disclosure, the mel-frequency spectrogram feature and the wavelet transform spectrogram feature are subjected to deep learning feature extraction to obtain deep learning features corresponding to the mel-frequency spectrogram feature and the wavelet transform spectrogram feature, respectively. The deep learning features corresponding to the mel-frequency spectrogram feature and the wavelet transform spectrogram feature, respectively, are fused to obtain a fusion feature, and the first road surface type of the road surface to be identified is obtained based on the fusion feature. Since the mel-frequency spectrogram feature and the wavelet transform spectrogram feature are subjected to feature extraction by the deep neural network corresponding thereto, respectively, more comprehensive and rich road surface state information is obtained from different angles, thereby further improving the accuracy of road surface type identification.
[0174] Fig. 10 is a diagram illustrating a method of fusing the statistical features corresponding to the first audio data into the embodiment shown in Fig. 8, and fusing the statistical features into the embodiment shown in Fig. 9, as shown in Figs. 10 and 11. The input module further performs S1113, i.e., extraction of the statistical features, in the embodiment shown in Fig. 9. The deep learning feature extraction module is configured to perform S112113, and the feature fusion module is configured to perform S11212”.
[0175] S1113: performing statistical feature extraction on the first audio data to obtain statistical features corresponding to the first audio data.
[0176] In some embodiments, the statistical features corresponding to the first audio data include a feature vector composed of mean, variance, standard deviation, median, quartile, maximum value, minimum value, root mean square, etc.
[0177] Mean:
[0178] Variance:
[0179] Standard deviation:
[0180] Median: Median{X1, X2, … Xn}. n
[0181] Quartile: Quartile{X1, X2, … Xn}. n
[0182] Maximum value: Max{X1, X2, … Xn}. n
[0183] Minimum value: Min{X1, X2, … Xn}. n
[0184] Root mean square:
[0185] where X i represents the i-th audio data, and n represents the number of audio data.
[0186] After obtaining the statistical features corresponding to the first audio data, the statistical features are input into a deep learning model for feature extraction to obtain deep learning features corresponding to the statistical features, as shown in S112113 in Fig. 10.
[0187] S112113: inputting the statistical features into a deep perception machine to obtain deep learning features corresponding to the statistical features.
[0188] Here, the deep perception machine is a multi-layer perception machine (MLP) including an input layer, three hidden layers, and an output layer. Here, the size of the input layer is set to the number of statistical features, the sizes of the three hidden layers are set to 64, 256, and 512 respectively, and the size of the output layer is set to the number of deep learning feature extraction, which is 96 in some embodiments. Each layer in the MLP is implemented through a fully connected layer, and a ReLU activation function is added after each fully connected layer to enhance the non-linear fitting capability of the network.
[0189] In some embodiments, the deep learning features corresponding to each of the at least two types of time-frequency graph features and the deep learning features corresponding to the statistical features are fused to obtain fused features, such as S11212”.
[0190] S11212”: The deep learning features corresponding to the statistical features, the deep learning features corresponding to the mel time-frequency graph features, and the deep learning features of the wavelet transform time-frequency graph features are fused to obtain fused features.
[0191] In some embodiments of the present disclosure, deep learning feature extraction is performed on the statistical features, the mel time-frequency graph features, and the wavelet time-frequency graph features to obtain deep learning features corresponding to the statistical features, the mel time-frequency graph features, and the wavelet time-frequency graph features respectively, and the deep learning features corresponding to the statistical features, the mel time-frequency graph features, and the wavelet time-frequency graph features obtained are fused to obtain fused features. Based on the fused features, the first road surface type of the road surface to be identified is obtained. Since the statistical features, the mel time-frequency graph features, and the wavelet time-frequency graph features are respectively extracted by the deep neural network corresponding thereto, more comprehensive and rich road surface state information is obtained from different angles, and therefore the accuracy of road surface type identification is further improved.
[0192] On the basis of the above embodiments, as shown in FIG. 12, the above method further includes: performing denoising processing on the first audio data before feature extraction, at this time, S111 can further include S1110 and S111’.
[0193] S1110: Denoising processing of environmental interference noise is performed on the first audio data to obtain denoised first audio data.
[0194] In some embodiments, since the frequency of environmental noise is generally greater than the frequency of tire noise, a low-pass filter can be used to perform environmental interference noise denoising processing on the first audio data to obtain denoised first audio data. Here, the frequency threshold of the low-pass filter can be determined according to the driving speed of the vehicle, for example, when the driving speed of the vehicle is less than or equal to 50 km / h, the frequency threshold of the low-pass filter is set to 2500 Hz; when the driving speed of the vehicle is greater than 50 km / h, the frequency threshold of the low-pass filter is set to 5000 Hz.
[0195] S111’:performing feature extraction on the denoised first audio data to obtain time-frequency map features corresponding to the first audio data.
[0196] The step of performing feature extraction on the denoised first audio data is described in detail in the foregoing embodiments, which will not be described here.
[0197] In some embodiments of the present disclosure, by performing environmental interference noise denoising processing on the first audio data to obtain denoised first audio data, and performing feature extraction on the denoised first audio data to obtain time-frequency map features corresponding to the first audio data, the quality of the first audio data is improved, and the road surface condition can be more accurately reflected, thus further improving the accuracy of road surface type identification.
[0198] The above embodiments are described by taking the collection of tire noise audio data of one tire as an example, and FIG. 13 is described by taking the collection of tire noise audio data of two tires as an example, and the road surface type identification is performed based on the tire noise audio data of each of the two tires. In this way, the road surface type identification result corresponding to each of the two tires can be obtained, and the final road surface type is determined by arbitration based on the road surface type identification results of the two tires.
[0199] For ease of description, the audio data of one tire is described as first audio data, the audio data of another tire is described as second audio data, the road surface type identification result based on the first audio data is described as first road surface type, the road surface type identification result based on the second audio data is described as second road surface type, and the final road surface type based on the first road surface type and the second road surface type is described as target road surface type. At this time, the identification method further includes S20 to S22.
[0200] S20: obtaining second audio data.
[0201] Here, the second audio data is the tire noise audio data of the second tire of the vehicle driving on the road surface to be identified.
[0202] S21: performing road type identification based on the second audio data to obtain a second road type of the to-be-identified road.
[0203] S22: obtaining a target road type of the to-be-identified road based on the first road type and the second road type.
[0204] In some embodiments, as shown in FIG. 14, S21 can further include S211 and S212:
[0205] S211: performing feature extraction on the second audio data to obtain a time-frequency graph feature corresponding to the second audio data.
[0206] S212: performing road type identification based on the time-frequency graph feature corresponding to the second audio data to obtain a second road type of the to-be-identified road.
[0207] Here, the detailed description of S20, S21, S211, and S212 can refer to the description of the process based on the first audio data processing in the foregoing embodiments, which will not be described herein again.
[0208] In some embodiments, as shown in FIG. 15, the identification method further includes S1501 to S1503.
[0209] S1501: determining whether the first road type and the second road type are the same, if yes, performing S1502, and if no, performing S1503.
[0210] S1502: determining that the first road type or the second road type is the target road type.
[0211] If the first road type and the second road type are the same, it is determined that the first road type or the second road type is the target road type.
[0212] S1503: determining the target road type of the to-be-identified road according to the confidence degrees corresponding to the first road type and the second road type, respectively.
[0213] If the first road type and the second road type are different, the target road type of the to-be-identified road is determined according to the confidence degrees corresponding to the first road type and the second road type, respectively.
[0214] In some embodiments, determining the target road type of the to-be-identified road according to the confidence degrees corresponding to the first road type and the second road type, respectively, can include: if one of the confidence degrees corresponding to the first road type and the second road type is greater than a preset threshold, and the other confidence degree is less than or equal to the preset threshold, determining the road type with the confidence degree greater than the preset threshold as the target road type.
[0215] If the confidence degrees corresponding to the first road surface type and the second road surface type are both greater than the preset threshold or both less than the preset threshold, a road surface type with a higher priority is determined as the target road surface type according to the priorities of the first road surface type and the second road surface type.
[0216] In some embodiments, taking the preset threshold as 0.9 for example, it is assumed that the confidence degree of the first road surface type is greater than 0.9 and the confidence degree of the second road surface type is less than 0.9 in the confidence degrees of the first road surface type and the second road surface type, and then the first road surface type is determined as the target road surface type. Conversely, if the confidence degree of the first road surface type is less than 0.9 and the confidence degree of the second road surface type is greater than 0.9 in the confidence degrees of the first road surface type and the second road surface type, the second road surface type is determined as the target road surface type. If the confidence degrees of the first road surface type and the second road surface type are both greater than 0.9 or both less than 0.9, a road surface type with a higher priority is determined as the target road surface type according to the priorities of the first road surface type and the second road surface type.
[0217] Here, the priorities of the first road surface type and the second road surface type are related to driving difficulties corresponding to the road surface types, and the greater the driving difficulty, the higher the priority. For example, the priority of a wet mud road is higher than the priority of a dry asphalt road.
[0218] In some other embodiments, determining the target road surface type of the to-be-identified road surface according to the confidence degrees corresponding to the first road surface type and the second road surface type can include: determining a road surface type with a greater confidence degree in the first road surface type and the second road surface type as the target road surface type of the to-be-identified road surface.
[0219] In some embodiments of the present disclosure, the target road surface type of the to-be-identified road surface is obtained based on the first road surface type and the second road surface type, that is, the audio data corresponding to the two tires are respectively subjected to road surface type identification, and the two identification results are arbitrated to determine the final target road surface type, thereby further improving the accuracy of road surface type identification.
[0220] In each of the above embodiments, the road surface type identification based on the time-frequency feature map features corresponding to the audio data can be performed by a target road surface type identification model. For example, the time-frequency feature corresponding to the first audio data is input into the target road surface type identification model to obtain the first road surface type corresponding to the to-be-identified road surface.
[0221] Here, the target road surface type identification model is obtained by training a candidate road surface type identification model based on training samples. In some embodiments, the training samples include audio data samples. The audio data samples include audio data content and road surface type labels corresponding to the audio data.
[0222] In some embodiments, as the number of collected training data increases, the target road surface type identification model can also be continuously optimized based on new training data. As shown in FIG. 16, the identification method can also include S1601 to S1612.
[0223] S1601: Obtain training data.
[0224] S1602: Construct a training data set.
[0225] S1603: Determine whether there is a target road surface type identification model. If yes, perform S1604; if no, perform S1605.
[0226] S1604: Determine whether there is an optimal model parameter. If yes, perform S1608; if no, perform S1611.
[0227] S1605: Build a candidate road surface type identification model.
[0228] S1606: Perform model training.
[0229] S1607: After the model converges, save the optimal model parameter.
[0230] S1608: Load the optimal model parameter.
[0231] S1609: Perform model training.
[0232] S1610: After the model converges, update the optimal model parameter.
[0233] S1611: Perform model training.
[0234] S1612: After the model converges, save the optimal model parameter.
[0235] As shown in FIG. 17, the road surface type identification system 1700 according to some embodiments of the present disclosure includes a cloud server 1701, a left rear wheel microphone 1702 for collecting first audio data of the left rear wheel, a right rear wheel microphone 1703 for collecting second audio data of the right rear wheel, a data collection module 1704 for collecting data collected by the microphones, a data preprocessing module 1705 for filtering out environmental interference noise and performing data enhancement processing on the tire noise audio data, a target road surface type identification module 1706, an arbitration module 1707, and a driving mode selection module 1708 for selecting a suitable driving mode according to the target road surface type. Here, the cloud server 1701 is used for model training and updating the target road surface type identification module 1706. The steps performed by the other modules are described above in the related description of the previous embodiments, and will not be described again here.
[0236] During the process of obtaining the training samples, at least one of the sample data can be subjected to environmental interference noise filtering processing or data enhancement processing. As shown in FIG. 18, the environmental interference noise filtering processing is performed to improve the quality of the sample data. The environmental interference noise filtering processing is mainly low-pass filtering processing, and details are described above, and will not be described again here. The data enhancement processing is performed to expand the sample data and increase the number of sample data. Here, the data enhancement processing includes at least one of the following: pitch enhancement processing, pitch enhancement processing, time shift processing, or time mask processing.
[0237] According to some embodiments of the present disclosure, the target road surface type identification model is obtained by training the candidate road surface type identification model based on the training samples, and the road surface type identification is performed based on the target road surface type identification model, thereby improving the accuracy of the road surface type identification. Through the expansion processing, the number of training samples is increased, and the robustness of the target road surface type identification model is enhanced.
[0238] In some embodiments, the candidate road surface type identification model is implemented by using Matlab, and the experimental equipment is a notebook computer with an external GeForce GTX 3080 graphics card. The memory of the graphics card (the memory in the graphics card) is 10 GB.
[0239] The training samples used are collected by experimental personnel in real vehicles. The training samples contain 10760 pieces of tire noise audio data, and contain 11 road surface types, which are dry asphalt, dry cement road surface, dry grassland, dry mud, dry sand, hard sand, wet asphalt, wet broken cement road surface, wet sand, wet soil road, and wet sandy gravel road surface.
[0240] First, statistical features, mel-frequency time-frequency feature and wavelet transform time-frequency feature of the fetal noise audio data are extracted respectively; then, three types of feature data sets are combined, and the label of the combined feature data set is the label of the original fetal noise audio; finally, the training set and the validation set are divided according to the ratio of 8:2.
[0241] In the training process of the candidate road surface type identification model, the number of iterations of some embodiments of the present disclosure is 25, the number of selected samples (batch_size) for one training is set to 256, the loss function is a cross-entropy loss function, and the optimizer is an Adam optimizer. In terms of learning rate, this embodiment adopts a gradually decreasing strategy, for example, the learning rate is initially set to 0.01, and the learning rate is reduced by 10 times every 10 iterations.
[0242] It can be understood that the loss of the candidate road surface type identification model is large when it starts training, and the learning rate needs to be set larger to speed up the convergence of the model. As the training process continues, the loss becomes smaller, and to prevent the candidate road surface type identification model from oscillating around the optimal value, the learning rate needs to be set smaller at this time. Such a processing method can make the model converge to the optimal value faster and reduce the training time of the model.
[0243] During the training of the model, the loss of the model decreases rapidly and the accuracy increases rapidly, as shown in FIG. 19A, after a preset number of iterations (such as 12 times), the loss tends to stabilize at about 0.05, as shown in FIG. 19B, after a preset number of iterations (such as 12 times), the accuracy of the validation set is about 0.98. In order to understand the accuracy of each ground surface category in the training sample in detail, as shown in FIG. 20, some embodiments of the present disclosure also draw a confusion matrix diagram of the validation set.
[0244] In some embodiments, as shown in FIG. 21, the road surface type identification device 2100 of some embodiments of the present disclosure includes an acquisition module 2101, an identification module 2102 and an arbitration module 2103. The acquisition module 2101 is configured to acquire first audio data, the first audio data being fetal noise audio data of a first tire of a vehicle driving on a to-be-identified road surface. The identification module 2102 is configured to obtain a first road surface type of the to-be-identified road surface based on the first audio data. The arbitration module 2103 is configured to obtain a target road surface type of the to-be-identified road surface based on the first road surface type and the second road surface type.
[0245] In some embodiments, the acquisition module 2101 is further configured to perform data enhancement processing on the audio data sample to obtain a plurality of audio data samples.
[0246] Herein, the data enhancement processing includes at least one of: pitch enhancement processing, tone enhancement processing, time translation processing, or time mask processing.
[0247] In some embodiments, the identification module 2102 is further configured to perform feature extraction on the first audio data to obtain time-frequency graph features corresponding to the first audio data; and perform road surface type identification based on the time-frequency graph features corresponding to the first audio data to obtain the first road surface type of the to-be-identified road surface.
[0248] In some embodiments, the identification module 2102 is further configured to perform deep learning feature extraction on the time-frequency graph features corresponding to the first audio data to obtain deep learning features corresponding to the time-frequency graph features; and obtain the first road surface type of the to-be-identified road surface based on the deep learning features.
[0249] In some embodiments, the time-frequency graph features corresponding to the first audio data include at least two types of time-frequency graph features.
[0250] The identification module 2102 is further configured to input the time-frequency graph features corresponding to the first audio data into a deep learning model to perform feature extraction to obtain deep learning features corresponding to the time-frequency graph features; perform fusion processing on the deep learning features corresponding to each type of time-frequency graph features in the obtained at least two types of time-frequency graph features to obtain fused features; and obtain the first road surface type of the to-be-identified road surface based on the fused features.
[0251] In some embodiments, the at least two types of time-frequency graph features include mel-frequency time-frequency graph features and wavelet transform time-frequency graph features.
[0252] The identification module 2102 is further configured to input the mel-frequency time-frequency graph features into a deep convolutional neural network to obtain deep learning features corresponding to the mel-frequency time-frequency graph features; and input the wavelet transform time-frequency graph features into a deep residual network to obtain deep learning features of the wavelet transform time-frequency graph.
[0253] In some embodiments, the identification module 2102 is further configured to perform short-time Fourier transform on the first audio data to obtain a frequency spectrum of the first audio data, perform filter processing on the frequency spectrum through a mel filter bank to obtain mel-frequency time-frequency graph features corresponding to the first audio data; and perform wavelet transform on the first audio data to obtain wavelet transform time-frequency graph features corresponding to the first audio data.
[0254] In some embodiments, the identification module 2102 is further configured to perform statistical feature extraction on the first audio data to obtain statistical features corresponding to the first audio data.
[0255] The identification module is configured to perform deep learning feature extraction on the statistical features to obtain deep learning features corresponding to the statistical features, and perform fusion processing on the deep learning features corresponding to each of the at least two types of time-frequency graph features and the deep learning features corresponding to the statistical features to obtain fusion features.
[0256] In some embodiments, the identification module 2102 is further configured to input the statistical features into a deep perception machine to obtain deep learning features corresponding to the statistical features.
[0257] In some embodiments, the identification module 2102 is further configured to perform noise removal processing on the first audio data to obtain denoised first audio data, and perform feature extraction on the denoised first audio data to obtain time-frequency graph features corresponding to the first audio data.
[0258] In some embodiments, the identification module 2102 is further configured to perform noise removal processing on the first audio data using a low-pass filter to obtain denoised first audio data.
[0259] In some embodiments, a frequency threshold of the low-pass filter is determined according to a driving speed of the vehicle.
[0260] In some embodiments, the acquisition module 2101 is further configured to acquire second audio data, where the second audio data is tire noise audio data of a second tire of the vehicle driving on the to-be-identified road surface.
[0261] The identification module 2102 is further configured to identify a road surface type based on the second audio data to obtain a second road surface type of the to-be-identified road surface.
[0262] In some embodiments, the identification module 2102 is further configured to perform feature extraction on the second audio data to obtain time-frequency graph features corresponding to the second audio data, and identify a road surface type based on the time-frequency graph features corresponding to the second audio data to obtain a second road surface type of the to-be-identified road surface.
[0263] The road surface type identification apparatus 2100 further includes an arbitration module 2103 configured to obtain a target road surface type of the to-be-identified road surface based on the first road surface type and the second road surface type.
[0264] In some embodiments, the arbitration module 2103 is configured to: if the first road surface type is the same as the second road surface type, determine the first road surface type or the second road surface type as the target road surface type; and if the first road surface type is different from the second road surface type, determine the target road surface type of the to-be-identified road surface according to the confidence degrees corresponding to the first road surface type and the second road surface type respectively.
[0265] In some embodiments, the arbitration module 2103 is further configured to: if one of the confidence degrees corresponding to the first road surface type and the second road surface type is greater than a preset threshold, and the other confidence degree is less than or equal to the preset threshold, determine the road surface type with the confidence degree greater than the preset threshold as the target road surface type.
[0266] If the confidence degrees corresponding to the first road surface type and the second road surface type are both greater than the preset threshold or both less than the preset threshold, determine the road surface type with the higher priority as the target road surface type according to the priorities of the first road surface type and the second road surface type.
[0267] In some embodiments, the arbitration module 2103 is further configured to: determine the road surface type with the higher confidence degree as the target road surface type of the to-be-identified road surface from among the first road surface type and the second road surface type.
[0268] In some embodiments, the identification module 2102 is further configured to: input the time-frequency graph feature corresponding to the first audio data into a target road surface type identification model to obtain the first road surface type of the to-be-identified road surface.
[0269] In some embodiments, the target road surface type identification model is obtained by training a candidate road surface type identification model based on training samples, where the training samples include audio data samples. The audio data samples include audio data content and a road surface type label corresponding to the audio data.
[0270] The road surface type identification apparatus of some embodiments of the present disclosure can be used to implement the technical solutions of the above method embodiments, and has similar implementation principles and technical effects, which will not be described here.
[0271] Some embodiments of the present disclosure further provide a controller, including a processor and a memory, where the memory stores programs or instructions executable on the processor, and the programs or instructions are executed by the processor to implement the steps of the road surface type identification method embodiments described above.
[0272] Some embodiments of the present disclosure further provide a vehicle, comprising: a processor and a memory, the memory storing programs or instructions executable on the processor, the programs or instructions being executed by the processor to implement the steps of the road surface type identification method embodiments described above.
[0273] Some embodiments of the present disclosure further provide a cloud server, comprising: a processor and a memory, the memory storing programs or instructions executable on the processor, the programs or instructions being executed by the processor to implement the steps of the road surface type identification method embodiments described above.
[0274] Some embodiments of the present disclosure further provide a computer readable storage medium, the computer readable storage medium storing programs or instructions, the programs or instructions being executed by a processor to implement the steps of the road surface type identification method embodiments described above.
[0275] Some embodiments of the present disclosure further provide a computer program product, the computer program product being executed by a processor of a vehicle or a cloud server to implement the steps of the road surface type identification method embodiments described above.
[0276] It should be noted that, in this document, the terms "comprising", "including", or any other variant thereof are intended to cover non-exclusive inclusions, so that processes, methods, articles, or devices that include a series of elements not only include those elements, but also include other elements not explicitly listed, or inherent to such processes, methods, articles, or devices. Without more limitations, the element defined by the statement "comprising a" does not exclude the presence of additional identical elements in the process, method, article, or device that includes the element. In addition, it should be noted that the scope of the methods and devices in some embodiments of the present disclosure is not limited to the order of performing functions as shown or discussed, but can also include performing functions in a substantially simultaneous manner or in reverse order according to the functions involved, for example, the described method can be performed in an order different from the described order, and various steps can also be added, omitted or combined. In addition, the features described with reference to certain examples can be combined in other examples.
[0277] From the above description of the embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be realized by means of computer software products and general hardware platforms as necessary, of course, they can also be realized by hardware. The computer software product is stored in a storage medium (such as a read-only memory (ROM), a random access memory (RAM), a magnetic disk, an optical disk, etc.), and includes a plurality of instructions for making a terminal or a network side device execute the method described in various embodiments of the present disclosure.
[0278] The embodiments of the present disclosure are described above with reference to the accompanying drawings, but the present disclosure is not limited to the specific embodiments described above, and the specific embodiments described above are merely illustrative, but not restrictive, and a person of ordinary skill in the art can make many forms of implementation under the inspiration of the present disclosure without departing from the purpose of the present disclosure and the scope protected by the claims, and these implementations all belong to the protection of the present disclosure.
Claims
1. A road surface type identification method, comprising: Acquire first audio data, wherein the first audio data is tire noise frequency data of the first tire of a vehicle traveling on the road surface to be identified; Based on the first audio data, the first road surface type of the road surface to be identified is obtained.
2. The method according to claim 1, wherein, The step of obtaining the first road surface type of the road surface to be identified based on the first audio data includes: Feature extraction is performed on the first audio data to obtain the time-frequency graph features corresponding to the first audio data; Based on the time-frequency map features corresponding to the first audio data, road surface type identification is performed to obtain the first road surface type of the road surface to be identified.
3. The method according to claim 2, wherein, The step of identifying the road surface type based on the time-frequency map features corresponding to the first audio data to obtain the first road surface type of the road surface to be identified includes: The time-frequency graph features corresponding to the first audio data are input into a deep learning model for feature extraction to obtain the deep learning features corresponding to the time-frequency graph features. Based on the deep learning features, the first road surface type of the road surface to be identified is obtained.
4. The method according to claim 3, wherein, The time-frequency graph features corresponding to the first audio data include at least two types of time-frequency graph features; The step of inputting the time-frequency graph features corresponding to the first audio data into a deep learning model for feature extraction to obtain the deep learning features corresponding to the time-frequency graph features includes: The at least two types of time-frequency map features are respectively input into a deep learning model for feature extraction, to obtain the deep learning features corresponding to each type of time-frequency map feature in the at least two types of time-frequency map features; The deep learning features corresponding to each type of time-frequency map feature are fused to obtain fused features; The process of obtaining the first road surface type of the road surface to be identified based on the deep learning features includes: Based on the fusion features, the first road surface type of the road surface to be identified is obtained.
5. The method according to claim 4, wherein, The at least two types of time-frequency plot features include: Mel time-frequency plot features and wavelet transform time-frequency plot features; The step of inputting the at least two types of time-frequency map features into a deep learning model for feature extraction, to obtain the deep learning feature corresponding to each type of time-frequency map feature, includes: The Mel time-frequency map features are input into a deep convolutional neural network to obtain the deep learning features corresponding to the Mel time-frequency map features; The wavelet transform time-frequency map features are input into a deep residual network to obtain the deep learning features of the wavelet transform time-frequency map.
6. The method according to claim 5, wherein, The step of extracting features from the first audio data to obtain the time-frequency graph features corresponding to the first audio data includes: Perform a short-time Fourier transform on the first audio data to obtain the spectrum of the first audio data, and then filter the spectrum using a Mel filter bank to obtain the Mel time-frequency plot features corresponding to the first audio data; Perform wavelet transform on the first audio data to obtain the wavelet transform time-frequency graph features corresponding to the first audio data.
7. The method according to claim 4, further comprising: Statistical features are extracted from the first audio data to obtain the statistical features corresponding to the first audio data; The statistical features are input into a deep learning model for feature extraction to obtain the deep learning features corresponding to the statistical features. The process of fusing the deep learning features corresponding to each of the at least two types of time-frequency map features to obtain fused features includes: The deep learning features corresponding to each type of time-frequency map feature and the deep learning features corresponding to the statistical features in the at least two types of time-frequency map features are fused to obtain fused features.
8. The method according to claim 7, wherein, The step of inputting the statistical features into a deep learning model for feature extraction to obtain the deep learning features corresponding to the statistical features includes: The statistical features are input into a deep perceptron to obtain the deep learning features corresponding to the statistical features.
9. The method according to claim 2, wherein, The step of extracting features from the first audio data to obtain the time-frequency graph features corresponding to the first audio data includes: The first audio data is subjected to denoising processing to remove environmental interference noise, resulting in denoised first audio data. Feature extraction is performed on the denoised first audio data to obtain the time-frequency graph features corresponding to the first audio data.
10. The method according to claim 9, wherein, The step of denoising the first audio data to obtain denoised first audio data includes: A low-pass filter is used to denoise the first audio data to remove environmental interference noise, resulting in denoised first audio data.
11. The method according to claim 10, wherein, The frequency threshold of the low-pass filter is determined based on the vehicle's speed.
12. The method according to any one of claims 1-11, further comprising: Acquire second audio data; wherein the second audio data is tire noise frequency data of the second tire of the vehicle traveling on the road surface to be identified; Based on the second audio data, road surface type identification is performed to obtain the second road surface type of the road surface to be identified; Based on the first road surface type and the second road surface type, the target road surface type of the road surface to be identified is obtained.
13. The method according to claim 12, wherein, The step of identifying the road surface type based on the second audio data to obtain the second road surface type of the road surface to be identified includes: Feature extraction is performed on the second audio data to obtain the time-frequency graph features corresponding to the second audio data; Based on the time-frequency map features corresponding to the second audio data, road surface type identification is performed to obtain the second road surface type of the road surface to be identified.
14. The method according to claim 12, wherein, The step of obtaining the target road type of the road to be identified based on the first road type and the second road type includes: If the first road surface type is the same as the second road surface type, then the first road surface type or the second road surface type is determined to be the target road surface type; If the first road surface type is different from the second road surface type, the target road surface type of the road surface to be identified is determined according to the confidence levels corresponding to the first road surface type and the second road surface type, respectively.
15. The method according to claim 14, wherein, The step of determining the target road type of the road to be identified based on the confidence levels corresponding to the first road type and the second road type includes: If, among the confidence levels corresponding to the first road surface type and the second road surface type, one confidence level is greater than a preset threshold and the other confidence level is less than or equal to the preset threshold, then the road surface type with a confidence level greater than the preset threshold is determined to be the target road surface type. If the confidence levels corresponding to the first road surface type and the second road surface type are both greater than the preset threshold or both are less than the preset threshold, then the road surface type with higher priority is determined as the target road surface type based on the priority of the first road surface type and the second road surface type.
16. The method of claim 14, wherein, The step of determining the target road type of the road to be identified based on the confidence levels corresponding to the first road type and the second road type includes: The road surface type with higher confidence among the first road surface type and the second road surface type is determined as the target road surface type of the road surface to be identified.
17. The method according to any one of claims 2-11, wherein, The step of identifying the road surface type based on the time-frequency map features corresponding to the first audio data to obtain the first road surface type of the road surface to be identified includes: The time-frequency map features corresponding to the first audio data are input into the target road surface type recognition model to obtain the first road surface type of the road surface to be identified.
18. The method according to claim 17, wherein, The target road surface type identification model is obtained by training the candidate road surface type identification model based on training samples; The training samples include: audio data samples; the audio data samples include: audio data content and road surface type labels corresponding to the audio data.
19. The method of claim 18, further comprising: By performing data augmentation processing on the audio data samples, multiple audio data samples are obtained; The data enhancement processing includes at least one of the following: tone enhancement processing, pitch enhancement processing, time shift processing, or time masking processing.
20. A road surface type identification device, comprising: The acquisition module is configured to acquire first audio data, which is the tire noise frequency data of the first tire of a vehicle traveling on the road surface to be identified. The recognition module is configured to obtain the first road surface type of the road surface to be recognized based on the first audio data; The arbitration module is configured to obtain the target road type of the road to be identified based on the first road type and the second road type.
21. A controller, comprising: A processor and a memory, the memory storing a program or instructions executable on the processor, the program or instructions, when executed by the processor, implementing the steps of the road surface type identification method according to any one of claims 1-19.
22. A vehicle comprising: A processor and a memory, the memory storing a program or instructions executable on the processor, the program or instructions, when executed by the processor, implementing the steps of the road surface type identification method according to any one of claims 1-19.
23. A cloud server, comprising: A processor and a memory, the memory storing a program or instructions executable on the processor, the program or instructions, when executed by the processor, implementing the steps of the road surface type identification method according to any one of claims 1-19.
24. A computer-readable storage medium, wherein, The readable storage medium stores a program or instructions that, when executed by a processor, implement the steps of the road surface type identification method according to any one of claims 1-19.
25. A computer program product, wherein, When the program product is executed by the processor of a vehicle or a cloud server, it implements the steps of the road surface type identification method according to any one of claims 1-19.
Citation Information
Patent Citations
System and method for classifying a road surface
CN105844211A
Systems, methods, and computer-readable storage media for a vehicle
CN114791731A
Road surface meteorological condition identification method and system based on road noise frequency analysis
CN115762565A
Computer-implemented method for machine learning of road markings using audio signals, control unit for automated driving functions, method and computer program for recognizing road markings
DE102019209634A1
Environmental condition monitoring for a vehicle
GB202308651D0