A method and device for learning semantic understanding of auditory information for autonomous driving of a vehicle
By collecting and training auditory information in a virtual environment and using deep learning models for semantic understanding, the problem of insufficient understanding of sound events in auditory perception systems in autonomous driving is solved, and the vehicle's perception and decision-making capabilities in complex traffic scenarios are improved.
Patent Information
- Application Number
- CN202510602094.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-12
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2045-05-12
AI Technical Summary
The existing auditory perception systems lack deep semantic understanding of sound events in autonomous driving, and cannot meet the needs of autonomous driving in complex traffic scenarios.
Through the virtual audio acquisition module, acquiring sound events in a virtual environment, generating a virtual audio data set, and training through deep learning models to obtain audio features and context features, combining multi-scale convolution and timing modeling networks to achieve semantic understanding of auditory information.
It improves the perception and understanding of auditory information by autonomous driving vehicles, can quickly identify complex sound events, and enhances the decision-making ability of the vehicle in complex environments.
Smart Images

Figure CN120126505B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of autonomous driving technology, and in particular to a method and device for learning semantic understanding of auditory information for autonomous driving of a vehicle. Background Art
[0002] In the field of autonomous driving, visual perception technologies (such as cameras and lidar) have made significant progress, enabling the identification of objects such as traffic signs, pedestrians, and vehicles. However, relying solely on visual perception has certain limitations. For example, in low light, inclement weather, or under occlusion, visual information may not be sufficient to support safe vehicle decision-making. Further incorporating auditory information can compensate for these shortcomings in visual perception.
[0003] However, existing auditory perception systems mainly focus on sound source localization and simple sound detection, lack deep semantic understanding of sound events, and cannot meet the needs of autonomous driving in complex traffic scenarios. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to provide a method and device for learning the semantic understanding of auditory information in vehicle autonomous driving, so as to improve the perception and understanding ability of the autonomous driving system of auditory information.
[0005] In order to solve the above technical problems, the technical solution adopted by the present invention is:
[0006] A method for learning semantic understanding of auditory information for autonomous driving of a vehicle, comprising:
[0007] The virtual audio acquisition module collects sound events in the virtual environment to obtain a virtual audio dataset;
[0008] Classifying the virtual audio data in the virtual audio dataset and generating pseudo labels according to different classifications;
[0009] Acquire scene information of target virtual audio data, and use the scene information as a scene semantic label of the target virtual audio data;
[0010] Using the virtual audio dataset with the pseudo labels and the scene semantic labels for deep learning model training to obtain an audio analysis model;
[0011] Acquire real-time audio data, and extract audio features of the real-time audio data through the feature extraction layer of the audio analysis model;
[0012] Obtaining a context feature corresponding to the audio feature in a time series, and generating an audio time feature according to the target audio feature and the context feature;
[0013] The audio time features are analyzed through the output layer of the audio analysis model to obtain the classification and scene corresponding to the real-time audio data.
[0014] In order to solve the above technical problems, another technical solution adopted by the present invention is:
[0015] A device for learning the semantic understanding of auditory information for autonomous driving of a vehicle comprises a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the device implements the various steps in the method for learning the semantic understanding of auditory information for autonomous driving of a vehicle as described above.
[0016] The beneficial effects of the present invention are: by collecting sound events in a virtual environment based on a virtual audio collection module, it is possible to simulate sound events encountered by a vehicle during actual driving. Compared with collecting sound events in real scenes, various sound events can be quickly simulated in a virtual environment to collect and generate a large number of virtual audio data sets, thereby improving the efficiency of data set collection; by classifying audio data, associating it with scene information and setting corresponding semantic labels, the deep learning model can learn the classification method of audio data and determine the scene of audio data; in the actual recognition process, by acquiring real-time audio data and combining it with its contextual features corresponding to the time series, and then analyzing it through an audio analysis model, it is possible to improve the recognition ability of complex sound events, thereby improving the perception and understanding ability of autonomous driving vehicles of auditory information. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 This is a flowchart of the steps of a method for learning semantic understanding of auditory information for autonomous driving of a vehicle according to an embodiment of the present invention;
[0018] Figure 2 Schematic diagram of an application scenario of a method for learning semantic understanding of auditory information for autonomous driving of a vehicle according to an embodiment of the present invention;
[0019] Figure 3 This is a flowchart of another step in a method for learning semantic understanding of auditory information for autonomous driving of a vehicle according to an embodiment of the present invention;
[0020] Figure 4 This is a structural diagram of a device for learning semantic understanding of auditory information for autonomous driving of a vehicle in an embodiment of the present invention. DETAILED DESCRIPTION
[0021] To illustrate the technical content, achieved objectives and effects of the present invention in detail, the following description is given in conjunction with the embodiments and accompanying drawings.
[0022] A method for learning semantic understanding of auditory information for autonomous driving of a vehicle, comprising:
[0023] The virtual audio acquisition module collects sound events in the virtual environment to obtain a virtual audio dataset;
[0024] Classifying the virtual audio data in the virtual audio dataset and generating pseudo labels according to different classifications;
[0025] Acquire scene information of target virtual audio data, and use the scene information as a scene semantic label of the target virtual audio data;
[0026] Using the virtual audio dataset with the pseudo labels and the scene semantic labels for deep learning model training to obtain an audio analysis model;
[0027] Acquire real-time audio data, and extract audio features of the real-time audio data through the feature extraction layer of the audio analysis model;
[0028] Obtaining a context feature corresponding to the audio feature in a time series, and generating an audio time feature according to the target audio feature and the context feature;
[0029] The audio time features are analyzed through the output layer of the audio analysis model to obtain the classification and scene corresponding to the real-time audio data.
[0030] From the above description, it can be seen that the beneficial effects of the present invention are: by collecting sound events based on a virtual audio acquisition module in a virtual environment, it is possible to simulate sound events encountered by a vehicle during actual driving. Compared with collecting sound events in real scenes, various sound events can be quickly simulated in a virtual environment to collect and generate a large number of virtual audio data sets, thereby improving the efficiency of data set collection; by classifying audio data, associating it with scene information, and setting corresponding semantic labels, the deep learning model can learn the classification method of audio data and determine the scene of audio data; in the actual recognition process, by obtaining real-time audio data and combining it with its contextual features corresponding to the time series, and then analyzing it through an audio analysis model, it is possible to improve the recognition ability of complex sound events, thereby improving the perception and understanding ability of autonomous driving vehicles of auditory information.
[0031] Furthermore, the collecting of sound events in the virtual environment by the virtual audio collection module to obtain a virtual audio data set includes:
[0032] Setting different traffic scenarios, weather conditions, time periods, and traffic densities in a virtual environment to obtain different scenario information;
[0033] Under target scene information, triggering a target sound event in the virtual environment through a script; the target sound event includes different types of virtual audio data;
[0034] The virtual audio data of the target sound event inside and outside the vehicle are collected through the interface of the virtual audio collection module to obtain the virtual audio data set.
[0035] From the above description, it can be seen that by adjusting the variables of scene information in the virtual environment and triggering different sound events under different scene information, the simulation of different sound events in the actual driving scene of the vehicle is achieved.
[0036] Furthermore, using the virtual audio dataset with the pseudo labels and the scene semantic labels for deep learning model training includes:
[0037] The feature extraction layer of the deep learning model includes three parallel convolutional layers with different convolution kernel sizes;
[0038] Extracting audio features from the virtual audio data respectively through each of the convolutional layers;
[0039] The audio features obtained by the three convolutional layers are concatenated in the channel dimension to obtain multi-scale audio features;
[0040] The deep learning model is trained based on the multi-scale audio features.
[0041] As can be seen from the above description, by setting up three parallel convolution layers with different convolution kernel sizes, convolution kernels of different sizes have different receptive fields, which can capture detailed features and global features. After splicing the feature information of three different scales, a more complete and richer feature representation is formed, thereby significantly improving the network's ability to capture sound features at different time scales. This enables the model to better understand the complexity of audio signals and improve its performance in application scenarios such as autonomous driving.
[0042] Furthermore, before using the virtual audio dataset with the pseudo labels and the scene semantic labels for deep learning model training, the method further includes:
[0043] Selecting first target virtual audio data from the virtual audio data set;
[0044] Performing data enhancement on the first target virtual audio data to obtain at least one enhanced audio data;
[0045] Selecting second target virtual audio data that is different from the first target virtual audio data from the virtual audio data set;
[0046] The using the labeled virtual audio dataset for deep learning model training includes:
[0047] Inputting the first target virtual audio data, the enhanced audio data, and the second target virtual audio data into the deep learning model respectively to obtain a first target feature vector, an enhanced feature vector, and a second target feature vector;
[0048] Combining the first target feature vector and the enhanced feature vector into a positive sample pair, and combining the second target feature vector and the second target feature vector into a negative sample pair;
[0049] A contrast loss function is constructed based on the positive sample pairs and the negative sample pairs to train the deep learning model.
[0050] From the above description, it can be seen that by performing data augmentation on audio data from the same sound source to obtain enhanced audio data and forming positive sample pairs, the purpose is to improve the model's performance in learning the same audio under different conditions; and obtaining audio data different from the target audio data and forming negative sample pairs so that the model can learn the ability to distinguish different audio content; and constructing a contrast loss function through positive sample pairs and negative sample pairs, the model can effectively extract the semantic features of audio samples, thereby improving its performance in autonomous driving tasks.
[0051] Furthermore, the contrast loss function includes:
[0052] ;
[0053] ;
[0054] Among them, z i is the first target feature vector; To enhance the feature vector; is the second target feature vector; is the similarity function; is a parameter; N is the number of the second target feature vector.
[0055] As can be seen from the above description, the similarity function can effectively describe the similarity between two feature vectors; and by constructing a loss function, the model can learn more robust semantic feature representations, thereby improving its ability to understand complex audio events; it effectively enhances the model's feature learning ability under unsupervised conditions, providing a high-quality feature foundation for subsequent semantic association and reasoning.
[0056] Furthermore, the using the virtual audio dataset with the pseudo labels and the scene semantic labels for deep learning model training also includes:
[0057] Construct audio event classification loss function;
[0058] Performing weight calculation on the audio event classification loss function and the contrast loss function to obtain a comprehensive loss function;
[0059] The deep learning model is trained according to the comprehensive loss function.
[0060] From the above description, we can see that by combining audio event classification loss and contrast loss, the accuracy and robustness of the model for semantic feature extraction are further improved.
[0061] Furthermore, the audio event classification loss function includes:
[0062] ;
[0063] The comprehensive loss function includes:
[0064] ;
[0065] in, is the comprehensive loss function; Loss function for audio event classification; is the contrast loss function; is the weight coefficient; is the audio feature; N is the number of audio features.
[0066] From the above description, we can see that setting the weight coefficient and , and adjust in real time according to the convergence of different tasks during training and , thereby achieving efficient collaborative optimization of multiple tasks by dynamically balancing the contributions of the two parts of the loss.
[0067] Furthermore, it also includes:
[0068] Acquire other modal data in real time;
[0069] Extracting modal features corresponding to target modal data through different modal feature models;
[0070] fusing the audio feature with all the modal features to obtain a joint feature;
[0071] The joint feature is analyzed to obtain the scene corresponding to the audio feature.
[0072] From the above description, we can see that by acquiring other modal data such as images, videos, or humidity, and obtaining the corresponding modal features, the modal features are combined with audio features, so that the model can capture the correlation information between audio data and other modal data, and improve the semantic understanding ability in complex scenarios.
[0073] Furthermore, the step of fusing the audio feature with all the modal features to obtain a joint feature includes:
[0074] Projecting the audio features and all the modal features to the same feature dimension;
[0075] Calculating an attention weight between each of the modal features and the audio feature through a dot product operation;
[0076] Performing weighted summation on all the modal features using the normalized attention weights to obtain a fusion feature;
[0077] The fusion feature is concatenated with the audio feature to obtain the joint feature.
[0078] From the above description, it can be seen that the above fusion method enables the audio features to be effectively fused with other modal features, thereby improving the model's processing efficiency for joint features.
[0079] Another embodiment of the present invention provides a device for learning the semantic understanding of auditory information for autonomous driving of a vehicle, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the processor implements the various steps in the method for learning the semantic understanding of auditory information for autonomous driving of a vehicle as described above.
[0080] The method and device for learning the semantic understanding of auditory information for autonomous driving provided by the present invention can be applied to autonomous driving scenarios, and are described below through specific implementation methods:
[0081] Example 1
[0082] Please refer to Figure 1 , a learning method for semantic understanding of auditory information for autonomous driving of a vehicle, comprising:
[0083] S1. Collect sound events in the virtual environment through a virtual audio acquisition module to obtain a virtual audio data set. In this embodiment, the GTA V virtual environment is combined with the DeepGTAV module to collect audio data of the virtual scene. Specifically:
[0084] S11. Setting different traffic scenarios, weather conditions, time periods, and traffic densities in the virtual environment to obtain different scenario information; for example, selecting representative traffic scenarios in Grand Theft Auto V, including urban roads, highways, intersections, tunnels, etc., to ensure scenario diversity; setting various weather conditions such as sunny, rainy, and foggy days; setting different time periods such as daytime, nighttime, and dusk; and setting different traffic densities such as peak hours and low traffic scenarios.
[0085] S12. Under the target scene information, trigger a target sound event in the virtual environment through a script; the target sound event includes different types of virtual audio data. For example, specific sound events such as sirens, horns, tire friction, and collisions can be triggered through script control to ensure that the data covers common traffic sounds.
[0086] S13. Collect the virtual audio data of the target sound event inside and outside the vehicle through the interface of the virtual audio acquisition module to obtain the virtual audio data set. For example, by installing the DeepGTAV module in GTA V, use the API interface provided by it to collect multimodal data. Use the extended function of DeepGTAV to capture environmental sounds through the audio engine built into the virtual application; when recording audio, ensure that the sampling rate is 44.1kHz or higher to ensure sound quality. Use a virtual microphone array to simulate multi-channel audio acquisition and record the spatial distribution information of the sound, such as the direction and distance of the sound source; use the extended script of DeepGTAV to simulate the sound environment inside and outside the vehicle, and collect the sound data inside and outside the cockpit respectively; use the microphone array to collect multi-channel original sound signals S i (t); perform noise reduction and enhancement processing on the signal, filter out the background noise N(t), and obtain the sound signal S i '(t):
[0087] .
[0088] Among them, the sound signal S is obtained i After '(t), the processed signal is subjected to short-time Fourier transform (STFT) to extract the time-frequency feature matrix:
[0089] ;
[0090] The features are then normalized, using the logarithmic amplitude normalization method to enhance the separability of small amplitude signals:
[0091] .
[0092] S2. Classify the virtual audio data in the virtual audio dataset and generate pseudo labels based on the different classifications. For example, multi-category semantic labels such as emergency alarm, horn, and tire slip can be manually annotated. Alternatively, a pre-trained model (such as the AudioSet pre-trained model) can be used to perform preliminary classification and generate pseudo labels.
[0093] S3. Obtain scene information of the target virtual audio data, and use the scene information as a scenario-based semantic label for the target virtual audio data; for example, associate specific scene information with each audio sample, such as an associated scene such as an "intersection" or "congested scene", thereby constructing a scenario-based semantic label.
[0094] S4. Using the virtual audio dataset with the pseudo-labels and the scene-based semantic labels for deep learning model training to obtain an audio analysis model. The deep learning model includes an audio feature extraction network and a time series modeling network.
[0095] (1) In this embodiment, the audio feature extraction network uses ResNet-18 as the backbone network, and designs a multi-scale convolution module to enhance the ability to capture sound features of different time scales to achieve feature extraction; by fusing features of different scales together, a richer and more robust feature representation can be provided, thereby improving the performance of the model; for example, the feature extraction layer includes three parallel convolution layers with different convolution kernel sizes; the audio features in the virtual audio data are extracted respectively through each of the convolution layers; the audio features obtained by the three convolution layers are spliced in the channel dimension to obtain multi-scale audio features; the deep learning model is trained based on the multi-scale audio features.
[0096] In one specific embodiment, the sizes of the three convolution kernels are 3×3, 5×5, and 7×7, respectively, with a stride of 1 and the same padding. For example, for a 7x7 convolution kernel, 3 pixels are padded on both sides of the input feature map to ensure that the temporal dimension of the output feature map is consistent with the input. Each convolution layer is followed by a batch normalization layer and a ReLU activation function to accelerate network convergence and improve nonlinear expression capabilities. The number of output channels of each convolution layer is 43, 42, and 43, respectively. After batch normalization and ReLU activation, the size of the concatenated feature map is (C1+C2+C3)×T×1, where T is the time step. Since C1+C2+C3=128, the final feature map has a size of 128×T×1. The last dimension 1 can usually be ignored, resulting in a final representation of 128×T. This 128×T feature map contains feature information from three convolution layers of different scales, achieving multi-scale feature fusion.
[0097] Convolution kernels of different sizes have different receptive fields. The 3x3 convolution kernel has a smaller receptive field and focuses primarily on very short-term local information in the audio signal. For example, it is very effective in capturing subtle changes in sounds, such as brief bursts of consonants and sharp changes in high-frequency noise, such as distinguishing "ba" from "pa." The 5x5 convolution kernel has a larger receptive field than the 3x3 convolution kernel and can capture features at longer time scales. This helps the network understand longer speech units such as phonemes and syllables and recognize the duration and frequency changes of vowels, such as distinguishing "ship" from "sheep." The 7x7 convolution kernel has the largest receptive field and can capture patterns in audio signals at longer time scales, such as pitch, rhythm, and speaker emotion; for example, it can distinguish between declarative sentences and interrogative sentences.
[0098] (2) The temporal modeling network adopts a bidirectional long short-term memory (Bi-LSTM) structure to capture the temporal context information of audio features. The network consists of two layers of bidirectional LSTM, each layer containing 256 hidden units to ensure the ability to model complex temporal dependencies. At the same time, to prevent overfitting, a Dropout (random deactivation) mechanism is introduced between the two layers of LSTM, and the Dropout rate is set to 0.5. The input of the bidirectional LSTM is the feature map output by the audio feature extraction network (CNN), with a feature dimension of 128×T. The network can simultaneously capture the forward and backward temporal dependencies of the audio signal, generate feature representations with global temporal context information, and provide support for subsequent semantic understanding and reasoning. The final output features contain rich temporal series information, which significantly improves the ability to recognize complex sound events.
[0099] In an optional embodiment, a contrastive learning mechanism is introduced in the feature extraction process. By constructing positive and negative sample pairs and designing a contrastive loss function, the model's ability to extract semantic features is enhanced. Specifically:
[0100] S41. Select first target virtual audio data from the virtual audio data set; for example, the first target virtual audio data is x.
[0101] S42. Perform data enhancement on the first target virtual audio data to obtain at least one enhanced audio data; for example, perform data enhancement on the first target virtual audio data x by time stretching, adding noise, or adjusting the volume to generate multiple variants to obtain multiple enhanced audio data x'.
[0102] S43, selecting second target virtual audio data different from the first target virtual audio data from the virtual audio data set; for example, the second target virtual audio data is x k , sample x k It should be significantly different from sample x in content, category or other characteristics.
[0103] S44: Input the first target virtual audio data, the enhanced audio data, and the second target virtual audio data into the deep learning model respectively to obtain a first target feature vector z i , enhanced feature vector z j and the second target feature vector z k .
[0104] S45, combining the first target feature vector and the enhanced feature vector into a positive sample pair, and combining the second target feature vector and the second target feature vector into a negative sample pair; that is, feature vector z i and z j Constitutes a positive sample pair, representing the performance of the same audio under different conditions. These enhanced samples are regarded as positive sample pairs, which are intended to help the model learn the performance of the same audio under different conditions, thereby enhancing its ability to extract semantic features. Feature vector z i and z k This constitutes a negative sample pair, which is used to reflect the audio features that are different from the positive sample, thereby providing the necessary comparison basis for contrastive learning.
[0105] S46. Construct a contrast loss function based on the positive sample pairs and the negative sample pairs to train the deep learning model; wherein the goal of contrastive learning is to maximize the similarity between the positive sample pairs and minimize the similarity between the positive samples and the negative samples; the specific form of the similarity function usually uses cosine similarity to measure the similarity between two feature vectors, which is defined as:
[0106] ;
[0107] Among them, the value range of cosine similarity is between [-1, 1]. The closer the value is to 1, the more similar the two vectors are, and the closer it is to -1, the less similar they are. is the norm of the eigenvector.
[0108] Construct positive and negative sample pairs and design a contrast loss function to enhance semantic feature extraction capabilities. The loss function is defined as follows:
[0109] ;
[0110] in, is the first target feature vector; To enhance the feature vector; is the second target feature vector; is the similarity function; is a parameter controlling the smoothness of the distribution; N is the number of second target feature vectors. By optimizing this loss function, the model can learn more robust semantic feature representations, thereby improving its understanding of complex audio events. This mechanism effectively enhances the model's feature learning capabilities under unsupervised conditions, providing a high-quality feature foundation for subsequent semantic association and reasoning.
[0111] In a further implementation, this embodiment further improves the accuracy and robustness of semantic feature extraction by jointly optimizing the audio event classification loss and the contrastive learning loss. The audio event classification loss function is as follows:
[0112] ;
[0113] By performing weight calculation on the audio event classification loss function and the contrast loss function, a comprehensive loss function is obtained:
[0114] ;
[0115] in, is the comprehensive loss function; Loss function for audio event classification; is the contrast loss function; is the weight coefficient, which is used to dynamically balance the contribution of the two parts of the loss. The weight coefficient value is adjusted in real time according to the convergence of different tasks during training, thereby achieving efficient collaborative optimization of multiple tasks; is the audio feature; N is the number of audio features.
[0116] The deep learning model is trained based on the comprehensive loss function. The training process includes a time series prediction task and a mask reconstruction task. In the time series prediction task, the audio sample is divided into multiple sub-segments, which are randomly shuffled and then fed into the model. The model is required to predict the correct time sequence of the sub-segments, thereby learning the temporal dependencies of the audio signal. In the mask reconstruction task, a portion of the input audio spectrogram is randomly masked, and the model is required to reconstruct the masked spectrogram. The local and global features of the audio signal are learned by minimizing the reconstruction error (e.g., mean square error).
[0117] S5. Acquire real-time audio data and extract audio features of the real-time audio data through the feature extraction layer of the audio analysis model; that is, obtain audio features through the above-mentioned audio feature extraction network. The real-time audio data is also analyzed by the sound source localization algorithm to determine the direction and distance of the sound source. Specifically:
[0118] S51. Data acquisition: Use multiple microphone arrays, such as linear arrays or circular arrays, to capture audio signals; the signal recorded by each microphone will vary depending on the direction and distance of the sound source.
[0119] S52. Signal preprocessing: Preprocess the collected audio signal, including denoising, gain adjustment, and sampling rate normalization, to improve the accuracy of subsequent analysis; for example, use a filter (such as a low-pass filter) to remove high-frequency noise.
[0120] S53. Feature extraction: Extract features from the preprocessed signal. The extracted time domain features include signal amplitude, signal energy, instantaneous frequency, etc.; the frequency domain features include spectrum, Mel-frequency cepstral coefficients (MFCC), short-time Fourier transform (STFT), etc.
[0121] S54. Time Delay Estimation: Infer the direction of the sound source by calculating the time difference (TDOA) between the sound waves received by different microphones. Use the cross-correlation function to calculate the cross-correlation between the signals from different microphones and find the delay corresponding to the maximum correlation value. Then, use peak detection to detect the peak in the cross-correlation function to determine the time delay.
[0122] S55. Direction estimation: Use triangulation to calculate the direction of the sound source based on the time delay and the geometric layout of the microphone array, such as calculating the azimuth and elevation angles.
[0123] S56. Sound Source Distance Calculation: This can be estimated by analyzing signal strength, such as sound pressure level, or by combining time delay and sound speed. The sound source localization algorithm can utilize particle filtering, a dynamic estimation method based on Bayesian theory that approximates the posterior probability distribution by generating and updating a large number of random particles. It is particularly well-suited for sound source localization in nonlinear and non-Gaussian noise environments. Using this sound source localization algorithm, the system can accurately calculate the direction and distance of the siren, providing critical support for subsequent traffic decisions.
[0124] S6. Obtain the contextual features corresponding to the audio features in the time series, and generate audio time features based on the target audio features and the contextual features; that is, obtain the corresponding contextual features through the above-mentioned timing modeling network and generate audio time features.
[0125] In this embodiment, the semantic embedding network converts audio features into high-dimensional semantic vectors to achieve semantic expression and feature enhancement of audio signals. For example, if the input audio features are represented as F(t), the semantic embedding network performs nonlinear mapping to generate a high-dimensional semantic vector e, which is calculated as follows:
[0126] ;
[0127] Among them, f embedRepresents the mapping function of the semantic embedding network; the network consists of multiple layers of fully connected layers and activation functions, which can capture the deep semantic information of audio features; through the semantic embedding network, audio features are converted into high-dimensional semantic vectors. The semantic vectors not only contain the time series characteristics of the audio signal, but also can express its semantic association information, providing high-quality feature representation for subsequent semantic reasoning and classification tasks.
[0128] At the same time, the audio embedding vector is combined with the context information to achieve semantic understanding of the audio event. The inference model takes the embedding vector e and the context information C as input to generate the semantic understanding result R of the event. Its calculation formula is as follows:
[0129] ;
[0130] Among them, f reason Represents the mapping function of the inference model. Contextual information includes temporal, spatial, environmental, or other modal features. The model jointly models the embedding vector and contextual information through a deep neural network, capturing the relationship between audio events and context, thereby generating accurate semantic understanding results.
[0131] This embodiment introduces a Graph Neural Network (GNN) into semantic feature modeling and employs a Graph Attention Network (GAT) to model and reason about the spatiotemporal relationships of audio events. In the graph structure, nodes represent sound events, and edges represent the spatiotemporal relationships between them. Through the graph attention mechanism, the network dynamically assigns weights between different nodes, thereby capturing global correlation information about sound events. The parameters of the graph attention network are configured as follows: 4 attention heads and 128 output dimensions. Each node aggregates information from neighboring nodes through a multi-head attention mechanism and updates it based on its own features, generating a high-dimensional feature representation that incorporates global spatiotemporal relationships. This network effectively models the complex spatiotemporal dependencies between sound events, enhancing semantic understanding and reasoning capabilities of audio events.
[0132] Furthermore, a dynamic memory network (DMN) is introduced into the semantic reasoning process to store and manage historical contextual information, thereby improving the semantic understanding of audio events. The DMN is designed and configured as follows: a memory capacity of 10, and a first-in-first-out (FIFO) strategy is used for memory updates. The DMN dynamically maintains the semantic content related to the current audio event by storing recent historical contextual information and provides support during the reasoning process. In practical applications, the DMN effectively captures the temporal dependencies and contextual associations of audio events. Through the FIFO update strategy, the DMN ensures the timeliness of the memory content while avoiding the storage of redundant information. This mechanism significantly enhances the model's ability to model long-term dependencies in complex audio scenes, providing more comprehensive contextual support for semantic reasoning and event classification. By combining audio signals and contextual information, the system can analyze the congestion level and dynamic changes in the current traffic environment, providing comprehensive environmental awareness for intelligent vehicle control.
[0133] S7. Analyze the audio temporal features through the output layer of the audio analysis model to obtain the classification and scenario corresponding to the real-time audio data. Through an efficient audio event classification model, the system can quickly and accurately detect and identify the sound of an ambulance siren, ensuring a timely response in emergency situations. For example, the system can identify the sound of an ambulance siren in real time. Upon detecting the sound of an ambulance siren, the system generates control instructions through a semantic reasoning model, automatically controlling the vehicle to slow down and evade to the right, providing priority passage for the ambulance.
[0134] In an optional embodiment, the multimodal fusion module uses a cross-modal attention mechanism to achieve the fusion of audio features with other modalities. This mechanism is based on scaled dot-product attention and dynamically adjusts the contribution of different modal features by calculating the correlation weights between audio features and features of other modalities. Specifically:
[0135] A1. Acquire other modal data in real time; for example, visual data such as images and videos, sensor data such as temperature, humidity, acceleration, and text.
[0136] A2. Modal features corresponding to the target modal data can be extracted using different modal feature models. For example, for visual data, features extracted from images or videos using convolutional neural networks (CNNs) or obtained through visual attention mechanisms can be used. When processing video or audio streams, time series features (such as inter-frame differences and temporal features of audio waveforms) can also serve as important modal features to help capture dynamic changes. For example, text features can be extracted using natural language processing (NLP) techniques, such as word embeddings, sentence embeddings, or contextual features generated by more complex models (such as BERT and GPT).
[0137] A3. Fusing the audio feature with all the modal features to obtain a joint feature. Specifically:
[0138] A31. Project the audio features and all the modal features to the same feature dimension; for example, project the audio features and image features to the same feature dimension d, and the dimension can be configured according to actual needs; the projection can be achieved through a linear transformation (such as a fully connected layer) to ensure that they have the same dimension d.
[0139] A32. Calculate the attention weight between each modal feature and the audio feature by performing a dot product operation. The specific calculation method is as follows:
[0140] ;
[0141] ;
[0142] Among them, W a and W o is the learned weight matrix.
[0143] A33. Perform weighted summation on all the modal features using the normalized attention weights to obtain fused features; for example, after normalizing the attention weights corresponding to the image features, time series features, and sensor features, perform weighted summation on the image features, time series features, and sensor features to obtain fused features.
[0144] A34. Concatenate the fused features with the audio features to obtain the joint features. Specifically, the weighted summed fused features are concatenated with the audio features along the channel dimension to form a joint feature representation containing multimodal information. This mechanism enables the system to effectively capture the correlation between audio and other modalities, improving semantic understanding in complex scenarios.
[0145] A4. Analyze the joint features to obtain the scene corresponding to the audio features.
[0146] Please refer to Figure 2 as well as Figure 3 In a specific embodiment, the entire system works in a complex traffic environment as follows Figure 3 As shown in the figure, starting from the "Start" node, the initial stage of sound source analysis begins. Environmental detection is performed to understand the external conditions and background noise of the sound source. Audio signals are collected through a microphone array to obtain the original data of the sound source. Feature extraction is performed on the collected audio signal to extract feature information that helps locate the sound source. Specific events in the audio signal are detected to identify the existence and type of the sound source. A model is constructed to describe the motion state and changes of the sound source. Sound source analysis and positioning are performed to calculate the direction and distance of the sound source. Based on the direction and distance of the sound source, the graph attention network combines multimodal information such as traffic flow and vehicle type to determine the lane traffic status, obtain the lane traffic status and the distribution of surrounding vehicles, and then make corresponding decisions, such as speed control and direction control. The process ends after completing all steps, and feedback and adjustments are provided as needed.
[0147] Example 2
[0148] Please refer to Figure 4 A device for learning the semantic understanding of auditory information for autonomous driving of a vehicle includes a memory, a processor, and a computer program stored in the memory and runnable on the processor. When the processor executes the computer program, each step of a method for learning the semantic understanding of auditory information for autonomous driving of a vehicle as described in Example 1 is implemented.
[0149] In summary, the method and device for learning the semantic understanding of auditory information for autonomous driving provided by the present invention improve the semantic recognition ability of autonomous driving vehicles for auditory signals through deep learning and self-supervised learning; and combine visual and environmental context information to achieve visual and auditory multimodal information fusion, enabling autonomous driving vehicles to perform situational reasoning more comprehensively; and enable autonomous driving vehicles to make decisions on driving strategies based on the recognition of auditory events and scenes.
[0150] The above descriptions are merely embodiments of the present invention and are not intended to limit the patent scope of the present invention. Any equivalent transformations made using the contents of the present invention's description and drawings, or directly or indirectly applied in related technical fields, are also included in the patent protection scope of the present invention.
Claims
1. A method for learning semantic understanding of auditory information for autonomous driving of a vehicle, characterized by: include: The virtual audio acquisition module collects sound events in the virtual environment to obtain a virtual audio dataset; Classifying the virtual audio data in the virtual audio dataset and generating pseudo labels according to different classifications; Acquire scene information of target virtual audio data, and use the scene information as a scene semantic label of the target virtual audio data; Using the virtual audio dataset with the pseudo labels and the scene semantic labels for deep learning model training to obtain an audio analysis model; Acquire real-time audio data, and extract audio features of the real-time audio data through the feature extraction layer of the audio analysis model; Obtaining context features corresponding to the audio features in a time series, and generating audio time features based on the audio features and the context features; Analyzing the audio time features through the output layer of the audio analysis model to obtain the classification and scene corresponding to the real-time audio data; Before using the virtual audio dataset with the pseudo labels and the scene semantic labels for deep learning model training, the method includes: Selecting first target virtual audio data from the virtual audio data set; Performing data enhancement on the first target virtual audio data to obtain at least one enhanced audio data; Selecting second target virtual audio data that is different from the first target virtual audio data from the virtual audio data set; Using the labeled virtual audio dataset for deep learning model training includes: Inputting the first target virtual audio data, the enhanced audio data, and the second target virtual audio data into the deep learning model respectively to obtain a first target feature vector, an enhanced feature vector, and a second target feature vector; Combining the first target feature vector and the enhanced feature vector into a positive sample pair, and combining the second target feature vector and the second target feature vector into a negative sample pair; A contrast loss function is constructed based on the positive sample pairs and the negative sample pairs to train the deep learning model.
2. The method for learning semantic understanding of auditory information for autonomous driving of a vehicle according to claim 1, characterized in that: The step of collecting sound events in the virtual environment by using the virtual audio collection module to obtain a virtual audio data set includes: Setting different traffic scenarios, weather conditions, time periods, and traffic densities in a virtual environment to obtain different scenario information; Under target scene information, triggering a target sound event in the virtual environment through a script; the target sound event includes different types of virtual audio data; The virtual audio data of the target sound event inside and outside the vehicle are collected through the interface of the virtual audio collection module to obtain the virtual audio data set.
3. The method for learning semantic understanding of auditory information for autonomous driving of a vehicle according to claim 1, characterized in that: The using the virtual audio dataset with the pseudo labels and the scene semantic labels for deep learning model training includes: The feature extraction layer of the deep learning model includes three parallel convolutional layers with different convolution kernel sizes; Extracting audio features from the virtual audio data respectively through each of the convolutional layers; The audio features obtained by the three convolutional layers are concatenated in the channel dimension to obtain multi-scale audio features; The deep learning model is trained based on the multi-scale audio features.
4. The method for learning semantic understanding of auditory information for autonomous driving of a vehicle according to claim 1, characterized in that: The contrast loss function includes: ; ; in, is the first target feature vector; To enhance the feature vector; is the second target feature vector; is the similarity function; is a parameter; N is the number of the second target feature vector.
5. The method for learning semantic understanding of auditory information for autonomous driving of a vehicle according to claim 1, characterized in that: The using the virtual audio dataset with the pseudo labels and the scene semantic labels for deep learning model training also includes: Construct audio event classification loss function; Performing weight calculation on the audio event classification loss function and the contrast loss function to obtain a comprehensive loss function; The deep learning model is trained according to the comprehensive loss function.
6. The method for learning semantic understanding of auditory information for autonomous driving of a vehicle according to claim 5, characterized in that: The audio event classification loss function includes: ; The comprehensive loss function includes: ; in, is the comprehensive loss function; Loss function for audio event classification; is the contrast loss function; is the weight coefficient; is the audio feature; M is the number of audio features.
7. The method for learning semantic understanding of auditory information for autonomous driving of a vehicle according to claim 1, characterized in that: Also includes: Acquire other modal data in real time; Extracting modal features corresponding to target modal data through different modal feature models; fusing the audio feature with all the modal features to obtain a joint feature; The joint feature is analyzed to obtain the scene corresponding to the audio feature.
8. The method for learning semantic understanding of auditory information for autonomous driving of a vehicle according to claim 7, characterized in that: The step of fusing the audio feature with all the modal features to obtain a joint feature includes: Projecting the audio features and all the modal features to the same feature dimension; Calculating an attention weight between each of the modal features and the audio feature through a dot product operation; Performing weighted summation on all the modal features using the normalized attention weights to obtain a fusion feature; The fusion feature is concatenated with the audio feature to obtain the joint feature.
9. A device for learning and understanding the semantics of auditory information for autonomous driving of a vehicle, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, it implements the various steps in the method for learning semantic understanding of auditory information for autonomous driving of a vehicle as described in any one of claims 1-8.
Citation Information
Patent Citations
Acoustic scene classification method and device and corresponding equipment
CN112446242A
Detection method for sensing driving scene and event based on sound
CN119132338A