Driving state recognition method, electronic equipment and storage medium
By combining feature fusion processing of audio, image, and vehicle status data, the problem of low accuracy in driver status recognition has been solved, achieving more efficient driver status recognition and improving driving safety.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-30
- Publication Date
- 2026-04-07
AI Technical Summary
Existing technologies have low accuracy in recognizing driver driving status, especially when identifying distracted driver behavior.
By simultaneously acquiring audio signals, vehicle status data, and image data from multiple perspectives, visual encoding vectors, posture feature vectors, audio feature vectors, and vehicle status feature vectors are extracted and fused to obtain multimodal interaction feature vectors, which are then used for driving status classification.
It improves the accuracy of driving status recognition, enabling more accurate identification of whether the driver is distracted, thus enhancing driving safety.
Smart Images

Figure CN121808593A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of vehicle and intelligent transportation technology, and in particular to a driving state recognition method, driving state recognition device, electronic device, computer-readable storage medium and computer program product. Background Technology
[0002] With the rapid development of intelligent assisted autonomous driving and ADAS (Advanced Driver Assistance Systems), vehicle safety systems are gradually acquiring capabilities in environmental perception, path planning, and automatic control. However, the driver remains a crucial decision-maker in vehicle control, and the driver's state is significantly correlated with vehicle and road safety. Numerous studies and traffic accident statistics indicate that driver distraction (such as using a mobile phone, talking to passengers, or thinking about problems) is one of the main causes of traffic accidents. Therefore, identifying the driver's state, such as the presence of distracted behavior, is an important auxiliary means to improve driving and road safety.
[0003] The methods used in related technologies to identify the driver's driving status suffer from low accuracy. Summary of the Invention
[0004] Therefore, it is necessary to provide a driving state recognition method, driving state recognition device, electronic device, computer-readable storage medium, and computer program product that can improve the accuracy of driving state recognition in response to the above-mentioned technical problems.
[0005] Firstly, this application provides a driving state recognition method, the method comprising:
[0006] Acquire driving-related data collected at the same moment during vehicle driving, including: audio signals, vehicle status data, and multiple image data obtained from image acquisition of the target object from multiple perspectives;
[0007] Image encoding is performed on multiple image data to obtain a visual encoding vector;
[0008] Driver posture information is extracted from multiple image datasets, and posture feature vectors are obtained based on the driver posture information;
[0009] The audio signal is subjected to feature extraction to obtain an audio feature vector;
[0010] The vehicle state data is processed to generate a feature vector to obtain a vehicle state feature vector.
[0011] The visual encoding vector, the posture feature vector, the audio feature vector, and the vehicle state feature vector are subjected to feature fusion processing to obtain a multimodal interaction feature vector;
[0012] The multimodal interaction feature vector is subjected to driving state classification processing to obtain the driving state category.
[0013] Secondly, this application also provides a driving state recognition device, the device comprising:
[0014] The data acquisition module is used to acquire driving-related data collected at the same time during the vehicle driving process. The driving-related data includes: audio signals, vehicle status data, and multiple image data obtained by acquiring images of the target object from multiple perspectives.
[0015] A visual feature extraction module is used to perform image encoding on multiple image data to obtain a visual encoding vector;
[0016] The posture feature extraction module is used to extract driver posture information from multiple image data and obtain posture feature vectors based on the driver posture information.
[0017] The audio feature extraction module is used to extract features from audio signals and obtain audio feature vectors.
[0018] The vehicle state feature extraction module is used to process vehicle state data to generate feature vectors and obtain vehicle state feature vectors.
[0019] The feature fusion module is used to perform feature fusion processing on the visual encoding vector, posture feature vector, audio encoding vector and vehicle state feature vector to obtain multimodal interaction feature vector;
[0020] The classification and recognition module is used to classify the multimodal interaction feature vectors to obtain the driving state category.
[0021] Thirdly, this application also provides an electronic device. The electronic device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the steps of the driving state recognition method in any of the above embodiments.
[0022] Fourthly, this application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program thereon, which, when executed by a processor, implements the steps of the driving state recognition method in any of the above embodiments.
[0023] Fifthly, this application also provides a computer program product. The computer program product includes a computer program that, when executed by a processor, implements the steps of the driving state recognition method in any of the above embodiments.
[0024] The aforementioned driving state recognition method, driving state recognition device, electronic device, storage medium, and computer program product simultaneously acquire audio signals, vehicle state data, and image data obtained from multiple perspectives of the target object during vehicle driving. While obtaining visual encoding vectors based on multiple image data, they also extract driver posture feature vectors based on these image data. Furthermore, by combining the visual encoding vectors and posture feature vectors extracted from the image data, the audio feature vectors extracted from the audio signals, and the vehicle state feature vectors extracted from the vehicle state data, driving state classification processing is performed to obtain driving state categories. Therefore, when recognizing the driver's driving state, the visual encoding vectors reflected in the image data, the driver's posture-related posture feature vectors in the image data, the audio features, and the vehicle state are all combined to jointly identify the driving state. This allows for the determination of the driver's driving state by combining visual presentation, driver posture, audio, and vehicle state, such as whether the driver is distracted, thus improving the accuracy of driving state recognition. Attached Figure Description
[0025] Figure 1 This is a schematic diagram illustrating an application scenario of the driving state recognition method in one embodiment.
[0026] Figure 2 This is a flowchart illustrating a driving state recognition method in one embodiment;
[0027] Figure 3 This is a schematic diagram of the process for obtaining multimodal interaction feature vectors in one embodiment;
[0028] Figure 4 This is a schematic diagram of the process for obtaining a training sample set in one embodiment;
[0029] Figure 5 This is a schematic diagram of the process for obtaining the training sample set in another embodiment;
[0030] Figure 6 This is a schematic diagram of the process for obtaining the driving state category in one embodiment;
[0031] Figure 7 This is a schematic diagram of the system architecture for an application scenario of the driving state recognition method in one embodiment.
[0032] Figure 8This is a flowchart illustrating the automatic generation, filtering, and label generation process of a sample image in an example.
[0033] Figure 9 This is an example of the workflow intent for classifying driving states using multimodal interaction feature vectors.
[0034] Figure 10 This is a schematic diagram of the driving state recognition device in one embodiment. Detailed Implementation
[0035] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0036] The embodiments of the technical solution of this application will now be described in detail with reference to the accompanying drawings. These embodiments are only used to more clearly illustrate the technical solution of this application and are therefore merely examples, and should not be used to limit the scope of protection of this application.
[0037] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains; the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the application; the terms “comprising” and “having”, and any variations thereof, in the specification, claims, and foregoing description of the drawings are intended to cover non-exclusive inclusion.
[0038] In the description of the embodiments of this application, technical terms such as "first" and "second" are used only to distinguish different objects and should not be construed as indicating or implying relative importance or implicitly specifying the number, specific order, or primary and secondary relationship of the indicated technical features. In the description of the embodiments of this application, "multiple" means two or more, unless otherwise explicitly defined.
[0039] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0040] In the description of the embodiments in this application, the term "and / or" is merely a description of the association relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. In addition, the character " / " in this document generally indicates that the related objects before and after are in an "or" relationship, and the term "multiple" refers to two or more (including two).
[0041] It should be noted that all information and data involved in this application (including but not limited to data used for analysis, stored data, and displayed data) are information and data authorized by the user or fully authorized by all parties, and the acquisition, transmission, storage, use, and processing of the relevant data comply with the relevant provisions of national laws and regulations. In the embodiments of this application, certain existing industry solutions such as software, components, and models may be mentioned. These should be considered exemplary, and their purpose is merely to illustrate the feasibility of implementing the technical solution of this application, but does not imply that the applicant has already used or necessarily used such a solution.
[0042] Currently, when identifying driving states, such as identifying and assessing whether a driver is distracted, the usual method is to capture images of the driver's face and analyze facial expressions based on the captured images to determine whether the driver is distracted or in a distracted state. However, observations of practical applications have revealed that this method of driving state identification suffers from low accuracy.
[0043] Research has revealed that when drivers are distracted, in addition to the recognition results of the visuals themselves, the driver's posture, the audio in the driving environment, and the vehicle's status differ from when the driver is not distracted. Furthermore, the visuals, driver's posture, audio in the driving environment, and vehicle status also differ under different distraction behaviors. Therefore, combining the visuals, driver's posture, audio in the driving environment, and vehicle status to classify and evaluate driving status can improve the accuracy of driving status recognition.
[0044] Accordingly, embodiments of this application provide a driving state recognition method, which can be applied to, for example... Figure 1The application environment shown can be related to scenarios such as in-vehicle driver monitoring systems (DMS), intelligent cockpit safety perception, driving behavior analysis, road safety management, and vehicle insurance risk assessment. Terminal 102 communicates with server 104 via a network. A data storage system can store the data that server 104 needs to process. The data storage system can be integrated onto server 104 or located in the cloud or on other network servers. Terminal 102 can collect driving-related data during vehicle operation and send this data to server 104, which then uses this data to identify the driver's driving state. In other examples, terminal 102 may collect driving-related data and then use this data to identify the driver's driving state. The terminal 102 can be, but is not limited to, any device capable of acquiring driving-related data, or acquiring driving-related data and being able to identify the driver's driving status based on that data. This includes, but is not limited to, in-vehicle terminals, personal computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices can include smart speakers, smart TVs, smart air conditioners, and smart in-vehicle devices. Portable wearable devices can include smartwatches, smart bracelets, and head-mounted devices. The server 104 can be implemented using a standalone server or a server cluster consisting of multiple servers.
[0045] In one embodiment, such as Figure 2 As shown, a driving state recognition method is provided, including the following steps:
[0046] Step S201: Acquire driving-related data collected at the same time during the vehicle driving process. The driving-related data includes: audio signals, vehicle status data, and multiple image data obtained from image acquisition of the target object from multiple perspectives.
[0047] Driving-related data refers to data directly or indirectly related to the driving process during vehicle operation. In this application's embodiments, this includes image data, audio signals, and vehicle status data.
[0048] Driving-related data at the same time refers to driving-related data with the same collection timestamp or the difference between collection timestamps being within the allowable time error range of the same collection period. That is, the collection timestamps of multiple image data, audio signal, and vehicle status data are all the same, or the errors between collection timestamps are all within the allowable time error range.
[0049] Image data refers to the acquired video frame images. A camera device with a shooting range covering the driver's position can be installed in the vehicle to capture video streams or images, obtaining corresponding image data. The target object is the driver, the user at the driver's position. The specific shooting angles for multiple perspectives are not limited; for example, cameras can be installed on the dashboard, rearview mirror, and right window to capture the driver from the dashboard view, rearview mirror view, and right window view respectively, obtaining multiple video frame images.
[0050] Audio signals refer to audio-related data that has been collected. Devices such as audio acquisition units can be installed in vehicles to collect audio signals within the vehicle. In the relevant examples of this application, the audio acquisition unit can be positioned relatively close to the driver's seat to collect the driver's audio signals as accurately as possible, thereby improving the accuracy of driving status recognition.
[0051] Vehicle status data refers to data related to the vehicle's status while driving. The method of obtaining vehicle status data is not limited; in the relevant examples, it can be obtained through the vehicle's CAN (Controller Area Network) bus, but is not limited to this. The specific data type of vehicle status data is not limited, and includes, but is not limited to, steering angle, throttle, brakes, lane departure, etc.
[0052] Step S202: Perform image encoding on multiple image data to obtain a visual encoding vector.
[0053] For multiple acquired image data sets, the image data can be encoded to obtain visual encoding vectors. These visual encoding vectors reflect the overall characteristics of the image data, thus providing a visual representation of the image data.
[0054] It is understandable that image encoding of multiple image data to obtain a visual encoding vector can include: separately encoding the multiple image data to obtain multiple image feature vectors, and then fusing these multiple image feature vectors to obtain the visual encoding vector. Accordingly, for multiple image data obtained from multiple different perspectives, each image data can be separately image encoded to obtain multiple image feature vectors. Based on these multiple image feature vectors, they can be fused together, and the resulting vector is the visual encoding vector.
[0055] By obtaining visual coding vectors from multiple image data acquired from multiple different angles, visual evaluation can be performed from multiple different angles, which can improve the accuracy of the obtained visual coding vectors and thus help to further improve the accuracy of driving state recognition.
[0056] Step S203: Extract driver posture information from image data and obtain posture feature vector based on driver posture information.
[0057] Furthermore, driver posture information can be extracted from the obtained image data to obtain posture feature vectors that reflect the driver's posture. These posture feature vectors reflect features related to the driver's posture within the corresponding image data, thus allowing the acquisition of posture feature vectors related to the driver's posture presented in the image data.
[0058] Step S204: Extract features from the audio signal to obtain audio feature vectors.
[0059] For the obtained audio signal, features can be extracted to obtain the corresponding audio feature vector. The method of feature extraction is not limited. In relevant examples, methods for extracting audio features and obtaining the audio feature vector may include: extracting the spectral features of the audio signal using Mel-frequency cepstral coefficients (MFCC); or generating audio feature vectors by using an audio feature encoder (e.g., an Audio Transformer encoder) to extract features from the spectral features.
[0060] Step S205: Perform feature vector generation processing on the vehicle state data to obtain the vehicle state feature vector.
[0061] Based on the obtained vehicle state data, a corresponding feature vector can be generated, thereby obtaining the vehicle state feature vector.
[0062] The method for generating feature vectors from vehicle state data is not limited. In relevant examples, it can involve normalizing multiple different types of vehicle state data, obtaining vehicle state data within a preset time window, and then encoding the vehicle state data within the preset time window using an encoder to obtain vehicle state feature vectors. The type of encoder used to encode the vehicle state data is not limited; some examples include, but are not limited to, lightweight MLPs (Multi-layer Perceptrons).
[0063] In other examples, for multiple different types of vehicle status data obtained, the vehicle status data within a preset time window can be normalized, and then the normalized vehicle status data can be encoded to obtain a vehicle status feature vector.
[0064] When determining the preset time window, it can be based on the time of obtaining multiple image data or audio signals, so that image data, audio signals and vehicle status data are obtained at the same time point or within the same time segment, and the time synchronization of image data, audio signals and vehicle status data is achieved, but it is not limited to this.
[0065] Step S206: Perform feature fusion processing on the visual encoding vector, posture feature vector, visual encoding vector and vehicle state feature vector to obtain multimodal interaction feature vector.
[0066] Based on the obtained visual encoding vector, posture feature vector, and vehicle state feature vector, feature fusion processing can be performed to obtain multimodal interaction feature vectors, which are then used for driving state classification.
[0067] Step S207: Perform driving state classification processing on the multimodal interaction feature vector to obtain the driving state category.
[0068] Based on the obtained multimodal interaction feature vectors, driving state classification can be performed on the multimodal interaction feature vectors to obtain the corresponding driving state categories. The specific categories of driving state categories can be different types of driver distraction, such as using a mobile phone, talking to passengers, thinking about problems, etc., but are not limited to these.
[0069] The driving state recognition method based on this embodiment simultaneously acquires audio signals, vehicle state data, and image data obtained from multiple perspectives of the target object during vehicle driving. While obtaining visual encoding vectors based on multiple image data, it also extracts driver posture feature vectors based on these image data. Furthermore, it combines the visual encoding vectors and posture feature vectors extracted from the image data, the audio feature vectors extracted from the audio signals, and the vehicle state feature vectors extracted from the vehicle state data to perform driving state classification processing and obtain driving state categories. Therefore, when recognizing the driver's driving state, it simultaneously combines the visual encoding vectors reflected in the image data, the driver's posture-related posture feature vectors in the image data, audio features, and vehicle state to jointly identify the driving state. This allows for the combined determination of the driver's driving state, such as whether the driver is distracted, by considering the visual presentation, driver posture, audio, and vehicle state, thereby improving the accuracy of driving state recognition.
[0070] In the process of encoding multiple image data to obtain visual encoding vectors, there are no restrictions on the method of encoding multiple image data separately to obtain multiple image feature vectors. In some examples, a pre-trained image encoder can be used to extract image features from multiple image data separately to obtain multiple image feature vectors.
[0071] In a specific example, a pre-trained image encoder can be used to encode multiple image data separately, obtain multiple image feature vectors, and then, based on the weight coefficients of the images at each angle in the image encoder, the image feature vectors corresponding to each angle are weighted and summed to obtain a visual encoding vector.
[0072] The type of image encoder obtained through pre-training is not limited, as long as it can extract image features to generate and obtain image feature vectors. In the relevant examples of this application, the pre-trained image encoder can be a Swing Transformer encoder, etc., but is not limited to this.
[0073] The method for training the image encoder for image feature extraction is not limited. In the relevant examples of this application, a self-supervised learning approach can be used to train the image encoder. Specific examples include the following methods for training the image encoder:
[0074] Obtain the sample input image, which includes sample images from multiple angles;
[0075] By randomly occluding sample images and reconstructing the occluded regions, the encoder to be trained is subjected to self-supervised learning, and after the self-supervised learning is completed, the pre-trained image encoder is obtained.
[0076] By randomly occluding sample images, the encoder to be trained reconstructs the image of the occluded region. The encoder is then trained by calculating the error between the reconstructed image region and the occluded region, or by calculating the error between the reconstructed image and the unoccluded sample image. This allows for the training of an image encoder capable of extracting image features through self-supervised learning. The random occlusion of sample images can involve any one or multiple sample images from the input image set; this embodiment does not impose specific limitations on this method.
[0077] In the training process of the encoder to be trained, the encoder can be trained by minimizing the reconstruction error, that is, minimizing the error between the sample input image and the reconstructed image. This error can be expressed as:
[0078] .
[0079] in, The input image is a sample image, which can be any one or more sample images from multiple viewpoints acquired at the same time. This indicates that for all samples in the training sample set D To find the expected value, we need to calculate the average loss across all samples. To obscure the operation, For the image encoder's encoding and decoding parameters, are trainable weights and biases. To reconstruct errors, The objective is to minimize the parameters. The function.
[0080] The encoder is trained by randomly occluding and reconstructing images, which essentially achieves pre-training using masked image modeling. This method can learn the latent features of images without setting labels for them, simplifying the training process.
[0081] Therefore, when extracting image features from image data to obtain image feature vectors, image feature extraction is performed using a pre-trained image encoder. This image encoder is trained through self-supervised learning by randomly occluding input sample images and reconstructing the occluded regions. This eliminates the need to label the input sample images, simplifying the training process. Furthermore, it can be implemented using a lightweight encoder, resulting in lower resource consumption and easier deployment. Specifically, when the input sample images include images from multiple angles, during the self-supervised learning process of the encoder to be trained, each sample image can be randomly occluded, and the occluded regions can be reconstructed to continue self-supervised learning training. The weight coefficients of the sample images at different angles are adjusted during training. Thus, the image encoder obtained after training can extract and recognize image feature vectors from sample images at different angles. It can also perform weighted summation of the image feature vectors corresponding to images at different angles based on the weight coefficients of the images at different angles to obtain a visual encoding vector.
[0082] Therefore, when training the image encoder, it is trained on sample images from multiple different angles. This enables the image encoder to learn the ability to extract image feature vectors from images at different angles. At the same time, it also learns and determines the weight coefficients of images at different angles. Then, when processing based on the trained image encoder, the image feature vectors of images at different angles can be weighted and summed based on the different weight coefficients of images at different angles. This results in the final visual encoding vector being determined by combining the differences between images at different angles, which can further improve the accuracy of the obtained visual encoding vector.
[0083] The method of fusing visual encoding vectors, pose feature vectors, audio feature vectors, and vehicle state feature vectors to obtain multimodal interaction feature vectors is not limited. In some examples, refer to... Figure 3 As shown, methods for obtaining multimodal interaction feature vectors may include:
[0084] Step S301: Perform linear mapping on the visual encoding vector, posture feature vector, audio feature vector, and vehicle state feature vector respectively to obtain the mapped visual encoding vector, mapped posture feature vector, mapped audio feature vector, and mapped vehicle state feature vector.
[0085] Through linear mapping, the visual encoding vector, posture feature vector, audio feature vector, and vehicle state feature vector can be transformed to the same dimension, that is, to obtain the mapped visual encoding vector, mapped posture feature vector, mapped audio feature vector, and mapped vehicle state feature vector with the same dimension.
[0086] Taking the linear mapping of a visual encoding vector into a mapped visual encoding vector as an example, the obtained mapped visual encoding vector can include a query vector, a key vector, and a value vector corresponding to the visual input, so that subsequent attention processing can be performed based on these query vector, key vector, and value vector. Again, using the visual encoding vector as an example, the process of linearly mapping the visual encoding vector can be expressed by the formula:
[0087]
[0088] in, , and These represent the query vector, key vector, and value vector, respectively. Represents the visual encoding vector. , and These represent the parameter matrices for the query vector, key vector, and value vector, respectively.
[0089] The specific method of linear mapping is not limited and can be any method already existing in related technologies. This application does not impose any specific restrictions on this.
[0090] Step S302: Perform cross-modal attention calculation on the mapped visual encoding vector, mapped posture feature vector, mapped audio feature vector, and mapped vehicle state feature vector to obtain multiple modal interaction features.
[0091] After obtaining the mapped visual encoding vector, mapped posture feature vector, mapped audio feature vector, and mapped vehicle state feature vector, cross-modal attention computation can be performed to obtain multiple modal interaction features.
[0092] Cross-modal attention computation refers to the process of selectively extracting relevant information from the feature vectors of another modality by paying attention to the feature vectors of another modality, in order to form a new integrated feature vector.
[0093] Taking the calculation of visual attention to pose as an example, and performing cross-modal attention calculation on the mapped visual encoding vector and the mapped pose feature vector, the modal interaction features obtained by visual attention to pose can be expressed by the formula:
[0094] .
[0095] in, Representing visual modalities attitude mode The attentional modal interaction features, also known as visual pose interaction features, The query vector represents the visual modality, that is, the query vector in the mapped visual encoding vector. and These represent the key vector and value vector of the attitude mode, respectively; that is, the key vector and value vector in the mapped attitude encoding vector. Used to calculate query vector With key vector The similarity is used to measure the degree of correlation between visual and pose information. This is a scaling factor used to prevent the dot product result from becoming too large, thus avoiding an excessively small gradient in the subsequent Softmax function. This is a normalization function. The specific method for cross-modal attention calculation is not limited; existing cross-modal attention calculation methods in related technologies can be used. This application's embodiments do not impose any restrictions on this.
[0096] In the relevant examples of this application, when performing cross-modal attention calculation, cross-modal attention can be calculated on the mapped feature vectors of any two modalities. For example, calculating visual attention to speech involves performing cross-modal attention calculation on the mapped visual encoding vector and the mapped audio feature vector to obtain visual-speech interaction features. Similarly, calculating visual attention to posture involves performing cross-modal attention calculation on the mapped visual encoding vector and the mapped posture feature vector to obtain visual-posture interaction features; calculating visual attention to a vehicle involves performing cross-modal attention calculation on the mapped visual encoding vector and the mapped vehicle state feature vector to obtain visual-vehicle interaction features; and calculating posture attention to speech involves performing cross-modal attention calculation on the mapped posture feature vector and the mapped speech feature vector to obtain posture-speech interaction features, and so on.
[0097] In this process, cross-modal attention can be calculated on any two of the mapped visual encoding vector, mapped posture feature vector, mapped audio feature vector, and mapped vehicle state feature vector to obtain multiple modal interaction features.
[0098] Step S303: Perform weighted summation on the interaction features of each modality to obtain the multimodal interaction feature vector.
[0099] Therefore, the obtained multimodal interaction features can be expressed by the following formula: .in, Represents the multimodal interaction feature vector. Representing the first of the four modalities: visual, pose, audio, or vehicle. The modal feature for the first Modal interaction features for attention of modal features.
[0100] Accordingly, when obtaining multimodal interaction feature vectors by performing feature fusion processing based on visual encoding vectors, posture feature vectors, audio feature vectors, and vehicle state feature vectors, the process involves linearly mapping each of these vectors. Then, cross-modal attention calculations are performed on the mapped visual encoding vectors, posture feature vectors, audio feature vectors, and vehicle state feature vectors to obtain multiple modal interaction features. Finally, a weighted summation of these modal interaction features is performed. This allows for the acquisition of the final multimodal interaction feature vectors based on the dependencies between feature vectors from different modalities, improving the accuracy of the final multimodal interaction feature vectors and further enhancing the accuracy of driving state recognition.
[0101] In the above embodiments, a trained driving state classification model can be used to process driving-related data to obtain multimodal interaction feature vectors, and then the multimodal interaction feature vectors can be classified to obtain driving state categories. Specifically, the trained driving state classification model can be used to encode multiple image data to obtain visual encoding vectors; extract driver posture information from multiple image data and obtain posture feature vectors based on the driver posture information; extract features from audio signals to obtain audio feature vectors; generate vehicle state feature vectors from vehicle state data; perform feature fusion processing on the visual encoding vector, posture feature vector, audio feature vector, and vehicle state feature vector to obtain multimodal interaction feature vectors; and then classify the multimodal interaction feature vectors to obtain driving state categories. Therefore, in step S207, when classifying the multimodal interaction feature vectors to obtain driving state categories, a trained driving state classification model can be used to process driving-related data to obtain multimodal interaction feature vectors, and then the multimodal interaction feature vectors can be classified to obtain driving state categories.
[0102] The specific model architecture of the driving state classification model is not limited, as long as it can extract features from the input driving-related data to obtain visual encoding vectors, posture feature vectors, audio feature vectors and vehicle state feature vectors, fuse them to obtain multimodal interaction feature vectors, and classify and identify the multimodal interaction feature vectors for output. Any possible deep learning model structure can be used to implement it.
[0103] The model architecture of a driving state classification model in a specific example can be as follows: Figure 9As shown, it includes an image encoder 901, a pose detection module 902, an audio encoder 903, an MLP encoder 904, a multi-head cross-modal attention module 905 connected to the image encoder 901, pose detection module 902, audio encoder 903, and MLP encoder 904, and a classifier 906 connected to the multi-head cross-modal attention module 905, wherein:
[0104] Image encoder 901 encodes image data from different perspectives to obtain image feature vectors from different perspectives, and fuses these feature vectors to obtain a visual encoding vector. The specific architecture of image encoder 901 is not limited; it can be an image encoder trained by randomly occluding and reconstructing images as described above. In some specific examples, the image encoder type can be a Swin Transformer encoder, while in other embodiments it can be a ViT (Vision Transformer) encoder, ConvNeXt (a pure convolutional network architecture), or CNN-LSTM (a hybrid model architecture that includes CNN (Convolutional Neural Network) and LSTM (Long Short-Term Memory)).
[0105] The pose detection module 902 extracts driver pose information from multiple image data and obtains pose feature vectors based on the driver pose information. The specific architecture of the pose detection module 902 is not limited. In some examples, it can be a hybrid model structure including DWpose (a whole-body pose estimation model) and a lightweight MLP (Multilayer Perceptron) model. In other examples, DWpose (a whole-body pose estimation model) can also be replaced by OpenPose (an open-source real-time multi-person pose estimation library), MediaPipe Pose (a real-time human keypoint detection model under the Google MediaPipe framework), or AlphaPose (an open-source multi-person pose estimation system), etc., but is not limited to these.
[0106] The audio encoder 903 extracts features from the input audio signal to obtain an audio feature vector. The specific architecture of the audio encoder 903 is not limited. Some examples may include MFCC (Mel-Frequency Cepstral Coefficients, a speech feature extraction method based on the characteristics of human hearing) and Audio Transformer encoder, but it is not limited to these.
[0107] The MLP encoder 904 performs feature vector generation processing on vehicle state data to obtain vehicle state feature vectors.
[0108] The multi-head cross-modal attention module 905, based on a multi-head cross-modal attention mechanism, performs feature fusion processing on visual encoding vectors, pose feature vectors, audio feature vectors, and vehicle state feature vectors to obtain multi-modal interaction feature vectors. The specific architecture of this multi-head cross-modal attention module 905 is not limited; some examples may use cross-modal attention (CMA), but it is not limited to these.
[0109] Classifier 906 performs driving state classification on multimodal interaction feature vectors to obtain driving state categories. The specific architecture of classifier 906 is not limited; in some examples, it can include a classification head that includes an MLP model and a Softmax (activation function).
[0110] Accordingly, when performing driving state classification processing on multimodal interaction feature vectors to obtain driving state categories, it is based on the driving state classification model obtained through training. This can effectively improve the efficiency of driving state recognition and help to perform stable and reliable classification of driving state recognition.
[0111] In the process of training the driving state classification model, the initial model is trained based on the acquired training sample set. The size of the training sample set affects the accuracy of the model training, and in practical applications, the extensive data annotation work also significantly impacts the efficiency of acquiring the training sample set. Therefore, in some embodiments, reference is made to... Figure 4 As shown, the methods for obtaining the training sample set for training the driving state classification model include:
[0112] Step S401: Obtain an initial sample set. Each initial sample in the initial sample set includes sample driving-related data, which includes: sample audio signals, sample vehicle status data, and sample image data from multiple angles.
[0113] The initial sample set is the raw sample set that can be used to train the driving state classification model. The driving-related data in each initial sample in the initial sample set can be obtained by collecting and acquiring sample image data, audio signals, and vehicle state data from multiple angles during the actual driving test by the test driver. Each initial sample can also contain a corresponding sample label, which includes the driving classification information corresponding to that initial sample. This driving classification information can be information determined manually, and it can be a classification label, a textual description of the driving classification category, etc., but is not limited to these.
[0114] Step S402: Extract driver posture key points from sample image data.
[0115] For the initial samples in the initial sample set, driver posture key points can be extracted from the sample image data of the initial samples. Driver posture key points are information related to the driver's posture, including but not limited to facial feature points, shoulder feature points, abdominal feature points, hand feature points, and foot feature points, as long as they can reflect the driver's posture characteristics.
[0116] There are no restrictions on the methods for extracting driver pose key points. For example, DWpose (a state-of-the-art (SOTA) generative model for key point detection) can be used to extract driver pose key points, but it is not limited to this.
[0117] In extracting key points of the driver's posture, one can arbitrarily select a sample image from multiple angles and extract the driver's posture information from that sample image.
[0118] Step S403: Use an image encoder to encode the sample image data to obtain the sample image feature vector.
[0119] For the sample image data, an image encoder is used to encode the sample image data to obtain the sample image feature vector. The extracted sample image feature vector reflects the overall characteristics of the sample image data.
[0120] The type of image encoder used to extract image features from sample image data is not limited. In this embodiment, the image encoder obtained by pre-training is used to encode the sample image data and obtain the sample image feature vector.
[0121] Step S404: Generate an attitude enhancement image based on the driver's attitude key points and the feature vector of the sample image.
[0122] The driver posture key points extracted in step S402 reflect the features related to driver posture in the sample image data. The sample image feature vector extracted in step S403 reflects the overall image features of the sample image data. Therefore, a posture enhancement image can be generated based on the driver posture key points and the sample image feature vector. This allows the generated posture enhancement image to retain the rationality of the driving posture determined based on the driver posture key points and to generate semantically rich images, such as posture enhancement images with rich and diverse lighting, angles, and action details.
[0123] There are no restrictions on the method used to generate pose-enhanced images. In some examples, new pose-enhanced images with consistent poses but diverse semantics can be generated based on Progressive Conditional Diffusion Models (PCDMs), which can be expressed by the formula: .in, The driver pose key points extracted from the sample image data can be represented as a structured skeleton diagram. This refers to the sample image feature vector extracted from the sample image data. This represents the model function corresponding to the asymptotic conditional diffusion model. This represents the generated pose-enhanced image.
[0124] This can be achieved by extracting driver posture key points from the same sample image data. and sample image feature vector This allows for the generation of high-quality images that meet posture requirements while preserving image details. In this embodiment, driver posture key points can also be extracted separately from different sample image data. and sample image feature vector This allows image details to be preserved while changing pose, and the resulting pose enhancement image can be used to improve image quality, which helps to further enhance the robustness and generalization ability of the model.
[0125] Step S405: Obtain the training sample set based on the initial sample set and pose-enhanced images.
[0126] Based on the obtained pose-enhanced images, a training sample set can be obtained using the pose-enhanced images and the initial sample set. It can be understood that the training samples in the generated training sample set include the initial samples from the initial sample set, as well as the training samples obtained based on the generated pose-enhanced images. The sample labels of the training samples obtained based on the pose-enhanced images can be set to be the same as the sample labels of the initial samples corresponding to that pose-enhanced image.
[0127] In the relevant examples of this application, the pose enhancement images can be further filtered, and a training sample set can be obtained based on the filtered pose enhancement images.
[0128] Accordingly, in some embodiments, reference Figure 5 As shown, a training sample set is obtained based on the initial sample set and pose-enhanced images, including:
[0129] Step S501: Recognize the pose-enhanced image using a visual language model to obtain the image semantics of the pose-enhanced image, and determine the matching degree between the image semantics and the label semantics of the sample labels of the initial samples.
[0130] Visual language models are large, multimodal models capable of processing both images (visual) and text (language) simultaneously. They enable the identification and analysis of the matching degree between images and text. The specific type of visual language model is not limited; for example, visual language models may include, but are not limited to, the CogVLM model (a basic visual language model).
[0131] By inputting the generated pose augmentation image and the corresponding sample label text (the text of the sample label of the initial sample or the text of the label semantics corresponding to the semantics of the sample label) into the visual language model, the visual language model can process the pose augmentation image and the sample label text and output the matching degree between the pose augmentation image and the sample label text.
[0132] In a specific example, the CogVLM model can be used to obtain the matching degree between the image semantics and the label semantics of the initial sample. The CogVLM model can perform deep fusion of linguistic features and visual features to achieve a deep understanding of images and text, and obtain the matching degree between image semantics and label semantics.
[0133] In a specific example, a visual language model is used to recognize the pose-enhanced image, obtain the image semantics of the pose-enhanced image, and determine the matching degree between the image semantics and the label semantics of the initial sample labels, including:
[0134] Image features of pose-enhanced images are extracted by a visual encoder in a visual language model, and then the image features are mapped to the same dimensional space as the text features to obtain image semantic text features.
[0135] The text encoder in the visual language model is used to encode the sample label text corresponding to the pose enhancement image to obtain the label text features.
[0136] Calculate the similarity between the semantic text features of the image and the text features of the label to obtain the matching degree between the semantics of the image and the semantics of the label.
[0137] Therefore, the process of obtaining the matching degree between pose-enhanced images and sample labels through visual language models can be expressed by the formula:
[0138] .
[0139] in, To enhance the pose of the image and the first The degree of match between each driving distraction label category The cosine similarity function is used. For the visual encoder of CogVLM, To enhance the pose of the image, For CogVLM text encoder, For the first Individual driving distraction tag categories The semantic matching prompts, such as the prompt "the driver is using a mobile phone," correspond to the label (driving category) "using a mobile phone." These semantic matching prompts describe the semantic concept represented by the label. The prompts can be predetermined, and there is a mapping relationship between them and the labels. During the classification process, the similarity between the semantic matching prompts and the labels is calculated to determine the feature's classification category. For example, the label corresponding to the semantic matching prompt with the highest similarity is taken as the feature's classification category.
[0140] The specific categories for driving distraction labels are not limited. Some examples may include, but are not limited to, visual distraction, manual distraction, cognitive distraction, and normal, but are not limited to these. When the driving distraction label is visual distraction, it indicates that the driver is shifting their gaze away from the road and traffic environment to other visual targets, such as looking away from the road, like looking at a mobile phone or the front passenger. When the driving distraction label is manual distraction, it indicates that the driver is taking one or both hands off the steering wheel, such as eating or adjusting the air conditioning. When the driving distraction label is cognitive distraction, it indicates that the driver is inattentive or distracted, such as daydreaming.
[0141] In other examples, the driving distraction label categories can be further subdivided. For example, driving distraction label categories may include: looking at a mobile phone, looking at passengers, making a phone call, eating with hands, and daydreaming, but are not limited to these. The specific subdivision of driving distraction label categories can be determined based on the level of detail in the classification of driving distraction.
[0142] Step S502: If the matching degree is greater than or equal to the matching degree threshold, retain the pose enhancement image; if the matching degree is less than the matching degree threshold, discard the pose enhancement image.
[0143] If the matching degree is greater than or equal to a preset matching degree threshold, it indicates that the semantics of the pose augmentation image is consistent with the semantics of the sample label, and thus the pose augmentation image is retained; otherwise, the pose augmentation image is discarded. The specific value of the matching degree threshold is not limited and can be determined based on the semantic matching accuracy between the image and the text. In some examples, the matching degree threshold can be set to 0.8, but it is not limited to this.
[0144] Step S503: Based on the initial sample set and the retained pose-enhanced images, obtain the training sample set.
[0145] Based on the initial sample set and the preserved pose-enhanced images, a training sample set can be obtained to train the driving state classification model.
[0146] Each training sample in the training sample set includes driving-related data, which includes: sample audio signals, sample vehicle state data, and multiple sample image data. Where the sample image data in the training sample is a generated pose-enhanced image, the corresponding sample audio signals and sample vehicle state data can be the sample audio signals and sample vehicle state data from the initial sample corresponding to the pose-enhanced image. In other examples, they can also be the sample audio signals and sample vehicle state data after data enhancement based on the sample audio signals and sample vehicle state data from the initial sample corresponding to the pose-enhanced image, but are not limited to these.
[0147] Accordingly, based on the obtained pose augmentation images, the generated pose augmentation images are further filtered by using a visual language model to measure the semantic matching degree between the generated pose augmentation images and the sample labels corresponding to the sample image data. This enables the automatic filtering and generation of samples corresponding to the generated pose augmentation images, which helps to improve the accuracy of the obtained training sample set, and in turn helps to improve the accuracy of the trained driving state classification model, and further helps to improve the accuracy of driving state recognition in practical applications.
[0148] Therefore, when obtaining the training sample set for training the driving state recognition model, based on the initial sample set, the driver posture key points are extracted from the sample image data in the initial sample set, and the sample image features obtained by image feature extraction are combined with the sample image data to generate posture enhancement images. The training sample set is then obtained based on the initial sample set and the generated posture enhancement images. This can effectively expand the effective samples and improve the convergence and stability of training the driving state classification model.
[0149] refer to Figure 6 As shown, in some embodiments, the driving state classification processing of the multimodal interaction feature vector in step S207 above to obtain the driving state category includes:
[0150] Step S601: Based on the driving state classification model obtained through training, perform driving state classification processing on the multimodal interaction feature vector to obtain the driving state recognition result of the current frame.
[0151] It is understandable that the driving state classification model obtained through training can encode multiple image data to obtain a visual encoding vector; extract driver posture information from multiple image data and obtain posture feature vectors based on the driver posture information; extract features from audio signals to obtain audio feature vectors; generate vehicle state feature vectors from vehicle state data; perform feature fusion processing on the visual encoding vector, posture feature vector, audio feature vector, and vehicle state feature vector to obtain a multimodal interaction feature vector; and perform driving state classification processing on the multimodal interaction feature vector to obtain the driving state category.
[0152] It is understandable that after obtaining the multimodal interaction feature vector, the multimodal interaction feature vector can be processed for driving state classification based on the output layer of the driving state classification model to obtain the driving state recognition result of the current frame. For example, the output layer may include a classification head and a post-processing part. The classification head is used to output the probability (also referred to as confidence) of the multimodal interaction feature vector mapped to each driving state category in this embodiment. The post-processing part determines and outputs the final classification prediction result, i.e., the identified driving state category, based on the probability of each driving state category.
[0153] The obtained current frame driving state recognition result can include the driving state category output by the post-processing part, or the confidence level output by the classification head (referred to as the first confidence level in this application embodiment). In the relevant embodiments of this application, the obtained current frame driving state recognition result includes the first confidence level of each driving state category. Accordingly, driving state classification processing is performed on the multimodal interaction feature vector to obtain the current frame driving state recognition result, which can be expressed by the formula: .in, Indicates the first The first confidence score of classifying the multimodal interaction feature vectors corresponding to the driving-related data of each frame into each driving state category. For activation function, For multimodal interaction feature vectors, This is the weight matrix. This is the bias vector.
[0154] Step S602: Obtain the driving state recognition results of adjacent frames by performing driving state classification processing on the driving-related data of adjacent frames according to the driving state classification model.
[0155] Driving-related data in adjacent frames refers to data that is temporally adjacent to the driving-related data in the current frame, and can be determined based on the timestamp of the acquired image data. When acquiring image data using a camera device, it is typically based on frame extraction from the video stream captured by the camera, and the data of one extracted video frame is used as the acquired image data. For example, in some examples, multiple video frames acquired from the video stream of a specific camera device can be sequentially denoted as follows: , … , .
[0156] Therefore, multiple driving-related data recorded in chronological order can be denoted as follows: , … , .in, , ... , These represent the corresponding audio signals. , ... , These represent the corresponding vehicle status data.
[0157] The order and number of adjacent frames are not limited. In some examples, the driving-related data of adjacent frames can be driving-related data from multiple earlier adjacent frames. The driving-related data at the current moment is used as... For example, assuming there are 3 adjacent frames, the driving-related data for adjacent frames includes: , and .
[0158] In other examples, the driving-related data of adjacent frames may include driving-related data from a first preset number of adjacent frames that are earlier in time, and driving-related data from a second preset number of adjacent frames that are later in time. The first and second preset numbers may be the same or different, and their specific values can be set based on the needs of the actual technical scenario. For example, they can be determined based on the acquisition frequency of the video captured by the camera device and the frame extraction frequency of the video stream captured by the camera device. Taking the driving-related data at the current moment as... Taking a first preset number of 3 and a second preset number of 2 as an example, the driving-related data of adjacent frames includes: , , , and .
[0159] The driving state classification model obtained through training can classify the driving-related data of each adjacent frame, thereby obtaining the driving state recognition results of each adjacent frame.
[0160] Step S603: Correct the current frame's driving state recognition result based on the driving state recognition results of adjacent frames to obtain the corrected driving state recognition result, and determine the driving state category based on the corrected driving state recognition result.
[0161] Based on the obtained driving state recognition results of adjacent frames, the driving state recognition results of the current frame can be corrected, and the driving state category can be determined based on the corrected driving state recognition results.
[0162] The method for correcting the current frame's driving state recognition result based on the obtained driving state recognition results of adjacent frames is not limited. In some examples, the current frame's driving state recognition result and the driving state recognition results of adjacent frames can include the output driving state category. In this case, the consistency between the current frame's driving state category and the driving state categories of adjacent frames can be compared. If they are inconsistent, the current frame's driving state category is modified to match the driving state categories of adjacent frames, thus obtaining the final determined driving state category. In this case, taking the driving-related data of adjacent frames including the driving-related data of the first preset number of adjacent frames with earlier times as an example, if the driving state categories of all adjacent frames are consistent, but the current frame's driving state category is inconsistent with them, it indicates that a sudden change has occurred in the driving state recognition, possibly resulting in a momentary misjudgment. Therefore, by modifying it to match the driving state category of adjacent frames, a smooth driving state recognition effect can be achieved.
[0163] In other examples of this application, the current frame driving state recognition result includes: the first confidence level of each driving state category corresponding to the multimodal interaction feature vector of the driving-related data of the current frame; the adjacent frame driving state recognition result includes: the second confidence level of each driving state category corresponding to the driving-related data of the adjacent frames.
[0164] It is understood that the driving-related data of adjacent frames can also be processed in the same way as described above to obtain the corresponding multimodal interaction feature vector, and the driving state classification processing of the multimodal interaction feature vector can be performed to obtain the driving state recognition result of the adjacent frame corresponding to the driving-related data of the adjacent frame. The driving state recognition result of the adjacent frame includes the second confidence of each driving state category.
[0165] Here, the first confidence level is the classification probability of assigning the multimodal interaction feature vector of driving-related data in the current frame to the corresponding driving state category. Similarly, the second confidence level is the classification probability of assigning the multimodal interaction feature vector of driving-related data in adjacent frames to the corresponding driving state category. It can be understood that different driving state categories can each have their own corresponding first and second confidence levels.
[0166] At this point, in step S603 above, the driving state recognition result of the current frame is corrected based on the driving state recognition results of adjacent frames to obtain the corrected driving state recognition result, and the driving state category is determined based on the corrected driving state recognition result, including:
[0167] Based on the second confidence level of each driving state category, the first confidence level of each driving state category is corrected to obtain the corrected confidence level corresponding to each driving state category;
[0168] The driving state category corresponding to the highest confidence among the corrected confidence scores for each driving state category is determined as the identified driving state category.
[0169] Accordingly, the first confidence level corresponding to each driving state category can be determined by classifying driving state data of the current frame, and then combined with the second confidence level corresponding to each driving state category of each adjacent frame determined by classifying driving state data of the adjacent frames. This process is used to correct the first confidence level corresponding to each driving state category of the current frame, thereby obtaining the corrected confidence level of each driving state category. The driving state category corresponding to the highest confidence level among the corrected confidence levels is then determined as the driving state category of the current frame.
[0170] It is understandable that when using a driving state classification model to classify the multimodal interaction feature vectors corresponding to driving-related data, the driving state classification model will obtain different confidence levels for different driving state categories. The confidence level corresponding to the driving state category reflects the credibility or probability of classifying the driving-related data into that driving state category. Therefore, in the existing methods, the driving state category with the highest confidence level is taken as the driving state category output by the driving state classification model.
[0171] In this embodiment, when processing the driving state data of each adjacent frame using the driving state classification model, the second confidence level corresponding to each driving state category is obtained. The first confidence level corresponding to each driving state category of the current frame's driving state data is then corrected to obtain the corrected confidence level. The driving state category corresponding to the current frame's driving state data is then determined by combining the corrected confidence level. This achieves the correction of the confidence level output by the driving state classification model based on a sliding time window, which can significantly reduce instantaneous misjudgments and improve prediction stability.
[0172] The width of the sliding time window is unlimited. The width of the sliding time window determines the number of adjacent frames. Taking the first preset number as 3 and the second preset number as 2 as an example, the width of the sliding time window can be 6, but it is not limited to this.
[0173] In some specific examples, the first confidence level of each driving state category can be weighted and summed with the second confidence level of each driving state category to obtain the corrected confidence level corresponding to each driving state category.
[0174] Taking a sliding time window with a width of 3, a first preset number of 1, and a second preset number of 1 as an example, the method for obtaining the corrected confidence level can be expressed by the formula:
[0175] .
[0176] in, For driving state classification models When classifying driving-related data at any given time (i.e., driving-related data for the current frame), the first... Confidence level for each driving state category For driving state classification models When classifying driving-related data at any given time, the first... Confidence level for each driving state category For driving state classification models When classifying driving-related data at any given time, the first... Confidence level for each driving state category The weighting coefficients represent the first confidence level of the driving-related data in the current frame. The weighting coefficient of the average second confidence level of driving-related data in adjacent frames.
[0177] Among them, the weighting coefficient The specific value can be determined by assessing the importance of driving-related data in the current frame compared to driving-related data in adjacent frames. Assuming that driving-related data in the current frame and driving-related data in adjacent frames are considered equally important, the weighting coefficient... The value can be 0.5. If the driving-related data in the current frame is more important than the driving-related data in adjacent frames, then the weighting coefficient... The value can be set to a value greater than 0.5, but is not limited to this, and can be determined based on actual technical needs.
[0178] Based on the embodiments described above, detailed examples are provided below.
[0179] The driving state recognition method of this application embodiment can be applied to scenarios where the driver's distracted driving behavior and the type of distraction are identified.
[0180] refer to Figure 7 As shown, the system architecture of the driving state recognition method in this application embodiment can include six architecture modules: a multimodal perception acquisition module 701, a self-supervised representation learning module 702, a posture guidance enhancement module 703, a semantic quality control module 704, a temporal confidence learning module 705, and a multimodal fusion module 706. These six modules together form a closed-loop system of data acquisition, self-learning, enhancement, optimization, and recognition.
[0181] The multimodal perception and acquisition module 701 collects multi-source driving state-related data through multiple perspectives of onboard cameras (e.g., dashboard view, rearview mirror view, right window view), microphones, and the vehicle's CAN bus. The acquired driving state-related data is synchronized using a unified timestamp to form an input data stream. .in These are multi-view video frames, i.e., image data from multiple perspectives. It is an audio signal. This is vehicle status data.
[0182] The image data obtained from the dashboard view captures the driver's frontal facial features, the image data obtained from the rearview mirror view covers changes in the driver's head posture, and the image data obtained from the right window view can identify the driver's side profile and gaze deviation. After the three views are fused, even if the driver's head turns to one side, monitoring can continue through other perspectives, ensuring continuous and effective detection.
[0183] The driving state data collected by the multimodal perception acquisition module 701 can be input into the self-supervised representation learning module 802 for representation learning.
[0184] The self-supervised representation learning module 702 performs self-supervised training on the image encoder (e.g., a lightweight Swing Transformer) based on the multi-view image data in the driving state data collected by the multi-modal perception acquisition module 701.
[0185] In a specific example, the self-supervised representation learning module 702 reconstructs the occluded regions by randomly occluding image patches of the input multi-view image data, extracting global and local semantic information. The self-supervised representation learning module 802 performs self-supervised training with the goal of minimizing reconstruction error. The image encoder obtained after training can be used for subsequent extraction of visual encoding vectors from images. Simultaneously, for input images from different viewpoints, weight coefficients for input images at different angles can be trained and obtained, which can be used in subsequent practical applications to perform a weighted summation of image feature vectors from different angles when obtaining visual encoding vectors.
[0186] The posture guidance enhancement module 703 uses the DWpose algorithm and other methods to extract driver posture key points (such as the posture information of face, shoulders, abdomen, hands and feet) from the input image. Based on the image encoder trained by the self-supervised representation learning module 702, it extracts the image feature vector of the input image. Then, using models such as Progressive Conditional Diffusion Models (PCDMs), it generates a new image (i.e., a posture-enhanced image) with consistent posture but diverse semantics based on the extracted driver posture key points and image feature vectors.
[0187] For the pose enhancement images generated by the pose guidance enhancement module 703, the semantic quality control module 704 automatically filters and generates labels for the pose enhancement images. The semantic quality control module 704 can evaluate the semantic matching degree between the generated pose enhancement images and the sample labels of the corresponding input images using a visual language model (e.g., the CogVLM model). If the matching degree is greater than or equal to the matching degree threshold, the pose enhancement image is retained, and the sample labels of the corresponding input images are determined as the labels of the pose enhancement images. If the matching degree is less than the matching degree threshold, the pose enhancement image is discarded, thereby realizing automatic sample filtering and label generation.
[0188] Accordingly, the pose guidance enhancement module 703 and the semantic quality control module 704 work together to achieve automatic generation, filtering, and label generation of sample images. Specifically, refer to... Figure 8 As shown:
[0189] In step S801, the pose guidance enhancement module 803 extracts the image feature vector of the input image based on the image encoder trained by the self-supervised representation learning module 702;
[0190] In step S802, the posture guidance enhancement module 703 uses the DWpose algorithm and other methods to extract key points of the driver's posture in the input image (such as the posture information of the face, shoulders, abdomen, hands and feet).
[0191] In step S803, the attitude guidance enhancement module 703 uses models such as Progressive Conditional Diffusion Models (PCDMs) to generate attitude enhancement images that are consistent in attitude but diverse in semantics based on the extracted driver attitude key points and image feature vectors.
[0192] In step S804, the semantic quality control module 704 can evaluate the matching degree of the generated pose-enhanced image with the label semantics of the sample labels of the corresponding sample images in the initial sample through a visual language model (e.g., the CogVLM model).
[0193] In step S805, the semantic quality control module 704 compares the obtained matching degree with the matching degree threshold. If the matching degree is greater than or equal to the matching degree threshold, then proceed to step S806; otherwise, proceed to step S807.
[0194] In step S806, the semantic quality control module 704 retains the pose enhancement image and determines the sample label of the corresponding sample image as the label of the pose enhancement image;
[0195] In step S807, the semantic quality control module 704 discards the pose enhancement image.
[0196] Based on the multiple sets of input data streams and corresponding driving classification labels obtained by the multimodal perception acquisition module 701, an initial sample set can be obtained. The pose enhancement images selected by the semantic quality control module 704 can then be used to obtain enhanced images. Based on these enhanced images, an enhanced sample set can be obtained. Therefore, combining the initial sample set obtained by the multimodal perception acquisition module 701 with the pose enhancement images selected by the semantic quality control module 704 forms a training sample set. The driving state classification model can then be trained using this training sample set.
[0197] In the process of training the driving state classification model and in the process of recognizing the driving state in the actual application scenario after training, the multimodal fusion module 706 can extract and fuse the features of the multimodal information of the acquired driving state data to obtain the multimodal interaction feature vector, and perform driving state classification processing on the multimodal interaction feature vector to obtain the confidence level corresponding to each driving state category.
[0198] Specifically, in combination Figure 9 As shown:
[0199] For the acquired multi-view (e.g., dashboard view, rearview mirror view, right window view) video frame data (i.e., image data from multiple different angles), the multimodal fusion module 706 uses the image encoder 901 trained by the self-supervised representation learning module 702 to extract image features from the image data of different viewpoints, obtaining image feature vectors for different viewpoints. Taking the multi-view including dashboard view, rearview mirror view, and right window view as an example, image feature vectors for three viewpoints can be generated. , , ,in, , , These represent image feature vectors from the dashboard view, rearview mirror view, and right window view, respectively. The image encoder 901 then performs a weighted summation of these feature vectors from different viewpoints to obtain a comprehensive visual feature (i.e., a visual encoding vector). .
[0200] For the collected multi-view video frame data (e.g., dashboard view, rearview mirror view, right window view), the multimodal fusion module 706 extracts the driver's two-dimensional key point coordinates from the video frame data through the attitude detection module 902. For example, the DWpose algorithm is used to extract the driver's attitude key point coordinates, and the extracted two-dimensional key point coordinates are converted into feature vectors. And through lightweight MLP models, etc., the feature vectors Encode the vector to obtain the pose feature vector. .
[0201] For the acquired audio signal, the multimodal fusion module 706 extracts the spectral features of the audio signal through Mel-frequency cepstral coefficients (MFCC), and encodes the spectral features through the audio encoder 903 to obtain the corresponding audio feature vector. .
[0202] For the acquired vehicle state data, the multimodal fusion module 706 normalizes and segments the raw data into temporal windows, and then encodes the normalized and temporal windowed data using a lightweight MLP encoder 904 to obtain the vehicle state feature vector. .
[0203] Subsequently, the multimodal fusion module 706 learns the dependencies between different modalities through the multi-head cross-modal attention (CMA) module 905 and performs feature fusion to obtain a multimodal interaction feature vector. Specifically:
[0204] The multimodal fusion module 706 performs linear mapping on each feature to obtain mapped feature vectors (including mapped visual encoding vectors, mapped posture feature vectors, mapped audio feature vectors, and mapped vehicle state feature vectors). Then, it forms multimodal interaction features for different mapped feature vectors, performs weighted summation on multiple multimodal interaction features to obtain multimodal interaction feature vectors, and uses classifier 906 to classify the multimodal interaction feature vectors to obtain the confidence scores of each driving state category.
[0205] The temporal confidence learning module 705 corrects the confidence of the driving state data of each frame by sliding time window to obtain corrected confidence. The driving state category corresponding to the highest corrected execution degree in the corrected confidence is determined as the recognized driving state category, so as to significantly reduce instantaneous misjudgment and improve the stability of inter-frame prediction.
[0206] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0207] Based on the same inventive concept, this application also provides a driving state recognition device for implementing the driving state recognition method described above. The solution provided by this device is similar to the implementation described in the above method; therefore, the specific limitations in one or more driving state recognition device embodiments provided below can be found in the limitations of the driving state recognition method described above, and will not be repeated here.
[0208] In one embodiment, reference Figure 10 As shown, a driving state recognition device is provided, comprising:
[0209] The data acquisition module 1001 is used to acquire driving-related data collected at the same time during the vehicle driving process. The driving-related data includes: audio signals, vehicle status data, and multiple image data obtained from image acquisition of the target object from multiple perspectives.
[0210] The visual feature extraction module 1002 is used to perform image encoding on multiple image data to obtain a visual encoding vector;
[0211] The posture feature extraction module 1003 is used to extract driver posture information from multiple image data and obtain posture feature vectors based on driver posture information.
[0212] The audio feature extraction module 1004 is used to extract features from the audio signal and obtain an audio feature vector.
[0213] The vehicle state feature extraction module 1005 is used to perform feature vector generation processing on vehicle state data to obtain vehicle state feature vectors.
[0214] The feature fusion module 1006 is used to perform feature fusion processing on the visual encoding vector, posture feature vector, visual encoding vector and vehicle state feature vector to obtain multimodal interaction feature vector;
[0215] The classification and recognition module 1007 is used to classify the driving state of the multimodal interaction feature vector to obtain the driving state category.
[0216] In some embodiments, the visual feature extraction module 1002 is used to use a pre-trained image encoder to encode multiple image data respectively to obtain multiple image feature vectors, and to perform weighted summation processing on the image feature vectors corresponding to each angle based on the weight coefficients of the images at each angle in the image encoder to obtain the visual encoding vector.
[0217] The device also includes: an image encoder training module for acquiring sample input images, which include sample images from multiple angles; performing self-supervised learning on the encoder to be trained by randomly occluding sample images and reconstructing the occluded areas, and obtaining the pre-trained image encoder after the self-supervised learning is completed.
[0218] In some embodiments, the feature fusion module 1006 is used to perform linear mapping on the visual encoding vector, posture feature vector, audio feature vector, and vehicle state feature vector respectively to obtain the mapped visual encoding vector, mapped posture feature vector, mapped audio feature vector, and mapped vehicle state feature vector; perform cross-modal attention calculation on the mapped visual encoding vector, mapped posture feature vector, mapped audio feature vector, and mapped vehicle state feature vector to obtain multiple multimodal interaction features; and perform weighted summation on each multimodal interaction feature to obtain a multimodal interaction feature vector.
[0219] In some embodiments, the classification and recognition module 1007 is used to perform driving state classification processing on the multimodal interaction feature vector using a trained driving state classification model to obtain the driving state category.
[0220] The device also includes: a sample set acquisition module, used to acquire an initial sample set, each initial sample in the initial sample set including sample driving-related data, including sample audio signals, sample vehicle state data, and sample image data acquired from multiple angles of the object; extracting driver posture key points from the sample image data; using an encoder to perform image encoding on the sample image data to obtain sample image feature vectors; generating posture enhancement images based on driver posture key points and sample image feature vectors; and obtaining a training sample set based on the initial sample set and posture enhancement images, the training sample set being used to train and obtain a driving state classification model.
[0221] In some embodiments, the sample set acquisition module is configured to recognize the pose enhancement image through a visual language model, obtain the image semantics of the pose enhancement image, and determine the matching degree between the image semantics and the label semantics of the sample labels of the initial samples; if the matching degree is greater than or equal to the matching degree threshold, the pose enhancement image is retained, and if the matching degree is less than the matching degree threshold, the pose enhancement image is discarded; and a training sample set is obtained based on the initial sample set and the retained pose enhancement images.
[0222] In some embodiments, the classification and recognition module 1007 is used to perform driving state classification processing on the multimodal interaction feature vector according to the obtained driving state classification model to obtain the driving state recognition result of the current frame; to obtain the driving state recognition result of the adjacent frame by performing driving state classification processing on the driving-related data of the adjacent frame according to the driving state classification model; to correct the driving state recognition result of the current frame based on the driving state recognition result of the adjacent frame to obtain the corrected driving state recognition result, and to determine the driving state category based on the corrected driving state recognition result.
[0223] In some embodiments, the current frame driving state recognition result includes: the first confidence level of each driving state category corresponding to the multimodal interaction feature vector; the adjacent frame driving state recognition result includes: the second confidence level of each driving state category corresponding to the driving-related data of the adjacent frames.
[0224] The classification and recognition module 1007 is used to correct the first confidence of each driving state category based on the second confidence of each driving state category, so as to obtain the corrected confidence of each driving state category; and to determine the driving state category corresponding to the highest confidence among the corrected confidence of each driving state category as the recognized driving state category.
[0225] In some embodiments, the classification and recognition module 1007 is used to perform a weighted summation of the first confidence level of each driving state category and the second confidence level of each driving state category to obtain the corrected confidence level corresponding to each driving state category.
[0226] Each module in the aforementioned driving state recognition device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the corresponding operations of each module.
[0227] In one embodiment, an electronic device is provided, which may be a computer device such as a terminal, including a processor, memory, communication interface, display screen, and input device connected via a system bus. The processor of the computer device provides computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The communication interface of the computer device is used for wired or wireless communication with an external terminal; wireless communication can be achieved through Wi-Fi, mobile cellular networks, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements a driving state recognition method. The display screen of the computer device may be an LCD screen or an e-ink display screen. The input device of the computer device may be a touch layer covering the display screen, or buttons, a trackball, or a touchpad located on the casing of the computer device, or an external keyboard, touchpad, or mouse, etc.
[0228] Those skilled in the art will understand that the structure of the above-described electronic device is only a partial structure related to the solution of this application and does not constitute a limitation on the electronic device to which the solution of this application is applied. The specific electronic device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.
[0229] In some embodiments, an electronic device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the driving state recognition method in any of the above embodiments.
[0230] In some embodiments, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps of the driving state recognition method in any of the above embodiments.
[0231] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps of the driving state recognition method in any of the above embodiments.
[0232] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments described above. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.
[0233] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0234] The above embodiments are merely illustrative of several implementation methods of this application, and their descriptions are relatively specific and detailed. However, they should not be construed as limiting the scope of this application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A driving state recognition method, characterized in that, The method includes: Acquire driving-related data collected at the same moment during vehicle driving, including: audio signals, vehicle status data, and multiple image data obtained from image acquisition of the target object from multiple perspectives; Image encoding is performed on multiple image data to obtain a visual encoding vector; Driver posture information is extracted from multiple image data sets, and a posture feature vector is obtained based on the driver posture information; The audio signal is subjected to feature extraction to obtain an audio feature vector; The vehicle state data is processed to generate a feature vector to obtain a vehicle state feature vector. The visual encoding vector, the posture feature vector, the audio feature vector, and the vehicle state feature vector are subjected to feature fusion processing to obtain a multimodal interaction feature vector; The multimodal interaction feature vector is subjected to driving state classification processing to obtain the driving state category.
2. The method according to claim 1, characterized in that, The step of encoding the multiple image data to obtain a visual encoding vector includes: A pre-trained image encoder is used to encode multiple image data separately to obtain multiple image feature vectors. Based on the weight coefficients of the images at each angle in the image encoder, the image feature vectors corresponding to each angle are weighted and summed to obtain the visual encoding vector. The methods for training the image encoder include: Acquire a sample input image, which includes sample images from multiple angles; The image encoder to be trained is self-supervised by randomly occluding the sample image and reconstructing the occluded area, and the image encoder is obtained after the self-supervised learning is completed.
3. The method according to claim 1, characterized in that, The step of performing feature fusion processing on the visual encoding vector, the pose feature vector, the audio feature vector, and the vehicle state feature vector to obtain a multimodal interaction feature vector includes: Linear mapping is performed on the visual encoding vector, the posture feature vector, the audio feature vector, and the vehicle state feature vector to obtain the mapped visual encoding vector, the mapped posture feature vector, the mapped audio feature vector, and the mapped vehicle state feature vector, respectively. Cross-modal attention calculation is performed on the mapped visual encoding vector, the mapped posture feature vector, the mapped audio feature vector, and the mapped vehicle state feature vector to obtain multiple modal interaction features; The multimodal interaction features are weighted and summed to obtain a multimodal interaction feature vector.
4. The method according to claim 1, characterized in that, The step of performing driving state classification processing on the multimodal interaction feature vector to obtain driving state categories includes: The driving-related data is processed using a driving state classification model obtained through training to obtain the multimodal interaction feature vector, and the driving state is classified using the multimodal interaction feature vector to obtain the driving state category. The methods for obtaining the training sample set for training the driving state classification model include: An initial sample set is obtained, and each initial sample in the initial sample set includes sample driving-related data, which includes: sample audio signals, sample vehicle status data, and sample image data obtained from multiple angles of the object. Extract key points of the driver's posture from the sample image data; The sample image data is encoded using an image encoder to obtain the sample image feature vector; Based on the driver's posture key points and the feature vector of the sample image, an posture enhancement image is generated; A training sample set is obtained based on the initial sample set and the pose enhancement image.
5. The method according to claim 4, characterized in that, The process of obtaining a training sample set based on the initial sample set and the pose-enhanced image includes: The pose-enhanced image is identified using a visual language model to obtain its image semantics, and the matching degree between the image semantics and the label semantics of the sample labels of the initial sample is determined. If the matching degree is greater than or equal to the matching degree threshold, the pose enhancement image is retained; if the matching degree is less than the matching degree threshold, the pose enhancement image is discarded. A training sample set is obtained based on the initial sample set and the retained pose-enhanced images.
6. The method according to claim 1, characterized in that, The step of performing driving state classification processing on the multimodal interaction feature vector to obtain driving state categories includes: Based on the driving state classification model obtained through training, the multimodal interaction feature vector is subjected to driving state classification processing to obtain the driving state recognition result of the current frame; Obtain the driving state identification results of adjacent frames by performing driving state classification processing on the driving-related data of adjacent frames according to the driving state classification model; The driving state recognition result of the current frame is corrected based on the driving state recognition result of the adjacent frames to obtain the corrected driving state recognition result, and the driving state category is determined based on the corrected driving state recognition result.
7. The method according to claim 6, characterized in that, The current frame driving state recognition result includes: the first confidence level of each driving state category corresponding to the multimodal interaction feature vector; the adjacent frame driving state recognition result includes: the second confidence level of each driving state category corresponding to the driving-related data of the adjacent frames. The step of correcting the current frame's driving state recognition result based on the driving state recognition results of adjacent frames to obtain a corrected driving state recognition result, and determining the driving state category based on the corrected driving state recognition result, includes: Based on the second confidence level of each driving state category, the first confidence level of each driving state category is corrected to obtain the corrected confidence level corresponding to each driving state category; The driving state category corresponding to the highest confidence among the corrected confidence scores for each driving state category is determined as the identified driving state category.
8. The method according to claim 7, characterized in that, The step of correcting the first confidence level of each driving state category based on the second confidence level of each driving state category to obtain the corrected confidence level corresponding to each driving state category includes: The first confidence level of each driving state category is weighted and summed with the second confidence level of each driving state category to obtain the corrected confidence level corresponding to each driving state category.
9. An electronic device comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 8.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 8.