A multimodal transformer-based driver assistance prompting method and device

Through the assisted driving prompt method based on multimodal transformer, the driver's emotional state and external environment information are integrated, and the problem of accurate driving prompts cannot be provided in the prior art is solved, thereby achieving safer and more personalized driving assistance.

CN119659663BActive Publication Date: 2025-05-16GUANGDONG UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510199795.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-05-16
Estimated Expiration
2045-02-24

AI Technical Summary

Technical Problem

The existing assisted driving prompt methods cannot effectively integrate the driver's emotional state and external environment information, resulting in the inability to provide accurate driving prompts.

Method used

The assisted driving prompt method based on multimodal transformer is adopted. By obtaining facial expression images and road conditions images, the Transformer self-attention module and cross attention module are used to generate assisted driving prompt results.

Benefits of technology

It achieves a comprehensive understanding and prediction of the driver's emotional state and road conditions, provides more accurate and timely driving tips, and ensures driving safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119659663B_ABST
    Figure CN119659663B_ABST
Patent Text Reader

Abstract

The present invention discloses a multimodal transformer-based assisted driving prompt method and device, which relates to the field of automobile driving technology and is used to solve the technical problem that the existing assisted driving prompt method cannot provide accurate driving prompts. The method includes using a multimodal transformer-based assisted driving prompt model to process the acquired multimodal images, namely, facial expression images and road condition images, and outputting assisted driving prompt results; wherein the multimodal transformer-based assisted driving prompt model includes an input module, a transformer self-attention module, a cross-attention module and a classification module.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of automobile driving technology, and in particular to a multimodal transformer-based auxiliary driving prompting method and device. Background Art

[0002] In recent years, with the rapid development of autonomous driving technology, people are paying more and more attention to the intelligent driving assistance systems of vehicles. Such systems are designed to improve driving safety and reduce the occurrence of traffic accidents. Although existing driving assistance systems can process certain road condition information, such as lane keeping and automatic braking, most of them ignore the important impact of the driver's emotional state on driving safety.

[0003] The driver's emotional state, such as tension, anger or distraction, can significantly increase the risk of traffic accidents. In addition, current systems are generally inefficient in processing multimodal data, especially when considering the driver's behavior and emotions in combination with external environmental information. Therefore, there is an urgent need for an intelligent system that can fully understand and predict the driving environment and driver status to provide more accurate and timely driving tips to ensure driving safety.

[0004] Existing assisted driving prompt methods still focus on the information processing of a single data source or a single modality, and have not yet been able to effectively integrate the driver's emotional state, facial expressions and other data to provide comprehensive driving assistance, resulting in the inability to provide accurate driving prompts. Summary of the invention

[0005] The present invention provides a multimodal transformer-based assisted driving prompt method and device, which are used to solve the technical problem that the existing assisted driving prompt method cannot provide accurate driving prompts.

[0006] The first aspect of the present invention provides an assisted driving prompt method based on a multimodal transformer, comprising:

[0007] Acquire a facial expression image and a road condition image, and input the facial expression image and the road condition image into a multimodal transformer-based assisted driving prompt model; the multimodal transformer-based assisted driving prompt model includes an input module, a Transformer self-attention module, a cross-attention module, and a classification module;

[0008] Using the input module to perform feature vectorization on the facial expression image and the road condition image respectively, to generate a plurality of expression target embeddings corresponding to the facial expression image and a plurality of road condition target embeddings corresponding to the road condition image;

[0009] Through the Transformer self-attention module, each of the expression target embeddings and each of the road condition target embeddings are respectively processed by a multi-head attention block, and the expression multi-head feature representation corresponding to each of the expression target embeddings and the road condition multi-head feature representation corresponding to each of the road condition target embeddings are output;

[0010] Taking each of the multi-head feature representations of expression and each of the multi-head feature representations of road conditions as inputs of the cross attention module, and outputting the cross feature representations of expression corresponding to each of the multi-head feature representations of expression and the cross feature representations of road conditions corresponding to each of the multi-head feature representations of road conditions;

[0011] The classification module is used to perform classification according to each of the expression cross feature representations, each of the expression multi-head feature representations, each of the road condition multi-head feature representations, and each of the road condition cross feature representations to generate an assisted driving prompt result.

[0012] Optionally, the adopting the input module to perform feature vectorization on the facial expression image and the road condition image respectively to generate a plurality of expression target embeddings corresponding to the facial expression image and a plurality of road condition target embeddings corresponding to the road condition image comprises:

[0013] Dividing the facial expression image and the road condition image respectively to generate a plurality of expression image blocks corresponding to the facial expression image and a plurality of road condition image blocks corresponding to the road condition image;

[0014] Flattening each of the expression image blocks and each of the road condition image blocks respectively, and outputting a one-dimensional expression vector corresponding to each of the expression image blocks and a one-dimensional road condition vector corresponding to each of the road condition image blocks;

[0015] Performing linear transformation on each of the one-dimensional expression vectors and each of the one-dimensional road condition vectors, respectively, and outputting expression mark embeddings corresponding to each of the one-dimensional expression vectors and road condition mark embeddings corresponding to each of the one-dimensional road condition vectors;

[0016] Adding each of the expression marker embeddings to the preset position embeddings corresponding to each of the expression marker embeddings, and outputting the expression target embeddings corresponding to each of the expression marker embeddings;

[0017] Each of the road condition mark embeddings is added to the preset position embedding corresponding to each of the road condition mark embeddings, and the road condition target embedding corresponding to each of the road condition mark embeddings is output.

[0018] Optionally, the Transformer self-attention module includes a plurality of Transformer autoencoders; the data processing process of the Transformer autoencoder is specifically as follows:

[0019] Performing a multi-head linear transformation on a plurality of first input vectors and a plurality of second input vectors input to the Transformer autoencoder to generate a plurality of query matrices, a plurality of key matrices, and a plurality of value matrices corresponding to each of the first input vectors, and a plurality of query matrices, a plurality of key matrices, and a plurality of value matrices corresponding to each of the second input vectors;

[0020] Calculating a plurality of first output self-attention heads corresponding to each of the first input vectors based on a plurality of query matrices, a plurality of key matrices, and a plurality of value matrices corresponding to each of the first input vectors;

[0021] Calculate a plurality of second output self-attention heads corresponding to each of the second input vectors based on a plurality of query matrices, a plurality of key matrices, and a plurality of value matrices corresponding to each of the second input vectors;

[0022] Respectively concatenate and linearly transform the multiple first output self-attention heads corresponding to each of the first input vectors and the multiple second output self-attention heads corresponding to each of the second input vectors, and output a first multi-head attention corresponding to each of the first input vectors and a second multi-head attention corresponding to each of the second input vectors;

[0023] Adding each of the first input vectors and the first multi-head attention corresponding to each of the first input vectors to generate a plurality of first residual features;

[0024] Adding each of the second input vectors and the second multi-head attention corresponding to each of the second input vectors to generate a plurality of second residual features;

[0025] Performing layer normalization on each of the first residual features and each of the second residual features respectively to generate a first normalized feature corresponding to each of the first residual features and a second normalized feature corresponding to each of the second residual features;

[0026] Using each of the first normalized features and each of the second normalized features as input of a feedforward network, and outputting a first target feature corresponding to each of the first normalized features and a second target feature corresponding to each of the second normalized features;

[0027] Adding each of the first normalized features and the first target features corresponding to each of the first normalized features to generate a plurality of first comprehensive features;

[0028] Adding each of the second normalized features and the second target features corresponding to each of the second normalized features to generate a plurality of second comprehensive features;

[0029] Layer normalization is performed on each of the first comprehensive features and each of the second comprehensive features to generate a first multi-head feature representation corresponding to each of the first comprehensive features and a second multi-head feature representation corresponding to each of the second comprehensive features.

[0030] Optionally, the cross attention module includes a first cross attention layer, a second cross attention layer, a third cross attention layer, and a fourth cross attention layer; the method of taking each of the expression multi-head feature representations and each of the road condition multi-head feature representations as inputs of the cross attention module, and outputting expression cross feature representations corresponding to each of the expression multi-head feature representations and road condition cross feature representations corresponding to each of the road condition multi-head feature representations, comprises:

[0031] Performing multi-head linear projection on each of the multi-head feature representations of the expression respectively to generate a plurality of expression query matrices, a plurality of expression key matrices and a plurality of expression value matrices corresponding to each of the multi-head feature representations of the expression;

[0032] Performing multi-head linear projection on each of the multi-head feature representations of the road condition to generate a plurality of road condition query matrices, a plurality of road condition key matrices and a plurality of road condition value matrices corresponding to each of the multi-head feature representations of the road condition;

[0033] Using a first cross attention layer to perform cross attention head calculation according to each of the expression key matrices, each of the expression value matrices, and each of the road condition query matrices, to determine a plurality of first expression cross multi-head attentions;

[0034] Performing cross-attention head calculation according to each of the expression query matrices, each of the road condition key matrices, and each of the road condition value matrices through a second cross-attention layer to determine a plurality of first road condition cross-multi-head attentions;

[0035] Respectively perform residual feedforward network block processing on each of the expression multi-head feature representations, each of the first expression cross-multi-head attentions, each of the road condition multi-head feature representations, and each of the first road condition cross-multi-head attentions to generate multiple initial expression feature representations and multiple initial road condition feature representations;

[0036] Performing linear projection on each of the initial expression feature representations and each of the initial road condition feature representations, respectively, and outputting an expression cross-query matrix corresponding to each of the initial expression feature representations and a road condition cross-query matrix corresponding to each of the initial road condition feature representations;

[0037] Using a third cross attention layer to perform cross attention head calculation according to each of the expression key matrices, each of the expression value matrices, and each of the road condition cross query matrices, to determine a plurality of second expression cross multi-head attentions;

[0038] Performing cross-attention head calculation according to each of the road condition key matrices, each of the road condition value matrices, and each of the expression cross-query matrices through a fourth cross-attention layer to determine a plurality of second road condition cross-multi-head attentions;

[0039] Each of the initial expression feature representations, each of the second expression cross-multi-head attentions, each of the initial road condition feature representations, and each of the second road condition cross-multi-head attentions is processed by a residual feedforward network block to generate multiple expression cross-feature representations and multiple road condition cross-feature representations.

[0040] Optionally, the performing residual feedforward network block processing on each of the initial expression feature representations, each of the second expression cross multi-head attentions, each of the initial road condition feature representations, and each of the second road condition cross multi-head attentions to generate multiple expression cross feature representations and multiple road condition cross feature representations includes:

[0041] Adding each of the initial expression feature representations and the second expression cross multi-head attention corresponding to each of the initial expression feature representations to generate a plurality of expression cross residual features;

[0042] Adding each of the initial road condition feature representations and the second road condition cross multi-head attention corresponding to each of the initial road condition feature representations to generate a plurality of road condition cross residual features;

[0043] Respectively performing layer normalization on each of the expression cross residual features and each of the road condition cross residual features to generate expression cross normalized features corresponding to each of the expression cross residual features and road condition cross normalized features corresponding to each of the road condition cross residual features;

[0044] Using each of the expression cross-normalized features and each of the road condition cross-normalized features as inputs of a feedforward network, and outputting expression cross-target features corresponding to each of the expression cross-normalized features and road condition cross-target features corresponding to each of the road condition cross-normalized features;

[0045] Adding each of the expression cross normalization features and the expression cross target features corresponding to each of the expression cross normalization features to generate a plurality of expression cross comprehensive features;

[0046] Adding each of the road condition cross normalization features and the road condition cross target features corresponding to each of the road condition cross normalization features to generate a plurality of road condition cross comprehensive features;

[0047] Each of the expression cross-comprehensive features and each of the road condition cross-comprehensive features is layer-normalized to generate an expression cross-feature representation corresponding to each of the expression cross-comprehensive features and a road condition cross-feature representation corresponding to each of the road condition cross-comprehensive features.

[0048] Optionally, the classification module includes a fully connected layer and a Softmax activation function layer; the classification module is used to perform classification according to each of the expression cross feature representations, each of the expression multi-head feature representations, each of the road condition multi-head feature representations, and each of the road condition cross feature representations to generate an assisted driving prompt result, including:

[0049] Adding each of the expression cross feature representations and each of the expression multi-head feature representations to generate a total input of facial expression data;

[0050] Adding each of the road condition multi-feature representations and each of the road condition cross-feature representations to generate a total input of road condition information data;

[0051] Concatenate the total facial expression data input and the total road condition information data input to determine an emotion classification input vector;

[0052] Using the emotion classification input vector as the input of the fully connected layer, and outputting the emotion fully connected feature vector;

[0053] A Softmax activation function layer is used to perform nonlinear mapping on the emotion fully connected feature vector to generate an emotion classification result;

[0054] Taking each of the multi-head feature representations of the road condition as the input of the fully connected layer, and outputting a road condition fully connected feature vector;

[0055] A Softmax activation function layer is used to perform nonlinear mapping on the road condition fully connected feature vector to generate a road condition classification result;

[0056] An assisted driving prompt result is generated according to the emotion classification result and the road condition classification result.

[0057] A second aspect of the present invention provides an auxiliary driving prompt device based on a multimodal transformer, comprising:

[0058] An acquisition module is used to acquire a facial expression image and a road condition image, and input the facial expression image and the road condition image into a multimodal transformer-based assisted driving prompt model; the multimodal transformer-based assisted driving prompt model includes an input module, a Transformer self-attention module, a cross-attention module and a classification module;

[0059] An adopting module, used for adopting the input module to perform feature vectorization on the facial expression image and the road condition image respectively, to generate a plurality of expression target embeddings corresponding to the facial expression image and a plurality of road condition target embeddings corresponding to the road condition image;

[0060] A multi-head attention processing module, used to perform multi-head attention block processing on each of the expression target embeddings and each of the road condition target embeddings through the Transformer self-attention module, and output expression multi-head feature representations corresponding to each of the expression target embeddings and road condition multi-head feature representations corresponding to each of the road condition target embeddings;

[0061] A cross-attention processing module, used to take each of the expression multi-head feature representations and each of the road condition multi-head feature representations as inputs of the cross-attention module, and output an expression cross-feature representation corresponding to each of the expression multi-head feature representations and a road condition cross-feature representation corresponding to each of the road condition multi-head feature representations;

[0062] A classification module is used to use the classification module to perform classification according to each of the above-mentioned expression cross-feature representations, each of the above-mentioned expression multi-head feature representations, each of the above-mentioned road condition multi-head feature representations, and each of the above-mentioned road condition cross-feature representations to generate an assisted driving prompt result.

[0063] A third aspect of the present invention provides a computer device, comprising a memory and a processor, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, the processor executes the steps of the multimodal transformer-based assisted driving prompt method as described in any one of the above items.

[0064] A fourth aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed, the steps of the multimodal transformer-based assisted driving prompt method as described in any one of the above items are implemented.

[0065] A fifth aspect of the present invention provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium, and the computer program includes program instructions, wherein when the program instructions are executed by a computer, the computer executes the steps of the multimodal transformer-based assisted driving prompt method as described in any one of the above items.

[0066] It can be seen from the above technical solutions that the present invention has the following advantages:

[0067] The above technical scheme of the present invention provides an assisted driving prompt method based on a multimodal transformer. First, a facial expression image and a road condition image are obtained, and the facial expression image and the road condition image are input into an assisted driving prompt model based on a multimodal transformer; the assisted driving prompt model based on a multimodal transformer includes an input module, a Transformer self-attention module, a cross-attention module and a classification module; then, the input module is used to perform feature vectorization on the facial expression image and the road condition image respectively, and a plurality of expression target embeddings corresponding to the facial expression image and a plurality of road condition target embeddings corresponding to the road condition image are generated; the Transformer self-attention module is used to perform multi-head attention block processing on each of the expression target embeddings and each of the road condition target embeddings, and the expression multi-head feature representation corresponding to each of the expression target embeddings, each The road condition target is embedded in the corresponding road condition multi-head feature representation; each of the expression multi-head feature representations and each of the road condition multi-head feature representations is used as the input of the cross-attention module, and the expression cross-feature representation corresponding to each of the expression multi-head feature representations and the road condition cross-feature representation corresponding to each of the road condition multi-head feature representations are output; finally, the classification module is used to classify each of the expression cross-feature representations, each of the expression multi-head feature representations, each of the road condition multi-head feature representations, and each of the road condition cross-feature representations to generate an assisted driving prompt result; based on the above scheme, the present invention uses an assisted driving prompt model based on a multimodal transformer to process the acquired multimodal images, namely facial expression images and road condition images, and outputs the assisted driving prompt result. The process comprehensively considers the driver's emotional state and road conditions, and can provide appropriate assistance and intervention according to the specific emotional state of each driver, thereby providing more accurate driving prompts. BRIEF DESCRIPTION OF THE DRAWINGS

[0068] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0069] Figure 1 A flowchart of a multimodal transformer-based assisted driving prompt method provided in Embodiment 1 of the present invention;

[0070] Figure 2 A schematic diagram of the structure of the multimodal transformer-based assisted driving prompt model provided in the first embodiment of the present invention;

[0071] Figure 3 A training framework diagram of the multimodal transformer-based assisted driving prompt model during initial operation provided in the first embodiment of the present invention;

[0072] Figure 4 A schematic diagram of adding a mark embedding and a position embedding provided in the first embodiment of the present invention;

[0073] Figure 5 A schematic diagram of the process flow of Part 1 data processing of multimodal multi-head cross attention provided in Example 1 of the present invention;

[0074] Figure 6 A structural block diagram of a multimodal transformer-based assisted driving prompt device provided in Example 2 of the present invention. DETAILED DESCRIPTION

[0075] The embodiments of the present invention provide a multimodal transformer-based assisted driving prompt method and device, which are used to solve the technical problem that the existing assisted driving prompt method cannot provide accurate driving prompts.

[0076] In order to make the purpose, features and advantages of the present invention more obvious and easy to understand, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described below are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0077] See also Figure 1 , Figure 1 A flowchart of the steps of a multimodal transformer-based assisted driving prompt method provided in Example 1 of the present invention.

[0078] The present invention provides a multimodal transformer-based assisted driving prompt method, comprising:

[0079] Step 101, obtaining facial expression images and road condition images, and inputting the facial expression images and road condition images into a multimodal transformer-based assisted driving prompt model; the multimodal transformer-based assisted driving prompt model includes an input module, a Transformer self-attention module, a cross-attention module and a classification module.

[0080] Please note that Figure 2The multimodal transformer-based assisted driving prompt model proposed in the present invention includes four parts: input module, transformer self-attention module, cross-attention module, and classification module; wherein, model training and practical application include two-modal input and two types of output. The two-modal input includes: facial expression pictures and road condition pictures (i.e., facial expression images and road condition images). The two types of output are: classification of emotions and classification of road conditions (i.e., emotion classification results and road condition classification results).

[0081] The overall data processing structure of the model is as follows: the road condition image is first input into a linear transformation layer, which converts the image into a series of vector representations. The converted vectors are then added with position information to maintain spatial continuity and contextual relationships. The processed vectors are input into the self-attention module of the Transformer, which is specifically processed for the road condition modality. The results are then output to three parts at the same time: the road condition classification part of the classification module, which is directly classified; the second is output to the cross-attention module, which interacts with the data of the facial expression modality to assist in emotion classification; the third is output to the emotion classification part of the classification module, which participates in the residual connection.

[0082] The facial expression images are also converted into multiple vectors through a dedicated linear transformation layer. These vectors are combined with position information to ensure that the model can understand the spatial relationship between the various parts of the image. These vectors are then input into the expression modality part of the self-attention module, and the results are then output to two parts at the same time. One is output to the cross-attention module, which is combined with the data of the road condition modality to participate in the emotion classification process; the other is output to the emotion classification part of the classification module to participate in the residual connection.

[0083] It is worth mentioning that in the initial use training stage of the multimodal transformer-based assisted driving prompt model, that is, when the model is initially running, it may not fully match the behavior patterns and reaction characteristics of individual drivers, resulting in errors in the judgment of the driver's emotions. In order to solve this problem, the present invention collects and stores the driver's behavior data under different driving conditions. Through the long-term accumulation and analysis of these data, it is possible to gradually learn and adapt to each driver's unique behavioral characteristics and emotional changes, thereby significantly improving the accuracy of emotion recognition, and adjusting the response strategy, so that the model is more in line with the actual driving situation.

[0084] Furthermore, in the case of multiple drivers using the same car, personalized models become the key to improving system performance. Storing specific data for each driver not only helps the system distinguish between different drivers, but also through detailed analysis of each data set, it can combine face recognition technology to customize model parameters for each driver. This customization includes adjusting the model's sensitivity, reaction time, and warning mechanism to adapt to the driving styles and habits of different drivers. For example, for experienced drivers, the system may be adjusted to a higher level of automated response, while for novice drivers, safety warnings and auxiliary guidance may be strengthened.

[0085] Through these measures, the driver assistance system can continuously learn and improve from actual applications, improve its performance in complex real-world environments, ensure the optimal combination of safety and personalized services, and thus greatly improve the driving experience and the overall effectiveness of the system. This data-driven approach is the cornerstone of achieving highly personalized and precise services, and is of great significance to improving the driver's driving experience.

[0086] Also, see Figure 3 ,For data storage: The present invention uses a postgresql database to store all collected multimodal data. These data include road images, facial expression images, and their corresponding labels and model output results. All data will be preprocessed and vectorized before being stored in the database. This step is to optimize storage efficiency and accelerate subsequent data retrieval and processing speed.

[0087] For vectorization processing: vectorization is the process of converting raw data into a format that can be directly used for model training. For example, facial expression images are converted into expression feature vectors by extracting key features through deep learning algorithms; road condition images are converted into road condition feature vectors by extracting key features through deep learning algorithms. These vectorized data not only reduce storage space usage, but also improve data processing efficiency.

[0088] For data backup and security: In order to ensure data security and prevent any accidental data loss, the present invention uses miniIO for multiple data backups. miniIO is an efficient object storage solution that supports high data availability and disaster recovery. This backup mechanism ensures that all critical data can be safely stored and quickly restored even under extreme conditions.

[0089] For data used for model training: The vectorized data stored in the PostgreSQL database will be used to continuously train and optimize the personalized Transformer model. By analyzing the long-term accumulated data, the model can learn more deeply about the behavior patterns and reaction characteristics of each driver, thereby customizing more accurate driving assistance strategies for each driver.

[0090] Step 102: Use an input module to perform feature vectorization on the facial expression image and the road condition image respectively, and generate multiple expression target embeddings corresponding to the facial expression image and multiple road condition target embeddings corresponding to the road condition image.

[0091] Specifically, step 102 may include the following sub-steps S21-S25:

[0092] Step S21, dividing the facial expression image and the road condition image respectively to generate a plurality of expression image blocks corresponding to the facial expression image and a plurality of road condition image blocks corresponding to the road condition image;

[0093] Step S22, flattening each expression image block and each road condition image block respectively, and outputting an expression one-dimensional vector corresponding to each expression image block and a road condition one-dimensional vector corresponding to each road condition image block;

[0094] Step S23, linearly transform each expression one-dimensional vector and each road condition one-dimensional vector, and output the expression mark embedding corresponding to each expression one-dimensional vector and the road condition mark embedding corresponding to each road condition one-dimensional vector;

[0095] Step S24, adding each expression mark embedding to the preset position embedding corresponding to each expression mark embedding, and outputting the expression target embedding corresponding to each expression mark embedding;

[0096] Step S25, adding each road condition mark embedding to the preset position embedding corresponding to each road condition mark embedding, and outputting the road condition target embedding corresponding to each road condition mark embedding.

[0097] It should be noted that in the tasks of emotion classification and road condition classification, the input data is in the form of images, so their feature extraction methods are highly similar. First, these images are divided into several fixed-size image blocks (patches), namely multiple expression image blocks and multiple road condition image blocks. Next, each image block is flattened, and the two-dimensional image block is converted into a one-dimensional vector, obtaining multiple expression one-dimensional vectors and multiple road condition one-dimensional vectors.

[0098] Furthermore, each flattened image patch vector (one-dimensional vector) is mapped to a high-dimensional space through a linear transformation. The vector generated by this process is called a token embedding, which represents the feature representation of the image patch. These token embeddings are further processed as input to the Transformer model (Transformer self-attention module) to extract higher-level features.

[0099] It is worth mentioning that see Figure 4After this step, in order to maintain the spatial order of the image blocks, the model will add a position encoding (preset position embedding) to each token embedding. At the same time, on the basis of all token embeddings, an additional special CLS Token (classification tag) is added. This Token does not represent any specific image block, but is used to summarize the representation of all information in subsequent classification tasks. Among them, the role of the position encoding is to provide the model with the relative position of each token in the image so that the model can understand the relationship between image blocks at different positions. For the calculation of each position encoding, the formula is as follows:

[0100] ;

[0101] ;

[0102] in, Encode the values ​​of the positions in the even-numbered dimensions; is the value of the position code on the odd dimension; position is the index of the position code, indicating the position of the current patch in the sequence; t is the index of the current dimension; n is the dimension of the position code vector.

[0103] Based on the above method, the model can identify the relative position relationship of each image block in the image, so as to better understand the spatial structure of the image.

[0104] Step 103: Perform multi-head attention block processing on each expression target embedding and each road condition target embedding through the Transformer self-attention module, and output the expression multi-head feature representation corresponding to each expression target embedding and the road condition multi-head feature representation corresponding to each road condition target embedding.

[0105] The Transformer self-attention module consists of multiple Transformer autoencoders.

[0106] Optionally, the data processing process of the Transformer autoencoder may include the following sub-steps S31-S:

[0107] Step S31, performing a multi-head linear transformation on the multiple first input vectors and the multiple second input vectors input to the Transformer autoencoder to generate multiple query matrices, multiple key matrices and multiple value matrices corresponding to each first input vector, and multiple query matrices, multiple key matrices and multiple value matrices corresponding to each second input vector;

[0108] Step S32, calculating a plurality of first output self-attention heads corresponding to each first input vector based on a plurality of query matrices, a plurality of key matrices, and a plurality of value matrices corresponding to each first input vector;

[0109] Step S33, calculating a plurality of second output self-attention heads corresponding to each second input vector based on a plurality of query matrices, a plurality of key matrices, and a plurality of value matrices corresponding to each second input vector;

[0110] Step S34, respectively concatenate and linearly transform the multiple first output self-attention heads corresponding to each first input vector and the multiple second output self-attention heads corresponding to each second input vector, and output the first multi-head attention heads corresponding to each first input vector and the second multi-head attention heads corresponding to each second input vector;

[0111] Step S35, adding each first input vector and the first multi-head attention corresponding to each first input vector to generate a plurality of first residual features;

[0112] Step S36, adding each second input vector and the second multi-head attention corresponding to each second input vector to generate a plurality of second residual features;

[0113] Step S37, respectively performing layer normalization on each first residual feature and each second residual feature to generate a first normalized feature corresponding to each first residual feature and a second normalized feature corresponding to each second residual feature;

[0114] Step S38: using each first normalized feature and each second normalized feature as input of a feedforward network, and outputting a first target feature corresponding to each first normalized feature and a second target feature corresponding to each second normalized feature;

[0115] Step S39: adding each first normalized feature and the first target feature corresponding to each first normalized feature to generate a plurality of first comprehensive features;

[0116] Step S310: Add each second normalized feature and the second target feature corresponding to each second normalized feature to generate a plurality of second comprehensive features;

[0117] Step S311: perform layer normalization on each first comprehensive feature and each second comprehensive feature respectively to generate a first multi-head feature representation corresponding to each first comprehensive feature and a second multi-head feature representation corresponding to each second comprehensive feature.

[0118] The first input vector and the second input vector are feature data input to the Transformer autoencoder. It can be understood that the first input vector and the second input vector can correspond to any feature data input to the Transformer autoencoder for data processing during model training or model application;

[0119] The data such as the first output self-attention head and the second output self-attention head are the intermediate feature data generated in the Transformer autoencoder;

[0120] The first multi-head feature representation and the second multi-head feature representation are feature data output by the Transformer autoencoder. It can be understood that they can correspond to any feature data output after the Transformer autoencoder performs data processing during model training or model application.

[0121] It should be noted that in the tasks of emotion classification and road condition classification, since the data types of both input data are images, their feature extraction methods are exactly the same, both using a multi-head attention model based on the self-attention mechanism, and each model contains 6 layers of encoders (Transformer autoencoders). In these encoder layers, the input of the first encoder is the embedding (token embeddings) of the image block plus the position encoding (i.e., expression target embedding, road condition target embedding), while the input of the remaining encoders is the output of the previous layer of Transformer (Transformer autoencoder). This design allows the network to gradually extract deep features of the image through a multi-layer structure.

[0122] Specifically, the input vector X is linearly projected and mapped to three different spaces, which are used to generate the query matrix (Query, Q), the key matrix (Key, K), and the value matrix (Value, V). Each matrix is ​​generated by performing matrix multiplication of the input vector with the corresponding learnable weight matrix. Specifically, the query, key, and value matrices are calculated using the following formulas:

[0123] ;

[0124] Among them, Q is the query matrix; K is the key matrix; V is the value matrix; is the learnable weight matrix corresponding to the query matrix; is the learnable weight matrix corresponding to the key matrix; is the learnable weight matrix corresponding to the value matrix.

[0125] Furthermore, the core of the self-attention mechanism is to calculate the similarity between the query (Q) and all keys (K). This similarity is achieved by calculating the dot product of the query and the key. The multi-head self-attention mechanism enhances the expressiveness of the model by fusing multiple self-attention heads together. Specifically, each head has independent linear transformations of queries, keys, and values. For example, the attention weight of the t-th self-attention head is calculated as:

[0126] ;

[0127] in, is the tth self-attention head; is the t-th query matrix; is the transpose of the t-th key matrix; is the scaling factor; is the t-th value matrix; is the similarity between the query and the key, divided by This is to ensure that the gradient remains stable during application. Next, the scores are normalized by the softmax function to obtain the normalized attention weights, which determine the importance of different values ​​(V) in the output.

[0128] Furthermore, after completing the calculation of each self-attention head (output self-attention head) corresponding to each input vector, all self-attention head outputs corresponding to each input vector are concatenated together to form a long vector, and a linear transformation is performed to obtain the final multi-head attention output, that is, the multi-head attention corresponding to each input vector. The formula is as follows:

[0129] ;

[0130] in, is the multi-head attention corresponding to the input vector X; For splicing; is the first self-attention head corresponding to the input vector X; are learnable parameters.

[0131] Furthermore, a residual connection is added to the output of each multi-head self-attention layer (Transformer autoencoder). The role of the residual connection is to add the input X and the output of the multi-head self-attention to ensure that the information is effectively transmitted in the deep network. Then, the result of the residual connection is normalized by layer normalization. The role of layer normalization is to make the mean of each sample 0 and the variance 1, thereby improving the stability of the network, accelerating the training process, and preventing internal covariate shift. This process is expressed by the following formula:

[0132] ;

[0133] Among them, Z is the normalized feature; X is the input vector; is the multi-head attention corresponding to the input vector X; Normalize the layer.

[0134] Furthermore, the next step is to pass through a feed-forward network containing a fully connected layer and an activation function. The network consists of two fully connected layers. The first fully connected layer maps the input to a higher dimension, and after the activation function, it is mapped back to the original dimension through the second fully connected layer. The output of FFN is also processed by residual connection and layer normalization. The formula is as follows:

[0135] ;

[0136] Among them, Output is a multi-head feature representation; FFN is a feed-forward network.

[0137] It is worth mentioning that in the Transformer self-attention module, there are differences in the output processing methods for facial expression data and road condition data to adapt to different application requirements.

[0138] For facial expression data, the model needs to pass the output of the last layer of Transformer self-attention encoder to the emotion classification part in the cross-attention module and the classification module, and the output is only connected to the residual in the classification module. This design enables the model to focus on extracting key features related to expressions and further strengthen the processing of these features through the cross-attention mechanism, thereby improving the accuracy of facial expression recognition.

[0139] For the road condition data, the model not only needs to send the output of the last layer of Transformer self-attention encoder to the cross attention module and the emotion classification part in the classification module, but also needs to pass the same output to the road condition classification part of the classification module. This step is added to enhance the classification ability of the model when processing road condition data, so that it can more effectively distinguish and identify different road conditions.

[0140] Such a processing strategy ensures that in the applications of emotion classification and road condition classification, the model can extract and utilize key information in a targeted manner, achieving higher task specialization and classification accuracy.

[0141] Step 104: Use the multi-head feature representations of each expression and the multi-head feature representations of each road condition as inputs of the cross-attention module, and output the expression cross-feature representation corresponding to each expression multi-head feature representation and the road condition cross-feature representation corresponding to each road condition multi-head feature representation.

[0142] The cross-attention module includes a first cross-attention layer, a second cross-attention layer, a third cross-attention layer, and a fourth cross-attention layer.

[0143] It should be noted that there is a close connection between human emotions and the environment they live in. Specifically, when observing different scenes, human emotions reflect the inner feelings and cognition of these scenes. For example, when seeing funny pictures, people usually feel happy and joyful; when facing disaster-related images, they may feel sad and sympathetic; and watching horror pictures may trigger fear and tension.

[0144] When observing road conditions, specific visual information can also trigger specific emotional responses. For example, when there are more pedestrians on the road, drivers usually become more cautious to avoid potential dangers; when there are more obstacles on the road, they may feel nervous and anxious, worrying about traffic accidents. These reactions are natural behaviors that humans have evolved to adapt to the environment and ensure safety.

[0145] Based on this understanding of the relationship between emotion and environment, the present invention cross-attentionally fuses the expression modality (facial expression) and the road condition modality (road condition observation). This fusion strategy is based on the recognition of the strong interaction between the two: changes in road conditions directly affect the driver's expression and emotional state. Through the cross-attention mechanism, the model can more accurately identify and understand the mutual influence of these two modal data when processing emotion classification tasks.

[0146] Specifically, the cross-attention module optimizes the accuracy of emotion recognition by learning the dynamic relationship between expression data and road condition data. This approach not only enhances the model's ability to understand complex emotions, but also enables it to provide more humane and contextualized response suggestions in practical applications such as driver assistance systems. Therefore, effectively integrating the data of these two modalities and using their interaction effects as the basis for emotion classification is a key step in achieving efficient emotion recognition.

[0147] Specifically, step 104 may include the following sub-steps S41-S49:

[0148] Step S41, performing multi-head linear projection on each expression multi-head feature representation respectively, generating multiple expression query matrices, multiple expression key matrices and multiple expression value matrices corresponding to each expression multi-head feature representation;

[0149] Step S42, performing multi-head linear projection on each road condition multi-head feature representation respectively, to generate multiple road condition query matrices, multiple road condition key matrices and multiple road condition value matrices corresponding to each road condition multi-head feature representation;

[0150] Step S43, using the first cross attention layer to perform cross attention head calculation according to each expression key matrix, each expression value matrix and each road condition query matrix, to determine multiple first expression cross attention heads;

[0151] Step S44, performing cross attention head calculation according to each expression query matrix, each road condition key matrix and each road condition value matrix through the second cross attention layer to determine a plurality of first road condition cross attention heads;

[0152] Step S45, performing residual feedforward network block processing on each expression multi-head feature representation, each first expression cross-attention head, each road condition multi-head feature representation, and each first road condition cross-attention head, respectively, to generate multiple initial expression feature representations and multiple initial road condition feature representations;

[0153] Step S46, linearly projecting each initial expression feature representation and each initial road condition feature representation, and outputting an expression cross-query matrix corresponding to each initial expression feature representation and a road condition cross-query matrix corresponding to each initial road condition feature representation;

[0154] Step S47, using the third cross attention layer to perform cross attention head calculation according to each expression key matrix, each expression value matrix and each road condition cross query matrix, to determine a plurality of second expression cross attention heads;

[0155] Step S48, performing cross-attention head calculation according to each road condition key matrix, each road condition value matrix and each expression cross-query matrix through the fourth cross-attention layer to determine a plurality of second road condition cross-attention heads;

[0156] Step S49, perform residual feedforward network block processing on each initial expression feature representation, each second expression cross-attention head, each initial road condition feature representation, and each second road condition cross-attention head to generate multiple expression cross-feature representations and multiple road condition cross-feature representations.

[0157] Further, step S49 may include the following sub-steps S491-S497:

[0158] Step S491, adding each initial expression feature representation and the second expression cross multi-head attention corresponding to each initial expression feature representation to generate multiple expression cross residual features;

[0159] Step S492, adding each initial road condition feature representation and the second road condition cross multi-head attention corresponding to each initial road condition feature representation to generate a plurality of road condition cross residual features;

[0160] Step S493, respectively performing layer normalization on each expression cross residual feature and each road condition cross residual feature to generate expression cross normalized features corresponding to each expression cross residual feature and road condition cross normalized features corresponding to each road condition cross residual feature;

[0161] Step S494, taking each expression cross-normalized feature and each road condition cross-normalized feature as input of a feedforward network, and outputting expression cross-target features corresponding to each expression cross-normalized feature and road condition cross-target features corresponding to each road condition cross-normalized feature;

[0162] Step S495, adding each expression cross normalization feature and the expression cross target feature corresponding to each expression cross normalization feature to generate a plurality of expression cross comprehensive features;

[0163] Step S496, adding each road condition cross normalization feature and the road condition cross target feature corresponding to each road condition cross normalization feature to generate a plurality of road condition cross comprehensive features;

[0164] Step S497, respectively perform layer normalization on each expression cross-comprehensive feature and each road condition cross-comprehensive feature to generate an expression cross-feature representation corresponding to each expression cross-comprehensive feature and a road condition cross-feature representation corresponding to each road condition cross-comprehensive feature.

[0165] It should be noted that in the cross-attention module of the model, there are designed the first cross-attention layer (multimodal multi-head cross-attention part1), the second cross-attention layer (multimodal multi-head cross-attention part2), the third cross-attention layer (multimodal multi-head cross-attention part3), and the fourth cross-attention layer (multimodal multi-head cross-attention part4). Therefore, the entire module includes a total of four cross-attention parts. Specifically, Part 2 and Part 3 of the multimodal multi-head cross-attention are specifically designed to process facial expression data, aiming to extract detailed features related to expressions. Part 1 and Part 4 of the multimodal multi-head cross-attention focus on the processing of road condition information, which helps the model to accurately understand and classify complex road conditions.

[0166] For further information, see Figure 5 ,Since the data processing processes of these four parts are basically the same in technology and methods, the main difference lies in the different data modalities processed. Therefore, this paper will explain the data processing process of this layer in detail through Part 1 of multimodal multi-head cross attention, providing a clear perspective for understanding the operation of the entire cross attention module.

[0167] First, the multimodal multi-head cross attention part1 has two input vectors, one is X (i.e., expression key matrix, expression value matrix) from the output of the facial expression part in the self-attention module, and the other is Y (i.e., road condition query matrix) from the output of the road condition information part in the self-attention module. The input vector X is linearly projected and mapped to two different spaces, which are used to generate key (Key, K) and value (Value, V) matrices respectively. The input vector Y is linearly projected and mapped to a space for generating queries (Query, Q). The generation of each matrix is ​​achieved by matrix multiplication of the input vector with the corresponding learnable weight matrix. Specifically, the query, key, and value matrices are calculated by the following formulas:

[0168] ;

[0169] Among them, Q is the query matrix; K is the key matrix; V is the value matrix; is the learnable weight matrix corresponding to the query matrix; is the learnable weight matrix corresponding to the key matrix; is the learnable weight matrix corresponding to the value matrix; Y is the output Y of the road condition information part in the self-attention module.

[0170] Furthermore, the cross-attention mechanism is similar to the self-attention mechanism, and the core is to calculate the similarity between the query (Q) and all keys (K). For example, the attention weight of the t-th cross-attention head is calculated as:

[0171] ;

[0172] in, is the t-th cross attention head; is the t-th query matrix; is the transpose of the t-th key matrix; is the scaling factor; is the t-th value matrix.

[0173] Furthermore, after completing the calculation of each cross-attention head corresponding to each multi-head feature representation, all the cross-attention head outputs corresponding to each multi-head feature representation are concatenated together to form a long vector, and the final multi-head attention output is obtained through a linear transformation to obtain the cross-multi-head attention corresponding to each multi-head feature representation. The formula is as follows:

[0174] ;

[0175] in, is the cross multi-head attention corresponding to the input vector X; For splicing; is the first cross attention head corresponding to the input vector X; are learnable parameters.

[0176] Furthermore, like the self-attention mechanism, a residual connection is added to the output of the cross-attention layer, that is, the residual feedforward network block processing. Specifically, the role of the residual connection is to add the input X and the output of the cross-attention to ensure that the information is effectively transmitted in the deep network. Then, the result of the residual connection is normalized by layer normalization. The role of layer normalization is to make the mean of each sample 0 and the variance 1, thereby improving the stability of the network, accelerating the training process, and preventing internal covariate shift. This process is expressed by the following formula:

[0177] ;

[0178] Among them, Z is the cross-normalized feature; X is the input vector; is the cross multi-head attention corresponding to the input vector X; Normalize the layer.

[0179] The next step is to pass through a feed-forward network containing a fully connected layer and an activation function. The network consists of two fully connected layers. The first fully connected layer maps the input to a higher dimension. After the activation function, it is mapped back to the original dimension through the second fully connected layer. The output of FFN is also processed by residual connection and layer normalization. The formula is as follows:

[0180] ;

[0181] Among them, Output is the cross feature representation; FFN is the feedforward network.

[0182] It is worth mentioning that the principles of the steps of performing residual feedforward network block processing on each expression multi-head feature representation, each first expression cross-multi-head attention, each road condition multi-head feature representation, and each first road condition cross-multi-head attention are consistent with the principles of the steps of performing residual feedforward network block processing on each initial expression feature representation, each second expression cross-multi-head attention, each initial road condition feature representation, and each second road condition cross-multi-head attention, and the present invention will not go into details.

[0183] Step 105: Use a classification module to perform classification according to the cross-feature representations of each expression, the multi-feature representations of each expression, the multi-feature representations of each road condition, and the cross-feature representations of each road condition, and generate an assisted driving prompt result.

[0184] The classification module includes a fully connected layer and a Softmax activation function layer.

[0185] Specifically, step 105 may include the following sub-steps S51-S58:

[0186] Step S51, adding each expression cross feature representation and each expression multi-head feature representation to generate a total input of facial expression data;

[0187] Step S52, adding the multiple feature representations of each road condition and the cross feature representations of each road condition to generate a total input of road condition information data;

[0188] Step S53, concatenating the total input of facial expression data and the total input of road condition information data to determine an emotion classification input vector;

[0189] Step S54, taking the emotion classification input vector as the input of the fully connected layer, and outputting the emotion fully connected feature vector;

[0190] Step S55, using a Softmax activation function layer to perform nonlinear mapping on the emotion fully connected feature vector to generate an emotion classification result;

[0191] Step S56, taking the multi-head feature representation of each road condition as the input of the fully connected layer, and outputting the road condition fully connected feature vector;

[0192] Step S57, using a Softmax activation function layer to perform nonlinear mapping on the road condition fully connected feature vector to generate a road condition classification result;

[0193] Step S58: Generate an assisted driving prompt result based on the emotion classification result and the road condition classification result.

[0194] It should be noted that the classification module is designed to consist of two parts: one for emotion classification and the other for road condition classification. For the emotion classification part, we first need to integrate the outputs from the cross-attention module and the self-attention module. Specifically, the output of Part 3 in the cross-attention module is Output of Part A in the self-attention module (Transformer self-attention module) that specifically processes facial expressions Perform addition operation to form comprehensive feature input of facial expression (total input of facial expression data):

[0195] ;

[0196] in, Total input for facial expression data; The output of Part 3 in the cross-attention module, including multiple cross-feature representations of expressions; It is the output of Part A in the Transformer self-attention module, including multiple expression multi-head feature representations.

[0197] Furthermore, in order to process the traffic information, the output of Part 4 in the cross-attention module Part B output for road condition information in the self-attention module The addition operation is also performed to obtain the total input of the traffic information (i.e. the total input of the traffic information data):

[0198] ;

[0199] in, It is the total input of traffic information data; The output of Part 4 in the cross-attention module, including multiple cross-feature representations of road conditions; It is the output of Part B in the Transformer self-attention module, including multiple road condition multi-head feature representations.

[0200] Furthermore, the total input of facial expression data Total input of traffic information data splicing to form the final input vector for emotion classification (i.e., the emotion classification input vector ):

[0201] ;

[0202] Furthermore, the vector It is then passed to the emotion classification part of the classification layer, where a fully connected layer and activation function are applied to convert the final feature vector into a prediction matrix (i.e., emotion classification results), which are used to classify emotional states.

[0203] Furthermore, compared to the complex data integration process of emotion classification, the processing of road condition classification is relatively simple. This part of the processing only involves converting the output of Part B of the self-attention training module for road condition information into Directly passed to the traffic classification part of the classification module. This direct pass ensures that the process from feature extraction to classification decision is as concise as possible to reduce information loss and speed up processing.

[0204] For example, in the Transformer-based multimodal large-model personal driving prompt framework, the model can not only perceive the environment outside the car, but also judge and respond to the driver's emotional state in real time by analyzing the driver's facial expressions, and give corresponding voice prompts. Assuming that the road condition information categories include "rainy day", "many roadblocks", "low light visibility", and "traffic jam", the following are some specific scenario analyses to show how the model is used in different driving environments:

[0205] Example 1: Drivers’ concerns about driving in rainy weather

[0206] Scenario analysis: When the camera outside the car detects that the external environment is rainy, the model analyzes the driver's facial expression and identifies the driver's possible worries.

[0207] System response: After the backend program receives the "worry" status output, combined with the road condition data showing "slippery in rainy days", the system determines that the driver's concerns are mainly caused by bad weather conditions.

[0208] Voice prompts (i.e. assisted driving prompts): In response to the current rainy conditions, the system will automatically provide the driver with a voice prompt: "Since the road may be slippery, it is recommended that you slow down to ensure driving safety." At the same time, the system will provide a recommended maximum driving speed to further guide safe driving.

[0209] Example 2: Driver anxiety caused by multiple roadblocks

[0210] Scenario analysis: When the external camera detects multiple obstacles on the road ahead, the model identifies the driver's possible anxiety by analyzing the driver's facial expressions and driving behavior.

[0211] System response: After the backend program receives the "anxiety" status output, further analysis shows that there are many roadblocks ahead, and it is determined that the reason for the driver's anxiety is the multiple obstacles ahead.

[0212] Voice prompt: The system will automatically provide a voice prompt: "Multiple obstacles are detected ahead, please slow down and be ready to stop at any time."

[0213] Example 3: Feeling of uneasiness while driving at night

[0214] Scenario analysis: When driving at night, if the driver shows an uneasy expression, the model combines this with the environmental data of night driving to determine the driver's uneasy feeling.

[0215] System response: Analysis results showed that the driver's anxiety might be related to the poor visibility at night, although there were no obvious obstacles ahead.

[0216] Voice prompt: The system will issue a voice prompt: "When driving at night, please turn on the high beam and maintain an appropriate speed to ensure safety."

[0217] Example 4: Driver frustration in traffic jams

[0218] Scenario analysis: During rush hour traffic jams, if the camera detects slow traffic ahead and the driver’s expression shows frustration or annoyance, the model will recognize this emotional state.

[0219] System Response: After the backend program confirms that traffic congestion is the cause of frustration, it triggers the corresponding voice output.

[0220] Voice prompt: The voice prompt provided by the system is: "Traffic congestion is detected ahead. It is recommended to be patient or use navigation to find an alternative route."

[0221] Through these examples, we can see that the Transformer-based multimodal large model (multimodal transformer-based assisted driving prompt model) can not only effectively identify and respond to the driver's emotional changes, but also provide specific and practical driving suggestions according to different driving environments, thereby improving driving safety and driving experience.

[0222] As a comparison of technical effects, we can refer to existing technologies. In recent years, with the rapid development of autonomous driving technology, people are paying more and more attention to the intelligent driving assistance system of vehicles. Such systems are designed to improve driving safety and reduce the occurrence of traffic accidents. Although existing driving assistance systems can process certain road conditions, such as lane keeping and automatic braking, most of them ignore the important impact of the driver's emotional state on driving safety.

[0223] Many research institutions and commercial companies around the world are constantly promoting the advancement of related technologies, striving to make breakthroughs in improving driving safety and vehicle automation. For example, Tesla's Autopilot system and Google's Waymo project have integrated a variety of sensors and artificial intelligence technologies to achieve high-precision perception and decision support of the vehicle environment. However, these systems mainly focus on processing the physical driving environment, and do not deeply interpret the driver's personal emotions and behaviors, so there are still limitations in fully understanding the driving environment and improving the personalized driving experience.

[0224] In China, with the rapid development of intelligent connected vehicle technology, many universities and research institutions are also actively exploring multimodal information fusion and processing technologies for driver assistance systems. Some research teams at home and abroad have achieved a series of results in vehicle environment perception, driving behavior analysis and safety assessment. Although these studies provide valuable theoretical and technical support, most studies still focus on single data source or single modality information processing, and have not yet been able to effectively integrate the driver's emotional state, facial expressions and voice data to provide comprehensive driving assistance.

[0225] The driver's emotional state, such as tension, anger or distraction, can significantly increase the risk of traffic accidents. In addition, current systems are generally inefficient in processing multimodal data, especially when considering the driver's behavior and emotions in combination with external environmental information. Therefore, there is an urgent need for an intelligent system that can fully understand and predict the driving environment and driver status to provide more accurate and timely driving tips to ensure driving safety.

[0226] Through long-term accumulated data training, the system can customize a unique model for each driver, further enhancing the system's personalized service capabilities. In existing driver assistance systems, most solutions adopt a "one-size-fits-all" strategy, that is, the feedback and prompts received by all drivers in similar situations are standardized. Although this approach is effective in dealing with routine situations, it ignores the differences in driving behavior, emotional reactions, and operating habits between individual drivers. For example, an experienced driver and a novice driver may have completely different reactions and the type of assistance they need when encountering an emergency.

[0227] In view of the above problems, the present invention proposes an assisted driving prompt method based on multimodal transformer, which can build a dynamically adjusted driving model (assisted driving prompt model based on multimodal transformer) based on deep learning and analysis of the driver's past behavior. Multimodal data such as the driver's facial expressions are collected through an integrated camera. These data are marked as different emotional states and behavioral patterns, such as tension, relaxation or anger, and are used to train a personalized transformer model.

[0228] In addition, through continuous data accumulation and analysis, it is possible to more accurately predict the possible reactions of individual drivers under specific road conditions. For example, if it detects that a driver often shows nervousness when encountering complex traffic conditions, it can provide reassuring information in advance or suggest more cautious driving strategies in similar situations, such as reducing speed, increasing distance from the vehicle in front, etc.

[0229] The present invention can also adjust the prompting method and content according to the driver's preferences and historical behavior data, which is more in line with the individual's acceptance and reaction methods. For example, for drivers who like to relax through music, their preferred music list can be automatically played when a high stress state is detected. For drivers who need to concentrate in silence, voice prompts can be reduced and necessary information can be provided through visual display.

[0230] Through this highly personalized service, the present invention not only improves the practicality and user satisfaction of the driving assistance system, but also has incomparable advantages in safety. It can provide the most appropriate assistance and intervention according to the specific needs of each driver, thereby effectively reducing the risk of accidents caused by individual differences among drivers.

[0231] In addition, the present invention uses PostgreSQL database for data storage and uses miniIO for data backup to ensure data security and reliability.

[0232] In summary, the technical solution of the present invention can effectively integrate multiple data sources and provide a new solution for the driving assistance system through advanced transformer model training, which can solve the shortcomings of the existing technology and ensure driving safety.

[0233] In an embodiment of the present invention, the present invention provides an assisted driving prompt method based on a multimodal transformer. First, a facial expression image and a road condition image are obtained, and the facial expression image and the road condition image are input into an assisted driving prompt model based on a multimodal transformer; the assisted driving prompt model based on a multimodal transformer includes an input module, a Transformer self-attention module, a cross-attention module and a classification module; then, the input module is used to perform feature vectorization on the facial expression image and the road condition image respectively, and a plurality of expression target embeddings corresponding to the facial expression image and a plurality of road condition target embeddings corresponding to the road condition image are generated; the Transformer self-attention module is used to perform multi-head attention block processing on each expression target embedding and each road condition target embedding respectively, and the expression multi-head feature representation corresponding to each expression target embedding, each road condition target embedding and the road condition target embedding are output. The road condition target is embedded in the corresponding road condition multi-head feature representation; each expression multi-head feature representation and each road condition multi-head feature representation are used as the input of the cross-attention module, and the expression cross-feature representation corresponding to each expression multi-head feature representation and the road condition cross-feature representation corresponding to each road condition multi-head feature representation are output; finally, a classification module is used to classify each expression cross-feature representation, each expression multi-head feature representation, each road condition multi-head feature representation, and each road condition cross-feature representation to generate an assisted driving prompt result; based on the above scheme, the present invention uses an assisted driving prompt model based on a multimodal transformer to process the acquired multimodal images, namely facial expression images and road condition images, and in the process of outputting the assisted driving prompt result, the driver's emotional state and road conditions are comprehensively considered, and can provide appropriate assistance and intervention according to the specific emotional state of each driver, thereby providing more accurate driving prompts.

[0234] See also Figure 6 , Figure 6A structural block diagram of a multimodal transformer-based assisted driving prompt device provided in Example 2 of the present invention.

[0235] The present invention provides a multimodal transformer-based auxiliary driving prompt device, comprising:

[0236] An acquisition module 601 is used to acquire a facial expression image and a road condition image, and input the facial expression image and the road condition image into a multimodal transformer-based assisted driving prompt model; the multimodal transformer-based assisted driving prompt model includes an input module, a Transformer self-attention module, a cross-attention module and a classification module;

[0237] Adopting module 602, for adopting input module to perform feature vectorization on facial expression image and road condition image respectively, and generating multiple expression target embeddings corresponding to facial expression image and multiple road condition target embeddings corresponding to road condition image;

[0238] A multi-head attention processing module 603 is used to perform multi-head attention block processing on each expression target embedding and each road condition target embedding respectively through the Transformer self-attention module, and output the expression multi-head feature representation corresponding to each expression target embedding and the road condition multi-head feature representation corresponding to each road condition target embedding;

[0239] A cross attention processing module 604 is used to take each expression multi-head feature representation and each road condition multi-head feature representation as input of a cross attention module, and output an expression cross feature representation corresponding to each expression multi-head feature representation and a road condition cross feature representation corresponding to each road condition multi-head feature representation;

[0240] The classification module 605 is used to use the classification module to classify according to the cross-feature representation of each expression, the multi-feature representation of each expression, the multi-feature representation of each road condition, and the cross-feature representation of each road condition to generate an assisted driving prompt result.

[0241] Further, module 602 is used to:

[0242] The facial expression image and the road condition image are divided respectively to generate a plurality of expression image blocks corresponding to the facial expression image and a plurality of road condition image blocks corresponding to the road condition image;

[0243] Flatten each expression image block and each road condition image block respectively, and output an expression one-dimensional vector corresponding to each expression image block and a road condition one-dimensional vector corresponding to each road condition image block;

[0244] Performing linear transformation on each expression one-dimensional vector and each road condition one-dimensional vector respectively, and outputting expression mark embedding corresponding to each expression one-dimensional vector and road condition mark embedding corresponding to each road condition one-dimensional vector;

[0245] Add each expression marker embedding to the preset position embedding corresponding to each expression marker embedding, and output the expression target embedding corresponding to each expression marker embedding;

[0246] Each road condition mark embedding is added to the preset position embedding corresponding to each road condition mark embedding, and the road condition target embedding corresponding to each road condition mark embedding is output.

[0247] Furthermore, the Transformer self-attention module includes multiple Transformer autoencoders; the data processing process of the Transformer autoencoder is specifically as follows:

[0248] Performing a multi-head linear transformation on the multiple first input vectors and the multiple second input vectors input to the Transformer autoencoder to generate multiple query matrices, multiple key matrices, and multiple value matrices corresponding to the first input vectors, and multiple query matrices, multiple key matrices, and multiple value matrices corresponding to the second input vectors;

[0249] Calculating a plurality of first output self-attention heads corresponding to each first input vector based on a plurality of query matrices, a plurality of key matrices, and a plurality of value matrices corresponding to each first input vector;

[0250] Calculating a plurality of second output self-attention heads corresponding to each second input vector based on a plurality of query matrices, a plurality of key matrices, and a plurality of value matrices corresponding to each second input vector;

[0251] Respectively concatenate and linearly transform the multiple first output self-attention heads corresponding to each first input vector and the multiple second output self-attention heads corresponding to each second input vector, and output the first multi-head attention heads corresponding to each first input vector and the second multi-head attention heads corresponding to each second input vector;

[0252] Adding each first input vector and the first multi-head attention corresponding to each first input vector to generate a plurality of first residual features;

[0253] Adding each second input vector and the second multi-head attention corresponding to each second input vector to generate a plurality of second residual features;

[0254] Performing layer normalization on each first residual feature and each second residual feature respectively to generate a first normalized feature corresponding to each first residual feature and a second normalized feature corresponding to each second residual feature;

[0255] Taking each first normalized feature and each second normalized feature as input of a feedforward network, and outputting a first target feature corresponding to each first normalized feature and a second target feature corresponding to each second normalized feature;

[0256] Adding each first normalized feature and a first target feature corresponding to each first normalized feature to generate a plurality of first comprehensive features;

[0257] Adding each second normalized feature and a second target feature corresponding to each second normalized feature to generate a plurality of second comprehensive features;

[0258] Layer normalization is performed on each first comprehensive feature and each second comprehensive feature to generate a first multi-head feature representation corresponding to each first comprehensive feature and a second multi-head feature representation corresponding to each second comprehensive feature.

[0259] Furthermore, the cross attention module includes a first cross attention layer, a second cross attention layer, a third cross attention layer, and a fourth cross attention layer; the cross attention processing module 604 includes:

[0260] The first submodule is used to perform multi-head linear projection on each expression multi-head feature representation, and generate multiple expression query matrices, multiple expression key matrices and multiple expression value matrices corresponding to each expression multi-head feature representation;

[0261] The second submodule is used to perform multi-head linear projection on each road condition multi-head feature representation, and generate multiple road condition query matrices, multiple road condition key matrices and multiple road condition value matrices corresponding to each road condition multi-head feature representation;

[0262] The third submodule is used to use the first cross attention layer to perform cross attention head calculation according to each expression key matrix, each expression value matrix and each road condition query matrix to determine multiple first expression cross multi-head attentions;

[0263] A fourth submodule is used to perform cross attention head calculation according to each expression query matrix, each road condition key matrix and each road condition value matrix through a second cross attention layer to determine multiple first road condition cross multi-head attentions;

[0264] A fifth submodule is used to perform residual feedforward network block processing on each expression multi-head feature representation, each first expression cross multi-head attention, each road condition multi-head feature representation, and each first road condition cross multi-head attention, to generate multiple initial expression feature representations and multiple initial road condition feature representations;

[0265] The sixth submodule is used to perform linear projection on each initial expression feature representation and each initial road condition feature representation, and output an expression cross-query matrix corresponding to each initial expression feature representation and a road condition cross-query matrix corresponding to each initial road condition feature representation;

[0266] A seventh submodule is used to use the third cross attention layer to perform cross attention head calculation according to each expression key matrix, each expression value matrix and each road condition cross query matrix to determine multiple second expression cross multi-head attentions;

[0267] An eighth submodule, configured to perform cross-attention head calculations according to each road condition key matrix, each road condition value matrix, and each expression cross-query matrix through a fourth cross-attention layer to determine a plurality of second road condition cross-multi-head attentions;

[0268] The ninth submodule is used to perform residual feedforward network block processing on each initial expression feature representation, each second expression cross-multi-head attention, each initial road condition feature representation, and each second road condition cross-multi-head attention, to generate multiple expression cross-feature representations and multiple road condition cross-feature representations.

[0269] Furthermore, the ninth submodule is specifically used for:

[0270] Adding each initial expression feature representation and the second expression cross multi-head attention corresponding to each initial expression feature representation to generate multiple expression cross residual features;

[0271] Adding each initial road condition feature representation and the second road condition cross multi-head attention corresponding to each initial road condition feature representation to generate multiple road condition cross residual features;

[0272] Performing layer normalization on each expression cross residual feature and each road condition cross residual feature respectively, generating expression cross normalized features corresponding to each expression cross residual feature and road condition cross normalized features corresponding to each road condition cross residual feature;

[0273] The cross-normalized features of each expression and the cross-normalized features of each road condition are respectively used as the input of the feedforward network, and the cross-target features of the expression corresponding to the cross-normalized features of each expression and the cross-target features of the road condition corresponding to the cross-normalized features of each road condition are output;

[0274] Adding each expression cross normalization feature and the expression cross target feature corresponding to each expression cross normalization feature to generate multiple expression cross comprehensive features;

[0275] Adding each road condition cross normalization feature and the road condition cross target feature corresponding to each road condition cross normalization feature to generate a plurality of road condition cross comprehensive features;

[0276] Layer normalization is performed on each expression cross-comprehensive feature and each road condition cross-comprehensive feature to generate an expression cross-feature representation corresponding to each expression cross-comprehensive feature and a road condition cross-feature representation corresponding to each road condition cross-comprehensive feature.

[0277] Furthermore, the classification module includes a fully connected layer and a Softmax activation function layer; the classification module 605 is specifically used for:

[0278] Adding the cross-feature representations of each expression and the multi-feature representations of each expression to generate a total input of facial expression data;

[0279] Adding the multi-head feature representations of each road condition and the cross-feature representations of each road condition to generate a total input of road condition information data;

[0280] Concatenate the total facial expression data input and the total road condition information data input to determine the emotion classification input vector;

[0281] The emotion classification input vector is used as the input of the fully connected layer, and the emotion fully connected feature vector is output;

[0282] The Softmax activation function layer is used to perform nonlinear mapping on the fully connected feature vector of emotions to generate emotion classification results;

[0283] The multi-head feature representation of each road condition is used as the input of the fully connected layer, and the road condition fully connected feature vector is output;

[0284] The Softmax activation function layer is used to perform nonlinear mapping on the fully connected feature vector of the road condition to generate the road condition classification result;

[0285] Based on the emotion classification results and road condition classification results, assisted driving prompt results are generated.

[0286] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described devices, modules and sub-modules can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0287] An embodiment of the present invention also provides a computer device, including a memory and a processor, wherein a computer program is stored in the memory; when the computer program is executed by the processor, the processor executes the steps of the multimodal transformer-based assisted driving prompt method as described in the first embodiment above.

[0288] An embodiment of the present invention further provides a computer-readable storage medium having a computer program / instruction stored thereon. When the computer program / instruction is executed by a processor, the steps of the multimodal transformer-based assisted driving prompt method as described in the first embodiment are implemented.

[0289] An embodiment of the present invention also provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the steps of the multimodal transformer-based assisted driving prompt method as described in the first embodiment above.

[0290] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0291] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0292] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features thereof may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. An assisted driving prompt method based on a multimodal transformer, characterized in that: include: Acquire a facial expression image and a road condition image, and input the facial expression image and the road condition image into a multimodal transformer-based assisted driving prompt model; the multimodal transformer-based assisted driving prompt model includes an input module, a Transformer self-attention module, a cross-attention module, and a classification module; Using the input module to perform feature vectorization on the facial expression image and the road condition image respectively, to generate a plurality of expression target embeddings corresponding to the facial expression image and a plurality of road condition target embeddings corresponding to the road condition image; Through the Transformer self-attention module, each of the expression target embeddings and each of the road condition target embeddings are respectively processed by a multi-head attention block, and the expression multi-head feature representation corresponding to each of the expression target embeddings and the road condition multi-head feature representation corresponding to each of the road condition target embeddings are output; Taking each of the multi-head feature representations of expression and each of the multi-head feature representations of road conditions as inputs of the cross attention module, and outputting the cross feature representations of expression corresponding to each of the multi-head feature representations of expression and the cross feature representations of road conditions corresponding to each of the multi-head feature representations of road conditions; The classification module is used to perform classification according to each of the expression cross feature representations, each of the expression multi-head feature representations, each of the road condition multi-head feature representations, and each of the road condition cross feature representations to generate an auxiliary driving prompt result; The Transformer self-attention module includes multiple Transformer autoencoders; the data processing process of the Transformer autoencoder is specifically as follows: Performing a multi-head linear transformation on a plurality of first input vectors and a plurality of second input vectors input to the Transformer autoencoder to generate a plurality of query matrices, a plurality of key matrices, and a plurality of value matrices corresponding to each of the first input vectors, and a plurality of query matrices, a plurality of key matrices, and a plurality of value matrices corresponding to each of the second input vectors; Calculating a plurality of first output self-attention heads corresponding to each of the first input vectors based on a plurality of query matrices, a plurality of key matrices, and a plurality of value matrices corresponding to each of the first input vectors; Calculate a plurality of second output self-attention heads corresponding to each of the second input vectors based on a plurality of query matrices, a plurality of key matrices, and a plurality of value matrices corresponding to each of the second input vectors; Respectively concatenate and linearly transform the multiple first output self-attention heads corresponding to each of the first input vectors and the multiple second output self-attention heads corresponding to each of the second input vectors, and output a first multi-head attention corresponding to each of the first input vectors and a second multi-head attention corresponding to each of the second input vectors; Adding each of the first input vectors and the first multi-head attention corresponding to each of the first input vectors to generate a plurality of first residual features; Adding each of the second input vectors and the second multi-head attention corresponding to each of the second input vectors to generate a plurality of second residual features; Performing layer normalization on each of the first residual features and each of the second residual features respectively to generate a first normalized feature corresponding to each of the first residual features and a second normalized feature corresponding to each of the second residual features; Using each of the first normalized features and each of the second normalized features as input of a feedforward network, and outputting a first target feature corresponding to each of the first normalized features and a second target feature corresponding to each of the second normalized features; Adding each of the first normalized features and the first target features corresponding to each of the first normalized features to generate a plurality of first comprehensive features; Adding each of the second normalized features and the second target features corresponding to each of the second normalized features to generate a plurality of second comprehensive features; Layer normalization is performed on each of the first comprehensive features and each of the second comprehensive features to generate a first multi-head feature representation corresponding to each of the first comprehensive features and a second multi-head feature representation corresponding to each of the second comprehensive features.

2. The multimodal transformer-based assisted driving prompt method according to claim 1, characterized in that: The step of using the input module to perform feature vectorization on the facial expression image and the road condition image respectively to generate a plurality of expression target embeddings corresponding to the facial expression image and a plurality of road condition target embeddings corresponding to the road condition image comprises: Dividing the facial expression image and the road condition image respectively to generate a plurality of expression image blocks corresponding to the facial expression image and a plurality of road condition image blocks corresponding to the road condition image; Flattening each of the expression image blocks and each of the road condition image blocks respectively, and outputting a one-dimensional expression vector corresponding to each of the expression image blocks and a one-dimensional road condition vector corresponding to each of the road condition image blocks; Performing linear transformation on each of the one-dimensional expression vectors and each of the one-dimensional road condition vectors, respectively, and outputting expression mark embeddings corresponding to each of the one-dimensional expression vectors and road condition mark embeddings corresponding to each of the one-dimensional road condition vectors; Adding each of the expression marker embeddings to the preset position embeddings corresponding to each of the expression marker embeddings, and outputting the expression target embeddings corresponding to each of the expression marker embeddings; Each of the road condition mark embeddings is added to the preset position embedding corresponding to each of the road condition mark embeddings, and the road condition target embedding corresponding to each of the road condition mark embeddings is output.

3. The multimodal transformer-based assisted driving prompt method according to claim 1, characterized in that: The cross attention module includes a first cross attention layer, a second cross attention layer, a third cross attention layer, and a fourth cross attention layer; the multi-head feature representations of each expression and the multi-head feature representations of each road condition are used as inputs of the cross attention module, and the cross feature representations of each expression corresponding to the multi-head feature representations of each expression and the cross feature representations of each road condition corresponding to the multi-head feature representations of each road condition are output, including: Performing multi-head linear projection on each of the multi-head feature representations of the expression respectively to generate a plurality of expression query matrices, a plurality of expression key matrices and a plurality of expression value matrices corresponding to each of the multi-head feature representations of the expression; Performing multi-head linear projection on each of the multi-head feature representations of the road condition to generate a plurality of road condition query matrices, a plurality of road condition key matrices and a plurality of road condition value matrices corresponding to each of the multi-head feature representations of the road condition; Using a first cross attention layer to perform cross attention head calculation according to each of the expression key matrices, each of the expression value matrices, and each of the road condition query matrices, to determine a plurality of first expression cross multi-head attentions; Performing cross-attention head calculation according to each of the expression query matrices, each of the road condition key matrices, and each of the road condition value matrices through a second cross-attention layer to determine a plurality of first road condition cross-multi-head attentions; Respectively perform residual feedforward network block processing on each of the expression multi-head feature representations, each of the first expression cross-multi-head attentions, each of the road condition multi-head feature representations, and each of the first road condition cross-multi-head attentions to generate multiple initial expression feature representations and multiple initial road condition feature representations; Performing linear projection on each of the initial expression feature representations and each of the initial road condition feature representations, respectively, and outputting an expression cross-query matrix corresponding to each of the initial expression feature representations and a road condition cross-query matrix corresponding to each of the initial road condition feature representations; Using a third cross attention layer to perform cross attention head calculation according to each of the expression key matrices, each of the expression value matrices, and each of the road condition cross query matrices, to determine a plurality of second expression cross multi-head attentions; Performing cross-attention head calculation according to each of the road condition key matrices, each of the road condition value matrices, and each of the expression cross-query matrices through a fourth cross-attention layer to determine a plurality of second road condition cross-multi-head attentions; Each of the initial expression feature representations, each of the second expression cross-multi-head attentions, each of the initial road condition feature representations, and each of the second road condition cross-multi-head attentions is processed by a residual feedforward network block to generate multiple expression cross-feature representations and multiple road condition cross-feature representations.

4. The multimodal transformer-based assisted driving prompt method according to claim 3, characterized in that: The method of performing residual feedforward network block processing on each of the initial expression feature representations, each of the second expression cross-multi-head attentions, each of the initial road condition feature representations, and each of the second road condition cross-multi-head attentions to generate multiple expression cross-feature representations and multiple road condition cross-feature representations includes: Adding each of the initial expression feature representations and the second expression cross multi-head attention corresponding to each of the initial expression feature representations to generate a plurality of expression cross residual features; Adding each of the initial road condition feature representations and the second road condition cross multi-head attention corresponding to each of the initial road condition feature representations to generate a plurality of road condition cross residual features; Respectively performing layer normalization on each of the expression cross residual features and each of the road condition cross residual features to generate expression cross normalized features corresponding to each of the expression cross residual features and road condition cross normalized features corresponding to each of the road condition cross residual features; Using each of the expression cross-normalized features and each of the road condition cross-normalized features as inputs of a feedforward network, and outputting expression cross-target features corresponding to each of the expression cross-normalized features and road condition cross-target features corresponding to each of the road condition cross-normalized features; Adding each of the expression cross normalization features and the expression cross target features corresponding to each of the expression cross normalization features to generate a plurality of expression cross comprehensive features; Adding each of the road condition cross normalization features and the road condition cross target features corresponding to each of the road condition cross normalization features to generate a plurality of road condition cross comprehensive features; Each of the expression cross-comprehensive features and each of the road condition cross-comprehensive features is layer-normalized to generate an expression cross-feature representation corresponding to each of the expression cross-comprehensive features and a road condition cross-feature representation corresponding to each of the road condition cross-comprehensive features.

5. The multimodal transformer-based assisted driving prompt method according to claim 1, characterized in that: The classification module includes a fully connected layer and a Softmax activation function layer; the classification module is used to classify according to each of the expression cross feature representations, each of the expression multi-head feature representations, each of the road condition multi-head feature representations, and each of the road condition cross feature representations to generate an assisted driving prompt result, including: Adding each of the expression cross feature representations and each of the expression multi-head feature representations to generate a total input of facial expression data; Adding each of the road condition multi-feature representations and each of the road condition cross-feature representations to generate a total input of road condition information data; Concatenate the total facial expression data input and the total road condition information data input to determine an emotion classification input vector; Using the emotion classification input vector as the input of the fully connected layer, and outputting the emotion fully connected feature vector; A Softmax activation function layer is used to perform nonlinear mapping on the emotion fully connected feature vector to generate an emotion classification result; Taking each of the multi-head feature representations of the road condition as the input of the fully connected layer, and outputting a road condition fully connected feature vector; A Softmax activation function layer is used to perform nonlinear mapping on the road condition fully connected feature vector to generate a road condition classification result; An assisted driving prompt result is generated according to the emotion classification result and the road condition classification result.

6. A multimodal transformer-based assisted driving prompt device, applied to the multimodal transformer-based assisted driving prompt method according to claim 1, characterized in that: include: An acquisition module is used to acquire a facial expression image and a road condition image, and input the facial expression image and the road condition image into a multimodal transformer-based assisted driving prompt model; the multimodal transformer-based assisted driving prompt model includes an input module, a Transformer self-attention module, a cross-attention module and a classification module; An adopting module, used for adopting the input module to perform feature vectorization on the facial expression image and the road condition image respectively, to generate a plurality of expression target embeddings corresponding to the facial expression image and a plurality of road condition target embeddings corresponding to the road condition image; A multi-head attention processing module, used to perform multi-head attention block processing on each of the expression target embeddings and each of the road condition target embeddings through the Transformer self-attention module, and output expression multi-head feature representations corresponding to each of the expression target embeddings and road condition multi-head feature representations corresponding to each of the road condition target embeddings; A cross-attention processing module, used to take each of the expression multi-head feature representations and each of the road condition multi-head feature representations as inputs of the cross-attention module, and output an expression cross-feature representation corresponding to each of the expression multi-head feature representations and a road condition cross-feature representation corresponding to each of the road condition multi-head feature representations; A classification module is used to use the classification module to perform classification according to each of the above-mentioned expression cross-feature representations, each of the above-mentioned expression multi-head feature representations, each of the above-mentioned road condition multi-head feature representations, and each of the above-mentioned road condition cross-feature representations to generate an assisted driving prompt result.

7. A computer device, characterized in that: It includes a memory and a processor, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, the processor executes the steps of the multimodal transformer-based assisted driving prompt method as described in any one of claims 1 to 5.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed, the multimodal transformer-based assisted driving prompt method as described in any one of claims 1 to 5 is implemented.

9. A computer program product, characterized in that The computer program product includes a computer program stored on a non-transitory computer-readable storage medium, and the computer program includes program instructions, wherein when the program instructions are executed by a computer, the computer executes the multimodal transformer-based assisted driving prompt method as described in any one of claims 1-5.

Citation Information

Patent Citations

  • ViT and StarGAN-based driver expression recognition method

    CN114005154A

  • Brain state determination method and device

    CN116204806A