Multimodal scenario risk determination method based on generative ai large language model

WO2025185005A8PCT designated stage Publication Date: 2025-10-02CENT SOUTH UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/098941
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-06
Filing Date
2024-06-13
Publication Date
2025-10-02

AI Technical Summary

Technical Problem

Existing traffic analysis methods rely too much on simplified models and limited data dimensions, and are unable to fully understand multimodal data sets, resulting in poor recognition efficiency, accuracy and robustness of risk assessment systems.

Method used

A multimodal scenario risk assessment method based on a generative AI large language model is adopted. The key features of multimodal data are extracted through the ALBEF algorithm. Combined with the Transformer model and the large language model, image-text comparative learning and cross-attention mechanism are used to fuse visual and text features, evaluate scenario risks and compare them with preset thresholds to identify high-risk scenarios.

Benefits of technology

It improves the ability to understand real-world scenarios, can quickly identify and flag high-risk scenarios, and provides more comprehensive risk assessment and real-time response capabilities for autonomous driving.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024098941_02102025_PF_FP_ABST
    Figure CN2024098941_02102025_PF_FP_ABST
Patent Text Reader

Abstract

The embodiments of the present disclosure belong to the technical field of data processing. Provided is a multimodal scenario risk determination method based on a generative AI large language model. The method specifically comprises: step 1, acquiring multimodal data to form a target data set, wherein the multimodal data comprises visual data and text data; step 2, using an ALBEF algorithm to extract key features corresponding to the target data set, and fusing the key features into a comprehensive scenario representation; and step 3, on the basis of a preset safety index and a large language model, evaluating a risk degree corresponding to the comprehensive scenario representation, comparing the risk degree with a risk threshold, and determining whether the scenario corresponding to the comprehensive scenario representation is a high-risk scenario. By means of the solution in the present disclosure, a high-risk scenario can be rapidly recognized and identified, so as to provide a basis for taking emergency measures, thereby enhancing the real-time response capability.
Need to check novelty before this filing date? Find Prior Art

Description

Multimodal scenario risk assessment method based on generative AI large language model Technical Field

[0001] The disclosed embodiments relate to the field of data processing technology, and in particular to a multimodal scenario risk assessment method based on a generative AI large language model. Background Art

[0002] Currently, scenario risk assessment plays a crucial role in intelligent transportation safety. It is not only the core of autonomous driving virtual simulation safety testing but also crucial for the development of autonomous driving test scenarios and the construction of high-risk scenario libraries. Accurate risk assessment is fundamental to ensuring road safety and guiding intelligent vehicles in making appropriate decisions.

[0003] However, traditional traffic analysis methods, due to their over-reliance on simplified models and limited data dimensions, often fail to fully understand multimodal datasets containing images, video feeds, and textual information. Consequently, existing risk assessment systems fail to fully utilize these heterogeneous data sources, which contain rich environmental information. Therefore, integrating heterogeneous data and deriving accurate risk assessments from them has become a major challenge that needs to be overcome in the field of intelligent connected vehicles.

[0004] It can be seen that there is an urgent need for a multimodal scenario risk assessment method based on a generative AI large language model with high recognition efficiency, accuracy and robustness.

[0005] Summary of the Invention

[0006] In view of this, the embodiments of the present disclosure provide a multimodal scenario risk assessment method based on a generative AI large language model, which at least partially solves the problems of poor recognition efficiency, accuracy and robustness in the existing technology.

[0007] The present disclosure provides a multimodal scenario risk assessment method based on a generative AI large language model, including:

[0008] Step 1: Acquire multimodal data to form a target data set, wherein the multimodal data includes visual data and text data;

[0009] Step 2: Apply the ALBEF algorithm to extract key features corresponding to the target dataset and fuse them into a comprehensive scene representation;

[0010] Step 3: Evaluate the risk level corresponding to the comprehensive scenario representation based on the preset safety indicators and the large language model and compare it with the risk threshold to determine whether the scenario corresponding to the comprehensive scenario representation is a high-risk scenario.

[0011] According to a specific implementation of the embodiment of the present disclosure, step 2 specifically includes:

[0012] In step 2.1, the visual data is divided into patches and then fed into the Transformer model through a linear transformation layer to extract visual features. In addition, the text data is processed through word segmentation and tokenization, and then processed by the text encoder to obtain text features.

[0013] Step 2.2: With image-text contrastive learning as the core, feature alignment is performed by downsampling and normalizing the refined feature vectors to form positive samples of visual features and text features.

[0014] In step 2.3, a multimodal encoder is used to integrate positive samples of visual features and text features, and output a comprehensive scene representation that comprehensively considers various features.

[0015] According to a specific implementation of the embodiment of the present disclosure, step 2.2 includes:

[0016] The normalized visual features and text features of each momentum encoder are denoted as g′ v (v′ cls ) and g′ w (w′ cls ), calculate and optimize the similarity scores between image to text and text to image in a normalized manner based on the softmax function:

[0017] Where s(I, T) = g v (v cls ) T g′ w (w′ cls ); s(T, I) = g w (w cls ) T g′ v (v′ cls ), τ represents the learnable temperature parameter, and defines y i2t (I) and y t2i (T) is the similarity of one-shot data of different modalities in the same scene, the probability of negative pair is 0, and the probability of positive pair is 1. The ITC loss is defined as the cross entropy H between p and y;

[0018] Using ITC as the optimization objective during training, related image-text pairs are brought close together in the embedding space, thus obtaining positive samples of visual and text features.

[0019] According to a specific implementation of the embodiment of the present disclosure, step 2.3 specifically includes:

[0020] In the multimodal encoder, visual features and textual features are fused through a cross-attention mechanism, with visual features as keys and values ​​and textual features as queries, and vice versa, and then trained on MLM and ITM tasks:

[0021] The model generates a joint feature representation as a comprehensive scene representation: u = MLP([v′ cls , w′ cls ]);

[0022] Among them, v′ cls =CrossAttn(v cls W); w′ cls =CrossAttn(w cls , V).

[0023] According to a specific implementation of the embodiment of the present disclosure, step 3 specifically includes:

[0024] Step 3.1, setting a time-based indicator, a deceleration-based indicator, and an energy-based indicator to form a preset safety indicator;

[0025] In step 3.2, the large language model is used to evaluate the risk level corresponding to the comprehensive scenario according to the preset safety indicators and compare it with the preset risk threshold. The high-risk scenarios are monitored and identified by combining the text prompt and the preset evaluation method.

[0026] According to a specific implementation method of an embodiment of the present disclosure, the Transformer model includes an encoder layer, multi-head attention, a feedforward fully connected network and a decoder layer, wherein the encoder layer includes a multi-head self-attention mechanism and a feedforward fully connected network, the multi-head attention contains multiple parallel self-attention heads, and the decoder includes multiple similar layers, each layer includes two multi-head self-attention modules and a feedforward fully connected network.

[0027] According to a specific implementation of the embodiment of the present disclosure, the specific processing flow of the Transformer model includes:

[0028] The input sequence is converted into an embedding vector, and for each even time step 2i, a sine function is used. To create an element in the position vector, where pos is the position index and d model is the embedding dimension, i is the dimension index, and for each odd time step 2i+1, the cosine function is used. To create the elements in the position vector, these position vectors are then added element-wise to the corresponding embedding vector to introduce position information;

[0029] The encoder takes as input an embedding vector with position information added, adds the output of the multi-head self-attention mechanism and the feed-forward fully connected network to its input, and then performs layer normalization.

[0030] The input is linearly transformed into three different matrices, which are used to generate the query vector Q, key vector K and value vector V respectively. Subsequently, the dot product is used to calculate the similarity score S = QK between Q and K. T , which represents how each position in the input pays attention to other positions. The score S is scaled by dividing by (d k is the key vector dimension), then apply the softmax function to the scaled scores and multiply the attention weights by the value vector V to obtain the attention function for a set of queries:

[0031] Execute the attention function in parallel on different projection dimensions to obtain the output of multi-head self-attention:

[0032] MultiHead(Q,K,V)=Concat(head1,...,head h )W o

[0033] in, h represents the number of multi-head attention;

[0034] The output of the self-attention module of each encoder layer is first added to the original input to form a residual connection, and then input to the layer normalization. The normalized output is used as the input of the feedforward network. After passing through the linear layer of the feedforward network and the nonlinear activation function ReLU: FFN(x) = max(0, xW1+b1)W2+b2, the network output is again subjected to the residual connection and layer normalization to obtain the final output of the encoder layer.

[0035] The first multi-head self-attention module uses a look-ahead mask, the second multi-head self-attention module uses the output of the encoder layer as the key and value, and the output of the first multi-head self-attention module of the decoder is used as the query. Each multi-head self-attention module output of the decoder is subjected to residual connection and layer normalization, and then input into the feedforward network. Finally, the output of the decoder is converted into the probability distribution of the predicted next word through an additional linear layer and a softmax layer.

[0036] According to a specific implementation of the embodiment of the present disclosure, the expression of the time-based indicator is

[0037] Where X represents the vehicle position (i represents the following vehicle, i-1 represents the leading vehicle), l represents the vehicle length, and v represents the speed.

[0038] According to a specific implementation of the embodiment of the present disclosure, the expression of the deceleration-based index is:

[0039] Where t represents the time interval;

[0040] Where MADR represents the maximum braking rate, Δt represents the time step, and t i and tf i denote the initial and final time steps, T i is the total travel time;

[0041] Among them, SSD L and SSD F Represents the stopping distance of the front vehicle and the rear vehicle, v F and v L Represent the speed of the moving vehicle and the preceding vehicle, t d represents the delay time, S represents the gap distance, d m Indicates the maximum deceleration.

[0042] According to a specific implementation of the embodiment of the present disclosure, the expression of the energy-based index is:

[0043] Among them, v1 and v2 represent the velocities of two vehicles with potential collision trajectories before the collision, and m1 and m2 represent the masses of the two vehicles.

[0044] The multimodal scenario risk determination scheme based on the generative AI large language model in the embodiment of the present disclosure includes: step 1, obtaining multimodal data to form a target data set, wherein the multimodal data includes visual data and text data; step 2, applying the ALBEF algorithm to extract key features corresponding to the target data set, and fusing them into a comprehensive scenario representation; step 3, evaluating the risk level corresponding to the comprehensive scenario representation based on preset safety indicators and the large language model and comparing it with the risk threshold, to determine whether the scenario corresponding to the comprehensive scenario representation is a high-risk scenario.

[0045] The beneficial effects of the embodiments of the present disclosure are as follows: through the scheme of the present disclosure, by combining text, image and video data, based on the Transformer algorithm and the large language model, the present invention makes full use of various types of data and improves the system's ability to understand real-world scenarios. This multimodal approach can capture rich information that a single data source cannot provide, thereby providing a more comprehensive perspective for risk assessment. At the same time, by calling a trained large language model, through pre-set SSM indicators and real-time monitoring, the present invention can quickly identify and mark high-risk scenarios, provide a basis for taking emergency measures, and thus enhance real-time response capabilities. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0047] FIG1 is a flow chart of a multimodal scenario risk assessment method based on a generative AI large language model according to an embodiment of the present disclosure;

[0048] FIG2 is a schematic diagram of a specific implementation process of a multimodal scenario risk determination method based on a generative AI large language model provided by an embodiment of the present disclosure;

[0049] FIG3 is a system framework diagram of a multimodal scenario risk determination method based on a generative AI large language model provided by an embodiment of the present disclosure;

[0050] FIG4 is a framework diagram of a multimodal data fusion module provided by an embodiment of the present disclosure;

[0051] FIG5 is a flow chart of a Transformer algorithm provided by an embodiment of the present disclosure. DETAILED DESCRIPTION

[0052] The embodiments of the present disclosure are described in detail below with reference to the accompanying drawings.

[0053] The following describes the embodiments of the present disclosure through specific examples, and those skilled in the art can easily understand other advantages and effects of the present disclosure from the contents disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all of the embodiments. The present disclosure can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present disclosure. It should be noted that, in the absence of conflict, the following embodiments and features in the embodiments can be combined with each other. Based on the embodiments in the present disclosure, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present disclosure.

[0054] It should be noted that various aspects of the embodiments within the scope of the appended claims are described below. It should be apparent that the aspects described herein can be embodied in a wide variety of forms, and any specific structure and / or function described herein is merely illustrative. Based on this disclosure, it should be understood by those skilled in the art that an aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number of aspects described herein can be used to implement an apparatus and / or practice a method. In addition, other structures and / or functionalities other than one or more of the aspects described herein can be used to implement this apparatus and / or practice this method.

[0055] It should also be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present disclosure. The illustrations only show components related to the present disclosure and are not drawn according to the number, shape and size of components in actual implementation. In actual implementation, the type, quantity and proportion of each component can be changed at will, and the component layout type may also be more complicated.

[0056] Additionally, in the following description, specific details are provided to provide a thorough understanding of the examples. However, one skilled in the art will appreciate that the aspects described can be practiced without these specific details.

[0057] The embodiments of the present disclosure provide a multimodal scenario risk determination method based on a generative AI large language model, which can be applied to the risk scenario determination process of autonomous driving scenarios.

[0058] See Figure 1, which is a flow chart of a multimodal scenario risk assessment method based on a generative AI large language model provided by an embodiment of the present disclosure. As shown in Figures 1 and 2, the method mainly includes the following steps:

[0059] Step 1: Acquire multimodal data to form a target data set, wherein the multimodal data includes visual data and text data;

[0060] In specific implementation, the system framework for the multimodal scenario risk assessment method based on a generative AI large model is shown in Figure 3. This method systematically collects video and still images from traffic cameras, drones, and vehicle-mounted cameras to capture a multidimensional perspective of the traffic environment. Furthermore, reports from traffic management centers and relevant descriptive text on social media are also collected. This systematic collection of multimodal data, including scene text descriptions, trajectory records, and dynamic recordings, builds a comprehensive dataset, ensuring that the dataset we obtain accurately and meticulously reflects real-world traffic conditions.

[0061] Step 2: Apply the ALBEF algorithm to extract key features corresponding to the target dataset and fuse them into a comprehensive scene representation;

[0062] Based on the above embodiment, step 2 specifically includes:

[0063] In step 2.1, the visual data is divided into patches and then fed into the Transformer model through a linear transformation layer to extract visual features. In addition, the text data is processed through word segmentation and tokenization, and then processed by the text encoder to obtain text features.

[0064] Step 2.2: With image-text contrastive learning as the core, feature alignment is performed by downsampling and normalizing the refined feature vectors to form positive samples of visual features and text features.

[0065] In step 2.3, a multimodal encoder is used to integrate positive samples of visual features and text features, and output a comprehensive scene representation that comprehensively considers various features.

[0066] Furthermore, the step 2.2 includes:

[0067] The normalized visual features and text features of each momentum encoder are denoted as g′ v (v′ cls ) and g′ w (w′ cls ), calculate and optimize the similarity scores between image to text and text to image in a normalized manner based on the softmax function:

[0068] Where s(I, T) = g v (v cls ) T g′ w (w′ cls); s(T, I) = g w (w cls ) T g′ v (v′ cls ), τ represents the learnable temperature parameter, and defines y i2t (I) and y t2i (T) is the similarity of one-shot data of different modalities in the same scene, the probability of negative pair is 0, and the probability of positive pair is 1. The ITC loss is defined as the cross entropy H between p and y;

[0069] Using ITC as the optimization objective during training, related image-text pairs are brought close together in the embedding space, thus obtaining positive samples of visual and text features.

[0070] Furthermore, step 2.3 specifically includes:

[0071] In the multimodal encoder, visual features and textual features are fused through a cross-attention mechanism, with visual features as keys and values ​​and textual features as queries, and vice versa, and then trained on MLM and ITM tasks:

[0072] The model generates a joint feature representation as a comprehensive scene representation: u = MLP([v′ cls , w′ cls ]);

[0073] Among them, v′ cls =CrossAttn(v cls W); w′ cls =CrossAttn(w cls , V).

[0074] Furthermore, the Transformer model includes an encoder layer, multi-head attention, a feedforward fully connected network and a decoder layer, wherein the encoder layer includes a multi-head self-attention mechanism and a feedforward fully connected network, the multi-head attention contains multiple parallel self-attention heads, and the decoder includes multiple similar layers, each layer includes two multi-head self-attention modules and a feedforward fully connected network.

[0075] Furthermore, the specific processing flow of the Transformer model includes:

[0076] The input sequence is converted into an embedding vector, and for each even time step 2i, a sine function is used. To create an element in the position vector, where pos is the position index and d modelis the embedding dimension, i is the dimension index, and for each odd time step 2i+1, the cosine function is used. To create the elements in the position vector, these position vectors are then added element-wise to the corresponding embedding vector to introduce position information;

[0077] The encoder takes as input an embedding vector with position information added, adds the output of the multi-head self-attention mechanism and the feed-forward fully connected network to its input, and then performs layer normalization.

[0078] The input is linearly transformed into three different matrices, which are used to generate the query vector Q, key vector K and value vector V respectively. Subsequently, the dot product is used to calculate the similarity score S = QK between Q and K. T , which represents how each position in the input pays attention to other positions. The score S is scaled by dividing by (d k is the key vector dimension), then apply the softmax function to the scaled scores and multiply the attention weights by the value vector V to obtain the attention function for a set of queries:

[0079] Execute the attention function in parallel on different projection dimensions to obtain the output of multi-head self-attention:

[0080] MultiHead(Q,K,V)=Concat(head1,...,head h )W o

[0081] in, h represents the number of multi-head attention;

[0082] The output of the self-attention module of each encoder layer is first added to the original input to form a residual connection, and then input to the layer normalization. The normalized output is used as the input of the feedforward network. After passing through the linear layer of the feedforward network and the nonlinear activation function ReLU: FFN(x) = max(0, xW1+b1)W2+b2, the network output is again subjected to the residual connection and layer normalization to obtain the final output of the encoder layer.

[0083] The first multi-head self-attention module uses a look-ahead mask, the second multi-head self-attention module uses the output of the encoder layer as the key and value, and the output of the first multi-head self-attention module of the decoder is used as the query. Each multi-head self-attention module output of the decoder is subjected to residual connection and layer normalization, and then input into the feedforward network. Finally, the output of the decoder is converted into the probability distribution of the predicted next word through an additional linear layer and a softmax layer.

[0084] In specific implementation, as shown in FIG4 , step 2 includes the following steps S21-S23:

[0085] S21. Multimodal data feature extraction. To efficiently encode multimodal data, images or video frames are divided into patches of equal size. Each patch is converted into an embedding vector through a linear transformation layer. These embedding vectors constitute the input sequence of the encoder. Since the Transformer infrastructure does not have the ability to capture sequence order, each embedding vector needs to be added with position information. Position encoding is usually generated using fixed sine and cosine functions and then added to the corresponding embedding vector. For each even time step 2i, a sine function is used to create an element in the position vector:

[0086] Where pos is the position index, d model is the embedding dimension and i is the dimension index.

[0087] For each odd time step 2i+1, use the cosine function to create the elements in the position vector:

[0088] These position vectors are added element-wise to the corresponding embedding vectors to incorporate position information. Similarly, text data undergoes word segmentation, is converted into a sequence of tokens, and then mapped to an embedding vector. Subsequently, each modality is processed by a modality-specific encoder.

[0089] As shown in Figure 5, the encoder accepts an embedding vector with position information incorporated into it as input. Each encoder layer consists of two main submodules: a multi-head self-attention mechanism and a feed-forward fully connected network. A residual connection surrounds these two submodules, meaning the output of each submodule is added to its input, followed by layer normalization.

[0090] The multi-head self-attention module contains several parallel self-attention heads. The input is first linearly transformed into three different matrices, which are used to generate the query vector Q, the key vector K, and the value vector V. Then, the dot product is used to calculate the similarity score S = QK between Q and K. T , which represents how each position in the input pays attention to other positions. The score S is scaled by dividing by (d k is the key vector dimension) to avoid gradient shrinkage caused by dot product. Then apply the softmax function to the scaled scores and multiply the attention weights by the value vector V to obtain the attention function for a set of queries:

[0091] Execute the attention function in parallel on different projection dimensions to obtain the output of multi-head self-attention:

[0092] MultiHead(Q,K,V)=Concat(head1,...,head h )W o

[0093] in, h represents the number of multi-head attention.

[0094] The output of the self-attention module in each encoder layer is then added to the original input to form a residual connection and then fed into the layer for normalization. The normalized output serves as the input to the feedforward network, passing through its linear layer and the nonlinear activation function ReLU: FFN(x) = max(0, xW1+b1)W2+b2. The network output then undergoes another residual connection and layer normalization to obtain the final output of the encoder layer.

[0095] The decoder also consists of multiple similar layers, each of which includes two multi-head self-attention modules and a feed-forward fully connected network. The first multi-head self-attention module uses a look-ahead mask to ensure that only the current and previous positions are used when predicting. The second multi-head self-attention module uses the output of the encoder layer as the key and value, and the output of the first multi-head self-attention module of the decoder as the query. Then, similar to the encoder, each multi-head self-attention module output of the decoder is residually connected and layer normalized before being input into the feed-forward network. Finally, the output of the decoder is converted into a probability distribution of the predicted next word through an additional linear layer and a softmax layer.

[0096] Furthermore, through stacked encoder and decoder layers, the model is able to learn the complex structure in the sequence and generate high-quality sequence output. The visual encoder uses the Transrormer framework to encode the input I and generate an embedded sequence {v cls ,v1,.…,v N}, where v cls Represents the global features of the image. The text encoder transforms the input text T into an embedding sequence {w cls w1,…,w M Both features will produce a CLS tag sequence with global information.

[0097] Specifically, the structure and operation steps of the Transformer model are as follows:

[0098] a. Input embedding and position encoding. First, the input sequence is converted into an embedding vector. For each even time step 2i, a sine function is used. To create an element in the position vector, where pos is the position index and d model is the embedding dimension, i is the dimension index. For each odd time step 2i+1, use the cosine function to create the elements in the position vector. These position vectors are then added element-wise to the corresponding embedding vector to introduce position information.

[0099] b. Encoder layer. The encoder accepts as input an embedding vector that incorporates position information. Each encoder layer consists of two main submodules: a multi-head self-attention mechanism and a feed-forward fully connected network. A residual connection surrounds these two submodules, meaning the output of each submodule is added to its input, followed by layer normalization.

[0100] c. Multi-head self-attention. This module contains several parallel self-attention heads. The input is first linearly transformed into three different matrices, which are used to generate the query vector Q, the key vector K, and the value vector V. Subsequently, the dot product is used to calculate the similarity score S = QK between Q and K. T , which represents how each position in the input pays attention to other positions. The score S is scaled by dividing by (d k is the key vector dimension) to avoid gradient shrinkage caused by dot product. Then apply the softmax function to the scaled scores and multiply the attention weights by the value vector V to obtain the attention function for a set of queries:

[0101] Execute the attention function in parallel on different projection dimensions to obtain the output of multi-head self-attention:

[0102] MultiHead(Q,K,V)=Concat(head1,...,head h )W o

[0103] in, h represents the number of multi-head attention.

[0104] d. Feedforward fully connected network. The self-attention module output of each encoder layer is first added to the original input to form a residual connection, and then input to the layer normalization. The normalized output serves as the input of the feedforward network. It passes through the linear layer of the feedforward network and the nonlinear activation function ReLU: FFN(x) = max(0, xW1+b1)W2+b2. The network output is again subjected to a residual connection and layer normalization to obtain the final output of the encoder layer.

[0105] e. Decoder layer. The decoder also consists of multiple similar layers, each of which includes two multi-head self-attention modules and a feed-forward fully connected network. The first multi-head self-attention module uses a look-ahead mask to ensure that only the current and previous positions are used when predicting. The second multi-head self-attention module uses the output of the encoder layer as the key and value, and the output of the first multi-head self-attention module of the decoder as the query. Then, similar to the encoder, each multi-head self-attention module output of the decoder is residually connected and layer normalized before being input into the feed-forward network. Finally, the output of the decoder is converted into a probability distribution of the predicted next word through an additional linear layer and a softmax layer.

[0106] Furthermore, through stacked encoder and decoder layers, the model is able to learn complex structures in sequences and generate high-quality text sequence outputs.

[0107] S22. When performing feature alignment, by implementing intra-batch and cross-batch comparative learning strategies, the generalization ability of the model is enhanced, and it can more accurately distinguish between positive and negative sample pairs.

[0108] The feature vector is refined through downsampling and normalization, and positive samples are formed. This stage focuses on image-text contrast learning, aiming to optimize the model to better capture the mutual correlation between single modalities and improve semantic level alignment. It learns a similarity function s = g v (v cls ) T gw (w cls ), so that the parallel image-text pairs have higher similarity scores. Where, g v and g w It is the low-dimensional representation of [CLS] embedded in the mapping after linear transformation.

[0109] The normalized features of each momentum encoder are recorded as g′ v (v′ cls ) and g′ w (w′ cls ), calculate and optimize the similarity scores between image to text and text to image in a normalized manner based on the softmax function:

[0110] Where s(I, T) = g v (v cls ) T g′ w (w′ cls ); s(T, I) = g w (w cls ) T g′ v (v′ cls ); τ represents the learnable temperature parameter. Define y i2t (I) and y t2i (T) is the similarity of one-shot data of different modalities in the same scene, the probability of negative pair is 0, and the probability of positive pair is 1. The ITC loss is defined as the cross entropy H between p and y:

[0111] Using ITC as the optimization objective during training forces related image-text pairs to be close in the embedding space, while unrelated pairs are kept far away.

[0112] S23, the multimodal encoder is responsible for integrating visual features and text features in the feature fusion stage, and outputting a unified scene representation that comprehensively considers various features.

[0113] In the multimodal encoder, visual and textual features are fused through a cross-attention mechanism. Visual features serve as keys and values, and textual features serve as queries, and vice versa, enabling the encoder to focus on relevant features. Subsequently, after training on MLM and ITM tasks:

[0114] The model generates a joint feature representation: u = MLP([v′ cls , w′ cls ]);

[0115] Among them, v′ cls =CrossAttn(c cls W); w′ cls =CrossAttn(w cls , V). It provides a powerful feature base for downstream tasks and provides richer semantic information for scene understanding.

[0116] Step 3: Evaluate the risk level corresponding to the comprehensive scenario representation based on the preset safety indicators and the large language model and compare it with the risk threshold to determine whether the scenario corresponding to the comprehensive scenario representation is a high-risk scenario.

[0117] Based on the above embodiment, step 3 specifically includes:

[0118] Step 3.1, setting a time-based indicator, a deceleration-based indicator, and an energy-based indicator to form a preset safety indicator;

[0119] In step 3.2, the large language model is used to evaluate the risk level corresponding to the comprehensive scenario according to the preset safety indicators and compare it with the preset risk threshold. The high-risk scenarios are monitored and identified by combining the text prompt and the preset evaluation method.

[0120] Furthermore, the expression of the time-based indicator is

[0121] Where X represents the vehicle position (i represents the following vehicle, i-1 represents the leading vehicle), l represents the vehicle length, and v represents the speed.

[0122] Furthermore, the expression of the deceleration-based index is:

[0123] Where t represents the time interval;

[0124] Where MADR represents the maximum braking rate, Δt represents the time step, and t i and tf i denote the initial and final time steps, T i is the total travel time;

[0125] Among them, SSD L and SSD F Represents the stopping distance of the front vehicle and the rear vehicle, v F and v L Represent the speed of the moving vehicle and the preceding vehicle, t d represents the delay time, S represents the gap distance, d m Indicates the maximum deceleration.

[0126] Furthermore, the expression of the energy-based index is

[0127] Among them, v1 and v2 represent the velocities of two vehicles with potential collision trajectories before the collision, and m1 and m2 represent the masses of the two vehicles.

[0128] In specific implementation, step 3 includes the following steps S31-S32:

[0129] S31. Using the features obtained from the multimodal data fusion process and based on multiple perspectives of safety assessment, this embodiment selects three types of SSM: "time-based", "deceleration-based" and "energy-based".

[0130] Among them, the time-based SSM uses the most common Time-to-Collision (TTC) to determine the time when the two vehicles collide while maintaining the collision path and speed difference:

[0131] Where X represents the vehicle position (i is the following vehicle, i-1 is the leading vehicle); l represents the vehicle length; and v represents the speed.

[0132] The deceleration-based SSM uses the Deceleration Rate to Avoid the Crash (DRAC) to calculate the minimum braking rate required for a vehicle to avoid a collision with another vehicle:

[0133] Here, t represents the time interval.

[0134] Use Crash Potential Index (CPI), taking into account the vehicle's braking ability and maximum deceleration rate:

[0135] Where MADR represents the maximum braking rate; Δt represents the time step; t i and tf i represent the initial and final time steps respectively; T i is the total travel time.

[0136] Distance-based SSM can also be thought of as deceleration-based SSM. The Rear-end Collision Risk Index (RCRI) identifies dangerous situations by comparing the stopping distances of the leading and following vehicles:

[0137] Among them, SSD 占 and SSD F Represents the stopping distance of the front vehicle and the rear vehicle, v F and v L Represent the speed of the moving vehicle and the preceding vehicle respectively; t d represents the delay time; S represents the gap distance; d m Indicates the maximum deceleration.

[0138] The energy-based SSM uses DeltaV to measure the speed change of road users due to collisions:

[0139] Among them, v1 and v2 represent the velocities of two vehicles with potential collision trajectories before the collision; m1 and m2 represent the masses of the two vehicles.

[0140] S32. Perform risk assessment analysis and import the feature values ​​obtained from multimodal fusion into the fine-tuned ChatGPT. Combined with the text prompt, specify an evaluation method (such as fuzzy comprehensive evaluation method, entropy weight method, etc., which are not specified in this invention) to monitor and identify high-risk scenarios.

[0141] Prompt in the example: (Procedural memory) You are a helpful assistant RoadSafety_Gpt (SSM threshold extraction comparison) Analyze the following surrogate safety measures extracted from a multimodal traffic scenario: [List of SSM features]. Compare each measure with the safety thresholds stored in the database and indicate any values ​​that exceed these thresholds. (Key SSM) The extracted Time-to-Collision (TTC) value is [x] seconds, and the Deceleration Rate is[y]m / s 2 .Compare these values ​​with the safety thresholds:TTC should not be less than[threshold TTC]seconds and Deceleration should not exceed[threshold Deceleration]m / s 2.Highlight any measurements that are outside safe limits. (Scenario Risk Evaluation) Apply an evaluation method to assess the risk level of the traffic scenario based on the previous comparisons. Use the [chosen evaluation method name, eg, Fuzzy Comprehensive Evaluation or Entropy Weight Method] to calculate the overall risk score and provide a brief explanation of your assessment. (Judgment) Based on the overall risk score obtained from the evaluation, classify the traffic scenario into low, medium, or high-risk categories.

[0142] The multimodal scenario risk assessment method based on the generative AI large language model provided in this embodiment, by combining text, image and video data, based on the Transformer algorithm and the large language model, the present invention makes full use of various types of data and improves the system's understanding of real-world scenarios. This multimodal method can capture rich information that a single data source cannot provide, thereby providing a more comprehensive perspective for risk assessment. At the same time, by calling the trained large language model, through pre-set SSM indicators and real-time monitoring, the present invention can quickly identify and mark high-risk scenarios, provide a basis for taking emergency measures, and thus enhance real-time response capabilities.

[0143] It should be understood that various parts of the present disclosure can be implemented in hardware, software, firmware, or a combination thereof.

[0144] The above description is merely a specific embodiment of the present disclosure, but the scope of protection of the present disclosure is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this disclosure should be included in the scope of protection of the present disclosure. Therefore, the scope of protection of the present disclosure should be based on the scope of protection of the claims.

Claims

1. A multimodal scenario risk assessment method based on a generative AI large language model, characterized by: include: Step 1: Acquire multimodal data to form a target data set, wherein the multimodal data includes visual data and text data; Step 2: Apply the ALBEF algorithm to extract key features corresponding to the target dataset and fuse them into a comprehensive scene representation; Step 3: Evaluate the risk level corresponding to the comprehensive scenario representation based on the preset safety indicators and the large language model and compare it with the risk threshold to determine whether the scenario corresponding to the comprehensive scenario representation is a high-risk scenario.

2. The method according to claim 1, characterized in that , the step 2 specifically includes: In step 2.1, the visual data is divided into patches and then fed into the Transformer model through a linear transformation layer to extract visual features. In addition, the text data is processed through word segmentation and tokenization, and then processed by the text encoder to obtain text features. Step 2.2: With image-text contrastive learning as the core, feature alignment is performed by downsampling and normalizing the refined feature vectors to form positive samples of visual features and text features. In step 2.3, a multimodal encoder is used to integrate positive samples of visual features and text features, and output a comprehensive scene representation that comprehensively considers various features.

3. The method according to claim 2, characterized in that , the step 2.2 includes: The normalized visual features and text features of each momentum encoder are denoted as g′ v (v′ cls ) and g′ w (w′ cls ), calculate and optimize the similarity scores between image to text and text to image in a normalized manner based on the softmax function: Where s(I, T) = g v (v cls ) T g′ w (w′ cls ); s(T, I) = g w (w cls ) T g′ v (v′ cls ), τ represents the learnable temperature Degree parameter, defining y i2t (I) and y t2i (T) is the similarity of one-shot data of different modalities in the same scene, the probability of negative pair is 0, and the probability of positive pair is 1. The ITC loss is defined as the cross entropy H between p and y; Using ITC as the optimization objective during training, related image-text pairs are brought close together in the embedding space, thus obtaining positive samples of visual and text features.

4. The method according to claim 3, characterized in that Step 2.3 specifically includes: In the multimodal encoder, visual features and textual features are fused through a cross-attention mechanism, with visual features as keys and values ​​and textual features as queries, and vice versa, and then trained on MLM and ITM tasks: The model generates a joint feature representation as a comprehensive scene representation: u = MLP([v′ cls , w′ cls ]); where, v′ cls = CrossAttn(v cls , W); w′ cls = CrossAttn(w cls , V).

5. The method according to claim 4, characterized in that The step 3 specifically includes: Step 3.1, setting a time-based indicator, a deceleration-based indicator, and an energy-based indicator to form a preset safety indicator; In step 3.2, the large language model is used to evaluate the risk level corresponding to the comprehensive scenario according to the preset safety indicators and compare it with the preset risk threshold. The high-risk scenarios are monitored and identified by combining the text prompt and the preset evaluation method.

6. The method according to claim 5, characterized in that The Transformer model includes an encoder layer, a multi-head attention, a feedforward fully connected network and a decoder layer, wherein the encoder layer includes a multi-head self-attention mechanism and a feedforward fully connected network, the multi-head attention includes multiple parallel self-attention heads, and the decoder includes multiple similar layers, each layer includes two multi-head self-attention modules and a feedforward fully connected network.

7. The method according to claim 6, characterized in that The specific processing flow of the Transformer model includes: The input sequence is converted into an embedding vector, and for each even time step 2i, a sine function is used. To create an element in the position vector, where pos is the position index and d model is the embedding dimension, i is the dimension index, and for each odd time step 2i+1, cosine function To create the elements in the position vector, these position vectors are then added element-wise to the corresponding embedding vector to introduce position information; The encoder takes as input an embedding vector with position information added, adds the output of the multi-head self-attention mechanism and the feed-forward fully connected network to its input, and then performs layer normalization. The input is linearly transformed into three different matrices, which are used to generate the query vector Q, key vector K and value vector V respectively. Subsequently, the dot product is used to calculate the similarity score S = QK between Q and K. T , which represents how each position in the input pays attention to other positions. The score S is scaled by dividing by (d k is the key vector dimension), then apply the softmax function to the scaled scores and multiply the attention weights by the value vector V to obtain the attention function for a set of queries: Execute the attention function in parallel on different projection dimensions to obtain the output of multi-head self-attention: MultiHead(Q,K,V)=Concat(head1,…,head h )W O in, h represents the number of multi-head attention; The output of the self-attention module of each encoder layer is first added to the original input to form a residual connection, and then input to the layer normalization. The normalized output is used as the input of the feedforward network. After passing through the linear layer of the feedforward network and the nonlinear activation function ReLU: FFN(x) = max(0, xW1+b1)W2+b2, the network output is again subjected to the residual connection and layer normalization to obtain the final output of the encoder layer. The first multi-head self-attention module uses a look-ahead mask, the second multi-head self-attention module uses the output of the encoder layer as the key and value, and the output of the first multi-head self-attention module of the decoder is used as the query. Each multi-head self-attention module output of the decoder is subjected to residual connection and layer normalization, and then input into the feedforward network. Finally, the output of the decoder is converted into the probability distribution of the predicted next word through an additional linear layer and a softmax layer.

8. The method according to claim 5, characterized in that , the expression of the time-based indicator is Where X represents the vehicle position (i represents the following vehicle, i-1 represents the leading vehicle), l represents the vehicle length, and v represents the speed.

9. The method according to claim 5, characterized in that , the expression of the deceleration-based index is Where t represents the time interval; Where MADR represents the maximum braking rate, Δt represents the time step, and t i and tf i denote the initial and final time steps, T i is the total travel time; Among them, SSD L and SSD F Represents the stopping distance of the front vehicle and the rear vehicle, v F and v L Represent the speed of the moving vehicle and the preceding vehicle, t d represents the delay time, S represents the gap distance, d m Indicates the maximum deceleration.

10. The method according to claim 5, characterized in that , the expression of the energy-based index is Among them, v1 and v2 represent the velocities of two vehicles with potential collision trajectories before the collision, and m1 and m2 represent the masses of the two vehicles.