Driver risk identification method and apparatus, terminal device and readable storage medium

By integrating video and audio features with a multimodal large language model, the problem of insufficient accuracy in driver risk identification in existing technologies is solved, enabling more accurate driver risk identification and safe driving intervention, thereby improving road traffic safety.

WO2025217803A1PCT designated stage Publication Date: 2025-10-23SHENZHEN STREAMING VIDEO TECH

Patent Information

Application Number
PCT/CN2024/087986
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-04-16
Publication Date
2025-10-23

AI Technical Summary

Technical Problem

Existing driver risk identification methods based on multimodal models have small parameters, simple structure, and limited multimodal feature alignment. They cannot accurately reflect the characteristics and needs of drivers under different driving conditions, resulting in insufficient recognition accuracy.

Method used

A driver risk identification method based on a multimodal large language model is adopted. By combining a video encoder, an audio encoder, a video adapter, an audio adapter, and a large language model, the powerful understanding and reasoning capabilities of the large language model are utilized to fuse the multimodal features of video and audio, thereby achieving accurate identification of driver risks.

Benefits of technology

It improves the accuracy and reliability of driver risk identification, can better identify the driver's risk characteristics and behavior patterns, promote safe driving, and improve road traffic safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024087986_23102025_PF_FP_ABST
    Figure CN2024087986_23102025_PF_FP_ABST
Patent Text Reader

Abstract

The present application is applicable to the technical field of vehicle safety, and provides a driver risk identification method and apparatus, a terminal device and a readable storage medium. The driver risk identification method comprises: acquiring a target video feature vector of video information of a driver, wherein the target video feature vector is aligned with an embedded space of a large language model; acquiring a target audio feature vector of audio information in a vehicle driven by the driver, wherein the target audio feature vector is aligned with the embedded space of the large language model; and inputting the target video feature vector and the target audio feature vector into the large language model to obtain a risk identification result of the driver. The present application can achieve driver risk identification.
Need to check novelty before this filing date? Find Prior Art

Description

Driver risk identification method and device, terminal equipment and readable storage medium TECHNICAL FIELD

[0001] The application belongs to the technical field of vehicle safety, and particularly relates to a driver risk identification method and device, terminal equipment and readable storage medium. BACKGROUND

[0002] With the increase of traffic flow and the increasing importance of road safety, driver risk identification has become an important link to ensure road safety. The risk identification scheme aims to accurately identify and respond to various potential risk factors during the driving process of the driver, such as driver distraction and fatigue. Accurate risk identification can help the driver take appropriate measures in time to avoid accidents and improve road safety. TECHNICAL PROBLEM

[0003] Embodiments of the application provide a driver risk identification method, device, terminal equipment and readable storage medium to realize driver risk identification. TECHNICAL SOLUTION

[0004] In a first aspect, the embodiments of the application provide a driver risk identification method, which comprises:

[0005] obtaining a target video feature vector of video information of a driver, the target video feature vector being aligned with an embedding space of a large language model;

[0006] obtaining a target audio feature vector of audio information in a vehicle driven by the driver, the target audio feature vector being aligned with the embedding space of the large language model;

[0007] inputting the target video feature vector and the target audio feature vector into the large language model to obtain a risk identification result of the driver.

[0008] In a second aspect, the embodiments of the application provide a driver risk identification device, which comprises:

[0009] a first obtaining module configured to obtain a target video feature vector of video information of a driver, the target video feature vector being aligned with an embedding space of a large language model;

[0010] a second obtaining module configured to obtain a target audio feature vector of audio information in a vehicle driven by the driver, the target audio feature vector being aligned with the embedding space of the large language model;

[0011] a risk identification module configured to input the target video feature vector and the target audio feature vector into the large language model to obtain a risk identification result of the driver.

[0012] In a third aspect, an embodiment of the present application provides a terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the steps of the driver risk identification method according to the first aspect when executing the computer program.

[0013] In a fourth aspect, an embodiment of the present application provides a computer readable storage medium, which stores a computer program, and the computer program, when executed by a processor, implements the steps of the driver risk identification method according to the first aspect.

[0014] In a fifth aspect, an embodiment of the present application provides a computer program product, which, when executed on a terminal device, causes the terminal device to perform the steps of the driver risk identification method according to the first aspect. Advantages

[0015] As can be seen from the above, the target video feature vector of the video information of the driver and the target audio feature vector of the audio information of the driver are obtained, and the target video feature vector and the target audio feature vector are aligned with the embedding space of the large language model, so that the target video feature vector and the target audio feature vector are input into the large language model to obtain the risk identification result of the driver. Based on the multi-modal features such as the target video feature vector and the target audio feature vector, and by using the extraordinary understanding and reasoning ability of the large language model, the driving behavior of the driver can be more accurately identified, that is, the driver risk identification is realized. BRIEF DESCRIPTION OF DRAWINGS

[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiments or prior art description will be briefly introduced as follows. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0017] FIG. 1 is an implementation flow diagram of a driver risk identification method according to an embodiment of the present application;

[0018] FIG. 2 is a processing flow example diagram of driver risk identification based on a multi-modal large language model;

[0019] FIG. 3 is a structural schematic diagram of a driver risk identification device according to an embodiment of the present application;

[0020] FIG. 4 is a structural schematic diagram of a terminal device according to an embodiment of the present application. Embodiments of the present application

[0021] In the following description, for purposes of explanation and not limitation, specific details are set forth such as particular architectures, techniques, etc. in order to provide a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application can be practiced in other embodiments that depart from these specific details. In other instances, detailed descriptions of well-known methods, devices, circuits, and

[0022] It is to be understood that the terminology "includes", "has", "holds", "contains" and / or "comprising", "including", "having" and their conjugates, as used herein, means "including but not limited to", and not to the exclusion of any other term or aspect.

[0023] It is also to be understood that the terminology "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items, as well as to the

[0024] As used in the description of the application and the appended claims, the term "if' can be interpreted to mean "when" or "upon" or "in response to determining" or "in response to detecting" depending on the context. Similarly, the phrase "if it is determined" or "if [a described condition or event] is detected" can be interpreted to mean "upon determining" or "in response to determining" or "upon detecting [the described condition or event]" or "in response to detecting [the described condition or event]" depending on the context.

[0025] In addition, the terms "first", "second", "third", etc. as used in the description of embodiments herein and throughout the claims (if any), are not used to connote any relative importance but are simply used to distinguish one element from another.

[0026] Reference throughout this specification to "one embodiment", "an embodiment", or "a specific embodiment", means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the present application. Thus, appearances of the phrases "in one embodiment", "in an embodiment", "in some embodiments", "in other embodiments", "in still other embodiments", and the like in the specification throughout, are not necessarily all referring to the same embodiment, unless otherwise noted or apparent from the context. The terms "including", "containing", "comprising", "having" and variations thereof herein are meant to be broad and encompass the terms "consisting of" and "consisting essentially of" unless otherwise noted.

[0027] Before the schemes of the present application are described, the nomenclature involved in the present schemes is explained for the convenience of the reader.

[0028] Currently, the driver risk identification method based on the multi-modal model has certain limitations. Such models have small parameter quantity, simple structure, limited multi-modal feature alignment, insufficient learned representation information, and insufficient generalization ability, and cannot capture complex data patterns and relationships. Therefore, the multi-modal model has certain limitations for driver risk identification, and cannot accurately reflect the characteristics and needs of drivers in different driving states. To solve this problem, the present application provides a driver risk identification method based on a multi-modal large language model. This method uses a powerful large language model as a brain to perform multi-modal tasks, which can better exploit multi-modal information, improve the effective capture of global information by the model, and enhance risk identification in different driving states. By adopting the multi-modal large language model scheme, the present application can more accurately identify the risk characteristics and behavior patterns of drivers, thereby better predicting and intervening in potential risk situations. This risk identification method based on a multi-modal large language model realizes more robust driver risk identification, guides drivers to drive safely, and promotes road traffic safety.

[0029] The driver risk identification method based on a multi-modal large language model provided by the present application has the following scientific, technical, and industrial values:

[0030] In view of the current situation that the precision identification and intelligent research and judgment ability of the driver risk identification method are insufficient, and effective prevention and control technology and equipment for high-risk traffic behaviors are lacking, the multi-modal information of drivers is focused as the research emphasis, and a mechanism, key technology, and integrated application closed-loop research framework is constructed. The risk monitoring and identification key technology research based on a multi-modal large language model is forward-looking, and has significant scientific value.

[0031] The following social benefits are provided:

[0032] Although the total number of traffic accidents and the number of deaths in China have been declining in recent years, both the death rate per 10,000 vehicles and the death rate per 1,000,000 kilometers are still much lower than those in developed countries. Traffic accidents account for a very high proportion of the total number of abnormal deaths in China, and have become one of the important unstable factors affecting social harmony and family happiness in China. In view of this, on October 9, 2016, the State Council Work Safety Office issued the “Guidelines for Implementing the Work of Controlling Major Accidents and Building a Dual Prevention Mechanism”, on December 9, 2016, the Central Committee of the Communist Party of China and the State Council issued the “Opinions on Promoting Reform and Development in the Field of Work Safety”, and in the same year, the Ministry of Transport issued the “Implementation Plan for Building a Dual Prevention System of Risk Classification and Control and Hidden Danger Governance in the Field of Work Safety in the Field of Work Safety”. Therefore, how to effectively monitor and control the risk of drivers has become one of the difficulties that need to be solved in the field of traffic safety in China at the present stage, and one of the key points for effectively improving the level of traffic safety.

[0033] The multimodal large language model in the present application aims to understand the visual content in the video information and the auditory content in the audio information using a large language model. Before using the multimodal large language model, a multimodal dataset for driver risk identification needs to be constructed and trained based on the designed multimodal large language model framework. Specifically, the multimodal large language model contains five components, namely, a video encoder, an audio encoder, a video adapter, an audio adapter, and a large language model.

[0034] Video encoder: In real application scenarios, traditional image encoders can only focus on the information in a single image and cannot fully perceive the driver state by utilizing the contextual information between the previous and subsequent frames in the video information. At present, there are many devices that support video recording during driving, such as Advanced Driving Assistance System (ADAS) cameras and Driver Monitor System (DMS) cameras, etc. Therefore, using a video encoder can better learn the state changes of the driver, identify the risks and intervene, and utilize the contextual information to identify some complex driving actions and scenarios. Therefore, driver risk identification algorithms based on video understanding will become a new research hotspot. For video-related tasks, the introduction of a video encoder can better utilize temporal and contextual information, improving the performance and efficiency of the model.

[0035] The video encoder can adopt the pre-trained visual module in BLIP-2, which consists of a ViT (Vision Transformer) and a Q-former. During training, the parameters are frozen. Assuming that a video information contains N frames of images, after passing through the video encoder, N two-dimensional embedding vectors, i.e., N video feature vectors, will be obtained, and N is an integer greater than 1.

[0036] Video adapter: Since the video feature vectors from the visual encoder do not consider any temporal information, the position embedding is further used as temporal information and applied to the representations from different frames. Then, the video feature vectors with position encoding information are input into a video-specific Q-former to obtain video encoding vectors. In addition, in order to adapt the video representation to the input of the large language model, a linear layer is added to convert the video encoding vectors into video query vectors, realizing the alignment of the video feature vectors and the language model space.

[0037] Audio encoder: In order to process the given audio information, an audio encoder is designed. Specifically, a pre-trained Imagebind is used as the audio encoder. First, K audio segments of 2 seconds are uniformly sampled from the audio information, K being an integer greater than 1, and then each audio segment is converted into a spectrogram using a mel-spectrogram transform, and the obtained spectrogram is input into the audio encoder to obtain the corresponding audio feature vector.

[0038] Audio adapter: Similar to the video adapter, the position encoding information representing the time information is first embedded into the audio feature vector from the audio encoder, and then the audio feature vector with the position encoding information is input into an audio-specific Q-former to obtain an audio encoding vector. Next, a linear layer is used to convert the audio encoding vector into an audio query vector, which is mapped into the embedding space of the large language model, realizing the alignment of the audio feature vector and the large language model space.

[0039] Large language model: The video query vector and the audio query vector are input into the large language model, and the model outputs a text description, which represents the risk identification result of the driver.

[0040] Before using the multi-modal large language model for driver risk identification, the video-related components and the audio-related components in the multi-modal large language model need to be separated and trained individually. Their training processes are basically the same, first pre-training them using large-scale data, and then fine-tuning them using high-quality instruction data sets to realize visual-linguistic alignment and auditory-linguistic alignment. After fine-tuning the video-related components and the audio-related components, the training of the multi-modal large language model is completed, and the trained multi-modal large language model is obtained.

[0041] It should be noted that the pre-training method of the multi-modal large language model is not limited in the present application. For example, a large-scale video caption data set can be used to pre-train the video-related components, and an audio text data set can be used to pre-train the audio-related components.

[0042] The trained multi-modal large language model is deployed to a terminal device, and the multi-modal large language model is processed in 8-bits, i.e., the number of bits in the operation process of the multi-modal large language model is reduced to 8 bits, to speed up the inference.

[0043] By applying the multi-modal large language model to the field of video understanding, the multi-modal features of video and audio are fused, and the extraordinary understanding and reasoning ability of the large language model is utilized, so that the multi-modal large language model can comprehensively perceive the state of the driver and more accurately identify the risk features and behavior patterns of the driver. The accuracy far exceeds that of traditional risk identification algorithms based on single-frame images, and the risk identification result conforms to human perception, and is more reliable and interpretable.

[0044] It should be understood that the size of the serial number of each step in the embodiment does not mean the order of execution, and the execution order of each process should be determined according to its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiment of the application.

[0045] In order to illustrate the technical solutions described in the present application, the following will be described by specific embodiments.

[0046] Referring to FIG. 1, it is an implementation flow diagram of a driver risk identification method provided by an embodiment of the present application, and the driver risk identification method is applied to a terminal device. As shown in FIG. 1, the driver risk identification method can include the following steps:

[0047] Step 101, obtaining a target video feature vector of video information of a driver.

[0048] The target video feature vector is a video feature vector compatible with the text input of the large language model, that is, the target video feature vector has the same dimension as the text input of the large language model. The target video feature vector is aligned with the embedding space of the large language model, which ensures high accuracy and understanding, so that the large language model can generate meaningful responses according to the video information.

[0049] The video information of the driver includes N frames of images, which can be a face image of the driver, or an image including the face and other body parts of the driver. N is an integer greater than 1.

[0050] Optionally, a video monitoring device located above the driving area can be used to collect the video information of the driver, and the video monitoring device sends the collected video information to the terminal device.

[0051] The terminal device is deployed with a trained multi-modal large language model, and the target video feature vector can be obtained through a video encoder and a video adapter in the multi-modal large language model. Specifically, the video information is input into the video encoder to obtain N first video feature vectors, each frame of image corresponding to a first video feature vector; the N first video feature vectors are input into the video adapter to obtain a target video feature vector.

[0052] The first video feature vector is a video feature vector output by the video encoder. By inputting the video information into the video encoder, the first video feature vector of each frame of image in the video information can be obtained. The first video feature vector is a two-dimensional embedding vector (i.e., a two-dimensional video feature vector).

[0053] The video encoder can better utilize the context information in the video information, and improve the performance and efficiency of the model.

[0054] There are many types of video monitoring devices (such as ADAS cameras, DMS cameras, etc.) installed on vehicles, and these video monitoring devices can all obtain video information of the driver during the driver's driving process. Therefore, using a video encoder can better learn the state changes of the driver (such as the driver changing from open eyes to closed eyes), thereby identifying the risks, and using contextual information to identify some complex driving actions and scenes, etc.

[0055] Of course, it can be understood that the pre-trained multi-modal large language model can also be deployed on the terminal device, and the pre-trained multi-modal large language model can be fine-tuned through the terminal device. Specifically, the video processing component and the audio processing component in the multi-modal large language model are fine-tuned separately.

[0056] In an optional embodiment, the fine-tuning scheme of the video processing component is as follows:

[0057] In the case of freezing the weight parameters of the video encoder and the large language model, based on the first instruction data set and the second instruction data set, the video adapter is fine-tuned to align the video features with the language features (i.e., visual language alignment), the first instruction data set includes a plurality of first video samples and a video description of each first video sample, and the second instruction data set includes a plurality of second video samples and a question and answer pair of each second video sample;

[0058] In the case of freezing the weight parameters of the video encoder, based on the first instruction data set and the second instruction data set, the video adapter and the large language model are fine-tuned at the same time to improve the multi-modal expression capability.

[0059] The terminal device can obtain a large amount of video information of different drivers driving vehicles, and one video information is a video sample. The obtained large amount of video samples are divided into a plurality of first video samples and a plurality of second video samples. The plurality of first video samples and the plurality of video samples can contain the same video samples or different video samples, which are not limited by the present application. The plurality refers to at least two.

[0060] The generation process of the first instruction data set is as follows:

[0061] First, the key frame of each first video sample is extracted; second, the obtained key frame and a randomly selected title are input into GPT4V to obtain a description of the key frame; and finally, each key frame, the title of the key frame, and the description are input into GPT4V, and GPT4V outputs a concise and spatiotemporal related video description. The video description is used to describe the corresponding first video sample. All the above first video samples and their video descriptions constitute the first instruction data set.

[0062] For example, one of the three titles below can be randomly selected to input into GPT4V.

[0063] Title 1: This is an infrared image of the driver. Briefly describe this image. Focus on the driver's facial expressions. Don't pay attention to anyone in the car other than the driver. Do not pay attention to the environment and objects in the car. Use a complete paragraph less than 100 words for your reply.

[0064] Title 2: This is an infrared image of the driver. Provide a concise depiction of this image. Focus on the driver's facial expressions. Don't pay attention to anyone in the car other than the driver. Do not pay attention to the environment and objects in the car. Use a complete paragraph less than 100 words for your reply.

[0065] Title 3: This is an infrared image of the driver. Summarize this image less than 100 words. Focus on the driver's facial expressions. Don't pay attention to anyone in the car other than the driver. Do not pay attention to the environment and objects in the car.

[0066] The generation process of the second instruction dataset is as follows:

[0067] For each second video sample, a question-answer pair is constructed, one question-answer pair includes one question and the corresponding answer of the question, and the answer will be labeled by a person watching the second video sample. For example, the question is "Whether the driver looks tired?" and the answer is "yes / no". All the above second video samples and their question-answer pairs constitute the second instruction dataset.

[0068] In an optional embodiment, the video adapter includes a first position embedding layer, a first Q-former and a first linear layer; N first video feature vectors are input into the video adapter to obtain a target video feature vector, including:

[0069] The N first video feature vectors are input into the first position embedding layer to obtain N second video feature vectors;

[0070] The N second video feature vectors are input into the first Q-former to obtain a third video feature vector;

[0071] The third video feature vector is input into the first linear layer to obtain the target video feature vector.

[0072] Since the first video feature vectors from the video encoder do not consider time information, after the first video feature vectors are input into the video adapter, the first position embedding layer in the video adapter can apply time information as position encoding information to representations from different frames, thereby obtaining second video feature vectors, which carry position encoding information. The number of second video feature vectors is the same as that of first video feature vectors, and the second video feature vectors increase position encoding information compared with the first video feature vectors.

[0073] The first Q-former is a video-specific Q-former for aggregating frame-level representations, aggregating N second video feature vectors into a third video feature vector. The third video feature vector is a one-dimensional video feature vector.

[0074] In order to adapt the third video feature vector to the input of the large language model, a linear layer (i.e., the first linear layer) is added to convert the third video feature vector into a target video feature vector, so as to project the third video feature vector into the same dimension as the text embedding of the large language model, and realize the alignment of the video feature and the space of the large language model.

[0075] Step 102, obtaining a target audio feature vector of audio information in a vehicle driven by a driver.

[0076] The target audio feature vector is an audio feature vector compatible with the text input of the large language model, that is, the target audio feature vector has the same dimension as the text input of the large language model. The target audio feature vector is aligned with the embedding space of the large language model, and this alignment ensures high accuracy and understanding, enabling the large language model to generate meaningful responses based on audio information.

[0077] The audio information in the vehicle can be audio information of all persons in the vehicle or audio information of the driver, which is not limited in the present application.

[0078] Optionally, an audio monitoring device (such as a microphone, an audio recorder, etc.) located in the vehicle can be used to collect the audio information in the vehicle, and the audio monitoring device sends the collected audio information to the terminal device. For example, an audio monitoring device located in front of the driving area is used to collect the video information of the driver.

[0079] For devices capable of recording video and audio simultaneously (i.e., video and audio recording devices), the video and audio recording device can be installed in the vehicle, and the installation position needs to ensure that the video and audio recording device can capture the driver. The audio and video information of the driver can be collected by the video and audio recording device, which includes the video information of the driver and the audio information in the vehicle.

[0080] It should be noted that in order to improve the accuracy of driver risk identification, the video information of the driver and the audio information in the vehicle driven by the driver need to be collected simultaneously, that is, the collection time of the video information is the same as the collection time of the audio information.

[0081] The terminal device is deployed with a trained multi-modal large language model, so that the target audio feature vector can be obtained through the audio encoder and the audio adapter in the multi-modal large language model. Specifically, K audio segments are sampled from the audio information, K is an integer greater than 1; the K audio segments are converted into K spectrograms respectively; the K spectrograms are input into the audio encoder to obtain K first audio feature vectors, one spectrogram corresponding to one first audio feature vector; the K first audio feature vectors are input into the audio adapter to obtain one target audio feature vector.

[0082] The first audio feature vector is an audio feature vector output by the audio encoder, and the first audio feature vector is a two-dimensional audio feature vector. The terminal device can uniformly sample K audio segments from the audio information, and the audio duration of each audio segment is a preset duration (e.g., 2 seconds). Then, the K audio segments are converted into corresponding spectrograms using the Mel-frequency spectrum transform to obtain K spectrograms. Each spectrogram is input into the audio encoder to obtain the corresponding first audio feature vector, and K first audio feature vectors are obtained.

[0083] In the case where the pre-trained multi-modal large language model is deployed on the terminal device, the fine-tuning scheme of the audio processing component is as follows:

[0084] In the case where the weight parameters of the audio encoder and the large language model are frozen, the audio adapter is fine-tuned based on the third instruction data set and the fourth instruction data set to align the audio features with the language features (i.e., auditory language alignment), the third instruction data set includes a plurality of first audio samples and an audio description of each first audio sample, and the fourth instruction data set includes a plurality of second audio samples and a question-answer pair of each second audio sample;

[0085] In the case where the weight parameters of the audio encoder are frozen, the audio adapter and the large language model are fine-tuned based on the third instruction data set and the fourth instruction data set to improve the multi-modal expression capability.

[0086] The terminal device can obtain audio information in the vehicle in addition to obtaining a large amount of video information of different drivers driving the vehicle, and one audio information is an audio sample. The obtained large amount of audio samples are divided into a plurality of first audio samples and a plurality of second audio samples. The first audio sample is an audio sample collected at the same time as the first video sample, and the video description of the first video sample can be used as the audio description of the corresponding first audio sample. The second audio sample is an audio sample collected at the same time as the second video sample, and the question-answer pair of the second video sample can be used as the question-answer pair of the corresponding second audio sample.

[0087] The third instruction data set described above includes all first audio samples and their audio descriptions. The fourth instruction data set described above includes all second audio samples and their question-answer pairs.

[0088] In an optional embodiment, the audio adapter includes a second position embedding layer, a second Q-former, and a second linear layer; K first audio feature vectors are input into the audio adapter to obtain a target audio feature vector, including:

[0089] The K first audio feature vectors are input into the second position embedding layer to obtain K second audio feature vectors;

[0090] The K second audio feature vectors are input into the second Q-former to obtain a third audio feature vector;

[0091] The third audio feature vector is input into the second linear layer to obtain the target audio feature vector.

[0092] Since the first audio feature vector from the audio encoder does not consider time information, after the first audio feature vector is input into the audio adapter, the second positional embedding layer in the audio adapter can apply time information as positional encoding information to the representation from different audio segments, thereby obtaining a second audio feature vector carrying the positional encoding information.

[0093] The second Q-former is an audio-specific Q-former for fusing second audio feature vectors of different audio segments, and fusing K second audio feature vectors into a third audio feature vector.

[0094] In order to adapt the third audio feature vector to the input of the large language model, a linear layer (i.e., a second linear layer) is added to convert the third audio feature vector into a target audio feature vector, so as to realize the alignment of the audio feature and the large language model space.

[0095] Step 103: inputting the target video feature vector and the target audio feature vector into the large language model to obtain a risk identification result of the driver.

[0096] The risk identification result is a piece of text indicating that there is a driving risk or no driving risk.

[0097] Since the target video feature vector and the target audio feature vector are both aligned with the embedding space of the large language model, the target video feature vector and the target audio feature vector can be input into the large language model as the input of the large language model. The large language model can output the risk identification result of the driver. As shown in FIG. 2, it is a processing flow example diagram of driver risk identification based on a multi-modal large language model. The audio-video information in FIG. 2 includes video information of the driver and audio information in the vehicle driven by the driver. The risk identification result in FIG. 2 is “Yes, the driver looks tired”, indicating that the driver has a driving risk.

[0098] If the risk identification result indicates that there is a driving risk, the terminal device sends an intervention instruction to the vehicle, and the intervention instruction is used to intervene in the driving behavior of the driver, so as to ensure the safe driving of the driver.

[0099] After receiving the intervention instruction, the vehicle can perform voice alarm to remind the driver to drive safely; or can reduce the speed of the vehicle driven by the driver to ensure the safe driving of the driver.

[0100] The embodiment is based on multi-modal features such as target video feature vectors and target audio feature vectors, and can more accurately identify the driving risk of the driver by using the excellent understanding and reasoning ability of the large language model, that is, the driver risk identification is realized.

[0101] Referring to FIG. 3, it is a structural schematic diagram of a driver risk identification device provided by an embodiment of the present application. For the convenience of description, only the parts related to the embodiment of the present application are shown.

[0102] The driver risk identification device comprises:

[0103] The first acquisition module 31 is configured to acquire a target video feature vector of video information of the driver, and the target video feature vector is aligned with an embedding space of the large language model.

[0104] The second acquisition module 32 is configured to acquire a target audio feature vector of audio information in a vehicle driven by the driver, and the target audio feature vector is aligned with the embedding space of the large language model.

[0105] The risk identification module 33 is configured to input the target video feature vector and the target audio feature vector into the large language model to obtain a risk identification result of the driver.

[0106] Optionally, the video information comprises N frames of images, N is an integer greater than 1, and the first acquisition module 31 comprises:

[0107] The first input unit is configured to input the video information into a video encoder to obtain N first video feature vectors, and each frame of image corresponds to a first video feature vector.

[0108] The second input unit is configured to input the N first video feature vectors into a video adapter to obtain a target video feature vector.

[0109] Optionally, the video adapter comprises a first position embedding layer, a first Q-former and a first linear layer, and the second input unit is specifically configured to:

[0110] input the N first video feature vectors into the first position embedding layer to obtain N second video feature vectors;

[0111] input the N second video feature vectors into the first Q-former to obtain a third video feature vector;

[0112] input the third video feature vector into the first linear layer to obtain the target video feature vector.

[0113] Optionally, the driver risk identification device further comprises:

[0114] The first fine-tuning module is configured to fine-tune the video adapter based on the first instruction data set and the second instruction data set in a case where the weight parameters of the frozen video encoder and the large language model are frozen, the first instruction data set comprising a plurality of first video samples and a video description of each first video sample, and the second instruction data set comprising a plurality of second video samples and a question and answer pair of each second video sample.

[0115] The second fine-tuning module is configured to fine-tune the video adapter and the large language model based on the first instruction data set and the second instruction data set in a case where the weight parameters of the frozen video encoder are frozen.

[0116] Optionally, the second acquisition module 32 comprises:

[0117] The audio sampling module is configured to sample K audio segments from the audio information, K being an integer greater than 1.

[0118] The audio conversion module is configured to convert the K audio segments into K spectrograms respectively, to obtain the K spectrograms.

[0119] The third input module is configured to input the K spectrograms into the audio encoder to obtain K first audio feature vectors, one spectrogram corresponding to one first audio feature vector.

[0120] The fourth input module is configured to input the K first audio feature vectors into the audio adapter to obtain a target audio feature vector.

[0121] Optionally, the audio adapter comprises a second position embedding layer, a second Q-former and a second linear layer; and the fourth input module is specifically configured to:

[0122] input the K first audio feature vectors into the second position embedding layer to obtain K second audio feature vectors;

[0123] input the K second audio feature vectors into the second Q-former to obtain a third audio feature vector;

[0124] input the third audio feature vector into the second linear layer to obtain the target audio feature vector.

[0125] Optionally, the driver risk identification device further comprises:

[0126] The third fine-tuning module is configured to fine-tune the audio adapter based on the third instruction data set and the fourth instruction data set in a case where the weight parameters of the frozen audio encoder and the large language model are frozen, the third instruction data set comprising a plurality of first audio samples and an audio description of each first audio sample, and the fourth instruction data set comprising a plurality of second audio samples and a question and answer pair of each second audio sample.

[0127] A fourth fine-tuning module is configured to fine-tune the audio adapter and the large language model based on the third instruction data set and a fourth instruction data set in a case where the weight parameters of the frozen audio encoder are frozen.

[0128] The driver risk identification apparatus provided by the embodiments of the present application can be applied in the foregoing method embodiments, and details can be referred to the description of the method embodiments, which will not be described herein.

[0129] Referring to FIG. 4, it is a structural schematic diagram of a terminal device according to an embodiment of the present application. As shown in FIG. 4, the terminal device 4 of this embodiment includes one or more processors 40 (only one processor is shown in the figure), a memory 41, and a computer program 42 stored in the memory 41 and executable on the processor 40. The processor 40 implements the steps in the above-mentioned driver risk identification method embodiments when executing the computer program 42.

[0130] The terminal device 4 can be a desktop computer, a notebook computer, a palm computer, a cloud server, and the like. The terminal device can include, but is not limited to, the processor 40 and the memory 41. Those skilled in the art can understand that FIG. 4 is only an example of the terminal device 4, and does not constitute a limitation on the terminal device 4, and can include more or fewer components than those shown, or combine certain components, or different components, for example, the terminal device can also include an input / output device, a network access device, a bus, and the like.

[0131] The processor 40 can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor.

[0132] The memory 41 can be an internal storage unit of the terminal device 4, for example, a hard disk or a memory of the terminal device 4. The memory 41 can also be an external storage device of the terminal device 4, for example, a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the terminal device 4. Further, the memory 41 can also include both the internal storage unit and the external storage device of the terminal device 4. The memory 41 is used to store the computer program and other programs and data required by the terminal device. The memory 41 can also be used to temporarily store data that has been output or will be output.

[0133] It can be clearly understood by those skilled in the art that, for the convenience and brevity of description, only the division of the above functional units and modules is exemplified, and in actual application, the above functions can be completed by different functional units and modules according to needs, that is, the internal structure of the apparatus is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit, and the integrated unit can be realized in the form of hardware or in the form of software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction, and do not limit the protection scope of the present application. The specific working process of the units and modules in the apparatus can refer to the corresponding process in the foregoing method embodiments, which will not be described here.

[0134] The embodiment of the present application also provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to realize the steps in each method embodiment.

[0135] The embodiment of the present application also provides a computer program product, which, when running on a terminal device, enables the terminal device to execute the steps in each method embodiment.

[0136] In the above embodiments, the description of each embodiment has its own emphasis, and the parts not described or recorded in detail in a certain embodiment can be referred to the related description of other embodiments.

[0137] Those skilled in the art can understand that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be realized by electronic hardware or a combination of computer software and electronic hardware. Whether the functions are realized in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to realize the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.

[0138] In the embodiments provided by the present application, it should be understood that the disclosed apparatus / terminal device and method can be implemented in other ways. For example, the apparatus / terminal device embodiments described above are merely schematic, and the division of the modules or units is merely a logical function division, and there can be another division manner in actual implementation. For example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual coupling or direct coupling or communication connection can be indirect coupling or communication connection through some interface, device or unit, and can be electrical, mechanical or in other forms.

[0139] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, i.e. can be located in one place, or can be distributed on a plurality of network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiments.

[0140] The integrated module / unit, if implemented in the form of a software functional unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on such understanding, all or part of the processes in the above-mentioned embodiment methods can also be completed by a computer program instructing related hardware, and the computer program can be stored in a computer readable storage medium. The computer program can implement the steps of each method embodiment when executed by a processor. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or some intermediate forms. The computer readable medium can include any entity or device capable of carrying the computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the contents included in the computer readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction, for example, in some jurisdictions, according to legislation and patent practice, the computer readable medium does not include electrical carrier signals and telecommunication signals.

[0141] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the protection scope of the present application.

Claims

1. A driver risk identification method characterized by, The driver risk identification method comprises: obtaining a target video feature vector of video information of a driver, the target video feature vector being aligned with an embedding space of a large language model; obtaining a target audio feature vector of audio information in a vehicle driven by the driver, the target audio feature vector being aligned with the embedding space of the large language model; inputting the target video feature vector and the target audio feature vector into the large language model to obtain a risk identification result of the driver.

2. The driver risk identification method according to claim 1, characterized in that, The video information comprises N frames of images, N being an integer greater than 1, and the target video feature vector of the video information of the driver comprises: inputting the video information into a video encoder to obtain N first video feature vectors, each frame of image corresponding to a first video feature vector; inputting the N first video feature vectors into a video adapter to obtain a target video feature vector.

3. The driver risk identification method according to claim 2, characterized in that, The video adapter comprises a first position embedding layer, a first Q-former and a first linear layer, and the inputting of the N first video feature vectors into the video adapter to obtain the target video feature vector comprises: inputting the N first video feature vectors into the first position embedding layer to obtain N second video feature vectors; inputting the N second video feature vectors into the first Q-former to obtain a third video feature vector; inputting the third video feature vector into the first linear layer to obtain the target video feature vector.

4. The driver risk identification method according to claim 2, characterized in that, Before the inputting of the N first video feature vectors into the video adapter, the method further comprises: fine-tuning the video adapter based on a first instruction data set and a second instruction data set in a case where the weight parameters of the video encoder and the large language model are frozen, the first instruction data set comprising a plurality of first video samples and a video description of each first video sample, and the second instruction data set comprising a plurality of second video samples and a question and answer pair of each second video sample; fine-tuning the video adapter and the large language model based on the first instruction data set and the second instruction data set in a case where the weight parameters of the video encoder are frozen.

5. The driver risk identification method according to any one of claims 1 to 4, characterized in that, The obtaining of the target audio feature vector of the audio information in the vehicle driven by the driver comprises: sampling K audio segments from the audio information, K being an integer greater than 1; converting the K audio segments into K spectrograms respectively to obtain K spectrograms; inputting the K spectrograms into an audio encoder to obtain K first audio feature vectors, one spectrogram corresponding to one first audio feature vector; inputting the K first audio feature vectors into an audio adapter to obtain a target audio feature vector.

6. The driver risk identification method according to claim 5, characterized in that The audio adapter comprises a second position embedding layer, a second Q-former and a second linear layer, and the inputting of the K first audio feature vectors into the audio adapter to obtain the target audio feature vector comprises: inputting the K first audio feature vectors into the second position embedding layer to obtain K second audio feature vectors; inputting the K second audio feature vectors into the second Q-former to obtain a third audio feature vector; and inputting the third audio feature vector into the second linear layer to obtain the target audio feature vector. inputting K second audio feature vectors into the second Q-former to obtain a third audio feature vector; inputting the third audio feature vector into the second linear layer to obtain the target audio feature vector.

7. The driver risk identification method according to claim 5, characterized in that, Before inputting K first audio feature vectors into the audio adapter, further comprising: freezing the weight parameters of the audio encoder and the large language model, fine-tuning the audio adapter based on a third instruction data set and a fourth instruction data set, the third instruction data set comprising a plurality of first audio samples and an audio description of each first audio sample, the fourth instruction data set comprising a plurality of second audio samples and a question and answer pair of each second audio sample; freezing the weight parameters of the audio encoder, fine-tuning the audio adapter and the large language model based on the third instruction data set and the fourth instruction data set.

8. A driver risk identification apparatus, characterized by, The driver risk identification device comprises: a first acquisition module configured to acquire a target video feature vector of video information of a driver, the target video feature vector being aligned with an embedding space of a large language model; a second acquisition module configured to acquire a target audio feature vector of audio information in a vehicle driven by the driver, the target audio feature vector being aligned with the embedding space of the large language model; a risk identification module configured to input the target video feature vector and the target audio feature vector into the large language model to obtain a risk identification result of the driver.

9. A terminal device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, The processor executes the computer program to implement the steps of the driver risk identification method according to any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, the computer program comprising instructions that, when executed by a computer, cause the computer to perform the method of any one of claims 1 to 9. The computer program is executed by the processor to implement the steps of the driver risk identification method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Dangerous driving behavior identification method and device, electronic equipment and storage medium

    CN114170585A

  • Behavior detection method and device, terminal equipment and storage medium

    CN114373189A

  • Traffic accident detection method and device, electronic equipment and medium

    CN116563801A

  • Deep synthesis audio detection method, system and product combined with large language model

    CN117577120A

  • Question and answer method and system based on multi-modal self-adaptive retrieval type enhanced large model

    CN117648429A

Cited By

  • Multi-modal behavior recognition method and device based on semantic augmentation and electronic equipment

    CN122116498A