Drunk driving early warning method and system based on large model and storage medium

By extracting feature and processing the driver's facial images, voice data and driving behavior data, the risk of drunk driving is judged in real time, solving the problem of real-time detection in the prior art, and improving driving safety and detection efficiency.

CN119975371APending Publication Date: 2025-05-13BEIJING UNISOUND INFORMATION TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510051399.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-13
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

In the prior art, alcohol detection cannot be carried out in real time, which reduces the efficiency of alcohol detection during vehicle driving.

Method used

By obtaining the driver's facial images, voice data and driving behavior data, the feature extraction and input into the pre-trained large model for self-attention mechanism processing, obtaining multimodal features and matching them with the drunk driving status characteristics, determining whether there is a risk of drunk driving, and if there is, early warning prompts and braking control are performed.

Benefits of technology

Real-time alcohol detection of drivers is realized, alcohol detection efficiency during vehicle driving is improved, and alcohol detection is effectively warned and controlled drunk driving behavior, improving driving safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119975371A_ABST
    Figure CN119975371A_ABST
Patent Text Reader

Abstract

The invention provides a drunk driving early warning method and system based on a large model and a storage medium, and the method comprises the steps: obtaining a face image, voice data and driving behavior data of a driver in a target vehicle, and carrying out the feature extraction of the face image, the voice data and the driving behavior data, facial features, voice features and behavior features are obtained; inputting the facial features, the voice features and the behavior features into a pre-trained large model for self-attention mechanism processing to obtain multi-modal features, and matching the multi-modal features with drunk driving state features to obtain feature similarity; if the feature similarity is larger than a similarity threshold value, it is judged that a drunk driving risk exists, drunk driving early warning prompting is conducted on the driver, and braking control is conducted on the target vehicle. According to the embodiment of the invention, alcohol detection can be carried out on the driver in real time, and the alcohol detection efficiency in the vehicle driving process is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of automobile technology, and in particular to a drunk driving warning method, system and storage medium based on a large model. Background Art

[0002] With the rapid development of economy and the continuous improvement of people's living standards, various types of vehicles have become indispensable means of transportation in people's lives. In order to ensure the safety of vehicle driving, the issue of alcohol detection of drivers has received more and more attention.

[0003] When investigating and punishing drunk driving, traffic police usually use a handheld alcohol tester to detect the alcohol value in the driver's exhaled breath. However, the handheld alcohol tester cannot perform alcohol detection in real time, which reduces the efficiency of alcohol detection during vehicle driving. Summary of the invention

[0004] The purpose of the embodiments of the present invention is to provide a drunk driving warning method, system and storage medium based on a large model to solve the problem that alcohol detection cannot be performed in real time in the prior art.

[0005] The embodiment of the present invention is implemented as follows: a drunk driving warning method based on a large model, the method comprising:

[0006] Acquire a facial image, voice data, and driving behavior data of a driver in a target vehicle, and perform feature extraction on the facial image, the voice data, and the driving behavior data to obtain facial features, voice features, and behavior features;

[0007] Input the facial features, the voice features, and the behavior features into a pre-trained large model for self-attention mechanism processing to obtain multimodal features, and match the multimodal features with the drunk driving status features to obtain feature similarity;

[0008] If the feature similarity is greater than a similarity threshold, it is determined that there is a risk of drunk driving, a drunk driving warning is issued to the driver, and the target vehicle is braked.

[0009] Preferably, feature extraction is performed on the facial image, the voice data and the driving behavior data to obtain facial features, voice features and behavior features, including:

[0010] Preprocessing the facial image and the voice data to obtain a preprocessed image and a preprocessed voice;

[0011] Performing convolution processing on the preprocessed image to obtain facial features, and extracting fundamental frequency, formant, speech rate and intonation in the preprocessed speech to obtain the speech features;

[0012] The driving parameter values ​​in the driving behavior data are normalized, and feature mapping is performed on the normalized driving parameter values ​​to obtain the behavior features.

[0013] Preferably, the facial features, the voice features and the behavioral features are input into a pre-trained large model for self-attention mechanism processing to obtain multimodal features, including:

[0014] Performing matrix conversion on the facial features, the voice features, and the behavior features according to the pre-trained large model to obtain a key matrix, a value matrix, and a query matrix;

[0015] Performing attention processing according to the key matrix, the value matrix, and the query matrix to obtain a self-attention feature, and performing multi-convolution processing on the self-attention feature to obtain a first convolution feature and a second convolution feature;

[0016] Weight processing is performed on the first convolution feature and the second convolution feature to obtain the multimodal feature.

[0017] Preferably, before inputting the facial features, the voice features and the behavior features into a pre-trained large model for self-attention processing to obtain multimodal features, the method further comprises:

[0018] Obtaining a drunk driving multimodal sample and a non-drunk driving multimodal sample, and inputting the drunk driving multimodal sample and the non-drunk driving multimodal sample into the large model for self-attention mechanism processing to obtain fusion features;

[0019] Decoding the fused features to obtain prediction features, and matching the prediction features with the drunk driving status features to obtain prediction similarity;

[0020] Loss calculation is performed on the fusion feature, the prediction feature and the prediction similarity to obtain a model loss, and parameters of the large model are updated according to the model loss until the large model converges to obtain the pre-trained large model.

[0021] Preferably, preprocessing the facial image and the voice data to obtain a preprocessed image and a preprocessed voice includes:

[0022] Performing grayscale processing on the facial image to obtain a grayscale image, and performing image filtering processing on the grayscale image to obtain a filtered image;

[0023] Performing key point recognition on the filtered image to obtain key point coordinates, and performing part cropping according to the key point coordinates to obtain the preprocessed image;

[0024] Performing pre-emphasis processing on the voice data to obtain pre-emphasized voice, and filtering the pre-emphasized voice to obtain filtered voice;

[0025] Endpoint detection is performed on the filtered speech, and speech screening is performed on the filtered speech according to the endpoint detection result to obtain the preprocessed speech.

[0026] Preferably, providing a drunk driving warning prompt to the driver and performing brake control on the target vehicle includes:

[0027] Control the warning light on the dashboard of the target vehicle to flash for warning, control the vehicle display screen to display the drunk driving warning information, and control the vehicle audio to give voice prompts;

[0028] Controlling the engine speed of the target vehicle to decrease to an idle state, and increasing the braking force until the target vehicle stops;

[0029] The vehicle ignition switch is locked until a preset unlocking signal is received, and then the vehicle ignition switch is turned on.

[0030] Preferably, performing loss calculation on the fusion feature, the prediction feature and the prediction similarity to obtain the model loss includes:

[0031] Calculating the similarity between the fused feature and the fused standard feature to obtain a first similarity, and determining a first loss according to the first similarity;

[0032] Calculating the similarity between the prediction feature and the prediction standard feature to obtain a second similarity, and determining a second loss according to the second similarity;

[0033] Calculating a difference between the predicted similarity and the standard similarity to obtain a third similarity, and determining a third loss according to the third similarity;

[0034] A weighted operation is performed on the first loss, the second loss, and the third loss to obtain the model loss.

[0035] Another object of an embodiment of the present invention is to provide a drunk driving warning system based on a large model, the system comprising:

[0036] A feature extraction module, used to obtain a facial image, voice data and driving behavior data of the driver in the target vehicle, and perform feature extraction on the facial image, the voice data and the driving behavior data to obtain facial features, voice features and behavior features;

[0037] A feature matching module, for inputting the facial features, the voice features and the behavior features into a pre-trained large model for self-attention mechanism processing to obtain multimodal features, and matching the multimodal features with the drunk driving status features to obtain feature similarity;

[0038] The early warning module is used to determine that there is a risk of drunk driving if the feature similarity is greater than a similarity threshold, issue a drunk driving early warning prompt to the driver, and perform braking control on the target vehicle.

[0039] Preferably, the feature extraction module is further used to: pre-process the facial image and the voice data to obtain a pre-processed image and a pre-processed voice;

[0040] Performing convolution processing on the preprocessed image to obtain facial features, and extracting fundamental frequency, formant, speech rate and intonation in the preprocessed speech to obtain the speech features;

[0041] The driving parameter values ​​in the driving behavior data are normalized, and feature mapping is performed on the normalized driving parameter values ​​to obtain the behavior features.

[0042] The embodiment of the present invention can effectively extract the driver's facial features, voice features and behavioral features by performing feature extraction on facial images, voice data and driving behavior data. The pre-trained large model can effectively perform self-attention mechanism processing on the facial features, voice features and behavioral features to obtain multimodal features. Based on the feature similarity between the multimodal features and the drunk driving status features, it can automatically determine whether the driver has a drunk driving risk. When the drunk driving risk is detected, the driver is given a drunk driving warning prompt and the target vehicle is braked, which effectively plays a drunk driving warning effect and improves the safety of vehicle driving. In this embodiment, the driver can be tested for alcohol in real time, which improves the efficiency of alcohol detection during vehicle driving. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 is a flow chart of a drunk driving warning method based on a large model provided by a first embodiment of the present invention;

[0044] Figure 2 is a structural schematic diagram of a drunk driving warning system based on a large model provided by a second embodiment of the present invention;

[0045] Figure 3 is a schematic diagram of the structure of a drunk driving prevention system based on a terminal-side large model provided in a third embodiment of the present invention;

[0046] Figure 4 It is a schematic diagram of the structure of a terminal device provided in the fourth embodiment of the present invention. DETAILED DESCRIPTION

[0047] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0048] In order to illustrate the technical solution of the present invention, a specific embodiment is provided below for illustration.

[0049] Embodiment 1

[0050] See also Figure 1 , is a flow chart of a drunk driving warning method based on a large model provided in the first embodiment of the present invention. The drunk driving warning method based on a large model can be applied to any device or system. The drunk driving warning method based on a large model includes the following steps:

[0051] Step S10, acquiring a facial image, voice data, and driving behavior data of a driver in a target vehicle, and performing feature extraction on the facial image, the voice data, and the driving behavior data to obtain facial features, voice features, and behavior features;

[0052] Among them, the target vehicle is equipped with a high-definition camera, a microphone array and a vehicle sensor group. The high-definition camera is installed at a suitable position in front of the vehicle's driving position, which can clearly capture the driver's facial expressions, eyes, head posture and other information, and output them in the form of image data. The frame rate is not less than 25 frames / second, and the image resolution reaches 1080p and above, ensuring that the details of the driver's facial features can be accurately extracted. The microphone array is distributed in multiple positions in the cab, which can accurately collect the driver's voice signal. It has a noise reduction function and can effectively filter out the ambient noise and other interfering sounds in the car. The sampling frequency is 16k to obtain clear and accurate voice data. The vehicle sensor group includes a steering wheel angle sensor, an accelerator pedal sensor, a brake pedal sensor, a vehicle speed sensor, etc. The steering wheel angle sensor can accurately measure the steering wheel's rotation angle and rotation speed; the accelerator pedal sensor and the brake pedal sensor can monitor the pedal's stepping depth and change rate in real time; the vehicle speed sensor accurately records the vehicle's driving speed, with an error range of less than 5km / h, and obtains comprehensive driving behavior data through sensors.

[0053] Optionally, feature extraction is performed on the facial image, the voice data, and the driving behavior data to obtain facial features, voice features, and behavior features, including:

[0054] Preprocessing the facial image and the voice data to obtain a preprocessed image and a preprocessed voice; wherein, by preprocessing the facial image and the voice data, noise in the facial image and the voice data can be effectively removed, thereby improving the quality of the facial image and the voice data;

[0055] The preprocessed image is convoluted to obtain facial features, and the fundamental frequency, formant, speech rate and intonation in the preprocessed speech are extracted to obtain the speech features; wherein, the preprocessed image is convoluted by using a convolutional neural network (CNN) algorithm based on deep learning to achieve a feature extraction effect on the preprocessed image, and facial features such as the degree of eye opening, pupil size changes, facial muscle movement features, etc. can be effectively extracted from the preprocessed image. Preferably, the fundamental frequency, formant, speech rate and intonation in the preprocessed speech can be extracted by using an algorithm such as Mel Frequency Cepstral Coefficient (MFCC) to obtain speech features;

[0056] The driving parameter values ​​in the driving behavior data are normalized, and the normalized driving parameter values ​​are feature mapped to obtain the behavior features; wherein, driving parameter values ​​such as the degree of frequent changes in the steering wheel angle, the abnormal frequency of stepping on the accelerator pedal and the brake pedal are extracted from the driving behavior data for normalization, and the driving parameter values ​​such as the degree of frequent changes in the steering wheel angle, the abnormal frequency of stepping on the accelerator pedal and the brake pedal are feature mapped according to a preset mapping relationship to obtain the behavior features. In this step, various types of driving behavior data collected by the vehicle sensor group are normalized, and data of different ranges and units are uniformly mapped to a specific numerical range, such as normalizing the steering wheel angle data to the [-1,1] interval, and normalizing the accelerator pedal and brake pedal data to the [0,1] interval, etc., so as to facilitate comprehensive analysis and processing by the large model on the end side.

[0057] Further, the facial image and the voice data are preprocessed to obtain a preprocessed image and a preprocessed voice, including:

[0058] Graying the facial image to obtain a gray image, and performing image filtering on the gray image to obtain a filtered image; wherein the facial image is grayed to convert the color image into a gray image to reduce the amount of data processing, and the gray image is image filtered to remove noise interference such as salt and pepper noise and Gaussian noise in the image by using Gaussian filtering or median filtering algorithm;

[0059] Performing key point recognition on the filtered image to obtain key point coordinates, and performing part cropping according to the key point coordinates to obtain the preprocessed image; wherein, the part cropping is performed according to the key point coordinates to extract the key area image of the driver's face, such as eyes, mouth and other parts, to facilitate subsequent feature extraction;

[0060] The speech data is pre-emphasized to obtain pre-emphasized speech, and the pre-emphasized speech is filtered to obtain filtered speech; wherein the speech data is pre-emphasized and filtered using a finite impulse response (FIR) or infinite impulse response (IIR) filter to remove low-frequency noise and other non-speech frequency band interference in the speech signal;

[0061] Endpoint detection is performed on the filtered speech, and speech screening is performed on the filtered speech according to the endpoint detection result to obtain the preprocessed speech; wherein, endpoint detection is performed on the filtered speech to accurately identify the valid speech segment in the speech signal, remove the silent part, and improve the accuracy of subsequent speech feature extraction.

[0062] Step S20, inputting the facial features, the voice features and the behavior features into a pre-trained large model for self-attention mechanism processing to obtain multimodal features, and matching the multimodal features with the drunk driving status features to obtain feature similarity;

[0063] Among them, the large model can adopt a large end-side model based on the Transformer architecture, which includes multiple encoder and decoder layers. The model input layer receives multimodal features from data preprocessing and feature extraction, and processes the multimodal features through the self-attention mechanism of the multi-layer Transformer encoder, which can automatically learn the correlation and importance weights between different modal features. For example, the large model can learn the strong correlation between the frequency of eye closure and the drunk driving status in facial features, as well as the potential connection between abnormal speech speed and drunk driving in voice features, and make a comprehensive judgment based on information such as unstable steering wheel control in driving behavior characteristics.

[0064] Optionally, the facial features, the voice features, and the behavioral features are input into a pre-trained large model for self-attention mechanism processing to obtain multimodal features, including:

[0065] Performing matrix conversion on the facial features, the voice features, and the behavior features according to the pre-trained large model to obtain a key matrix, a value matrix, and a query matrix;

[0066] Attention processing is performed according to the key matrix, the value matrix and the query matrix to obtain self-attention features, multi-convolution processing is performed on the self-attention features to obtain first convolution features and second convolution features, and weight processing is performed on the first convolution features and the second convolution features to obtain the multimodal features; wherein, through the self-attention mechanism, the large model focuses on key information in different modal features, such as eye status in facial features, changes in speech speed and intonation in voice features, and abnormal vehicle control in driving behavior features.

[0067] Furthermore, before inputting the facial features, the voice features and the behavior features into the pre-trained large model for self-attention processing to obtain the multimodal features, the method further includes:

[0068] Obtaining a drunk driving multimodal sample and a non-drunk driving multimodal sample, and inputting the drunk driving multimodal sample and the non-drunk driving multimodal sample into the large model for self-attention mechanism processing to obtain fusion features;

[0069] Decoding the fused features to obtain prediction features, and matching the prediction features with the drunk driving status features to obtain prediction similarity;

[0070] Performing loss calculation on the fusion feature, the prediction feature and the prediction similarity to obtain a model loss, and updating parameters of the large model according to the model loss until the large model converges to obtain the pre-trained large model;

[0071] Among them, during the large model training stage, a large number of multimodal data samples of drivers including drunk driving and non-drunk driving are collected and trained using supervised learning. The loss function uses a combination of cross entropy loss function and regularization term to prevent overfitting of the large model. The large model parameters are continuously adjusted through the back propagation algorithm to optimize the model performance so that the large model can accurately distinguish between drunk driving and non-drunk driving. After the large model training is completed, it is deployed on the vehicle-side equipment, such as the on-board intelligent control unit (ECU) or a dedicated edge computing device, to ensure that fast data analysis and decision-making can be performed locally in the vehicle without relying on cloud network connection, ensuring the real-time data processing and data privacy.

[0072] Furthermore, loss calculation is performed on the fusion feature, the prediction feature and the prediction similarity to obtain a model loss, including: calculating the similarity between the fusion feature and the fusion standard feature to obtain a first similarity, and determining a first loss based on the first similarity; calculating the similarity between the prediction feature and the prediction standard feature to obtain a second similarity, and determining a second loss based on the second similarity; calculating the difference between the prediction similarity and the standard similarity to obtain a third similarity, and determining a third loss based on the third similarity; performing a weighted operation on the first loss, the second loss and the third loss to obtain the model loss; wherein, during the weighted operation, the weighting coefficients of the first loss, the second loss and the third loss can be set according to demand.

[0073] Step S30: if the feature similarity is greater than the similarity threshold, it is determined that there is a risk of drunk driving, a drunk driving warning is given to the driver, and braking control is performed on the target vehicle;

[0074] Among them, the similarity threshold can be set according to needs. When the feature similarity is greater than the similarity threshold, it is determined that the driver is suspected of drunk driving and enters the warning and control steps; if the feature similarity is less than or equal to the similarity threshold, it returns to continue executing the data collection and drunk driving warning steps.

[0075] Optionally, the driver is given a drunk driving warning prompt and the target vehicle is braked, including: controlling the warning light on the dashboard of the target vehicle to flash as a warning, controlling the on-board display screen to display the drunk driving warning information, and controlling the vehicle audio to give voice prompts; controlling the engine speed of the target vehicle to reduce to idle state, and increasing the braking force until the target vehicle stops; locking the vehicle ignition switch until a preset unlocking signal is received, and then turning on the vehicle ignition switch.

[0076] When the driver is suspected of drunk driving, the warning light on the dashboard in the car flashes at a specific frequency, such as the red warning light flashing three times per second; at the same time, a striking drunk driving warning message pops up on the car's display screen, with the text displaying "Possible drunk driving behavior detected, please stop immediately!", and a sharp alarm sounds through the vehicle's audio system, lasting for no less than 10 seconds, to attract the driver's high attention.

[0077] At the same time as the warning is issued, the vehicle is controlled. The vehicle's power output is limited, and the engine speed is gradually reduced to the idle state, so that the vehicle cannot accelerate; the braking force is gradually increased to decelerate the vehicle smoothly until the vehicle stops completely. After the vehicle stops, the vehicle's ignition system is automatically locked to prevent the driver from starting the vehicle again until the driver is confirmed to be not drunk driving through legal alcohol testing methods or other legal unlocking measures are taken.

[0078] In this embodiment, by extracting features from facial images, voice data and driving behavior data, the driver's facial features, voice features and behavior features can be effectively extracted. The pre-trained large model can effectively process the facial features, voice features and behavior features with a self-attention mechanism to obtain multimodal features. Based on the feature similarity between the multimodal features and the drunk driving status features, it can automatically determine whether the driver has a risk of drunk driving. When the risk of drunk driving is detected, the driver is given a drunk driving warning prompt and the target vehicle is braked, which effectively plays a role in the drunk driving warning effect and improves the safety of vehicle driving. In this embodiment, the driver can be tested for alcohol in real time, which improves the efficiency of alcohol detection during vehicle driving. By collecting multimodal data, the driver's relevant information can be fully obtained, avoiding the limitations of a single data source; a large model that integrates deep learning algorithms and drunk driving related knowledge graphs is used for comprehensive analysis to improve the accuracy of drunk driving risk judgment; a warning signal can be issued in a timely manner to achieve early intervention in drunk driving behavior, effectively solving the problems of passive detection, untimely warning and difficulty in comprehensive monitoring in existing drunk driving prevention technologies, and greatly improving the effect and efficiency of drunk driving prevention.

[0079] Embodiment 2

[0080] See also Figure 2 , is a schematic diagram of the structure of a drunk driving warning system 100 based on a large model provided in a second embodiment of the present invention, including:

[0081] The feature extraction module 10 is used to obtain the facial image, voice data and driving behavior data of the driver in the target vehicle, and perform feature extraction on the facial image, the voice data and the driving behavior data to obtain facial features, voice features and behavior features.

[0082] Optionally, the feature extraction module 10 is further used to: pre-process the facial image and the voice data to obtain a pre-processed image and a pre-processed voice;

[0083] Performing convolution processing on the preprocessed image to obtain facial features, and extracting fundamental frequency, formant, speech rate and intonation in the preprocessed speech to obtain the speech features;

[0084] The driving parameter values ​​in the driving behavior data are normalized, and feature mapping is performed on the normalized driving parameter values ​​to obtain the behavior features.

[0085] Furthermore, the feature extraction module 10 is also used to: perform grayscale processing on the facial image to obtain a grayscale image, and perform image filtering processing on the grayscale image to obtain a filtered image;

[0086] Performing key point recognition on the filtered image to obtain key point coordinates, and performing part cropping according to the key point coordinates to obtain the preprocessed image;

[0087] Performing pre-emphasis processing on the voice data to obtain pre-emphasized voice, and filtering the pre-emphasized voice to obtain filtered voice;

[0088] Endpoint detection is performed on the filtered speech, and speech screening is performed on the filtered speech according to the endpoint detection result to obtain the preprocessed speech.

[0089] The feature matching module 11 is used to input the facial features, the voice features and the behavior features into a pre-trained large model for self-attention mechanism processing to obtain multimodal features, and match the multimodal features with the drunk driving status features to obtain feature similarity.

[0090] Optionally, the feature matching module 11 is further used to: perform matrix conversion on the facial features, the voice features and the behavior features according to the pre-trained large model to obtain a key matrix, a value matrix and a query matrix;

[0091] Performing attention processing according to the key matrix, the value matrix, and the query matrix to obtain a self-attention feature, and performing multi-convolution processing on the self-attention feature to obtain a first convolution feature and a second convolution feature;

[0092] Weight processing is performed on the first convolution feature and the second convolution feature to obtain the multimodal feature.

[0093] Furthermore, the feature matching module 11 is also used to: obtain a drunk driving multimodal sample and a non-drunk driving multimodal sample, and input the drunk driving multimodal sample and the non-drunk driving multimodal sample into the large model for self-attention mechanism processing to obtain a fusion feature;

[0094] Decoding the fused features to obtain prediction features, and matching the prediction features with the drunk driving status features to obtain prediction similarity;

[0095] Loss calculation is performed on the fusion feature, the prediction feature and the prediction similarity to obtain a model loss, and parameters of the large model are updated according to the model loss until the large model converges to obtain the pre-trained large model.

[0096] Furthermore, the feature fusion module 11 is further used to: calculate the similarity between the fused feature and the fused standard feature to obtain a first similarity, and determine a first loss according to the first similarity;

[0097] Calculating the similarity between the prediction feature and the prediction standard feature to obtain a second similarity, and determining a second loss according to the second similarity;

[0098] Calculating a difference between the predicted similarity and the standard similarity to obtain a third similarity, and determining a third loss according to the third similarity;

[0099] A weighted operation is performed on the first loss, the second loss, and the third loss to obtain the model loss.

[0100] The warning module 12 is used to determine that there is a risk of drunk driving if the feature similarity is greater than a similarity threshold, issue a drunk driving warning to the driver, and perform braking control on the target vehicle.

[0101] Optionally, the warning module 12 is further used to: control the warning light on the dashboard of the target vehicle to flash for warning, control the vehicle display screen to display the drunk driving warning information, and control the vehicle audio to give voice prompts;

[0102] Controlling the engine speed of the target vehicle to decrease to an idle state, and increasing the braking force until the target vehicle stops;

[0103] The vehicle ignition switch is locked until a preset unlocking signal is received, and then the vehicle ignition switch is turned on.

[0104] In this embodiment, by performing feature extraction on facial images, voice data and driving behavior data, the driver's facial features, voice features and behavioral features can be effectively extracted. The pre-trained large model can effectively perform self-attention mechanism processing on the facial features, voice features and behavioral features to obtain multimodal features. Based on the feature similarity between the multimodal features and the drunk driving status features, it can automatically determine whether the driver has a drunk driving risk. When a drunk driving risk is detected, the driver is given a drunk driving warning prompt and the target vehicle is braked, which effectively plays a drunk driving warning effect and improves the safety of vehicle driving. In this embodiment, the driver can be tested for alcohol in real time, which improves the efficiency of alcohol detection during vehicle driving.

[0105] Embodiment 3

[0106] See also Figure 3 , is a schematic diagram of the structure of a drunk driving prevention system 200 based on a terminal-side large model provided in a third embodiment of the present invention, including:

[0107] Multimodal data acquisition unit 201, data preprocessing and feature extraction unit 202, end-side large model core processing unit 203 and early warning and control execution unit 204.

[0108] The multimodal data acquisition unit 201 includes a high-definition camera, a microphone array and a vehicle sensor group. The data preprocessing and feature extraction unit 202 includes an image preprocessing module, a voice preprocessing module, a driving behavior data preprocessing module and a feature extraction module. The warning and control execution unit 204 includes a warning module and a vehicle control module.

[0109] The methods for preventing drunk driving based on the big model on the client side include:

[0110] 1. Data collection steps:

[0111] After the system is started, the multimodal data acquisition unit 201 starts working. The high-definition camera continuously captures the driver's facial image at a set frame rate, the microphone array collects the driver's voice signal in real time, and the vehicle sensor group synchronously records the driving behavior data. The collected data is marked and cached in chronological order to form a continuous data stream, waiting for subsequent processing.

[0112] 2. Data preprocessing and feature extraction steps:

[0113] The collected image data, voice data, and driving behavior data are preprocessed and feature extracted. After graying, filtering, and cropping, the image data is used to extract facial features by the feature extraction module; the voice data is pre-emphasized, filtered, and endpoint detected to extract voice features; the driving behavior data is normalized to extract driving behavior features. The extracted multimodal feature data is organized into a unified data format so that it can be input into the large model on the end for analysis.

[0114] 3. Model analysis and judgment steps:

[0115] The multimodal feature data after preprocessing and feature extraction is input into the core processing unit 203 of the large model on the end side. The large model on the end side performs a comprehensive analysis of the input feature data based on its trained parameters and internal complex neural network structure. Through the self-attention mechanism, the model focuses on key information in different modal features, such as the state of the eyes in facial features, changes in speech speed and intonation in voice features, and abnormal vehicle control in driving behavior features, and calculates the degree of match between these feature combinations and the drunk driving status. The model outputs a drunk driving suspicion probability value. When the probability value exceeds a preset threshold (for example, 0.6), it is determined that the driver is suspected of drunk driving and enters the early warning and control step; if the probability value does not exceed the threshold, the data collection and analysis process continues.

[0116] 4. Early warning and control steps:

[0117] When the client-side big model determines that the driver is suspected of drunk driving, the warning and control execution unit 204 is activated. The warning module sends out visual and auditory warning signals to remind the driver to stop driving. At the same time, the vehicle control module limits the power of the vehicle and performs deceleration and braking operations until the vehicle stops safely and the ignition system is locked, effectively preventing traffic accidents caused by drunk driving.

[0118] In this embodiment, multimodal data is collected by the end-side data collection module, so that relevant information of the driver can be fully obtained, avoiding the limitations of a single data source; the end-side large model module uses a large model that integrates deep learning algorithms and drunk driving related knowledge graphs for comprehensive analysis, thereby improving the accuracy of drunk driving risk judgment; the early warning module can issue early warning signals in a timely manner to achieve early intervention in drunk driving behavior; the end-side model update module can continuously optimize the model to adapt to different driving scenarios and driver groups, effectively solving the problems of passive detection, untimely warning, and difficulty in comprehensive monitoring in existing drunk driving prevention technologies, greatly improving the effectiveness and efficiency of drunk driving prevention.

[0119] Embodiment 4

[0120] Figure 4 2 is a block diagram of a terminal device 2 provided in the fourth embodiment of the present application. Figure 4 As shown, the terminal device 2 of this embodiment includes: a processor 20, a memory 21, and a computer program 22 stored in the memory 21 and executable on the processor 20, such as a program of a drunk driving warning method based on a large model. When the processor 20 executes the computer program 22, the steps in each embodiment of the drunk driving warning method based on a large model are implemented.

[0121] Exemplarily, the computer program 22 may be divided into one or more modules, which are stored in the memory 21 and executed by the processor 20 to complete the present application. The one or more modules may be a series of computer program instruction segments capable of completing specific functions, which are used to describe the execution process of the computer program 22 in the terminal device 2. The terminal device may include, but is not limited to, a processor 20 and a memory 21.

[0122] The processor 20 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.

[0123] The memory 21 may be an internal storage unit of the terminal device 2, such as a hard disk or memory of the terminal device 2. The memory 21 may also be an external storage device of the terminal device 2, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the terminal device 2. Further, the memory 21 may also include both an internal storage unit and an external storage device of the terminal device 2. The memory 21 is used to store the computer program and other programs and data required by the terminal device. The memory 21 may also be used to temporarily store data that has been output or is to be output.

[0124] In addition, each functional module in each embodiment of the present application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of software functional unit.

[0125] If the integrated module is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Among them, the computer-readable storage medium can be non-volatile or volatile. Based on this understanding, the present application implements all or part of the processes in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and the computer program can implement the steps of the above-mentioned various method embodiments when executed by the processor. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable storage medium may include: any entity or device capable of carrying computer program code, recording medium, U disk, mobile hard disk, disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electric carrier signal, telecommunication signal and software distribution medium, etc. It should be noted that the content contained in computer-readable storage media can be appropriately increased or decreased according to the requirements of legislation and patent practices in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practices, computer-readable storage media do not include electrical carrier signals and telecommunication signals.

[0126] The embodiments described above are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, a person skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. Such modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.

Claims

1. A drunk driving warning method based on a large model, characterized in that: The method comprises: Acquire a facial image, voice data, and driving behavior data of a driver in a target vehicle, and perform feature extraction on the facial image, the voice data, and the driving behavior data to obtain facial features, voice features, and behavior features; Input the facial features, the voice features, and the behavior features into a pre-trained large model for self-attention mechanism processing to obtain multimodal features, and match the multimodal features with the drunk driving status features to obtain feature similarity; If the feature similarity is greater than a similarity threshold, it is determined that there is a risk of drunk driving, a drunk driving warning is issued to the driver, and the target vehicle is braked.

2. The drunk driving warning method based on a large model as claimed in claim 1 is characterized in that: Feature extraction is performed on the facial image, the voice data, and the driving behavior data to obtain facial features, voice features, and behavior features, including: Preprocessing the facial image and the voice data to obtain a preprocessed image and a preprocessed voice; Performing convolution processing on the preprocessed image to obtain facial features, and extracting fundamental frequency, formant, speech rate and intonation in the preprocessed speech to obtain the speech features; The driving parameter values ​​in the driving behavior data are normalized, and feature mapping is performed on the normalized driving parameter values ​​to obtain the behavior features.

3. The drunk driving warning method based on a large model as claimed in claim 1 is characterized in that: The facial features, the voice features, and the behavior features are input into a pre-trained large model for self-attention mechanism processing to obtain multimodal features, including: Performing matrix conversion on the facial features, the voice features, and the behavior features according to the pre-trained large model to obtain a key matrix, a value matrix, and a query matrix; Performing attention processing according to the key matrix, the value matrix, and the query matrix to obtain a self-attention feature, and performing multi-convolution processing on the self-attention feature to obtain a first convolution feature and a second convolution feature; Weight processing is performed on the first convolution feature and the second convolution feature to obtain the multimodal feature.

4. The drunk driving warning method based on a large model as claimed in claim 1 is characterized in that: Before inputting the facial features, the voice features and the behavior features into the pre-trained large model for self-attention processing to obtain multimodal features, the method further includes: Obtaining a drunk driving multimodal sample and a non-drunk driving multimodal sample, and inputting the drunk driving multimodal sample and the non-drunk driving multimodal sample into the large model for self-attention mechanism processing to obtain fusion features; Decoding the fused features to obtain prediction features, and matching the prediction features with the drunk driving status features to obtain prediction similarity; Loss calculation is performed on the fusion feature, the prediction feature and the prediction similarity to obtain a model loss, and parameters of the large model are updated according to the model loss until the large model converges to obtain the pre-trained large model.

5. The drunk driving warning method based on a large model as claimed in claim 2 is characterized in that: Preprocessing the facial image and the voice data to obtain a preprocessed image and a preprocessed voice includes: Performing grayscale processing on the facial image to obtain a grayscale image, and performing image filtering processing on the grayscale image to obtain a filtered image; Performing key point recognition on the filtered image to obtain key point coordinates, and performing part cropping according to the key point coordinates to obtain the preprocessed image; Performing pre-emphasis processing on the voice data to obtain pre-emphasized voice, and filtering the pre-emphasized voice to obtain filtered voice; Endpoint detection is performed on the filtered speech, and speech screening is performed on the filtered speech according to the endpoint detection result to obtain the preprocessed speech.

6. The drunk driving warning method based on a large model as claimed in claim 1 is characterized in that: Providing a drunk driving warning to the driver and performing brake control on the target vehicle, including: Control the warning light on the dashboard of the target vehicle to flash for warning, control the vehicle display screen to display the drunk driving warning information, and control the vehicle audio to give voice prompts; Controlling the engine speed of the target vehicle to decrease to an idle state, and increasing the braking force until the target vehicle stops; The vehicle ignition switch is locked until a preset unlocking signal is received, and then the vehicle ignition switch is turned on.

7. The drunk driving warning method based on a large model as claimed in claim 4 is characterized in that: The loss calculation is performed on the fusion feature, the prediction feature and the prediction similarity to obtain the model loss, including: Calculating the similarity between the fused feature and the fused standard feature to obtain a first similarity, and determining a first loss according to the first similarity; Calculating the similarity between the prediction feature and the prediction standard feature to obtain a second similarity, and determining a second loss according to the second similarity; Calculating a difference between the predicted similarity and the standard similarity to obtain a third similarity, and determining a third loss according to the third similarity; A weighted operation is performed on the first loss, the second loss, and the third loss to obtain the model loss.

8. A drunk driving warning system based on a large model, characterized in that: The system comprises: A feature extraction module, used to obtain a facial image, voice data and driving behavior data of the driver in the target vehicle, and perform feature extraction on the facial image, the voice data and the driving behavior data to obtain facial features, voice features and behavior features; A feature matching module, for inputting the facial features, the voice features and the behavior features into a pre-trained large model for self-attention mechanism processing to obtain multimodal features, and matching the multimodal features with the drunk driving status features to obtain feature similarity; The early warning module is used to determine that there is a risk of drunk driving if the feature similarity is greater than a similarity threshold, issue a drunk driving early warning prompt to the driver, and perform braking control on the target vehicle.

9. The drunk driving warning system based on a large model as claimed in claim 8, characterized in that: The feature extraction module is also used for: Preprocessing the facial image and the voice data to obtain a preprocessed image and a preprocessed voice; Performing convolution processing on the preprocessed image to obtain facial features, and extracting fundamental frequency, formant, speech rate and intonation in the preprocessed speech to obtain the speech features; The driving parameter values ​​in the driving behavior data are normalized, and feature mapping is performed on the normalized driving parameter values ​​to obtain the behavior features.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Abnormal driving behavior detection method and system, electronic equipment and storage medium

    CN120747926A