A driving risk early warning method and system based on a large language model
By combining a large language model with vehicle sensors and eye-tracking glasses to obtain information about the driving environment and driver status, personalized voice and image warnings are generated, which solves the problem of insufficient warnings of driving risks when the driver's attention is distracted, and improves driving safety and reaction speed.
Patent Information
- Application Number
- CN202410808777.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-21
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2044-06-21
AI Technical Summary
Existing technologies are insufficient in addressing driving risk warning issues. They fail to effectively address driver distraction, especially in non-emergency situations, and cannot provide timely and personalized risk warnings when drivers are distracted.
The system employs a large language model combined with vehicle sensors and eye-tracking glasses to acquire information about the driving environment and driver status. Pedestrian recognition and segmentation are performed using the Dense Prediction Transformer and Segment Anything Model. The YOLO algorithm is used to identify the positions of vehicles and pedestrians. Combined with the driver's gaze point, voice and image warning information is generated to provide personalized risk alerts.
It enables timely and effective personalized risk warnings when the driver's attention is distracted, improving driving safety and reaction speed, and enhancing the driver's ability to perceive potential risks through a combination of voice and vision.
Smart Images

Figure CN118701092B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the field of driving risk prediction, and particularly relates to a driving risk early warning method and system based on a large language model. BACKGROUND
[0002] The development of autonomous driving technology has provided the possibility for drivers to supervise the driving system while performing other tasks. Modern cars have been able to perform a series of driving tasks while allowing drivers to perform other work or leisure activities during vehicle operation, significantly improving the convenience and comfort of driving. This multi-task processing capability, while bringing convenience to drivers, has introduced new challenges, especially in ensuring that drivers can take over control in a timely and effective manner when necessary.
[0003] Existing assisted driving systems mostly focus on immediate reactions to emergency situations, issuing warnings to prompt drivers to take immediate action, but lack detailed management of driver attention allocation in non-emergency situations. This is particularly important in the context of autonomous driving, as drivers may become distracted due to over-reliance on automated systems, resulting in delayed reactions when human intervention is truly needed. These systems often fail to consider the cognitive characteristics of drivers and potential risks in the future, leading to a lack of sufficient preparation and alertness in complex or unexpected situations.
[0004] Traditional risk warning systems use risk fields, Bayesian networks, and other models. With the advent of artificial intelligence, new solutions have emerged for risk detection. Deep learning methods learn the relative depth of objects in images or estimate the depth information of objects in indoor scenes by training neural networks. Interactive methods such as voice prompts and interface interactions
[0005] Despite the significant progress made in risk detection, current risk warning systems still cannot meet the needs of comprehensive risk presentation. This gap is mainly due to two factors: first, the unpredictability and complexity of real-world driving environments, making it challenging to handle boundary conditions using software simulation tests; second, driving automation provides opportunities for distraction for drivers, and driving state affects the effectiveness of warning. Existing research attempts to apply persuasion theory to adjust the state of drivers. However, these methods do not consider the current driving task and the state of driver distraction, and it is unclear when the system should persuade the driver and how the driver will react to persuasion. SUMMARY
[0006] The present application provides a driving risk early warning method based on a large language model, which detects the distraction behavior of drivers, judges the risks in the future period of time, generates persuasive voice prompts and intuitive visual interfaces, and realizes risk early warning.
[0007] The embodiment of the application provides a driving risk early warning method based on a large language model, which comprises the following steps:
[0008] S1, pedestrian risk information, driver distraction behavior state information and prompt words obtained in real time are input into a large language model, the prompt words comprising risk judgment reference information, persuasion strategies and persuasion principles;
[0009] S2, whether early warning is needed is determined based on the pedestrian risk information, the driver distraction behavior state information and the risk judgment reference information through the large language model;
[0010] S3, if early warning is needed, voice data and image data for early warning of the driver are output based on the persuasion strategies and the persuasion principles through the large language model, the voice data is converted into prompt sound for reminding the driver, corresponding images are displayed based on the image data through a head-up display system, and driving risk early warning is realized based on the prompt sound and the displayed images.
[0011] Preferably, the pedestrian risk information acquisition method comprises the following steps:
[0012] S21, road environment information is acquired through a vehicle sensor, the road environment information comprising traffic flow, pedestrians, road conditions, illumination and weather;
[0013] S22, the road environment information is input into a Dense Prediction Transformer model to obtain pedestrian distance information, a safety distance label set based on traffic engineering principles is acquired, the pedestrian distance information is compared with the safety distance label, and the pedestrian distance information is divided into corresponding early warning areas;
[0014] The road environment information is input into a pixel-level segmentation of a Segment Anything Model and a Grounding DINO to distinguish sidewalks and lanes, and position coordinates of pedestrians in a scene are obtained based on obtained position coordinates of the pedestrians, the sidewalks and the lanes and corresponding masks, so that it is determined whether the pedestrians are on the sidewalks or the lanes;
[0015] S23, the corresponding early warning areas and the position coordinates of the pedestrians in the scene obtained in step S22 are converted into natural language, and the natural language is the pedestrian risk information.
[0016] Preferably, the driver distraction behavior state information acquisition method comprises the following steps:
[0017] A camera built in the eye movement glasses captures scene images and fixation points in a field of view of the driver in real time;
[0018] The YOLO visual algorithm is used to identify objects in the scene image, and the recognition result is compared with the fixation point for coincidence degree, and when the coincidence degree reaches the set threshold, the driver distraction behavior state information is obtained.
[0019] Preferably, the risk judgment reference information is used to provide a reference for the risk judgment of the large language model, and the risk judgment reference information includes an accident probability ranking of the driver distraction behavior and specific risk factors that need to be paid attention to in the driving scene.
[0020] Preferably, the persuasion strategy includes state feedback, risk emphasis, reliable prompt, and social relationship.
[0021] The state feedback provides environmental feedback to attract the attention of the driver;
[0022] The risk emphasis provides past cases to warn potential consequences;
[0023] The reliable prompt provides information about the needs and interests of the driver to adjust the driving state of the driver;
[0024] The social relationship points out the problems of the driver and guides the behavior of the driver by using the social relationship.
[0025] Preferably, the persuasion principle is used as a restriction condition for the output of the large language model.
[0026] Preferably, the prompt word further includes a role and an output format.
[0027] The role is used to define the role of the large language model as a driving assistant and perform corresponding duties.
[0028] The output format is used to specify the format of the output content of the large language model.
[0029] Preferably, the corresponding image is displayed based on the image data through the head-up display system, and the corresponding image includes basic driving data display, action danger warning sign, and eye movement guide sign.
[0030] The basic driving data display includes basic driving data such as speed, time, and navigation direction.
[0031] The eye movement guide sign is a plurality of arc segments, the center of the arc segment moves with the fixation point of the driver, the arc segment is located at the center position between the fixation point and the corresponding action danger warning sign, the center point of the arc segment is located on the line connecting the fixation point and the corresponding action danger warning sign, and the color of the arc segment is consistent with the color of the corresponding action danger warning sign.
[0032] Preferably, the action danger warning sign includes a road danger sign and a roadside danger sign, wherein:
[0033] A method for obtaining a road danger sign, comprising:
[0034] A road vehicle area and a corresponding center position in the road environment information are dynamically identified using a YOLO algorithm, a first sign is set based on the road vehicle area and the corresponding center position, a driver's gaze point receiving an eye movement eye output is obtained in real time, and when the gaze point falls within the road vehicle area, the size of the first sign is adjusted to obtain a road danger sign;
[0035] A method for obtaining a roadside danger sign, comprising:
[0036] A moving roadside pedestrian and / or non-motor vehicle in the road environment information is identified, a YOLO algorithm is used to dynamically identify a roadside pedestrian and / or non-motor vehicle contour and a corresponding center position, a third sign is set based on the roadside pedestrian and / or non-motor vehicle contour and the corresponding center position, a driver's gaze point receiving an eye movement eye output is obtained in real time, and when the gaze point falls within the contour of the roadside pedestrian and / or non-motor vehicle, the size and shape of the third sign are adjusted to obtain a roadside danger sign.
[0037] In another aspect, the present application also provides a driving risk warning system based on a large language model, comprising:
[0038] An input module for inputting road environment information, driver distraction behavior state information and prompt words into a large language model, wherein the prompt words include risk judgment reference information, persuasion strategies and persuasion principles;
[0039] A processing module for judging whether warning is needed through the large language model based on the road environment information, the driver distraction behavior state information and the risk judgment reference information;
[0040] An output module for outputting voice data and image data for warning the driver through the large language model based on the persuasion strategies and the persuasion principles if warning is needed, converting the voice data into a prompt sound for reminding the driver, displaying a corresponding image through a head-up display system based on the image data, and realizing driving risk warning based on the prompt sound and the displayed image.
[0041] Compared with the prior art, the present application has the following advantages:
[0042] 1. The present application inputs real-time obtained road environment information and driver distraction behavior state information into a large language model, performs early warning judgment and obtains early warning information through the large language model, so as to fully consider the specific driving task and the distraction state of the driver, so as to provide personalized risk warning at the appropriate time by using appropriate persuasion strategy, compared with a single warning system, the driver's attention can be more efficiently guided, so as to improve the reaction speed and driving safety.
[0043] 2. The present application combines voice and visual interaction modalities to provide more rich and intuitive early warning information, and enhances the driver's perception ability of potential risks. BRIEF DESCRIPTION OF DRAWINGS
[0044] Figure 1 FIG. 1 is a schematic diagram of the overall process of the driving risk early warning system based on the large language model of the embodiment of the present application;
[0045] Figure 2 FIG. 4 is an example diagram of the large language model prompt of the embodiment of the present application;
[0046] Figure 3 FIG. 6 is an example schematic diagram of the adaptive head-up display interface of the embodiment of the present application;
[0047] Figure 3 (a) is a risk object map provided by the road danger sign and the roadside danger sign of the specific embodiment of the present application;
[0048] Figure 3 (b) is a change diagram of the rear interface of the driver looking at the front vehicle;
[0049] Figure 3 (c) is a change diagram of the rear interface of the driver looking at the roadside danger sign guided by the existing eye movement guide sign. DETAILED DESCRIPTION
[0050] The present application will be described in detail below in conjunction with specific embodiments. The following embodiments will help those skilled in the art to further understand the present application, but do not limit the present application in any form. It should be pointed out that those skilled in the art can make several changes and improvements without departing from the concept of the present application. These all belong to the protection scope of the present application.
[0051] In view of the insufficient early warning for the driver distraction in the prior art, the specific embodiment of the present application obtains the risk warning for the driver by using the large language model based on the obtained driver analysis behavior state information and road environment information, so as to early warn the driver at the appropriate time by using the appropriate persuasion strategy.
[0052] As Figure 1As shown, the embodiment of the present application provides a driving risk early warning method based on a large language model, which comprises:
[0053] (1) Obtain road environment information and driver distraction behavior state information.
[0054] In an embodiment, the present embodiment obtains road environment information using a vehicle-mounted sensor, which includes traffic flow, pedestrians, road conditions, lighting, and weather.
[0055] In order to obtain more road environment information from the vehicle-mounted camera video, a Dense Prediction Transformer (DPT) model is deployed to predict the depth of each pixel in the current image to obtain pedestrian distance information. Combined with the "safe distance" label set based on traffic engineering principles, pedestrians are divided into the following categories to provide more context information and warnings for the LLM. The pedestrian distance information is divided into corresponding warning areas, which include:
[0056] 1. High alert zone (0-5 meters): very close;
[0057] 2. Warning area (5-10 meters): nearby;
[0058] 3. Warning area (10-15 meters): medium;
[0059] 4. Other: very far away.
[0060] The Segment Anything Model (SAM) and Grounding DINO are used to distinguish between sidewalks and roads to obtain the location of pedestrians and the type of surface they are currently on (i.e., on a sidewalk or on a road). Using the distance (coordinates of pedestrians and roads) and the mask (segmented image) output by SAM, based on the obtained location coordinates of pedestrians, sidewalks, and lanes, and the corresponding mask, the location coordinates of pedestrians in the scene are determined to determine whether the pedestrians are on a sidewalk or on a road, which can affect the judgment of the large language model on whether to warn. The image is divided into "lower left, near front, lower right, upper left, far front, and upper right" six parts to determine the position of the pedestrian. This process converts the position information of the pedestrian on the image into position information that is easy for the LLM to understand, i.e., aggregates the data obtained by each of the above modules into natural language, i.e., pedestrian risk information, thereby reducing the complexity of the input information.
[0061] In an embodiment, the present embodiment provides specific steps for obtaining driver distraction behavior state information:
[0062] The scene image and fixation point in the driver's field of view are captured in real time by the camera built in the eye movement glasses, the objects in the scene image are identified by using the YOLO visual algorithm, the coincidence degree of the identification result and the fixation point is compared, and the driver's distraction behavior state information is obtained when the coincidence degree reaches the set threshold. The eye movement glasses used in the embodiment are Tobii Pro Glasses 3 glasses, and the driver's distraction behaviors include using a mobile phone, adjusting in-vehicle equipment, eating or taking an object, etc. By using the eye movement tracking technology, the visual interface elements are automatically adjusted according to the attention distribution of the driver, and unnecessary interference is reduced.
[0063] (2) The road environment information, the driver's distraction behavior state information and the prompt word obtained in real time are input into the large language model to determine whether a warning is needed, and the prompt word for determining whether to warn includes [role] [reference information] [output format] part to comprehensively evaluate the road risk and auxiliary tasks to determine whether a warning is needed.
[0064] (3) If a warning is needed, the warning information obtained by the large language model based on the prompt word is a voice data and image data for warning the driver, and the prompt word for obtaining the warning information includes [role] [reference information] [output format] [persuasion strategy] [persuasion principle], in an embodiment, the voice data is a persuasion text, and the image data is a potential risk judgment of road objects, if a warning is not needed, steps (1) and (2) are repeated.
[0065] In a specific embodiment, the prompt word input into the large language model provided by the embodiment includes the following parts as shown in Figure 2
[0066] [Role]: It prompts the role and responsibility of the driving assistant that the large language model needs to play.
[0067] [Reference information]: It gives the reference for the large language model to make risk judgments, including the accident probability ranking of the driver's distraction behavior and the specific risk factors that need to be paid attention to in the driving scene.
[0068] [Persuasion strategy]: Four persuasion strategies are provided, including state feedback, risk emphasis, reliable suggestion and social contact.
[0069] [Persuasion principle]: It provides the guiding principle of simplifying and clearly conveying information.
[0070] [Output format]: It specifies the format of the content output by the large language model.
[0071] Where the driver distraction behavior accident probability ranking comes from "Visual-Manual NHTSA Driver Distraction Guidelines for In-Vehicle Electronic Devices," (April 26, 2013. https: / / www.federalregister.gov / documents / 2013 / 04 / 26 / 2013-09883 / visual-manual-nhtsa-driver-distraction-guidelines-for-in-vehicle-electronic-devices).
[0072] Persuasion strategies include:
[0073] 1. "State Feedback": Provide timely environmental feedback to attract the driver's attention;
[0074] 2. "Highlight the risk": Provide past cases to warn of potential consequences;
[0075] 3. "Reliable tips": Provide information suitable for the driver's needs and interests to adjust the driver's driving state;
[0076] 4. "Social relations": Point out the driver's problems and use social relations to guide the driver's behavior.
[0077] Persuasion principles include:
[0078] 1. Keep it simple and neat, don't use rhetorical questions, and reduce behavior recommendations;
[0079] 2. Don't describe what happened, but say what the driver did;
[0080] 3. Avoid overemphasizing consequences;
[0081] 4. Colloquial, avoid excessive seriousness, and keep the tone consistent with people's daily conversations.
[0082] (4) The specific embodiment of the present application converts voice data into prompt sound for reminding the driver and performs voice output, displays the corresponding image through the head-up display system based on the image data, obtains the adaptive HUD interface, and realizes the driving risk warning based on the prompt sound and the displayed image.
[0083] In a specific embodiment, the persuasion content is converted into voice by artificial intelligence, the artificial intelligence is a Baidu voice synthesis API, and the potential risk is prompted through the head-up display interface, the head-up display interface is composed of three parts, as shown in Figure 3 (a)- Figure 3 (b) shown:
[0084] 1. Basic driving data display: Basic driving data display is shown in a fixed position at the bottom of the interface, which shows basic driving data including speed, time and navigation direction, providing continuous driving information for the driver;
[0085] 2. Dynamic danger warning signs: As shown in Figure 3 (a), road danger signs and roadside danger signs mark risk objects, and eye movement guide signs guide the line of sight through an arc. The warning signs in the interface are divided into two types, which are used to indicate potential dangers on the road and on the roadside respectively:
[0086] (1) Road danger signs: used to indicate the risk of vehicles on the road. According to the YOLO algorithm, the rectangular area where the vehicle is located is dynamically identified, with length l road , width w road , and center position (x road , y road ). The sign is designed as a red hollow equilateral triangle with a side length of b = 0.5w road , and when the driver's gaze point falls within a certain vehicle area, the sign automatically adjusts the side length to 0.6b to reduce visual interference, as shown in Figure 3 (b), after the driver's gaze falls on the danger on the front vehicle, the size of the triangle sign on the vehicle becomes smaller, and the eye movement guide sign corresponding to the vehicle disappears.
[0087] (2) Roadside danger signs: These signs are used to indicate the presence of risk pedestrians and non-motor vehicles on the roadside, and move with the risk source. The sign uses four straight-line frames to frame a yellow rectangular outline, with length l side , width w side , and center position (x side , y side ) determined by the YOLO algorithm segmentation. When the driver's line of sight enters the area where the sign is located, the sign becomes a yellow solid equilateral triangle with a side length of 0.6w side to optimize the driver's risk perception. As shown in Figure 3 Figure 3 (c), the driver is guided to pay attention to the roadside danger sign by the existing eye movement guide sign, and the boundary box of the sign disappears, replaced by a small triangle, and the corresponding eye movement guide sign disappears.
[0088] 3. Eye movement guide sign: The eye movement guide sign is an arc segment with a radius r and a length of 1 / 8 of the circumference, and the center dynamically follows the driver's gaze point (x gaze , y gaze ) movement. The arc indicates the driver's gaze point and the center position of the potential danger object The color of the guide sign matches the corresponding danger sign to guide the driver's line of sight to directly notice the potential threat and make a response in advance.
[0089] In another aspect, the embodiments of the present application also provide a driving risk warning system based on a large language model, comprising:
[0090] An input module is configured to input road environment information, driver distraction behavior state information, and a prompt word into the large language model, wherein the prompt word includes risk judgment reference information, a persuasion strategy, and a persuasion principle.
[0091] A processing module is configured to judge whether a warning is needed based on the road environment information, the driver distraction behavior state information, and the risk judgment reference information through the large language model.
[0092] An output module is configured to output voice data and image data for warning the driver based on the persuasion strategy and the persuasion principle through the large language model if a warning is needed, convert the voice data into a prompt sound for reminding the driver, display a corresponding image based on the image data through a head-up display system, and implement driving risk warning based on the prompt sound and the displayed image.
Claims
1. A driving risk early warning method based on a large language model, characterized in that, The method comprises the following steps: S1, inputting real-time obtained pedestrian risk information, driver distraction behavior state information and prompt words into a large language model, wherein the prompt words comprise risk judgment reference information, persuasion strategy and persuasion principle; S2, judging whether early warning is needed based on the pedestrian risk information, the driver distraction behavior state information and the risk judgment reference information through the large language model; S3, if early warning is needed, outputting voice data and image data for early warning of the driver based on the persuasion strategy and the persuasion principle through the large language model, converting the voice data into prompt sound for reminding the driver, displaying corresponding images based on the image data through a head-up display system, and realizing driving risk early warning based on the prompt sound and the displayed images. The method for obtaining the pedestrian risk information comprises the following steps: S21, obtaining road environment information through a vehicle sensor, wherein the road environment information comprises traffic flow, pedestrians, road conditions, illumination and weather; S22, inputting the road environment information into a Dense Prediction Transformer model to obtain pedestrian distance information, obtaining a safety distance label set based on traffic engineering principles, comparing the pedestrian distance information with the safety distance label, and dividing the pedestrian distance information into corresponding early warning areas; inputting the road environment information into a pixel-level segmentation of a Segment Anything Model and a Grounding DINO to distinguish a sidewalk and a lane, obtaining position coordinates of the pedestrians in the scene based on obtained position coordinates of the pedestrians, the sidewalk and the lane, and corresponding masks, and determining whether the pedestrians are on the sidewalk or the lane; S23, converting the corresponding early warning areas and the position coordinates of the pedestrians in the scene obtained in step S22 into natural language, wherein the natural language is the pedestrian risk information; displaying corresponding images based on the image data through a head-up display system, wherein the corresponding images comprise basic driving data display, action danger warning signs and eye movement guide signs; The basic driving data display comprises basic driving data such as speed, time and navigation direction. The eye movement guide sign comprises a plurality of arc segments, the center of the arc segment moves with the gaze point of the driver, the arc segment is located at the center position between the gaze point and the corresponding action danger warning sign, the center point of the arc segment is located on the line connecting the gaze point and the corresponding action danger warning sign, and the color of the arc segment is consistent with the color of the corresponding action danger warning sign. 2.The driving risk pre-warning method based on large language model according to claim 1, characterized in that, The method for obtaining the driver distraction behavior state information comprises the following steps: capturing scene images and gaze points in the field of view of the driver through a camera built in the eye movement glasses in real time; adopting a YOLO visual algorithm to identify objects in the scene images, comparing the identification results with the gaze points in terms of coincidence degree, and obtaining the driver distraction behavior state information when the coincidence degree reaches a set threshold. 3.The driving risk pre-warning method based on large language model according to claim 1, characterized in that, The risk judgment reference information is used to provide a reference for risk judgment of the large language model, and comprises an accident probability ranking of the driver distraction behavior and specific risk factors that need to be paid attention to in the driving scene. 4.The driving risk pre-warning method based on large language model according to claim 1, characterized in that, The persuasion strategy includes state feedback, risk emphasis, reliable cues and social relationships. The state feedback provides environmental feedback to attract the driver's attention. The risk emphasis provides past cases to warn of potential consequences. The reliable cues provide information about the driver's needs and interests to adjust the driver's driving state. The social relationship points out the driver's problem and uses social relationships to guide the driver's behavior. 5.The driving risk pre-warning method based on large language model according to claim 1, characterized in that, The persuasion principle is used to limit the output of voice data by the large language model. 6.The driving risk pre-warning method based on large language model according to claim 1, characterized in that, The prompt word also includes a role and an output format. The role is used to define the large language model as a driving assistant role and perform corresponding duties. The output format is used to specify the format of the large language model output content. 7.The driving risk pre-warning method based on large language model according to claim 1, characterized in that, The action danger warning sign includes a road danger sign and a roadside danger sign, wherein: The method for obtaining a road danger sign includes: Using the YOLO algorithm to dynamically identify the on-road vehicle area and the corresponding center position in the road environment information, setting a first sign based on the on-road vehicle area and the corresponding center position, obtaining the driver's gaze point output by receiving eye movement in real time, and adjusting the size of the first sign to obtain the road danger sign when the gaze point falls within the on-road vehicle area. The method for obtaining a roadside danger sign includes: Identifying moving roadside pedestrians and / or non-motor vehicles in the road environment information, using the YOLO algorithm to dynamically identify the outline and corresponding center position of the roadside pedestrians and / or non-motor vehicles, setting a third sign based on the outline and corresponding center position of the roadside pedestrians and / or non-motor vehicles, obtaining the driver's gaze point output by receiving eye movement in real time, and adjusting the size and shape of the third sign to obtain the roadside danger sign when the gaze point falls within the outline of the roadside pedestrians and / or non-motor vehicles.
8. The driving risk early warning system based on a large language model is applied to the driving risk early warning method based on a large language model according to any one of claims 1-7, characterized in that, It includes: An input module for inputting road environment information, driver distraction behavior state information and prompt words into a large language model, the prompt words including risk judgment reference information, persuasion strategies and persuasion principles; A processing module for determining whether a warning is needed based on road environment information, driver distraction behavior state information and risk judgment reference information through a large language model; An output module for outputting voice data and image data for warning drivers through a large language model based on persuasion strategies and persuasion principles if a warning is needed, converting the voice data into a prompt sound to remind the driver, displaying the corresponding image based on the image data through a heads-up display system, and implementing driving risk warning based on the prompt sound and displayed image.
Citation Information
Patent Citations
Information control device
CN110998688A
Head-up display method and device and computer readable storage medium
CN116524013A