Intelligent driving system and method and electronic equipment
Through a multimodal large language model and dynamic safety perception module, the intelligent driving system generates personalized teaching strategies and adaptive safety response strategies, solving the problems of teaching mismatch and low perception accuracy caused by unified standards and static sensors in existing systems, and improving teaching efficiency and safety.
Patent Information
- Application Number
- CN202511183353.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-22
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2045-08-22
AI Technical Summary
Existing intelligent driving systems, due to their use of unified standards and static sensor fusion strategies, are unable to adapt to differences in students' abilities, resulting in the inability to accurately match teaching strategies. Furthermore, the perception accuracy is low in complex scenarios, leading to the risk of misbraking or missed detection.
A multimodal large language model is used to generate personalized teaching strategies. Combined with meteorological, multi-dimensional vehicle-mounted sensors and visual data, scenario risks are dynamically assessed, adaptive safety response strategies are generated, and multimodal interactive feedback is provided through the behavior management module.
It achieves personalized matching of teaching strategies, improves teaching efficiency and safety, reduces the risks of false braking and missed detection, and improves the system's perception accuracy and robustness in complex environments.
Smart Images

Figure CN120681155A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of intelligent driving technology, and in particular to an intelligent driving system, method, and electronic device. Background Art
[0002] In the areas of driving education and human-machine co-driving, there is an urgent need to build an intelligent system that can collaboratively achieve precise driver development and safe management in complex scenarios. This system must generate personalized teaching strategies based on the differences in students' driving abilities, rather than using a uniform training standard. Furthermore, it must optimize multi-source sensor data fusion mechanisms in complex scenarios such as sudden inclement weather and unconventional obstacles, generating safe driving decisions that adapt to environmental disturbances.
[0003] Currently, representative technical solutions typically employ a driving behavior scoring system based on a preset rule base, combined with a static sensor fusion strategy. This solution applies the same standardized scoring model to all learners, for example, monitoring fixed thresholds such as steering wheel vibration frequency and following distance to trigger text alerts and display deduction notices on the in-car screen.
[0004] However, the existing solutions use a static rule base and fixed sensor weights, which makes it impossible for teaching strategies to adapt to differences in students' abilities. This leads to frequent conflicts such as missed high-risk operations or excessive low-risk alerts. In interference scenarios such as heavy rain and backlight, the fixed fusion weights of sensors significantly reduce perception accuracy, leading to the risk of false braking or missed detections, thus causing poor performance of existing intelligent driving systems. Summary of the Invention
[0005] The purpose of this application is to provide an intelligent driving system, method and electronic equipment to solve the problem of poor operation performance of intelligent driving systems in the prior art.
[0006] To solve the above technical problems, in a first aspect, the present application provides an intelligent driving system, including a teaching strategy module, a safety perception module and a behavior management module; The teaching strategy module is used to generate personalized teaching strategies adapted to students' abilities based on multi-dimensional student data through a multimodal large language model; The safety perception module is used to integrate meteorological, on-board multi-dimensional sensor and visual data to assess the risk level of different scene types by constructing road friction coefficient and slippery trend, analyzing scene visibility, and performing behavior clustering and motion vector analysis on dynamic targets. Based on the assessment results and the preset intervention level, a safety response strategy is generated, wherein the safety response strategy includes one or any combination of speed control, avoidance path, braking force and steering compensation; The behavior management module is used to identify driving behavior data and generate multimodal interactive feedback in combination with physiological data and status data to guide and correct driving behavior.
[0007] Optionally, the scene type includes a weather change scene; The safety perception module is specifically used to collect meteorological data and vehicle sensor data, the meteorological data including rainfall intensity and fog density, and the vehicle sensor data including vehicle speed and acceleration; Dynamically binding the meteorological data to the vehicle position to generate space meteorological parameters; Acquire a scene image through polarization imaging processing and perform an optical descattering operation to generate a defogging image; Analyzing the visibility level based on the defogging image and inferring the road slipperiness trend by combining the spatial meteorological parameters with the visibility level; Based on the intervention degree corresponding to the weather change scenario and the slippery road surface change trend, a safety response strategy matching the environmental interference is generated. The safety response strategy includes a speed control curve linked to the slippery road surface change trend and a control strategy for the minimum following distance threshold.
[0008] Optionally, the scene type includes a human-machine hybrid scene; The safety perception module is specifically used to collect pedestrian behavior data and vehicle-mounted camera data, wherein the pedestrian behavior data includes pedestrian movement trajectory and density data; Based on the collected pedestrian behavior data, a cluster analysis is performed on pedestrian behavior patterns to classify pedestrians into stationary, walking, and running groups. Processing the vehicle camera data using binocular stereo vision technology, generating a depth map by parallax calculation to determine the distances between the pedestrians in the stationary group, the walking group, and the running group and the vehicle, and outputting a three-dimensional scene feature map containing pedestrian classification information; Based on the three-dimensional scene feature map, the position, target type and motion vector of pedestrians and non-motor vehicles in each classification group are detected, wherein the motion vector of the non-motor vehicle includes a speed direction change feature; Assessing the environmental risk level based on the motion vector change rate of the pedestrians in the walking group and the running group, the speed direction change characteristics of the non-motor vehicle, and the distance between the non-motor vehicle and the vehicle; Based on the degree of intervention corresponding to the human-machine hybrid scenario, the stationary group of pedestrians is selectively included in the risk calculation scope, and a safety response strategy matching the environmental interference is generated according to the environmental risk level. The safety response strategy includes an avoidance path and speed control strategy linked to the environmental risk level and matching the environmental interference.
[0009] Optionally, the scene type includes a two-way road scene; The safety perception module is specifically used to obtain road surface humidity data, road surface texture data and vehicle inertia measurement data to construct a road friction coefficient curve; Analyzing the road surface texture by an ultrasonic sensor, determining the road surface micromorphology to calculate a roughness index, and outputting a slipperiness feature vector; Based on the road friction coefficient curve and the slipperiness characteristic vector, predicting a short-term road condition change trend to output a slipperiness risk level; Based on the intervention level corresponding to the two-way road scenario, combined with the slippery risk level and vehicle interference data from the oncoming lane, a safety response strategy that matches the environmental interference is generated. The safety response strategy includes an adaptive braking force and steering compensation strategy linked to the road slippery risk level for anti-skid control and safe passing.
[0010] Optionally, the safety perception module is also used to adjust the weights corresponding to multiple sensors to compensate for environmental interference by fusing lidar data, inertial measurement unit data and synchronous positioning and mapping data, and generate a compensated safety response strategy.
[0011] Optionally, the security perception module includes: a fusion perception module, a fusion positioning module and a decision planning module; The safety perception module is specifically used to perform the following process when fusing lidar data, inertial measurement unit data, and synchronous positioning and mapping data, adjusting the weights corresponding to multiple sensors to compensate for environmental interference, and generating a compensated safety response strategy: Send weight configuration instructions to the fusion perception module and fusion positioning module to adjust the fusion weights of lidar data, camera data, millimeter wave radar data, inertial measurement unit data, and simultaneous positioning and mapping data; The fusion perception module responds to the weight configuration instruction, utilizes the pre-calibrated lidar-camera-millimeter-wave radar parameters, fuses the weighted lidar point cloud data, camera image data and millimeter-wave radar point track data, and performs target detection, recognition and trajectory prediction through a deep learning model; The fusion positioning module responds to the weight configuration instruction and fuses the weighted lidar data, inertial measurement unit data and high-precision point cloud map built based on the laser inertial odometry SLAM framework in areas where satellite positioning fails or is interfered with, to achieve centimeter-level vehicle positioning through point cloud matching; The decision-making planning module receives the target information output by the fusion perception module, the high-precision location information output by the fusion positioning module, the scene type and the degree of intervention, and applies target tracking, trajectory prediction and collision detection algorithms to generate a compensated safety response strategy.
[0012] Optionally, the multi-dimensional student data includes learning ability data, driving habit data, and error type data; The teaching strategy module is specifically used to perform combined processing of the learning ability data, driving habit data and error type data through a multimodal large language model to form a student state feature set; Based on the student status feature set, a path generation method is used to construct a personalized teaching path including a basic training stage, a scenario simulation stage and a comprehensive assessment stage; According to the time distribution and content configuration of the personalized teaching path, a personalized teaching strategy is output.
[0013] Optionally, the teaching strategy module is specifically configured to perform the following process when performing the combined processing of the learning ability data, driving habit data, and error type data by a multimodal large language model to form a student state feature set: The attention concentration index in the learning ability data is mapped to the first-dimensional text descriptor to generate the first-category language feature; the steering wheel turning frequency index in the driving habit data is mapped to the second-dimensional operation descriptor to generate the second-category language feature; the operation error code in the error type data is mapped to the third-dimensional scene descriptor to generate the third-category language feature; Sequentially concatenating the first, second, and third language features through a lightweight pure language model, converting the concatenated features into structured language units using a rule converter, and outputting a text descriptor set containing multiple labels; Through the end-to-end multimodal vertical model, the original sensor data stream, the gaze trajectory timing signal output by the eye tracker, the angle change waveform output by the steering wheel sensor, and the erroneous operation video clips captured by the vehicle camera are synchronously received; In a unified embedding space, the gaze trajectory timing signal is encoded into a first multidimensional data group, the angle change waveform is encoded into a second multidimensional data group, and the incorrect operation video clip is encoded into a third multidimensional data group; the first multidimensional data group, the second multidimensional data group and the third multidimensional data group are fused to generate a multimodal feature vector; the text descriptor set and the multimodal feature vector are aligned along the time axis, and a student state feature set is generated through a feature binder.
[0014] Optionally, a teaching execution monitoring module is further included, which is used to perform the following processes during the execution of the personalized teaching strategy: During the execution of the basic training phase, the scenario simulation phase, or the comprehensive assessment phase, multimodal data including the trainee's voice commands, images, and text input data are collected through a multimodal data fusion engine; Processing the multimodal data using a deep learning algorithm to identify user operation intentions; Perceive the interior environment and external driving scene through camera and radar sensors; Combined with the user's operation intention, the in-vehicle environment and the external driving scene information, a cognitive state data set representing the trainee's current cognitive state and scene interaction needs is generated and updated.
[0015] Optionally, a teaching strategy adjustment module is further included, configured to execute the following progressive process based on the cognitive state data set generated by the teaching execution monitoring module: generating and executing an interaction mode switching instruction based on current vehicle speed information in the cognitive state data set, wherein the interaction mode switching instruction controls the interaction interface to enable portrait display and voice-first interaction at high speeds, and to enable landscape display and touch interaction at low speeds; Generate and execute a teaching content adjustment instruction based on the user operation intention information and cognitive state information in the cognitive state data set, wherein the teaching content adjustment instruction adjusts the detail level, presentation method or strength of auxiliary information of the teaching prompt; Based on the evaluation results of the student's operation proficiency and error patterns of the cognitive state data set, a teaching mode switching instruction is generated and executed, and the screen is linked to display the teaching prompt content corresponding to the mode; Synchronously call the map engine service to load a high-resolution static map and display the vehicle's position.
[0016] Optionally, a teaching task collaborative execution module is further included, which is used to execute the following progressive process based on the interactive mode switching instruction, teaching content adjustment instruction and teaching mode switching instruction generated by the teaching strategy adjustment module, and the cognitive state data set generated by the teaching execution monitoring module: By combining the user operation intention information in the cognitive state data set with the current teaching steps recorded by the system, the decision planning engine analyzes the contextual semantics of the student's instructions in the teaching-related interaction; Coordinate the execution order and output logic of the teaching guidance tasks, safety control tasks, and route guidance tasks that need to be executed in parallel based on the current teaching strategy defined by the interactive mode switching instructions, teaching content adjustment instructions, and teaching mode switching instructions generated by the teaching strategy adjustment module, and a predefined task priority algorithm, to avoid task conflicts; During the coordinated execution of tasks, the perception engine service is called to obtain identification information of key perception targets around the vehicle; The coordinated task instructions are executed in a linked manner, and the identification information of the key perception targets obtained is presented on the interactive interface.
[0017] Optionally, an emotional service intervention module is further included, which is used to generate student emotional state data by analyzing the student's facial image and voice data collected by the teaching execution monitoring module through emotion recognition technology; Based on the student's emotional state data, a soothing service instruction is generated and triggered. The instruction controls the teaching system to perform operations, including switching the voice interaction tone to a softer mode and providing targeted encouraging guidance content; Combined with the cognitive state data set and emotional state data generated by the teaching execution monitoring module, predict the operational errors that may occur to students in the current teaching scenario; Based on the prediction results, predictive operation reminder information is generated and pushed to the interactive terminal.
[0018] Optionally, the teaching strategy module is also used to simultaneously collect the student's heart rate signal and eye tracking signal output by the wearable physiological monitoring device, and the steering wheel angle signal and throttle opening signal output by the vehicle operating device; convert the heart rate signal into a learning stress level indicator, and convert the eye tracking signal into a line of sight focus time indicator, to jointly form the student's learning ability data; convert the steering wheel angle signal into a steering frequency indicator, and convert the throttle opening signal into an acceleration smoothness indicator, to jointly form the student's driving habit data; and combine the number of operating errors and error type codes in historical training records to form the student's error type data.
[0019] Optionally, the safety perception module is further configured to obtain environmental data output by an environmental perception device; determine a scene type based on a scene feature identifier in the environmental data, wherein the scene types include a weather change scene, a human-machine mixed scene, and a two-way road scene; When the scene type is a weather change scene, the rainfall intensity and fog density parameters in the real-time meteorological data are collected and combined with the moving object density index in the spatial obstacle distribution data to generate the weather change scene type code; When the scene type is a human-machine mixed scene, the pedestrian movement trajectory features and non-motor vehicle motion vectors are extracted, and the trajectory intersection density and relative speed change rate are calculated to generate the human-machine mixed scene type code; When the scene type is a two-way road scene, the distance parameters and relative speed parameters of the vehicles in the opposite lane are obtained to identify the lane line curvature change characteristics and generate the two-way road scene type code; Synchronously collect vehicle speed and road curvature values from vehicle driving status data; Based on the scene type code, vehicle speed value and road curvature value, the intervention degree corresponding to each scene is calculated.
[0020] Optionally, the output end of the safety perception module is connected to the input end of the behavior management module. The behavior management module is used to perform the following process when identifying driving behavior data and generating multimodal interactive feedback in combination with physiological data and status data to guide and correct driving behavior: Obtain torque change data output by the steering wheel torque sensor and facial muscle displacement data collected by the vehicle camera; Converting the torque change data into a steering wheel operation force level, and converting the facial muscle displacement data into a mouth corner lift amplitude and a brow wrinkle degree; Based on the deviation between the steering wheel operation force level and the preset standard force range, the behavioral conflict coefficient is generated in combination with the upward range of the mouth corners and the degree of frowning between the eyebrows; When the behavior conflict coefficient exceeds a first threshold, the voice feedback generator is triggered to perform the following operations: matching a first voice segment in a steering operation description phrase library according to the steering wheel operation force level; matching a second voice segment in a emotional state description phrase library according to the mouth corner upward angle and the eyebrow wrinkling degree; and combining the first voice segment and the second voice segment into a primary correction instruction; After outputting the primary correction instruction, continuously monitoring the change trends of the torque change data and the facial muscle displacement data, and if the rate of decrease of the behavior conflict coefficient is lower than a second threshold, extracting related follow-up question segments from a multi-round dialogue template library; The associated question segment and the correction operation demonstration segment are combined to generate a secondary correction instruction, and a multi-round interactive voice sequence including the primary correction instruction and the secondary correction instruction is output.
[0021] In a second aspect, the present application provides an intelligent driving method, comprising: Generate personalized teaching strategies tailored to students' abilities based on multi-dimensional student data through a multimodal large language model; In the process of implementing the personalized teaching strategy, the system integrates meteorological, on-board multi-dimensional sensor and visual data to evaluate the risk level of different scene types by constructing the road friction coefficient and slippery trend, analyzing scene visibility, and performing behavioral clustering and motion vector analysis on dynamic targets. Based on the evaluation results and the preset intervention level, a safety response strategy is generated, wherein the safety response strategy includes one or any combination of speed control, avoidance path, braking force and steering compensation; In the process of executing the safety response strategy, driving behavior data is identified, and multimodal interactive feedback is generated in combination with physiological data and status data to guide and correct driving behavior.
[0022] In a third aspect, the present application provides an electronic device, comprising: memory for storing computer programs; A processor is used to implement the steps of the intelligent driving method as described in the second aspect above when executing the computer program.
[0023] The teaching strategy module in the intelligent driving system provided in this application dynamically generates and adjusts teaching strategies based on multi-dimensional student data (such as operational proficiency, reaction speed, and error types) using a multimodal large language model. This allows the teaching content and difficulty to be precisely matched to the driving ability levels of different students. This shifts the traditional unified teaching model, resolves conflicts caused by differences in ability, and significantly improves teaching efficiency and safety.
[0024] The safety perception module innovatively integrates meteorological data, multi-dimensional on-board sensors (such as radar and cameras), and visual data. By dynamically analyzing road friction coefficients and slippery trends, accurately analyzing environmental visibility, and performing behavioral clustering and motion vector prediction for surrounding dynamic targets (such as vehicles and pedestrians), it achieves a refined and dynamic risk assessment for different scenarios. Based on this assessment, the module adaptively generates optimal safety response strategies, including precise speed control recommendations, safe avoidance paths, and necessary braking and steering compensation. This overcomes the limitations of static rule bases and fixed sensor weights in complex interference scenarios such as heavy rain, fog, and backlight, significantly improving the accuracy and robustness of environmental perception. This directly reduces the risks of "false braking" (braking when it should not) and "missed detection" (failure to respond to a warning or braking warning) caused by perception errors, significantly enhancing the system's safety assurance capabilities in adverse conditions.
[0025] The Behavior Management Module continuously identifies driving operation data and conducts a comprehensive analysis based on the student's physiological state (e.g., fatigue and tension). Based on this analysis, the module provides timely, appropriate, and easy-to-understand guidance and corrections to the student through various interactive modes, including voice prompts, visual alerts, and tactile feedback (e.g., steering wheel vibration). This module not only provides immediate feedback on operational issues but also translates the risk assessment results of the Safety Perception Module and the guidance provided by the Teaching Strategy Module into intuitive interactive instructions, forming a complete "perception-decision-feedback" closed loop. This effectively guides students to develop correct driving habits and proactively intervenes before potential dangers arise. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] In order to more clearly illustrate the embodiments of the present application or the technical solutions of the prior art, the following is a brief introduction to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0027] Figure 1 This is a system architecture diagram of an intelligent driving system provided by one embodiment of the present application; Figure 2 This is a system architecture diagram of a teaching strategy module provided by an embodiment of the present application; Figure 3 This is a flow chart of a security perception module provided by one embodiment of the present application; Figure 4 This is a flow chart of calculating a behavior conflict coefficient by a behavior management module provided by one embodiment of the present application; Figure 5 This is a flowchart of a teaching execution monitoring module provided by an embodiment of the present application; Figure 6 This is a flow chart of a teaching strategy adjustment module provided by an embodiment of the present application; Figure 7 This is a schematic diagram of a scenario in which a map engine service switches to a teaching mode according to an embodiment of the present application; Figure 8 This is a flowchart of a teaching task collaborative execution module provided by an embodiment of the present application; Figure 9 This is a flowchart of an intelligent driving method provided by an embodiment of the present application. DETAILED DESCRIPTION
[0028] Research has revealed significant limitations in current intelligent driving teaching and assistance systems. At the instructional level, these systems generally employ uniform standards for training and assessment of all learners, failing to adapt to individual differences in driving ability. This one-size-fits-all approach leads to critical issues: for less capable learners, the system may fail to promptly identify high-risk maneuvers; for more capable learners, it may trigger unnecessary alerts for low-risk behaviors, creating a conflict between safety monitoring and instructional guidance. Regarding environmental perception, the system's mechanism for fusing and processing meteorological, onboard sensor (e.g., radar, camera), and visual data is relatively rigid in complex interference scenarios such as heavy rain, fog, and backlight, often relying on pre-set fixed rules and sensor weightings. This static fusion strategy significantly reduces perception accuracy in changing environments, making it prone to misjudgments. For example, the system fails to promptly warn or brake on slippery roads (missed detection) or causes unnecessary sudden braking in safe situations (false triggering), severely limiting system reliability and safety.
[0029] To address the above technical bottlenecks, the present invention proposes an innovative intelligent driving processing system. The core of the system consists of three intelligent collaborative modules: Teaching strategy module: Utilizes advanced multimodal large language model technology to deeply analyze students' multi-dimensional data (such as operating habits and proficiency), dynamically generate customized teaching strategies for students of different abilities, and completely change the unified standard model.
[0030] Safety Perception Module: Targeting complex environments (such as inclement weather and unconventional obstacles), this module innovatively integrates meteorological information, multi-dimensional onboard sensor data, and visual information. This module calculates road friction and slippery risk in real time, accurately analyzes scene visibility, and analyzes the behavioral trends and movement directions of surrounding dynamic targets (vehicles and pedestrians), enabling a refined, graded assessment of risks across different scenarios. Based on this assessment, the module adaptively generates a comprehensive safety response strategy, including speed adjustment, avoidance path planning, braking force, and steering compensation.
[0031] Behavior Management Module: Real-time monitoring of driving operations, combined with the student's physiological state (such as fatigue and attention), provides immediate and accurate behavior guidance and correction through various interactive methods such as voice, visual prompts or tactile feedback.
[0032] This application directly addresses the contradiction between "high-risk omissions" and "low-risk excessive warnings" caused by unified standards through the personalized generation capabilities of the teaching strategy module, ensuring that teaching truly matches the individual level of students. Through the dynamic multi-source data fusion and in-depth scene analysis (friction coefficient, visibility, target behavior prediction) of the safety perception module, it effectively overcomes the problem of low sensor fusion accuracy in interference environments caused by static rules, significantly reducing the risks of "false braking" and "missed detection." The multimodal feedback mechanism of the behavior management module further consolidates the teaching effect and the timeliness of safety intervention. The three modules work together to improve the adaptability, safety, and overall effectiveness of the intelligent driving teaching system.
[0033] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below in conjunction with the accompanying drawings and specific embodiments. Obviously, the embodiments described are only a part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without making any creative efforts are within the scope of protection of the present application.
[0034] Figure 1 FIG1 shows an architecture diagram of an intelligent driving system provided by an embodiment of the present application. Figure 1 As shown, the system includes a teaching strategy module 11, a safety perception module 12 and a behavior management module 13; The teaching strategy module 11 is configured to generate personalized teaching strategies adapted to the students' abilities based on multi-dimensional student data through a multimodal large language model; The multidimensional student data includes learning ability data, driving habit data, and error type data. Specifically, multidimensional student data refers to data representing information such as the student's attention and cognitive ability during the learning process. Learning ability data refers to information that quantifies the student's level of focus, such as an attention concentration index. Driving habit data refers to information that records the student's driving behavior patterns, such as a steering wheel turning frequency index. Error type data refers to information that categorizes the student's operational errors, such as operational error codes.
[0035] The teaching strategy module 11 is specifically used to perform the following process: combining and processing the learning ability data, driving habit data and error type data through a multimodal large language model to form a student state feature set; based on the student state feature set, using a path generation method to construct a personalized teaching path including a basic training stage, a scenario simulation stage and a comprehensive assessment stage; outputting a personalized teaching strategy based on the time distribution and content configuration of the personalized teaching path.
[0036] As a possible solution, before the teaching strategy module 11 "combines and processes the learning ability data, driving habit data and error type data through a multimodal large language model to form a student state feature set", the student's heart rate signal and eye tracking signal output by the wearable physiological monitoring device, as well as the steering wheel angle signal and throttle opening signal output by the vehicle operating device can be collected; the student's heart rate signal is converted into a learning stress level indicator, and the eye tracking signal is converted into a gaze focus time indicator, to jointly form the student's learning ability data.
[0037] Specifically, wearable physiological monitoring equipment is used to collect the trainees' original heart rate signals and eye tracking signals at a specific frequency. The heart rate signal is preprocessed: a low-pass filter is used to remove motion artifacts, the standard deviation SDNN of the RR interval within a time window is calculated, and the result is mapped to a learning stress level index of 0-100 points based on the calibration pressure curve. The eye tracking signal is processed synchronously: the pupil coordinates are identified to determine the state of stay in key areas such as the instrument panel and rearview mirror, blinking and instantaneous saccade interference are eliminated, the effective focus time within every 5 minutes is accumulated, and the result is converted into a 0-100 point line of sight focus index in proportion. Finally, the two indicators are input into the fusion module, and a comprehensive learning ability data package is generated according to the stress level weight and the focus time weight, and stored in the central database.
[0038] The steering wheel angle signal is converted into a steering frequency index, and the throttle opening signal is converted into an acceleration smoothness index, which together form the student's driving habit data; combined with the number of operating errors and error type codes in historical training records, the student's error type data is formed.
[0039] Specifically, steering wheel angle and throttle opening signals are acquired from the vehicle's CAN bus at a specific frequency. Steering wheel signals undergo threshold detection and processing: A valid steering action is recorded when the angle changes continuously exceed ±5°. The number of turns per minute is counted to form a steering frequency index. The throttle opening rate of change is calculated every second: the standard deviation σ within a specific time window is taken and, using the formula: 10-2 × max(σ-2, 0), an acceleration smoothness index (scaled from 1 to 10) is generated. These indicators are combined into JSON-formatted driving habit data. Simultaneously, a historical training database is accessed: all student operation records from the past 30 days are extracted, the total number of errors is counted, and, based on a pre-set error coding table (e.g., code 1 for accidental emergency brake application and code 2 for understeer), the frequency of each error type is categorized and aggregated to form a structured error type dataset. Driving habit data and error type data are bound and stored in the student behavior archive system.
[0040] As a possible implementation scheme, the teaching strategy module 11 is specifically configured to perform the following process when executing the process of "combining and processing the learning ability data, driving habit data, and error type data through a multimodal large language model to form a student state feature set": In the first stage, the attention concentration index in the learning ability data is mapped to the first-dimensional text descriptor to generate the first type of language features; the steering wheel turning frequency index in the driving habit data is mapped to the second-dimensional operation descriptor to generate the second type of language features; the operation error code in the error type data is mapped to the third-dimensional scene descriptor to generate the third type of language features; the first type of language features, the second type of language features and the third type of language features are sequentially spliced through a lightweight pure language model, and the spliced features are converted into structured language units using a rule converter, and a text descriptor set containing multiple labels is output, among which the attention concentration index is converted into a focus level label, the steering frequency index is converted into an operation rhythm label, and the operation error code is converted into an error pattern label; The first dimension, text descriptors, refers to the "person" dimension text descriptors, which are text description features mapped from the attention concentration index. The second dimension, operation descriptors, refers to the "vehicle" dimension operation descriptors, which are operational behavior features mapped from the steering wheel turning frequency index. The third dimension, scene descriptors, refers to the "field" dimension scene descriptors, which are scene context features mapped from the operational error codes.
[0041] A lightweight pure language model is an AI model used to process text feature sequences, while a rule converter is a component that converts text splicing features into structured units.
[0042] The text descriptor set refers to a feature set that includes a focus level label, an operation rhythm label, and an error pattern label. Among them, the focus level label refers to a quantitative indicator that represents the student's focus level, including three levels: (high / medium / low). The operation rhythm label refers to a quantitative indicator that represents the steering wheel operation speed, including (quick / smooth / slow). The error pattern label refers to a classification indicator that represents the error type, including (oversteering / delayed reaction / lack of observation). For example, in a driving teaching simulation system, a student's learning ability data shows that the concentration index is 0.85, indicating high concentration. The steering wheel turning frequency index in the driving habit data is 10 times per minute, indicating a medium operating speed. The operation error code in the error type data is 001, indicating a common steering error.
[0043] In the first phase of the present embodiment, the teaching strategy module 11 uses a lightweight pure language model fine-tuned from an open-source model. It uses an algorithm to abstract three-dimensional features. First, the attention concentration indicator in the learning ability data is mapped to a "person" dimension text descriptor to generate language features representing the level of focus. Simultaneously, the steering wheel turning frequency indicator in the driving habit data is mapped to a "vehicle" dimension operation descriptor to generate language features describing the operation rhythm. Finally, the operation error codes in the error type data are mapped to a "field" dimension scene descriptor to generate scene features reflecting the error pattern. These three types of language features are sequentially concatenated and fed into a rule converter for structured processing. This process combines algorithmic automatic labeling with manually labeled samples to supervise and fine-tune the model. Ultimately, the concentration value is converted into high / medium / low level labels, the steering frequency value is mapped into rapid / smooth / slow rhythm labels, and the operation error codes are parsed into pattern labels such as oversteering / delayed reaction / missed observation. The output is a set of text descriptors containing three-dimensional structured descriptions, providing a linguistic data foundation for teaching strategy generation.
[0044] Furthermore, entering the second stage, the teaching strategy module 11 is also used to synchronously receive the original sensor data stream, the gaze trajectory timing signal output by the eye tracker, the angle change waveform output by the steering wheel sensor, and the incorrect operation video clip captured by the vehicle camera through an end-to-end multimodal vertical model; encode the gaze trajectory timing signal into a first multidimensional data group, encode the angle change waveform into a second multidimensional data group, and encode the incorrect operation video clip into a third multidimensional data group in a unified embedding space; fuse the first multidimensional data group, the second multidimensional data group and the third multidimensional data group to generate a multimodal feature vector; align the text descriptor set with the multimodal feature vector according to the time axis, and generate a student state feature set through a feature binder.
[0045] Among them, the first multidimensional data group corresponds to the label index output by the lightweight model, the second multidimensional data group corresponds to the feature channel output by the end-to-end model, and the third multidimensional data group is bound to the time node of the teaching stage.
[0046] In the embodiments of this application, the end-to-end multimodal vertical model refers to a comprehensive model that processes data from multiple sensors. The raw sensor data stream refers to the data sequence collected from the device. The gaze trajectory timing signal refers to the trainee's gaze movement data output by the eye tracker. The steering angle change waveform refers to the angle change data output by the steering wheel sensor. The incorrect operation video clip refers to the recording of incorrect actions captured by the vehicle's onboard camera.
[0047] The unified embedding space refers to the shared space after data encoding. The first multidimensional data set is the vector representation of the encoded gaze trajectory time series signal. The second multidimensional data set is the vector representation of the encoded angle change waveform. The third multidimensional data set is the vector representation of the encoded error operation video clip. The multimodal feature vector is the feature representation obtained by fusing multiple data sets. The feature binder is a tool for aligning text and sensor features.
[0048] like Figure 2 As shown, the second-stage teaching strategy module 11 simultaneously receives three types of raw data through an end-to-end multimodal vertical model: gaze trajectory time series signals output by the eye tracker, steering angle waveforms captured by the steering wheel sensor, and video clips of incorrect operation captured by the vehicle camera. These data are encoded separately in a unified embedding space: the gaze trajectory is converted into a first multidimensional data set representing attention in the "person" dimension, the steering angle waveform is extracted into a second multidimensional data set describing the operational state in the "vehicle" dimension, and the video clips are parsed by a deep network into a third multidimensional data set depicting the scene context in the "field" dimension. After the three data sets are fused to generate a multimodal feature vector, they are precisely aligned with the set of text descriptors output by the first stage. A feature binder is used to associate the focus level label index generated by the lightweight model with the first multidimensional data set, the operation rhythm label channel is mapped to the second multidimensional data set, and the error mode label node is bound to the third multidimensional data set, ultimately forming a spatiotemporally synchronized feature set of the student's state.
[0049] As a possible solution, in the process of the teaching strategy module 11 "based on the student state feature set, using a path generation method to construct a personalized teaching path including a basic training stage, a scenario simulation stage and a comprehensive assessment stage; outputting a personalized teaching strategy according to the time distribution and content configuration of the personalized teaching path", specifically by using a weighted decision tree algorithm based on the concentration level label, operation rhythm label, error pattern label and multimodal feature vector in the student state feature set to construct a three-stage teaching path including a basic training stage, a scenario simulation stage and a comprehensive assessment stage.
[0050] Among them, the process of constructing a three-stage teaching path first generates a basic ability value by calculating the weighted coefficients of the concentration level and the operation rhythm. The concentration level is assigned a weight of 0.5, 0.3, and 0.2 according to the three levels of high, medium and low, accounting for 60% respectively. The operation rhythm is assigned a weight of 0.3, 0.6, and 0.3 according to the three levels of fast, smooth and slow, accounting for 40%. When the value is lower than 0.3, the direction control and throttle linear basic unit is configured. Between 0.3 and 0.6, the steering acceleration and deceleration coordinated training module is started. When it exceeds 0.6, it jumps directly to the scene simulation stage. At the same time, the time proportion of this stage is reversed based on the total class hours. For example, the basic ability value of 0.42 in the total 20 hours of class time corresponds to 11.6 hours of training allocation.
[0051] Secondly, after entering the scenario simulation stage, the training content is automatically generated based on the error mode label. Oversteering triggers the 50-meter radius curve speed limit training, reaction delay activates the vehicle-following emergency braking module, and the missing configuration of the three-way intersection scanning unit is observed. The scenario difficulty parameters are adjusted in combination with the operation rhythm label. The rapid type scenario is simulated by reducing the speed by 30%, and the slow type scenario extends the response time window by 20%.
[0052] Finally, the comprehensive assessment phase begins. This phase generates a customized route based on defect repair logic. To address oversteer defects, a 50-meter sharp curve with a speed limit is embedded. Furthermore, a sudden pedestrian crossing test point is added to observe missing issues. This results in an assessment plan encompassing core repair items such as cornering speed control and pedestrian response. Finally, a structured teaching strategy is output through the timeline alignment engine.
[0053] In the intelligent driving system, the safety perception module 12 is used to integrate meteorological, on-board multi-dimensional sensor and visual data. By constructing the road friction coefficient and slippery trend, analyzing scene visibility, and performing behavior clustering and motion vector analysis on dynamic targets, it can assess the risk level of different scene types. Based on the assessment results and the preset intervention level, it generates an adaptive safety response strategy covering speed control, avoidance path, braking force and steering compensation.
[0054] In an embodiment of the present application, before generating a safety response strategy that matches the environmental interference based on the scene type and the preset intervention level, it is first necessary to obtain environmental data output by the environmental perception device (such as various sensors, specifically set according to needs); determine the scene type based on the scene feature identifier in the environmental data, wherein the scene types include weather change scenes, human-machine mixed scenes and two-way road scenes, and the scene feature identifier includes identification information possessed by each scene. For example, in a two-way road scene, it is necessary to detect identification information such as two-way lanes and vehicles, so as to determine that the current scene type is a two-way road scene.
[0055] In a specific embodiment, when the scene type is a weather change scene, the safety perception module 12 is specifically used to collect rainfall intensity and fog density parameters in real-time meteorological data, and combine the moving object density index in the spatial obstacle distribution data to generate a weather change scene type code.
[0056] Rainfall intensity refers to the amount of rain falling to the ground per unit time, typically measured in millimeters per hour, and is used to quantify the magnitude of real-time precipitation. Fog density refers to the degree to which visibility is reduced by tiny water droplets or ice crystals suspended in the air. Higher density indicates thicker fog and more restricted vision. The moving object density indicator is derived from spatial obstacle distribution data and refers to the number or concentration of moving obstacles (such as vehicles and pedestrians) within a specific area. The classification identifier that represents the current complex weather and dynamic obstacle conditions is the weather change scene type code.
[0057] In another specific embodiment, when the scene type is a human-machine mixed scene, the safety perception module 12 is specifically used to extract pedestrian movement trajectory characteristics and non-motor vehicle motion vectors, calculate the trajectory intersection density and relative speed change rate, and generate a human-machine mixed scene type code.
[0058] Among them, pedestrian trajectory characteristics refer to the sequence of continuous location points left by pedestrians when they move in a specific area, which is captured and depicted by technical means. Non-motor vehicle motion vectors refer to data describing the motion status of non-motor vehicles such as bicycles and electric bicycles. It contains information on the direction and speed of movement of non-motor vehicles at a specific moment, and together constitute a vector representing its instantaneous motion trend. Trajectory intersection density refers to the frequency with which pedestrian trajectories and non-motor vehicle trajectories intersect or approach potential conflict areas within a set time and space range. The relative speed change rate refers to the speed of the relative movement speed change between pedestrians and adjacent non-motor vehicles. The classification identifier that can identify the complexity of pedestrian-non-motor vehicle interactions and the probability of potential conflict in the current environment is the human-machine hybrid scene type coding.
[0059] In another specific embodiment, when the scene type is a two-way road scene, the lane curvature change characteristics are identified by obtaining the distance parameters and relative speed parameters of the vehicles in the opposite lane to generate the two-way road scene type code.
[0060] The distance parameter for vehicles in the oncoming lane refers to the straight-line distance between the vehicle and the nearest vehicle in the opposite lane, and is used to measure the distance between the two vehicles. The relative speed parameter for vehicles in the oncoming lane refers to the difference in speed between the vehicle and the vehicle in the oncoming lane along the road. The lane line curvature variation characteristic refers to the property of the curvature of the road boundary marking line that changes over time or space. The classification identifier that can identify the complexity of traffic flow interaction and the potential risk level in the current two-way road environment is the two-way road scene type code.
[0061] After determining the scene type codes for each scenario, it is also necessary to synchronously collect the vehicle speed value and road curvature value in the vehicle driving status data, and calculate the intervention degree corresponding to each scenario based on the scene type codes, vehicle speed values and road curvature values of each scenario.
[0062] Specifically, according to the scene type indicated by the scene type code, the corresponding intervention calculation mode is selected: Model 1: Calculate the degree of intervention under weather change scenarios.
[0063] Specifically, the distributed tire sensing unit synchronously collects multi-axis tire pressure values and calculates the maximum dynamic difference of each axial tire pressure value per unit time; the road adhesion coefficient is extracted from the scene type code, and the axial dynamic difference is multiplied by the adhesion coefficient to generate the axial grip risk value. Establish a matching relationship between the pressure fluctuation time series and the curve curvature change: when the axial grip risk value reaches its peak before the curve curvature changes, it is marked as an advance event; when the axial grip risk value reaches its peak after the curve curvature changes, it is marked as a lag event; determine the time series risk coefficient based on the proportional relationship between the advance event and the lag event ; Input the axial grip risk value, vehicle speed value, curve radian value and timing risk coefficient into the intervention synthesizer; In the intervention synthesizer, the following is performed: the axial grip risk value is combined with the vehicle speed coefficient Multiply the first intermediate value and convert the curve curvature coefficient Multiply the second intermediate value by the time series risk coefficient, add the first intermediate value and the second intermediate value to get the total risk value, and multiply the total risk value by the weather hazard weight in the scenario type code Get the corresponding intervention level :
[0064] Mode 2: Calculate the degree of intervention in a human-machine hybrid scenario.
[0065] Specifically, the trajectory intersection density and relative speed change rate are extracted from the scene type code; the pedestrian-non-motor vehicle interaction risk value is calculated as: (trajectory intersection density × relative speed change rate) / safety distance margin; a risk value distribution heat map is established and the density of high-risk areas is identified. ; Inputting the high-risk area density, vehicle speed value and road curvature value into the intervention synthesizer; In the intervention synthesizer, the following steps are executed: multiplying the high-risk area density by the vehicle speed coefficient to obtain an area risk value, and converting the road curvature value into a curvature risk coefficient. , superimpose the regional risk value and the curvature risk value to obtain a mixed risk value, and multiply the mixed risk value by the human-computer interaction weight in the scene type coding Degree of intervention :
[0066] Mode 3: Calculate the degree of intervention in a two-way road scenario.
[0067] Specifically, the oncoming vehicle distance parameter and relative speed parameter are extracted from the scene type code; the meeting time margin is calculated. :(distance to the oncoming vehicle / relative speed)×lane width coefficient; identify the lane line curvature change rate and generate a curve risk coefficient; combine the time margin value, the curve risk coefficient, and the vehicle speed value Input intervention synthesizer; The intervention synthesizer performs the following steps: converting the time margin value into a time risk index, multiplying the curve risk coefficient by the vehicle speed value to obtain a spatial risk index, superimposing the time risk index and the spatial risk index to obtain a meeting risk value, and multiplying the risk value by the meeting risk weight in the scene type code. Degree of intervention :
[0068] After determining the intervention level for each scenario, you can generate a security response strategy that matches the environmental interference based on the scenario type and the preset intervention level. The following are solutions for generating security response strategies that match the environmental interference for different scenario types: As an optional solution, the scene type includes a weather change scene; The safety perception module 12 is specifically used to collect meteorological data and vehicle sensor data, the meteorological data including rainfall intensity and fog density, and the vehicle sensor data including vehicle speed and acceleration; dynamically bind the meteorological data to the vehicle position to generate spatial meteorological parameters; obtain scene images through polarization imaging processing, and perform optical descattering operations to generate defogged images; analyze visibility levels based on the defogged images, and combine the spatial meteorological parameters with the visibility levels to infer the trend of road slippery changes; and generate a safety response strategy that matches the environmental interference based on the degree of intervention corresponding to the weather change scenario and the trend of road slippery changes, the safety response strategy including a speed control curve linked to the trend of road slippery changes and a control strategy for a minimum following distance threshold.
[0069] In the above scheme, rainfall intensity refers to the amount of rain falling to the ground per unit time, typically measured in millimeters per hour, and is used to quantify rainfall intensity. Fog density refers to the concentration of tiny water droplets or ice crystals suspended in the air; higher concentrations indicate denser fog and lower visibility. Spatial meteorological parameters are comprehensive parameters that dynamically link real-time meteorological data, such as rainfall intensity and fog density, with the vehicle's current location via the Global Positioning System (GPS), reflecting the meteorological conditions surrounding the vehicle. Dehazed images are sharpened images generated by capturing the original scene using polarization imaging technology and then processing them using an optical descattering algorithm to effectively eliminate fog and haze interference. Visibility levels are visually recognizable distance classifications derived from dehazed images, indicating the maximum range at which the driver can clearly observe objects ahead. Road slipperiness trend refers to the dynamic change direction of the road friction coefficient, predicted by combining meteorological data and visibility information, and is used to assess the increasing or decreasing risk of slipperiness. Speed control curves are optimized curves that dynamically adjust vehicle speed based on real-time risk, ensuring that speed matches environmental risk. The minimum following distance threshold is the minimum safe following distance a vehicle should maintain under specific conditions, calculated based on dynamic risk. The safety response strategy is a comprehensive control solution that integrates the speed control curve and the minimum following distance threshold to proactively address environmental disturbances.
[0070] In an embodiment of the present application, first, real-time rainfall intensity data, such as 15 mm per hour, and fog density data, such as 0.2 grams per cubic meter, are collected through the on-board meteorological sensor, and at the same time, vehicle speed data of 80 kilometers per hour and acceleration data of 0.5 meters per square second are obtained from the vehicle bus. Secondly, the global positioning system is used to dynamically associate meteorological parameters with the latitude and longitude coordinates of the vehicle to generate spatial meteorological parameters, such as spatiotemporal labels of high rainfall intensity and medium fog density on the K50 section of the highway. Then, the polarization imaging camera is used to capture the haze scattering image of the scene in front, and the DehazeNet convolutional neural network is applied to perform optical descattering processing, and a dehazed image is output, such as a clear display of lane lines within 200 meters. The Canny edge detection algorithm is then used to analyze the maximum visible distance of 200 meters in the dehazed image and map it to visibility level three. The spatial meteorological parameters and visibility levels are then input into the long and short-term memory network model to deduce the trend of slippery road conditions, such as the friction coefficient. In the next 5 minutes, it will drop by 20%, i.e. .
[0071] Assume that the intervention level calculated by extraction model 1 =0.8, and the comprehensive risk value is calculated by combining it with the slippery descent rate of 20% through the risk fusion function: .
[0072] Speed control curve: For example, based on the initial vehicle speed v_0 = 80 kilometers per hour, the speed reduction per second is calculated according to the comprehensive risk value of 0.96: , Generate a linear speed control function: , After 10 seconds, the target vehicle speed drops to 40 kilometers per hour.
[0073] Minimum following distance threshold: Based on the vehicle dynamics model, substituting the wet friction coefficient and intervention level 0.8: The degree of intervention, where v = 22.22 meters per second, g = 9.8 meters per second squared, is calculated to be 150 meters.
[0074] The final safety response strategy was formed: implement a speed control curve with a 10-second linear deceleration to 40 kilometers per hour, and maintain a minimum following distance of 150 meters until the weather sensor detects that the risk has subsided.
[0075] As another optional solution, the scene type includes a human-machine hybrid scene; The safety perception module 12 is specifically used to collect pedestrian behavior data and vehicle camera data, wherein the pedestrian behavior data includes pedestrian movement trajectory and density data; based on the collected pedestrian behavior data, perform cluster analysis on pedestrian behavior patterns to classify pedestrians into a stationary group, a walking group, and a running group; use binocular stereo vision technology to process the vehicle camera data, generate a depth map through parallax calculation to determine the distance between pedestrians and vehicles in the stationary group, the walking group, and the running group, and output a three-dimensional scene feature map containing pedestrian classification information; based on the three-dimensional scene feature map, detect the positions of pedestrians and non-motor vehicles in each classification group The method includes the following steps: evaluating the environmental risk level based on the motion vector change rate of the pedestrians in the walking group and the running group, the speed direction change characteristics of the non-motor vehicle, and the distance between the non-motor vehicle and the vehicle; selectively including the pedestrians in the stationary group in the risk calculation range based on the degree of intervention corresponding to the human-machine hybrid scenario, and generating a safety response strategy that matches the environmental interference based on the environmental risk level, the safety response strategy including an avoidance path and a speed control strategy that are linked to the environmental risk level.
[0076] In the above scheme, pedestrian trajectory refers to the spatial position sequence formed by the continuous movement of pedestrians. Density data refers to the statistical value of the number of pedestrians in a unit area. Pedestrian behavior pattern clustering refers to the classification of pedestrian movement status into groups using machine learning algorithms: a stationary group (no change in position), a walking group (constant speed movement), and a running group (accelerated movement). Binocular stereo vision refers to a technology that uses two cameras to simulate the principle of human eye parallax. A depth map refers to a spatial distance heat map generated through parallax calculation. A three-dimensional scene feature map refers to a three-dimensional environment model that integrates a depth map with pedestrian classification information. A motion vector refers to the composite vector of the speed and direction of a target object. A speed and direction change feature refers to the angular rate of deflection of the non-motor vehicle's motion direction. An environmental risk level refers to a quantitative indicator of danger based on the dynamics of moving objects and distance calculations. An avoidance path refers to a planned trajectory for vehicles to avoid high-risk areas. A speed control strategy refers to a decision-making plan for adjusting vehicle speed based on risk dynamics.
[0077] In an embodiment of the present application, first, a laser radar and a camera are used to collect pedestrian movement trajectories, such as a sequence of position coordinates per second and density data, such as 2 people per square meter. Secondly, the DBSCAN clustering algorithm is used to classify pedestrian behavior patterns: the stationary group is divided based on a position change threshold of 0.1 meters per second, the walking group is divided into a speed of 1-2 meters per second, and the running group is divided into a speed>2 meters per second. For example, 5 stationary pedestrians, 3 walking pedestrians, and 2 running pedestrians are detected. Then, the scene image is captured by a binocular camera, and the SGBM semi-global matching algorithm is applied to calculate the parallax of the left and right views, and a depth map is generated, such as the distance value between pedestrians and vehicles is marked. Then, the pedestrian classification information is integrated to output a three-dimensional scene feature map, such as the pedestrians in the running group are 30 meters away from the vehicle and the direction angle is 60 degrees.
[0078] Target detection is then performed based on this feature map. YOLOv5 is used to identify non-motorized vehicles, such as bicycles, and their motion vectors are calculated. For example, the speed direction changes by 15 degrees per second. Risk assessment is performed by calculating the rate of change of the running group's motion vector: the speed direction changes by 30 degrees within 0.5 seconds. Combined with a distance of 30 meters, the risk formula is used: , Non-motor vehicle risk value: (distance 20 meters), total environmental risk level: 1.5+0.9=2.4 (high risk).
[0079] Then call the intervention level calculated by mode 2 , selective inclusion into the static group: when Ignore the static group and only calculate the dynamic target risk. Based on the risk level 2.4 and the vehicle position, use the A* algorithm to plan the avoidance path: for example, the initial coordinates are (0,0), the target point is (50,10), and the high-risk area is bypassed: the path point sequence , lateral offset .
[0080] Speed control: basic speed , speed reduction ratio Risk Level , control strategy: .
[0081] The final safety response strategy was generated: executing an avoidance path with a lateral offset of 3.6 meters while reducing the vehicle speed to 22 kilometers per hour.
[0082] As another optional solution, the scene type includes a two-way road scene; The safety perception module 12 is specifically used to obtain road surface humidity data, road surface texture data and vehicle inertia measurement data to construct a road friction coefficient curve; analyze the road surface texture through an ultrasonic sensor, determine the road surface micromorphology to calculate the roughness index, and output a slippery characteristic vector; based on the road friction coefficient curve and the slippery characteristic vector, predict the short-term road state change trend to output a slippery risk level; according to the intervention degree corresponding to the two-way road scenario, combined with the slippery risk level and the vehicle interference data of the opposite lane, generate a safety response strategy that matches the environmental interference, the safety response strategy including an adaptive braking force and steering compensation strategy linked to the road slippery risk level, for anti-skid control and safe passing.
[0083] In the above scheme, road surface moisture data refers to the percentage of moisture in the road surface. Road surface texture data refers to the microscopic roughness and concavity of the road surface acquired through ultrasonic scanning. Vehicle inertial measurement data refers to motion state parameters including accelerometer and gyroscope outputs. Road friction coefficient curve refers to a graph depicting how tire-road friction changes with moisture and texture. Roughness index refers to a numerical metric quantifying the degree of microscopic road surface roughness. Slipperiness feature vector refers to a multidimensional representation of slipperiness risk formed by combining roughness and moisture. Short-term road condition trend refers to the direction of friction coefficient decay predicted over the next few seconds based on real-time data. Slipperiness risk level refers to a quantitative hazard classification based on the friction decay trend. Oncoming lane vehicle interference data refers to the relative distance and speed difference between the vehicle and the oncoming vehicle. Adaptive braking force refers to the braking force dynamically adjusted based on slipperiness risk. Steering compensation strategy refers to the steering angle correction applied to offset sideslip risk.
[0084] In this embodiment, a road surface humidity sensor is used to collect surface humidity data, for example, 80%. An ultrasonic sensor is used to scan the road texture, for example, to detect an average bump depth of 0.25 mm. A longitudinal acceleration of 0.3 g and a yaw rate of 5 degrees per second are obtained from the vehicle's inertial measurement unit. The roughness index is then calculated:
[0085] Combined with a humidity of 80%, a slipperiness feature vector is generated, for example, [0.8, 1.25]. The historical friction coefficient data and the current slipperiness feature vector are then input into the long-short-term memory network model to predict short-term road condition trends. For example, if the friction coefficient \mu decreases from 0.6 to 0.45 in the next three seconds, the output slipperiness risk level is 0.7 (range 0-1).
[0086] Then call the intervention degree calculated by model three , and integrate the interference data of the opposite lane: the distance of the opposite vehicle Meters, relative speed km / h (traveling in opposite directions), meeting time margin: .
[0087] Finally, the adaptive braking force is calculated according to the formula: basic braking force , compensation coefficient , final braking force: The steering compensation strategy is based on the vehicle dynamics model:
[0088] Vehicle weight , lateral acceleration , tire cornering stiffness , Generate a safety response strategy: Apply 426N braking force (a 42% increase) during the oncoming vehicle, while increasing the steering wheel compensation angle by 2.3 degrees.
[0089] For the different scenario types corresponding to the above three optional solutions, when the environmental detection device simultaneously identifies multiple scenarios, such as the coexistence of heavy rain weather and human-machine mixed scenarios, the system can set the weights of various scenarios according to needs based on the weighted summation algorithm, thereby obtaining a comprehensive safety response strategy.
[0090] For example, first calculate the comprehensive priority through the risk weight fusion algorithm: the risk of human-machine mixed scenarios such as pedestrian running , giving the highest weight That is, personal safety takes priority, and weather change scenarios such as rainfall risks Weight , comprehensive risk value .
[0091] Then, dynamic arbitration of strategies is performed: the avoidance path is set to take precedence over the speed reduction command, such as human-machine strategy > weather strategy, and the vehicle speed is taken as the weighted average: , the following distance takes the strict maximum value: .
[0092] Finally, a hybrid human-machine strategy generates an avoidance path with a lateral offset of 3.6 meters. The vehicle's speed is then reduced to 30 km / h, and a 150-meter distance between vehicles is maintained, using a weather strategy. Exit conditions, such as pedestrians leaving or rain easing, are monitored in real time. Weight allocation and conflict resolution rules are used to reduce the accident rate associated with the hybrid strategy.
[0093] The above are all examples. The specific process of determining the comprehensive security response strategy under different hybrid scenarios can be set according to needs. In addition, the security response strategy plan under different scenario types can be set according to needs.
[0094] After determining the safety response strategy corresponding to each scenario type, the safety perception module 12 is further used to adjust the weights corresponding to multiple sensors to compensate for environmental interference by fusing lidar data, inertial measurement unit data and synchronous positioning and mapping data, and generate a compensated safety response strategy.
[0095] Specifically, the safety perception module 12 includes: a fusion perception module 121, a fusion positioning module 122 and a decision planning module 123; wherein, the safety perception module 12 is used to perform the following when performing the following operations: Figure 3 The process shown: Send weight configuration instructions to the fusion perception module 121 and the fusion positioning module 122 to adjust the fusion weights of the lidar data, camera data, millimeter wave radar data, inertial measurement unit data, and synchronous positioning and mapping data; The fusion perception module 121 responds to the weight configuration instruction, uses the pre-calibrated lidar-camera-millimeter wave radar parameters, fuses the weighted lidar point cloud data, camera image data and millimeter wave radar point track data, and performs target detection, recognition and trajectory prediction through a deep learning model; The fusion positioning module 122 responds to the weight configuration instruction and fuses the weighted lidar data, the inertial measurement unit data and the high-precision point cloud map constructed based on the laser inertial odometry SLAM framework in areas where satellite positioning fails or is interfered with, to achieve centimeter-level vehicle positioning through point cloud matching; The decision-making planning module 123 receives the target information output by the fusion perception module 121, the high-precision location information output by the fusion positioning module, the scene type and the degree of intervention, and applies target tracking, trajectory prediction and collision detection algorithms to generate a compensated safety response strategy.
[0096] The laser radar data and camera data are obtained by the vehicle-mounted LiDAR sensor and multi-view vehicle-mounted camera respectively, and the millimeter wave radar data is obtained by the 77GHz millimeter wave radar array. The specific data acquisition method is explained. The fusion perception module 121 processes the weighted laser radar point cloud using the CenterPoint++ model and outputs the obstacle center point heat map (such as Figure 2 The red box marks the location of pedestrians), and the YOLOv7 model is used to analyze the compensated image data to accurately identify traffic signs and non-motor vehicle targets.
[0097] To cope with environmental interference, the safety perception module 12 dynamically adjusts the sensor weights: in heavy rain scenarios, the lidar weight is reduced from the baseline value of 0.4 to 0.2 (precipitation interferes with point cloud quality), while the millimeter-wave radar weight is increased from 0.3 to 0.5; in backlit scenarios, the camera weight is reduced to 0.3, and the polarization filter algorithm is activated simultaneously to compensate for image quality. After multi-source data fusion, a centimeter-level accurate three-dimensional grid map is constructed ( Figure 5 road model), providing a reliable environmental representation for decision making.
[0098] The fusion positioning module 122 responds to the weight configuration instruction and fuses the weighted lidar data, the inertial measurement unit data and the high-precision point cloud map constructed based on the laser inertial odometry SLAM framework in areas where satellite positioning fails or is interfered with, to achieve centimeter-level vehicle positioning through point cloud matching; In areas where satellite signals are lost, the fusion positioning module 122 activates the LIO-SAM algorithm to construct a point cloud map. The inertial measurement unit provides initial six-degree-of-freedom pose values, and the laser point cloud optimizes the pose using an iterative closest point algorithm. During the real-time positioning phase, a normal distribution transformation algorithm is used to match the current point cloud frame with the pre-built map, outputting vehicle coordinates with an error of less than ±5 cm. A multi-source fusion algorithm integrates real-time positioning data using a weighted formula: Dynamically weighting is applied to RTK satellite positioning, laser SLAM, and IMU dead reckoning.
[0099] The decision-making planning module 123 receives the target information output by the fusion perception module 121, the high-precision location information output by the fusion positioning module, the scene type and the degree of intervention, and applies target tracking, trajectory prediction and collision detection algorithms to generate a compensated safety response strategy.
[0100] Among them, the decision-making planning module 123 is based on path planning and scene simulation of high-precision maps, supports complex teaching scenarios, and adopts target tracking, trajectory prediction, and collision detection algorithms to realize vehicle deceleration, lane changing and overtaking, and emergency braking decisions.
[0101] Based on the received anti-interference processed target trajectory and centimeter-level vehicle posture data, this module implements dynamic compensation for safety strategies in three scenarios. The following are specific examples: To address weather change scenarios, a longitudinal acceleration jitter suppression module is added to basic linear deceleration strategies, such as 80km / h → 50km / h, to strictly limit the acceleration fluctuation amplitude to less than 0.1 times the acceleration of gravity.
[0102] For mixed human-machine scenarios, when generating the avoidance path, a curvature smoothing constraint algorithm is embedded, such as a lateral offset of 3.6 meters, to force the path curvature change rate to be less than 0.05 radians per second to avoid sideslip caused by sudden turns.
[0103] In two-way road scenarios, steering wheel compensation angle control, such as 2.3 degrees of synchronous fusion of IMU real-time yaw rate feedback, can compress execution delay to less than 50 milliseconds through closed-loop control.
[0104] This compensation mechanism significantly improves the robustness of the strategy in interference environments through dual verification of the preload dynamics model and real-time sensor feedback. The emergency braking false trigger rate is reduced to 1.2 times per thousand kilometers, while ensuring the smoothness of teaching operations.
[0105] In the intelligent driving system, the output of the safety perception module 12 is connected to the input of the behavior management module 13 to transmit the safety response strategy for behavior recognition; The behavior management module 13 is used to perform the following process when identifying driving behavior data and generating multimodal interactive feedback in combination with physiological data and status data to guide and correct driving behavior: The system obtains torque change data output by a steering wheel torque sensor and facial muscle displacement data collected by an on-board camera; converts the torque change data into a steering wheel operation force level, and simultaneously converts the facial muscle displacement data into a mouth corner lift amplitude and a brow wrinkle degree; generates a behavior conflict coefficient based on a deviation between the steering wheel operation force level and a preset standard force range, in combination with the mouth corner lift amplitude and brow wrinkle degree; and triggers a speech feedback generator to perform the following operations when the behavior conflict coefficient exceeds a first threshold: matches a first speech segment in a steering operation description phrase library based on the steering wheel operation force level; matches a second speech segment in a emotional state description phrase library based on the mouth corner lift amplitude and brow wrinkle degree; combines the first speech segment and the second speech segment into a primary correction instruction; and continuously monitors the changing trends of the torque change data and facial muscle displacement data after outputting the primary correction instruction. If the rate of decrease of the behavior conflict coefficient is lower than a second threshold, extracts a related follow-up question segment from a multi-round dialogue template library; combines the related follow-up question segment with a correction operation demonstration segment to generate a secondary correction instruction, and outputs a multi-round interactive speech sequence containing the primary correction instruction and the secondary correction instruction.
[0106] Among them, the torque change data output by the steering wheel torque sensor refers to the change information of the torque or force detected on the vehicle steering wheel, reflecting the fluctuation of the force of the driver's operation of the steering wheel; the facial muscle displacement data collected by the on-board camera is the position movement information of the key points of the driver's face captured by the in-vehicle camera, which is used to evaluate the change of expression; the steering wheel operation force level is an indicator for quantitatively grading the torque change data, indicating the intensity level of the operation force; the amplitude of the upward movement of the corners of the mouth refers to the displacement distance of the corners of the mouth upward, reflecting the degree of smile or joy; the degree of frown between the eyebrows refers to the depth of contraction of the area between the eyebrows, indicating the degree of tension or frowning.
[0107] Among them, the behavior conflict coefficient is an indicator value calculated based on the deviation between the steering wheel operation force level and the preset standard force range, combined with the amplitude of the mouth corners and the degree of frowning between the eyebrows, which is used to measure the inconsistency in driving behavior; the first threshold is the critical value of the set behavior conflict coefficient, which is used to trigger subsequent operations; the voice feedback generator is a component in the system, responsible for generating voice feedback; the steering operation description phrase library is a database that stores phrases related to steering wheel operations; the first voice segment is a short voice segment matched from the library, describing the steering operation; the emotional state description phrase library is a database that stores phrases related to emotion description; the second voice segment is matched from the library A short voice clip is matched to describe the emotional state; the primary corrective instruction is the preliminary guidance voice content formed by combining the first and second voice clips; the multi-round dialogue template library is a database that stores dialogue templates; the associated question clip is a follow-up question voice clip extracted from the library; the correction operation demonstration clip is a demonstration voice clip that shows the correct driving operation; the secondary corrective instruction is the advanced guidance voice content formed by combining the associated question clip and the correction operation demonstration clip; the multi-round interactive voice sequence is the final output multi-round voice sequence containing the primary corrective instruction and the secondary corrective instruction; the second threshold is the critical value of the set behavior conflict coefficient decrease rate, which is used to determine whether additional instructions are needed.
[0108] In this embodiment of the present application, the system first collects a sequence of torque change data from a steering wheel torque sensor. A normalization algorithm is used to compress the raw data to a range of 0 to 100. This data is then converted to a steering wheel force level based on a preset force range mapping table. For example, a torque value of 1.2 Nm is normalized to 85, corresponding to a "strong" level. Simultaneously, the onboard camera uses the Dlib feature point detection algorithm to track 68 key facial points. Using a Euclidean distance calculation engine, the system outputs the degree of upward movement of the mouth corners and the degree of frowning between the brows. A 3mm Y-axis displacement of the mouth corners is detected as a smile, while a 2mm contraction of the brow distance is detected as a frown.
[0109] Next, the behavior conflict coefficient generation algorithm is started, calculating the absolute difference between the current force level and the midpoint of the standard force range of 50, and weighted fusion expression feature values. The specific formula is: , Where β is set to 1.0, representing the weight of the degree of frown between the eyebrows (M), and α is set to 0.5, representing the weight of the upward angle of the mouth corners (S). When the coefficient C exceeds the preset threshold of 10.0, dual-channel voice matching is triggered. The steering operation description phrase library uses a key-value mapping algorithm to index the first voice segment, "Please operate the steering wheel gently," based on the "strong" level. The emotional state description phrase library uses a range matching algorithm, and a frown between the eyebrows of 2 mm triggers the second voice segment, "You look nervous." These two segments are then spliced together using a timeline algorithm to generate preliminary correction instructions.
[0110] The subsequent behavioral conflict coefficient is sampled once per second, and the rate of decline is analyzed through linear regression: If the slope calculated at three consecutive sampling points is less than 0.5 coefficient per second, the dialogue template library activates the semantic association algorithm to extract the follow-up question segment "Do you need help?" and combines it with the pre-stored demonstration segment "Lightly grip the steering wheel edge" to form a secondary correction instruction, ultimately forming a complete multi-round interaction sequence. For example, when the initial force level is 85, the frown between the eyebrows is 2 mm, and the corners of the mouth are raised 1 mm: , If the coefficient drops less than 7.5 within 15 seconds after the voice message is triggered, additional instructions will be given.
[0111] The behavior management module 13 generates a behavior conflict coefficient based on the deviation between the steering wheel operation force level and the preset standard force range, combined with the upward range of the mouth corners and the degree of frown between the eyebrows, and specifically performs the following steps: Figure 4 The process shown: The original current signal generated by the vocal cord vibration is collected through the throat contact sensor, and the number of fluctuation segments with a frequency exceeding a preset Hz in the original current signal is extracted as a fast jitter count; the ratio of the fast jitter count to the total number of signal segments is calculated to generate a sound fluctuation value, the steering wheel operation force level is subtracted from the median of the standard force range to obtain a force deviation value, and the force deviation value is multiplied by the sound fluctuation value to generate an operation tension parameter; the mouth corner position coordinates and the glabella area coordinates output by the facial recognition device are obtained; the Euclidean distance between the mouth corner position coordinates and the reference smile position is calculated to generate a droop index, and the glabella area coordinates per unit time are measured. The contraction displacement generates a shrinkage strength index, and the droop index is added to the shrinkage strength index to generate an expression conflict value; when it is detected that the fast jitter count value is greater than zero, the three-source synthesizer multiplies the operation tension parameter with the expression conflict value to obtain a first intermediate product, and then multiplies the first intermediate product with the sound fluctuation value to obtain the behavior conflict coefficient; when it is not detected that the fast jitter count value is greater than zero, the three-source synthesizer multiplies the operation tension parameter with the expression conflict value to obtain a second intermediate product, and then multiplies the second intermediate product with a predefined fixed compensation factor to obtain the behavior conflict coefficient.
[0112] The laryngeal contact sensor collects the raw current signal generated by vocal cord vibration. The rapid jitter count extracts the number of fluctuation segments in the signal with a frequency exceeding a preset hertz. The sound fluctuation value is the ratio of the rapid jitter count to the total number of signal segments. The force deviation value is calculated by subtracting the median of the standard force range from the steering wheel force level. The operational tension parameter is generated by multiplying the force deviation value by the sound fluctuation value. The facial recognition device outputs the coordinates of the mouth corner position and the glabella area. The droop index is generated by calculating the Euclidean distance between the mouth corner position coordinates and the baseline smile position. The wrinkle intensity index is the amount of contraction displacement of the glabella area coordinates per unit time. The expression conflict value is obtained by adding the droop index and the wrinkle intensity index. The three-source synthesizer is configured to: when the rapid jitter count value is greater than zero, the operational tension parameter, the expression conflict value, and the sound fluctuation value are multiplied in sequence to generate the behavioral conflict coefficient; otherwise, the product of the operational tension parameter and the expression conflict value is multiplied by a fixed compensation factor to generate the behavioral conflict coefficient.
[0113] In the embodiment of the present application, after the throat contact sensor collects the original current signal, it first applies a bandpass filtering algorithm to extract the fluctuation segments whose frequency exceeds a preset threshold (such as 200Hz), and uses zero-crossing detection technology to count the number of these segments to generate a fast jitter count. ; Perform division operation by proportional calculator Generate sound fluctuation values, where is the total number of signal fragments.
[0114] At the same time, the steering wheel sensor outputs the operating force level L, and the subtractor calculates the difference between it and the pre-stored standard median value. The algebraic difference is used to obtain the velocity deviation value, and the multiplier is used to multiply it with the sound fluctuation value to generate the operation tension parameter. The facial recognition device uses a feature point tracking algorithm to obtain the coordinates of the mouth corners. and glabella coordinates The geometry processor calculates the mouth corner to the reference smile position through vector operations The Euclidean distance of the sag index is generated At the same time, the differential processor derives the displacement modulus of the glabella area according to the time series coordinate data to generate the wrinkle intensity index ; The adder sums the two and outputs the expression conflict value .
[0115] Three-source synthesizer start-up condition branch logic: first detect Calculate the behavior conflict coefficient; otherwise, pass the secondary multiplier and compensator according to Calculation. For example, when hour ,like but ,when Pixels, Pixel time , and finally generate the conflict coefficient .
[0116] Furthermore, the system also includes a teaching execution monitoring module 14, which is used to perform the following operations during the execution of the personalized teaching strategy: Figure 5 The process shown is as follows: during the execution of the basic training stage, the scenario simulation stage or the comprehensive assessment stage, multimodal data including the trainee's voice commands, images and text input data are collected through a multimodal data fusion engine; the multimodal data are processed using a deep learning algorithm to identify the user's operation intention; the environmental information inside the vehicle and the external driving scene are perceived through sensors such as cameras and radars; and a cognitive state data set representing the trainee's current cognitive state and scene interaction needs is generated and updated based on the user's operation intention, the in-vehicle environment and the external driving scene information.
[0117] The basic training phase refers to the initial training portion of the teaching path. The scenario simulation phase refers to the simulated driving portion of the teaching path. The comprehensive assessment phase refers to the test and evaluation portion of the teaching path.
[0118] A multimodal data fusion engine is a system that integrates voice, image, and text data. Multimodal data refers to a collection of information consisting of voice commands, images (including facial expressions and driving posture), and text input. User intent refers to the student's action goals during interaction. The in-vehicle environment refers to information about the vehicle's internal state, such as the student's location and focus of attention. The external driving scene refers to information about traffic conditions outside the vehicle, such as current speed and road type. The cognitive state data set refers to data that represents the student's cognitive needs and scene interactions, including their level of understanding and attention.
[0119] In the embodiment of the present application, during the basic training phase (the initial training portion of the teaching path), the teaching execution monitoring module 14 focuses on the student's theoretical knowledge learning and basic operation training. This process uses a multimodal data fusion engine to collect the student's voice commands (such as verbal answers to questions), image data (including facial expressions and driving posture images to capture the student's attention distribution in a static environment), and text input data (such as simulated operation logs) in real time. Next, deep learning algorithms (such as multimodal data fusion engines such as convolutional neural networks and natural language processing models) are used to process this multimodal data to identify the user's operational intentions (for example, whether the student is trying to understand traffic rules or perform simple control operations). At the same time, cameras and radar sensors are used to perceive environmental information within the vehicle (such as the student's sitting position and whether their attention focus deviates from the teaching interface) as well as the external driving scene (although the external scene is relatively simple at this stage, such as simulating a stationary road type). Finally, the module combines the recognized user operation intentions, the in-vehicle environment, and the external scene information to generate and update the cognitive status data set: for example, if the trainee's facial expression shows confusion and the voice command is vague, combined with the distracted focus in the car, the system will update the "level of understanding" to low and the "attention concentration" to weak to reflect cognitive needs; this data set is fed back to the teaching strategy module in real time to adjust the basic training content (such as repeating theoretical modules or simplifying operation steps).
[0120] During the scenario simulation phase (the simulated driving portion of the instructional path), the instructional execution monitoring module 14 focuses on the student's interactions within the dynamic simulated environment. This process begins by using a multimodal data fusion engine to collect richer voice commands (such as the student's verbal commands during the simulated driving), image data (including real-time facial expressions and driving posture to monitor emotional changes under stress), and text input data (such as simulated steering wheel or pedal inputs). Deep learning algorithms (such as recurrent neural networks and intent recognition models) then process this data to identify the user's intended action (for example, the student's intended goal when changing lanes or avoiding obstacles). Simultaneously, cameras and radar sensors perceive the in-vehicle environment (such as the student's body position and whether their attention is focused on the simulated instrument panel) and the external driving scene (such as the current simulated vehicle speed, road type, and traffic flow information). The module combines the user's operational intentions, the in-vehicle environment, and external scene information to generate and update a cognitive status data set. For example, when a complex intersection appears in the external scene, if the student's voice command is hesitant and the image shows a tense posture, the system will update the "understanding level" to medium and the "attention concentration" to high but easily distracted to capture the scene interaction needs (such as the need for more lane change practice). This data set is dynamically updated to optimize the difficulty of scene simulation or provide instant feedback.
[0121] During the comprehensive assessment phase (the test and evaluation component of the teaching path), the teaching execution monitoring module 14 emphasizes the assessment of the student's actual driving ability. This process uses a multimodal data fusion engine to intensively collect voice commands (such as the student's decision-making statements during the assessment), image data (including high-precision facial expressions and driving posture images to assess cognitive load under pressure), and text input data (such as assessment log entries). Deep learning algorithms (such as ensemble learning models and intent classifiers) process this data to identify the user's operational intent (for example, the student's response goals in an emergency situation). Simultaneously, cameras and radar sensors perceive the in-vehicle environment (such as whether the student's attention is firmly focused on key controls) and the external driving scene (such as actual vehicle speed, road type, and real-time traffic events). The module combines user operational intent, in-vehicle environment, and external scene information to generate and update a cognitive status data set. For example, in an external highway scenario, if the student's operational intent is incorrect and the image indicates distraction, the system will update the "understanding level" to low and the "attention concentration" to fluctuating to quantify the interaction requirements of the scenario (such as revealing knowledge gaps). This data set serves as the evaluation output and is used to generate the final assessment report and subsequently adjust the teaching strategy.
[0122] The above process ensures that the cognitive state data set can accurately represent the students' cognitive state and scenario requirements in real time at each stage, supporting the dynamic optimization of personalized teaching paths.
[0123] Furthermore, the system also includes a teaching strategy adjustment module 15 for executing the following based on the cognitive state data set generated by the teaching execution monitoring module 14: Figure 6 The progressive process shown is as follows: based on the current vehicle speed information in the cognitive state data set, an interactive mode switching instruction is generated and executed, and the interactive mode switching instruction controls the interactive interface to enable vertical screen display and voice priority interaction at high speed, and enable horizontal screen display and touch interaction at low speed; based on the user operation intention information and cognitive state information in the cognitive state data set, a teaching content adjustment instruction is generated and executed, and the teaching content adjustment instruction adjusts the detail level, presentation method or intensity of auxiliary information of the teaching prompts; based on the evaluation results of the student's operation proficiency and error mode of the cognitive state data set, a teaching mode switching instruction is generated and executed, and the screen is linked to display the teaching prompt content corresponding to the mode; the map engine service is synchronously called to load a high-resolution static map and display the vehicle posture.
[0124] Among them, the cognitive state data set refers to a dynamic data set that includes quantitative indicators such as student understanding, concentration, and operational proficiency. The interactive mode switching instruction refers to the control signal that automatically adjusts the human-computer interaction form according to the current vehicle speed, realizing the switching between high-speed vertical screen voice interaction and low-speed horizontal screen touch interaction. The teaching content adjustment instruction refers to the decision-making instruction that dynamically adjusts the level of detail of the teaching prompts, such as simplified or enhanced version, the presentation method, such as text / 3D animation, and the auxiliary intensity, such as the prompt frequency, based on the student's cognitive state. The teaching mode switching instruction refers to the control instruction for switching between basic teaching / intensive training / independent practice modes based on the operational proficiency assessment results. The map engine service refers to the background system that provides centimeter-level high-precision static maps and real-time rendering of vehicle posture.
[0125] The map engine service is used to load high-resolution static maps in all scenarios and display vehicle postures; personalized scenario teaching, automatically switching between "primary teaching" and "advanced teaching" modes according to changes in driving ability, such as Figure 7 As shown, the linked screen prompts content.
[0126] In one embodiment of the present application, the teaching strategy adjustment module performs a progressive adjustment process based on the cognitive state data set generated by the teaching execution monitoring module. First, the interaction mode is dynamically switched based on the current vehicle speed: when the vehicle speed reaches or exceeds 60 kilometers per hour, the system automatically activates the portrait display interface and prioritizes voice interaction, such as simplifying instrument information and providing a voice broadcast of road conditions every 20 seconds in a rainstorm on a highway. When the vehicle speed drops below 60 kilometers per hour, the system switches to landscape mode and activates touch interaction, such as expanding the 3D steering animation teaching panel and enlarging the touch buttons by 50% for easier operation during rainy and foggy urban road training.
[0127] Secondly, the teaching content is adjusted based on the user's operational intent and cognitive state. For students whose comprehension level falls below 0.5, enhanced prompts are triggered, showcasing tire track mechanics animations and step-by-step voice explanations in a rainstorm reversing scenario. For rule-based intent, graphic and text comparison cards are used, such as a comparison chart of braking distances on slippery roads. The prompt frequency is automatically doubled when attention level drops below 0.6. The teaching mode is then switched based on the proficiency assessment results. When a student's proficiency level reaches 0.8 and their error rate falls below 2 per minute in continuous cornering practice, the system switches to independent practice mode and hides the basic guidance layer. If the proficiency level falls below 0.4 or the error rate exceeds 5 per minute, the enhanced training mode is activated, dynamically overlaying semi-transparent guidance on the screen, such as the real-time marking of the steering angle safety threshold for meeting on a rainy night. Finally, the map engine is synchronously called to load centimeter-level static maps and render the vehicle's position in real time. For example, during complex driving school tests, the deviation between the current training path and the vehicle's heading angle is highlighted.
[0128] Furthermore, the system also includes a teaching task collaborative execution module 16 for executing the following based on the interactive mode switching instructions, teaching content adjustment instructions and teaching mode switching instructions generated by the teaching strategy adjustment module 15 and the cognitive state data set generated by the teaching execution monitoring module 14: Figure 8 The progressive process shown: Through the decision-making planning engine, the contextual semantics of the trainees' instructions in teaching-related interactions are analyzed by combining the user operation intention information in the cognitive state data set with the current teaching steps recorded by the system; based on the current teaching strategy defined by the interactive mode switching instructions, teaching content adjustment instructions and teaching mode switching instructions generated by the teaching strategy adjustment module and the predefined task priority algorithm, the execution order and output logic of the teaching guidance tasks, safety control tasks and route guidance tasks that need to be executed in parallel are coordinated to avoid task conflicts; in the process of coordinating the execution of tasks, the perception engine service is called to obtain the identification information of the key perception targets around the vehicle; the coordinated task instructions are executed in a linked manner, and the obtained identification information of the key perception targets is presented on the interactive interface.
[0129] Among them, the decision-making planning engine refers to the arbitration system that coordinates conflicts between parallel tasks; contextual semantic analysis refers to the algorithm that analyzes the association between instructions and teaching steps; the task priority algorithm refers to the rules that define the execution weights of teaching, safety, and navigation tasks; the teaching guidance task refers to the functional instruction that outputs operation prompts to students; the safety control task refers to the protection instruction that triggers operation intervention; the route guidance task refers to the path guidance instruction that provides spatial navigation; the perception engine service refers to the system that calls the environmental perception data interface; the key perception target refers to objects around the vehicle such as pedestrians and vehicle traffic lights; the identification information refers to the position, speed, and trajectory data of the perception target; the decision-making planning engine analyzes the operation intention and the current teaching step; the engine executes contextual semantic analysis to infer the logical association of instructions; calls the task priority algorithm to coordinate the execution order of the three types of tasks; the weight of the safety control task is always higher than that of the teaching guidance task; the perception engine service requests the environmental perception data stream; the fusion sensor captures the identification information of the key perception target; the task arbitration result synchronously outputs the teaching prompts and safety warnings; the interactive interface superimposes the target location mark and navigation instructions; non-urgent teaching instructions are immediately frozen when the safety task is triggered; AR navigation arrows and pedestrian outline boxes are rendered to the interface.
[0130] The decision-making and planning engine is a central arbitration system that coordinates the parallel execution of multiple tasks, including instruction, safety, and navigation. It resolves command conflicts through contextual semantic analysis and a prioritization algorithm. Contextual semantic analysis analyzes the logical relationship between user instructions and the current instructional step, for example, interpreting the ambiguous instruction "speak louder" as increasing the voice broadcast volume. The task prioritization algorithm is a rule base that defines the execution weights of three types of tasks: instructional guidance, safety control, and route guidance. Safety control always takes precedence. Instructional guidance tasks provide student operational prompts, such as displaying "Please release the parking brake." Safety control tasks trigger emergency intervention protection commands, such as automatic braking and freezing the accelerator. Route guidance tasks provide spatial navigation path commands, such as generating AR arrows. The perception engine service is an interface system for accessing environmental perception data. Key perception targets include dynamic objects such as pedestrians and vehicles around the vehicle. Identification information refers to data such as the target's location, speed, and trajectory.
[0131] In the embodiment of the present application, contextual semantic analysis is first used to associate user instructions with the current teaching step. For example, when a student asks "What should I do next" in the "Release the Brake" step, the system analyzes the operational guidance intent based on the vehicle status (handbrake applied, engine off) and automatically generates the teaching task "Prompt to Release the Handbrake Process." Secondly, the task priority algorithm is called to arbitrate multi-task conflicts. For example, if the vehicle is detected in an unsafe area, the safety control task weight W_s = 1.0 forcibly overrides the teaching task weight W_t = 0.64, immediately freezing the "Handbrake Operation" teaching prompt and activating the safety warning voice "Do not start, there are pedestrians nearby." At the same time, the perception engine service is requested to fuse sensor data. For example, the millimeter-wave radar identifies a pedestrian approaching at 0.5m / s within 3 meters, and the lidar locates its coordinates (X: 2.8, Y: -1.5). The final coordinated output is: maintain the automatic brake lock, replace the original operational prompt with the safety instruction "Please wait for the pedestrian to leave", and render the pedestrian outline in red and the escape route arrow in green on the AR interface.
[0132] Furthermore, the system also includes an emotional service intervention module 17, which is used to generate student emotional state data based on the student facial image and voice data collected by the teaching execution monitoring module 14 through emotion recognition technology analysis; based on the student emotional state data, generate and trigger soothing service instructions, which control the teaching system to perform operations, including switching the voice interaction tone to a softer mode and providing targeted encouraging guidance content; combined with the cognitive state data set and emotional state data generated by the teaching execution monitoring module, predict operational errors that may occur to students in the current teaching scenario; based on the prediction results, generate and push predictive operation reminder information to the interactive terminal.
[0133] In the above scheme, the emotional service intervention module refers to a system that generates emotional state data by analyzing students' facial expressions and voice features through emotion recognition technology. The emotional state data includes quantitative indicators such as tension and frustration; the soothing service instructions refer to the interactive strategy control signals triggered by the emotional state, including switching to a soft voice tone or pushing encouraging content; the predictive operation reminder refers to the proactive warning information generated after combining cognitive state and emotional data to predict potential operational errors.
[0134] In the embodiment of the present application, firstly, the facial image of the trainee is captured in real time by the in-car camera, such as the facial features of frowning and drooping mouth corners during rainstorm curve training, and at the same time, the microphone collects trembling voice, such as the trainee growling "the steering wheel is too heavy", and the convolutional neural network is used to analyze the facial muscle activity to calculate the tension index of 0.82. The long-term and short-term memory voiceprint model is used to analyze the voice fundamental frequency fluctuation to output the frustration probability of 0.76, and generate a quantitative emotional state data set containing the two dimensions of tension and frustration; secondly, based on the data, the soothing service instruction is immediately triggered: the voice interaction system switches to a soft mode, for example, the harsh instruction "brake immediately" is converted into a gentle and gentle "Please brake lightly, we will take it slowly", and simultaneously push targeted encouragement content. For example, based on the trainee's operating proficiency of 0.3, a graphic prompt is generated, "You have mastered 80% of the steering skills, just keep the angle and you can pass!"; then the understanding degree 0.4 in the cognitive state data and the current emotional data are integrated, and the time series prediction model is input to predict the risk of operational errors. For example, the probability of understeering in heavy rain is 91%; finally, based on the prediction results, predictive operation reminders are generated and pushed to the AR interactive terminal. For example, a dynamic yellow arrow and text warning "It is easy to understeer in curves, it is recommended to increase the steering wheel by 10°" are projected on the windshield, completing the closed loop from emotion recognition to active intervention.
[0135] Based on the same concept, the embodiment of the present application provides an intelligent driving method. Figure 9 A flowchart of an intelligent driving method provided in an embodiment of the present application is shown in FIG. Figure 9 As shown, the method includes: S901. Generate personalized teaching strategies tailored to students' abilities based on multi-dimensional student data using a multimodal large language model. S902: During the execution of the personalized teaching strategy, the vehicle integrates meteorological, multi-dimensional onboard sensor, and visual data to assess the risk level of different scenario types by constructing road friction coefficients and slippery trends, analyzing scene visibility, and performing behavioral clustering and motion vector analysis on dynamic targets. Based on the assessment results and the preset intervention level, an adaptive safety response strategy encompassing speed control, avoidance paths, braking force, and steering compensation is generated. S903. During the execution of the safety response strategy, the driving behavior data is identified, and multimodal interactive feedback is generated in combination with the physiological data and the status data to guide and correct the driving behavior.
[0136] In the above scheme, multi-dimensional student data refers to a collection of information such as historical driving records, physiological characteristics, and cognitive status; personalized teaching strategies refer to phased training plans (basic / scenario / assessment) that adapt to the differences in students' abilities; complex scenarios refer to preset high-risk environments such as weather changes and human-machine mixing; safety response strategies refer to dynamic plans such as speed control and avoidance paths that are linked to environmental interference; and multimodal interactive feedback refers to real-time behavior guidance signals that integrate voice, vision, and touch.
[0137] In an embodiment of the present application, a multimodal large language model is first used to analyze the student's historical operation data, physiological indicators, and cognitive assessment results to generate a personalized teaching strategy. For example, a three-stage training plan (basic training → rainstorm scenario simulation → slippery road assessment) is customized for students with weak vehicle-following ability in rainy days. Secondly, the driving environment is monitored in real time during execution. When a specific complex scenario is detected, such as a rainfall intensity exceeding 20 mm per hour, the safety perception module identifies the weather change scenario type and calls a preset intervention level to generate a safety response strategy. For example, the vehicle speed is linearly reduced from 80 kilometers per hour to 50 kilometers per hour and the following distance is increased to 150 meters in conjunction with the slippery road trend. Finally, when executing the safety strategy, the student's driving behavior data, such as sudden braking operations, and physiological data, such as a sudden increase in heart rate to 120 beats per minute, are simultaneously identified, triggering multimodal interactive feedback including a soft voice prompt "brake slowly", a green brake gradient bar projected on the AR windshield, and 3 Hz vibration feedback on the steering wheel, forming a closed-loop teaching optimization of "strategy generation-scenario intervention-behavior correction".
[0138] The present application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any of the above-mentioned intelligent driving methods when executing the computer program.
[0139] The present application also provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps of any one of the above-described intelligent driving methods are implemented.
[0140] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, a read-only memory, a random access memory, a mobile hard disk, a magnetic disk, or an optical disk.
[0141] An embodiment of the present invention further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps in any of the above-mentioned intelligent driving method embodiments are implemented.
[0142] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present invention.
[0143] The above is a detailed introduction to an intelligent driving system, method and electronic device provided by the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method and core ideas of the present application. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the present application.
Claims
1. An intelligent driving processing system, characterized in that: Includes teaching strategy module, safety perception module and behavior management module; The teaching strategy module is used to generate personalized teaching strategies adapted to students' abilities based on multi-dimensional student data through a multimodal large language model; The safety perception module is used to integrate meteorological, on-board multi-dimensional sensor and visual data to assess the risk level of different scene types by constructing road friction coefficient and slippery trend, analyzing scene visibility, and performing behavior clustering and motion vector analysis on dynamic targets. Based on the assessment results and the preset intervention level, a safety response strategy is generated, wherein the safety response strategy includes one or any combination of speed control, avoidance path, braking force and steering compensation; The behavior management module is used to identify driving behavior data and generate multimodal interactive feedback in combination with physiological data and status data to guide and correct driving behavior.
2. The system according to claim 1, wherein: The scene types include weather change scenes; The safety perception module is specifically used to collect meteorological data and vehicle sensor data, the meteorological data including rainfall intensity and fog density, and the vehicle sensor data including vehicle speed and acceleration; Dynamically binding the meteorological data to the vehicle position to generate space meteorological parameters; Acquire a scene image through polarization imaging processing and perform an optical descattering operation to generate a defogging image; Analyzing the visibility level based on the defogging image and inferring the road slipperiness trend by combining the spatial meteorological parameters with the visibility level; Based on the intervention degree corresponding to the weather change scenario and the slippery road surface change trend, a safety response strategy matching the environmental interference is generated. The safety response strategy includes a speed control curve linked to the slippery road surface change trend and a control strategy for the minimum following distance threshold.
3. The system according to claim 1, wherein: The scenario types include human-machine hybrid scenarios; The safety perception module is specifically used to collect pedestrian behavior data and vehicle-mounted camera data, wherein the pedestrian behavior data includes pedestrian movement trajectory and density data; Based on the collected pedestrian behavior data, a cluster analysis is performed on pedestrian behavior patterns to classify pedestrians into stationary, walking, and running groups. Processing the vehicle camera data using binocular stereo vision technology, generating a depth map by parallax calculation to determine the distances between the pedestrians in the stationary group, the walking group, and the running group and the vehicle, and outputting a three-dimensional scene feature map containing pedestrian classification information; Based on the three-dimensional scene feature map, the position, target type and motion vector of pedestrians and non-motor vehicles in each classification group are detected, wherein the motion vector of the non-motor vehicle includes a speed direction change feature; Assessing the environmental risk level based on the motion vector change rate of the pedestrians in the walking group and the running group, the speed direction change characteristics of the non-motor vehicle, and the distance between the non-motor vehicle and the vehicle; Based on the degree of intervention corresponding to the human-machine hybrid scenario, the stationary group of pedestrians is selectively included in the risk calculation scope, and a safety response strategy matching the environmental interference is generated according to the environmental risk level. The safety response strategy includes an avoidance path and speed control strategy linked to the environmental risk level and matching the environmental interference.
4. The system according to claim 1, wherein: The scene type includes a two-way road scene; The safety perception module is specifically used to obtain road surface humidity data, road surface texture data and vehicle inertia measurement data to construct a road friction coefficient curve; Analyzing the road surface texture by an ultrasonic sensor, determining the road surface micromorphology to calculate a roughness index, and outputting a slipperiness feature vector; Based on the road friction coefficient curve and the slipperiness characteristic vector, predicting a short-term road condition change trend to output a slipperiness risk level; Based on the intervention level corresponding to the two-way road scenario, combined with the slippery risk level and vehicle interference data from the oncoming lane, a safety response strategy that matches the environmental interference is generated. The safety response strategy includes an adaptive braking force and steering compensation strategy linked to the road slippery risk level for anti-skid control and safe passing.
5. The system according to any one of claims 2 to 4, characterized in that: The safety perception module is also used to adjust the weights corresponding to multiple sensors to compensate for environmental interference by fusing lidar data, inertial measurement unit data and synchronous positioning and mapping data, and generate a compensated safety response strategy.
6. The system according to claim 5, characterized in that The security perception module includes: a fusion perception module, a fusion positioning module and a decision-making planning module; The safety perception module is specifically used to perform the following process when fusing lidar data, inertial measurement unit data, and synchronous positioning and mapping data, adjusting the weights corresponding to multiple sensors to compensate for environmental interference, and generating a compensated safety response strategy: Send weight configuration instructions to the fusion perception module and fusion positioning module to adjust the fusion weights of lidar data, camera data, millimeter wave radar data, inertial measurement unit data, and simultaneous positioning and mapping data; The fusion perception module responds to the weight configuration instruction, utilizes the pre-calibrated lidar-camera-millimeter-wave radar parameters, fuses the weighted lidar point cloud data, camera image data and millimeter-wave radar point track data, and performs target detection, recognition and trajectory prediction through a deep learning model; The fusion positioning module responds to the weight configuration instruction and fuses the weighted lidar data, inertial measurement unit data and high-precision point cloud map built based on the laser inertial odometry SLAM framework in areas where satellite positioning fails or is interfered with, to achieve centimeter-level vehicle positioning through point cloud matching; The decision-making planning module receives the target information output by the fusion perception module, the high-precision location information output by the fusion positioning module, the scene type and the degree of intervention, and applies target tracking, trajectory prediction and collision detection algorithms to generate a compensated safety response strategy.
7. The system according to claim 1, wherein: The multi-dimensional student data includes learning ability data, driving habit data, and error type data; The teaching strategy module is specifically used to perform combined processing of the learning ability data, driving habit data and error type data through a multimodal large language model to form a student state feature set; Based on the student status feature set, a path generation method is used to construct a personalized teaching path including a basic training stage, a scenario simulation stage and a comprehensive assessment stage; According to the time distribution and content configuration of the personalized teaching path, a personalized teaching strategy is output.
8. The system according to claim 7, characterized in that The teaching strategy module is specifically used to perform the following process when performing combined processing of the learning ability data, driving habit data, and error type data using a multimodal large language model to form a student state feature set: The attention concentration index in the learning ability data is mapped to the first-dimensional text descriptor to generate the first-category language feature; the steering wheel turning frequency index in the driving habit data is mapped to the second-dimensional operation descriptor to generate the second-category language feature; the operation error code in the error type data is mapped to the third-dimensional scene descriptor to generate the third-category language feature; Sequentially concatenating the first, second, and third language features through a lightweight pure language model, converting the concatenated features into structured language units using a rule converter, and outputting a text descriptor set containing multiple labels; Through the end-to-end multimodal vertical model, the original sensor data stream, the gaze trajectory timing signal output by the eye tracker, the angle change waveform output by the steering wheel sensor, and the erroneous operation video clips captured by the vehicle camera are synchronously received; In a unified embedding space, the gaze trajectory timing signal is encoded into a first multidimensional data group, the angle change waveform is encoded into a second multidimensional data group, and the incorrect operation video clip is encoded into a third multidimensional data group; the first multidimensional data group, the second multidimensional data group and the third multidimensional data group are fused to generate a multimodal feature vector; the text descriptor set and the multimodal feature vector are aligned along the time axis, and a student state feature set is generated through a feature binder.
9. The system according to claim 7, wherein: The system also includes a teaching execution monitoring module, which is used to perform the following processes during the execution of the personalized teaching strategy: During the execution of the basic training phase, the scenario simulation phase, or the comprehensive assessment phase, multimodal data including the trainee's voice commands, images, and text input data are collected through a multimodal data fusion engine; Processing the multimodal data using a deep learning algorithm to identify user operation intentions; Perceive the interior environment and external driving scene through camera and radar sensors; Combined with the user's operation intention, the in-vehicle environment and the external driving scene information, a cognitive state data set representing the trainee's current cognitive state and scene interaction needs is generated and updated.
10. The system according to claim 9, characterized in that The teaching strategy adjustment module is further included, which is used to execute the following progressive process based on the cognitive state data set generated by the teaching execution monitoring module: generating and executing an interaction mode switching instruction based on current vehicle speed information in the cognitive state data set, wherein the interaction mode switching instruction controls the interaction interface to enable portrait display and voice-first interaction at high speeds, and to enable landscape display and touch interaction at low speeds; Generate and execute a teaching content adjustment instruction based on the user operation intention information and cognitive state information in the cognitive state data set, wherein the teaching content adjustment instruction adjusts the detail level, presentation method or strength of auxiliary information of the teaching prompt; Based on the evaluation results of the student's operation proficiency and error patterns of the cognitive state data set, a teaching mode switching instruction is generated and executed, and the screen is linked to display the teaching prompt content corresponding to the mode; Synchronously call the map engine service to load a high-resolution static map and display the vehicle's position.
11. The system according to claim 10, wherein: The system further includes a teaching task collaborative execution module, which is configured to execute the following progressive process based on the interactive mode switching instructions, teaching content adjustment instructions, and teaching mode switching instructions generated by the teaching strategy adjustment module, and the cognitive state data set generated by the teaching execution monitoring module: By combining the user operation intention information in the cognitive state data set with the current teaching steps recorded by the system, the decision planning engine analyzes the contextual semantics of the student's instructions in the teaching-related interaction; Coordinate the execution order and output logic of the teaching guidance tasks, safety control tasks, and route guidance tasks that need to be executed in parallel based on the current teaching strategy defined by the interactive mode switching instructions, teaching content adjustment instructions, and teaching mode switching instructions generated by the teaching strategy adjustment module, and a predefined task priority algorithm, to avoid task conflicts; During the coordinated execution of tasks, the perception engine service is called to obtain identification information of key perception targets around the vehicle; The coordinated task instructions are executed in a linked manner, and the identification information of the key perception targets obtained is presented on the interactive interface.
12. The system according to claim 9, wherein: It also includes an emotional service intervention module for generating student emotional state data by analyzing the student's facial image and voice data collected by the teaching execution monitoring module through emotion recognition technology; Based on the student's emotional state data, a soothing service instruction is generated and triggered. The instruction controls the teaching system to perform operations, including switching the voice interaction tone to a softer mode and providing targeted encouraging guidance content; Combined with the cognitive state data set and emotional state data generated by the teaching execution monitoring module, predict the operational errors that may occur to students in the current teaching scenario; Based on the prediction results, predictive operation reminder information is generated and pushed to the interactive terminal.
13. The system according to claim 7, wherein: The teaching strategy module is also used to simultaneously collect the student's heart rate signal and eye tracking signal output by the wearable physiological monitoring device, as well as the steering wheel angle signal and throttle opening signal output by the vehicle operating device; convert the heart rate signal into a learning stress level indicator, and convert the eye tracking signal into a line of sight focus time indicator, to jointly form the student's learning ability data; convert the steering wheel angle signal into a steering frequency indicator, and convert the throttle opening signal into an acceleration smoothness indicator, to jointly form the student's driving habit data; and combine the number of operating errors and error type codes in historical training records to form the student's error type data.
14. The system according to claim 1, wherein: The safety perception module is further configured to obtain environmental data output by the environmental perception device; determine a scene type based on scene feature identifiers in the environmental data, wherein the scene types include weather change scenes, human-machine mixed scenes, and two-way road scenes; When the scene type is a weather change scene, the rainfall intensity and fog density parameters in the real-time meteorological data are collected and combined with the moving object density index in the spatial obstacle distribution data to generate the weather change scene type code; When the scene type is a human-machine mixed scene, the pedestrian movement trajectory features and non-motor vehicle motion vectors are extracted, and the trajectory intersection density and relative speed change rate are calculated to generate the human-machine mixed scene type code; When the scene type is a two-way road scene, the distance parameters and relative speed parameters of the vehicles in the opposite lane are obtained to identify the lane line curvature change characteristics and generate the two-way road scene type code; Synchronously collect vehicle speed and road curvature values from vehicle driving status data; Based on the scene type code, vehicle speed value and road curvature value, the intervention degree corresponding to each scene is calculated.
15. The system according to claim 1, wherein: The output of the safety perception module is connected to the input of the behavior management module. The behavior management module is used to perform the following process when identifying driving behavior data and generating multimodal interactive feedback in combination with physiological data and status data to guide and correct driving behavior: Obtain torque change data output by the steering wheel torque sensor and facial muscle displacement data collected by the vehicle camera; Converting the torque change data into a steering wheel operation force level, and converting the facial muscle displacement data into a mouth corner lift amplitude and a brow wrinkle degree; Based on the deviation between the steering wheel operation force level and the preset standard force range, the behavioral conflict coefficient is generated in combination with the upward range of the mouth corners and the degree of frowning between the eyebrows; When the behavior conflict coefficient exceeds a first threshold, the voice feedback generator is triggered to perform the following operations: matching a first voice segment in a steering operation description phrase library according to the steering wheel operation force level; matching a second voice segment in a emotional state description phrase library according to the mouth corner upward angle and the eyebrow wrinkling degree; and combining the first voice segment and the second voice segment into a primary correction instruction; After outputting the primary correction instruction, continuously monitoring the change trends of the torque change data and the facial muscle displacement data, and if the rate of decrease of the behavior conflict coefficient is lower than a second threshold, extracting related follow-up question segments from a multi-round dialogue template library; The associated question segment and the correction operation demonstration segment are combined to generate a secondary correction instruction, and a multi-round interactive voice sequence including the primary correction instruction and the secondary correction instruction is output.
16. An intelligent driving method, characterized in that: include: Generate personalized teaching strategies tailored to students' abilities based on multi-dimensional student data through a multimodal large language model; In the process of implementing the personalized teaching strategy, the system integrates meteorological, on-board multi-dimensional sensor and visual data to evaluate the risk level of different scene types by constructing the road friction coefficient and slippery trend, analyzing scene visibility, and performing behavioral clustering and motion vector analysis on dynamic targets. Based on the evaluation results and the preset intervention level, a safety response strategy is generated, wherein the safety response strategy includes one or any combination of speed control, avoidance path, braking force and steering compensation; In the process of executing the safety response strategy, driving behavior data is identified, and multimodal interactive feedback is generated in combination with physiological data and status data to guide and correct driving behavior.
17. An electronic device, characterized in that: include: memory for storing computer programs; A processor is configured to implement the steps of the intelligent driving method according to claim 16 when executing the computer program.
Citation Information
Patent Citations
Early warning method and system for non-vehicle target
CN112185146A
Roundabout entrance vehicle passing sequence decision-making system for intelligent network connection vehicles
CN115424445A
Vehicle control method and device, electronic equipment and vehicle
CN116061825A
Driving anti-skid and braking anti-lock performance improving method based on tire pressure
CN118358299A
Adaptive driving behavior evaluation and training system based on multi-modal large model
CN119513809A
Cited By
Traffic persuasion system integrating edge computing and multi-sensor interface and terminal
CN121075133A
Traffic guidance system and terminal integrating edge computing and multi-sensor interface
CN121075133B
Safety early warning communication method and system among multiple ships
CN121330955A
Driving training robot driving behavior analysis and correction system based on artificial intelligence
CN122090697A
An artificial intelligence-based driving robot driving behavior analysis and correction system
CN122090697B