Behavior pre-judgment control system based on multi-modal incremental learning and dynamic security verification

By constructing a spatiotemporal interaction graph and performing multi-level verification through a behavior prediction and control system based on multimodal incremental learning and dynamic safety verification, the system solves the problems of false triggering and false execution of dangerous actions in complex driving scenarios, and achieves higher accuracy and safety in driver behavior prediction and control.

CN121912982APending Publication Date: 2026-04-24RIVOTEK TECH (JIANGSU) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
RIVOTEK TECH (JIANGSU) CO LTD
Filing Date
2026-03-26
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Existing technologies struggle to reliably suppress false triggering and dangerous erroneous actions in complex driving scenarios during driver behavior prediction and control. They also lack a unified, dynamic, and closed-loop verification mechanism, which affects the safety and reliability of vehicle control.

Method used

A behavior prediction and control system based on multimodal incremental learning and dynamic safety verification is adopted. Through data acquisition, graph construction, multimodal incremental learning, dynamic safety verification, safety execution and learning closed-loop modules, a spatiotemporal interaction graph is constructed, candidate control actions are output and multi-level verification is performed, including environmental risk, driver intention, model reliability and interpretable consistency verification, providing revocable interaction and redundant control, and realizing closed-loop update.

Benefits of technology

It effectively suppresses misjudgment and release in complex scenarios, improves the accuracy and reliability of driver behavior prediction and control, reduces the false trigger rate and the probability of dangerous actions, and enhances the interpretability and overall operational safety of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121912982A_ABST
    Figure CN121912982A_ABST
Patent Text Reader

Abstract

The invention discloses a behavior pre-judgment control system based on multi-mode incremental learning and dynamic safety verification, and relates to the technical field of vehicle intelligent driving control. According to the technical scheme, the system is characterized by comprising a data acquisition module, a graph construction module, a multi-modal incremental learning module, a dynamic security verification module, a security execution module, a redundancy control module and a learning closed-loop module; the system constructs a space-time interaction diagram based on multi-modal data, outputs candidate control actions, candidate control action confidence coefficients, concept vectors and scene identifiers, performs dynamic security verification in combination with concept consistency scores and historical false triggering rates, and executes a control instruction after verification is passed. And performing closed-loop increment updating based on the revocation event or the takeover event. According to the scheme, the false triggering rate and the dangerous action execution probability can be reduced, and the interpretability, the verifiability and the operation safety of driver behavior pre-judgment control are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent driving control technology for vehicles, and more specifically, to a behavior prediction and control system based on multimodal incremental learning and dynamic safety verification. Background Technology

[0002] With the development of intelligent driving and driver assistance technologies, utilizing onboard sensors to acquire vehicle operation data, driver interaction data, and external environment perception data, and then predicting potential driver behaviors based on this information, has become an important technological direction in the field of vehicle control. Especially in scenarios such as lane changing, overtaking, and yielding, the system often needs to provide candidate control actions or assist in executing control commands in advance based on the driver's state and the surrounding traffic environment to improve driving efficiency and driving coordination.

[0003] In existing technologies, driver behavior prediction and control typically focus on directly generating candidate actions based on perceived information and model output, or making execution decisions mainly based on a single confidence level or fixed rule thresholds. While these approaches can achieve behavior prediction and control assistance to some extent, they often lack a unified, dynamic, and closed-loop verification mechanism to determine whether the current candidate control action has sufficient interpretable evidence, whether such an action belongs to a historically prone-to-mistake scenario in the current situation, and whether the control command continuously meets safety requirements during execution.

[0004] In this context, the core problem with existing technologies lies in their inability to reliably suppress false triggers and dangerous erroneous actions in driver behavior prediction and control under complex driving scenarios. If the system directly releases actions based solely on the model's high confidence output, or fails to dynamically adjust the decision threshold based on historical false triggers, its own interpretability consistency, and runtime monitoring results, it is prone to outputting inappropriate control actions when the driver is not yet ready to take over, environmental risks are not fully eliminated, or there are anomalies in the execution chain. This can negatively impact the safety and reliability of vehicle control. Summary of the Invention

[0005] To address the shortcomings of existing technologies, the present invention aims to provide a behavior prediction and control system based on multimodal incremental learning and dynamic security verification.

[0006] To achieve the above objectives, the present invention provides the following technical solution: a behavior prediction and control system based on multimodal incremental learning and dynamic security verification, comprising a data acquisition module, a graph construction module, a multimodal incremental learning module, a dynamic security verification module, a security execution module, a redundant control module, and a learning closed-loop module;

[0007] The data acquisition module is used to collect multimodal data including vehicle operation data, driver interaction data, and external environment perception data. The graph construction module is used to construct a spatiotemporal interaction graph based on the multimodal data. The multimodal incremental learning module is used to encode the spatiotemporal interaction graph to obtain a scene embedding vector, and output candidate control actions, candidate control action confidence, concept vector, and scene identifier based on the same scene embedding vector. The dynamic safety verification module is used to receive the candidate control actions, the candidate control action confidence, the concept vector, the scene identifier, and the multimodal data, calculate the concept consistency score based on the concept vector, and perform threshold adaptive verification of the candidate control action confidence based on the historical false trigger rate corresponding to the scene identifier to obtain a verification result. The verification result includes execution permission, execution blocking, or only prompting not to execute.

[0008] The safety execution module is used to generate control commands and issue them to the vehicle execution object to execute the candidate control action when the verification result is an execution permission, and to provide the revocable interaction of the control commands so as to trigger the revocation event of the control commands when the revocation conditions are met; the redundant control module is used to independently monitor the control commands of the safety execution module, and to take over when the preset takeover conditions are met so as to trigger the takeover event of the control commands, and block or replace the control commands.

[0009] The learning closed-loop module is used to update the historical false trigger rate based on the execution results of the control commands and the revocation or takeover events of the control commands, and to generate feedback information. This feedback information is then converted into feedback correction samples for the multimodal incremental learning module to perform incremental updates. Through this technical solution, driver behavior prediction, dynamic safety verification, execution control, and post-operation feedback updates can be incorporated into a single closed-loop framework.

[0010] Furthermore, the multimodal incremental learning module includes a spatiotemporal graph neural network unit, a concept bottleneck unit, a behavior prediction unit, and a scene identifier generation unit. The spatiotemporal graph neural network unit encodes graph features of the spatiotemporal interaction graph to obtain a scene embedding vector. The concept bottleneck unit outputs a concept vector based on the scene embedding vector. The behavior prediction unit outputs a candidate control action and its confidence level based on the scene embedding vector and the concept vector. The scene identifier generation unit outputs a scene identifier based on the scene embedding vector. The concept bottleneck unit, the behavior prediction unit, and the scene identifier generation unit share the same scene embedding vector and output them in parallel, ensuring that the candidate control action, the candidate control action confidence level, the concept vector, and the scene identifier correspond to the same driving scenario. The concept vector represents interpretable evidence for the candidate control action, and the scene identifier represents the current driving scenario category. Thus, action judgment, action interpretation, and scenario category information can be obtained simultaneously based on the same scenario representation.

[0011] Furthermore, the dynamic safety verification module includes an environmental risk verification unit, a driver willingness and takeover capability verification unit, a model reliability verification unit, an interpretable consistency verification unit, and a runtime monitoring verification unit. The environmental risk verification unit generates a risk envelope based on the vehicle's external environment perception data and determines whether the predicted trajectory of the candidate control action intersects with the risk envelope. The driver willingness and takeover capability verification unit determines whether the driver is in a takeover capability state based on the driver interaction data and checks for any opposition signals within a preset confirmation window. The model reliability verification unit performs reliability determination based on the confidence level of the candidate control action and the historical false trigger rate corresponding to the scene identifier. The interpretable consistency verification unit performs consistency determination based on the concept consistency score. The runtime monitoring verification unit receives the monitoring decision result output by the redundant control module and verifies the candidate control action and the control command based on the monitoring decision result. When all verification units pass, the verification result is execution permission; when any verification unit fails, the verification result is execution blocking or only a prompt not to execute. This hierarchical verification mechanism avoids relying solely on a single model output to directly release control actions.

[0012] Furthermore, the interpretable consistency verification unit calculates a concept consistency score based on the concept vector, which characterizes the degree of consistency between the current candidate control action and the interpretable evidence under the current scene identifier. The concept consistency score can be calculated based on the deviation of each concept component in the concept vector from its corresponding expected interval, wherein the expected interval is determined by the combination of the candidate control action and the scene identifier. The interpretable consistency verification unit determines failure if the concept consistency score is less than a preset consistency threshold, otherwise it determines success. This setting allows for constraints on whether the action results output by the model have sufficiently reasonable conceptual evidence.

[0013] Furthermore, the environmental risk verification unit calculates an environmental risk index for each obstacle in the external environment perception data, and generates the risk envelope based on the environmental risk index of each obstacle. The environmental risk index of an obstacle is related to the distance between the obstacle and the vehicle, the approach speed component of the obstacle relative to the vehicle, and scale parameters. The environmental risk verification unit determines the longitudinal and lateral expansion amounts based on the predicted occupied area of ​​each obstacle and the corresponding environmental risk index, generates corresponding local risk sub-envelopes around the predicted occupied area of ​​each obstacle, and obtains the risk envelope by taking the union of each local risk sub-envelope. When the predicted trajectory corresponding to the candidate control action intersects with the risk envelope, the environmental risk verification unit determines it as a failure; otherwise, it is a pass. Thus, environmental risk can be extended from single-point detection to envelope risk determination for future trajectories.

[0014] Furthermore, the learning closed-loop module updates the historical false trigger rate for each scenario identifier. The historical false trigger rate corresponding to each scenario identifier is determined by the number of false triggers determined by the control command cancellation event and / or takeover event under that scenario identifier, the number of triggers in which the candidate control action is granted execution permission under that scenario identifier, and a smoothing coefficient. The model reliability verification unit performs threshold adaptive verification of the confidence level of the candidate control action based on the historical false trigger rate. The confidence level verification threshold under the scenario identifier is jointly limited by a base threshold, an adjustment coefficient, a preset minimum confidence level threshold, and a preset maximum confidence level threshold. Through this setting, the candidate control actions under different scenarios can be dynamically adjusted according to the historical false trigger situation based on the threshold.

[0015] Furthermore, the learning closed-loop module assigns sample weights to the feedback correction samples, and the multimodal incremental learning module performs incremental updates based on the feedback correction samples and their corresponding sample weights. The sample weights are related to weight coefficients, event-addition coefficients, and a comprehensive reliability score, which is obtained by fusing the confidence verification score and the concept consistency score. Specifically, the sample weight corresponding to a revocation event is greater than the sample weight corresponding to a takeover event, and the sample weight corresponding to a takeover event is greater than the sample weight formed solely based on the control command execution results. By introducing a sample weight mechanism, feedback samples that better reflect system defects can play a greater corrective role in incremental updates.

[0016] Furthermore, the multimodal incremental learning module includes a two-stage distillation update mechanism, comprising behavioral distillation and concept distillation. Behavioral distillation constrains the distribution of candidate control actions output by the student model to remain consistent with that of the teacher model. Concept distillation constrains the concept vectors output by the student model to remain consistent with those of the teacher model, and incremental updates are performed on the student model parameters based on the total update loss. The total update loss includes a weighted task loss term determined based on the feedback correction samples and their corresponding weights, as well as distillation constraint terms related to the candidate control action distributions, concept vectors, and scene embedding vectors of the teacher and student models. This configuration allows for the correction of new errors using feedback correction samples while maintaining the model's stability within the original scenario.

[0017] Furthermore, the safety execution module includes a haptic feedback component and a cancellation decision logic. The haptic feedback component provides haptic feedback to the driver when the control command is issued. The cancellation decision logic triggers a cancellation event for the control command and cancels the control command when it detects that the driver's applied counterforce exceeds a preset force threshold or detects an opposition signal. The logic then sends the scene identifier, candidate control action, and confidence level of the candidate control action corresponding to the cancellation event to the learning closed-loop module to form feedback information. This provides the driver with the ability to directly veto commands during system execution.

[0018] Furthermore, the redundant control module includes a main control unit and a backup control unit; the main control unit is used to run the dynamic safety verification module and the safety execution module to execute the main control link when the preset takeover conditions are not met, and to generate and issue the control command based on the verification result; the backup control unit operates independently of the main control unit and is used to generate a monitoring decision result based on the candidate control action, the control command, and the vehicle execution object status and output it to the runtime monitoring verification unit;

[0019] The monitoring and decision result is used to characterize whether at least one of the following situations exists: the candidate control action is inconsistent with the control command, the control command exceeds the preset safety constraints of the vehicle execution object, the state of the vehicle execution object does not correspond to the control command, the main control link times out, or communication between the main control unit and the vehicle execution object is abnormal. When the monitoring and decision result indicates that at least one of the above situations exists, the preset takeover condition is determined to be met. When the preset takeover condition is met, the backup control unit completes the takeover within a preset takeover time, triggers the takeover event of the control command, and blocks or replaces the control command. This redundant takeover mechanism can further improve the operational security of the control execution phase.

[0020] Compared with the prior art, the present invention has the following beneficial effects:

[0021] 1. In this invention, by setting up a data acquisition module, a graph construction module, a multimodal incremental learning module, and a dynamic safety verification module, a spatiotemporal interaction graph is first constructed based on vehicle operation data, driver interaction data, and external environment perception data. Then, the multimodal incremental learning module outputs candidate control actions, candidate control action confidence scores, concept vectors, and scene identifiers. Finally, the dynamic safety verification module dynamically verifies the candidate control actions by combining the concept consistency score and the historical false trigger rate corresponding to the scene identifier. Therefore, control is no longer driven solely by the confidence score of a single model; instead, candidate control actions are simultaneously subject to interpretable evidence verification and scene reliability verification. This more effectively suppresses misjudgment and release in complex scenarios, improving the accuracy and reliability of driver behavior prediction control.

[0022] 2. This invention, by setting up a safe execution module, a redundant control module, and a learning closed-loop module, provides revocable interaction for control commands after successful verification. When preset takeover conditions are met, the redundant control module independently takes over, simultaneously converting revocation or takeover events into feedback correction samples for incremental updates by the multimodal incremental learning module. Using these techniques, the system can not only promptly block unsafe actions during the control execution phase but also continuously correct model outputs in scenarios prone to false triggering based on actual operational results. Therefore, it can reduce the false trigger rate and the probability of dangerous actions during long-term operation, and improve the system's interpretability, verifiability, and overall operational safety. Attached Figure Description

[0023] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this invention, illustrate exemplary embodiments of the invention and are used to explain the invention, but do not constitute an undue limitation of the invention. In the drawings: Figure 1 This is a flowchart of the process of the present invention. Detailed Implementation

[0024] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.

[0025] Furthermore, the technical solutions of the various embodiments of the present invention can be combined with each other, but only if they are based on the ability of those skilled in the art to implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such combination of technical solutions does not exist and is not within the scope of protection claimed by the present invention.

[0026] In this embodiment, a behavior prediction and control system based on multimodal incremental learning and dynamic safety verification is installed within the vehicle domain controller. This system includes a data acquisition module, a graph construction module, a multimodal incremental learning module, a dynamic safety verification module, a safety execution module, a redundant control module, and a learning closed-loop module. The data acquisition module connects to the vehicle's CAN bus, vehicle Ethernet, cameras, millimeter-wave radar, driver monitoring camera, steering wheel torque sensor, pedal position sensor, and actuator status interface to acquire multimodal data such as vehicle operation data, driver interaction data, and external environment perception data. Vehicle operation data includes at least vehicle speed, longitudinal acceleration, yaw rate, steering wheel angle, steering wheel angular velocity, brake pedal opening, accelerator pedal opening, turn signal status, and current gear. Driver interaction data includes at least steering wheel grip status, steering wheel counterforce, driver's gaze direction, eyelid opening and closing status, head orientation, voice objection signal, and braking intervention signal. External environment perception data includes at least lane lines in the current and adjacent lanes, vehicles ahead and behind, vehicles in adjacent lanes, guardrails, cones, pedestrians, static obstacles, and road curvature information. To ensure that data from different sources can be used for decision-making at the same time, the data acquisition module synchronizes the data from each sensor with a uniform time step of 100ms, and uses nearest neighbor interpolation and zero-order hold to fill in missing frames, and performs amplitude limiting processing on outliers.

[0027] In this embodiment, the graph construction module constructs a spatiotemporal interaction graph based on multimodal data. The nodes of the spatiotemporal interaction graph include at least the vehicle node, the preceding vehicle node, the following vehicle node, the preceding and following vehicle nodes in the left adjacent lane, the preceding and following vehicle nodes in the right adjacent lane, lane boundary nodes, obstacle nodes, and driver state nodes. The vehicle node includes the vehicle's speed, attitude, and control state characteristics; the vehicle node includes relative distance, relative speed, relative azimuth angle, and predicted trajectory segments; the lane boundary node includes lane width, curvature, lane line type, and drivability; the driver state node includes grip state, gaze focus area, turn signal trigger state, and opposition signal state. It also includes spatial relationship edges and temporal evolution edges. Spatial relationship edges represent the relative positional relationship between the vehicle and surrounding targets, as well as the coupling relationship between the driver state and the vehicle's control state. Temporal evolution edges connect the states of the same node across multiple consecutive time steps, thus forming a spatiotemporal interaction graph sequence containing 2 seconds of historical data and a sampling frequency of 10Hz. By using this graph construction method, the subsequent model can simultaneously utilize four types of information: "who is nearby", "who is approaching", "whether the driver intends to change lanes", and "how these states have changed in the last 2 seconds".

[0028] In this embodiment, the multimodal incremental learning module includes a spatiotemporal graph neural network unit, a concept bottleneck unit, a behavior prediction unit, and a scene label generation unit. The spatiotemporal graph neural network unit adopts a structure of two layers of graph attention network connected in series with a single layer of gated recurrent unit. Each node has an initial feature dimension of 32, and the hidden dimensions of the two graph attention networks are 64 and 128, respectively. The gated recurrent unit outputs a 128-dimensional scene embedding vector. The concept bottleneck unit receives the 128-dimensional scene embedding vector and outputs an 8-dimensional concept vector. These 8 concepts are: target lane forward clearance sufficiency, target lane rear-end collision risk, current lane congestion level, lane line clarity, driver's active lane change intention strength, driver takeover capability, road adhesion stability, and lateral maneuvering space sufficiency. Each concept component is normalized to between 0 and 1. The behavior prediction unit receives the concatenation result of the scene embedding vector and the concept vector, and outputs candidate control actions and their confidence levels. In this embodiment, the candidate control actions are taken from a set of five actions: lane keeping, changing lanes to the left, changing lanes to the right, slowing down to yield, and prompting but not executing. The scene identifier generation unit outputs scene identifiers based on the scene embedding vector. In this embodiment, the scene identifiers include at least high-speed following scenarios, high-speed overtaking scenarios, ramp merging scenarios, congested low-speed scenarios, and curved road scenarios. Since the concept bottleneck unit, behavior prediction unit, and scene identifier generation unit share the same scene embedding vector and output them in parallel, the candidate control action, candidate control action confidence, concept vector, and scene identifier naturally correspond to the same driving scenario.

[0029] The dynamic safety verification module executes in the logical order of environmental risk verification unit, driver willingness and takeover capability verification unit, model reliability verification unit, interpretable consistency verification unit, and runtime monitoring verification unit.

[0030] The environmental risk verification unit calculates an environmental risk index for each obstacle in the vehicle's external environment perception data and generates a risk envelope based on the environmental risk index. Specifically, for the first obstacle... For each obstacle, the distance between the obstacle and the vehicle is first determined based on external environmental perception data. Safety reference distance The approach speed component of the obstacle relative to the vehicle. and scale parameters and The corresponding environmental risk index is calculated according to the following formula:

[0031]

[0032] in, The distance between the obstacle and the vehicle. This is a safety baseline distance; This represents the approach velocity component of the obstacle relative to the vehicle, and is positive only when the obstacle is approaching the vehicle, and zero when it is moving away from the vehicle. and This is a scale parameter. Therefore, it can be seen that when the distance between the obstacle and the vehicle decreases, or the approach speed of the obstacle towards the vehicle increases, the corresponding environmental risk index... Increase; when the obstacle is far away, or is not approaching the vehicle, the corresponding environmental risk index increases. The environmental risk verification unit determines the longitudinal expansion and lateral expansion based on the predicted occupied area of ​​each obstacle and the corresponding environmental risk index. It generates a corresponding local risk sub-envelope around the predicted occupied area of ​​each obstacle and obtains the risk envelope by taking the union of the local risk sub-envelopes. When the predicted trajectory corresponding to the candidate control action intersects with the risk envelope, the environmental risk verification unit determines that it has failed; otherwise, it has passed.

[0033] In one specific embodiment, if an obstacle is detected to the left front, the distance between the obstacle and the vehicle is taken. meters, safety reference distance Meter, scale parameter =8 meters m / s; if the obstacle is approaching the vehicle, and the approach velocity component If the speed is meters per second, then its environmental risk index is:

[0034]

[0035] This indicates that the obstacle poses a high risk to the current candidate control action. In this embodiment, the longitudinal and lateral expansion amounts can be further determined based on the environmental risk index; for example, the longitudinal expansion amount can be taken as... Meters, lateral expansion measurement Meters, then the corresponding longitudinal expansion is meters, lateral expansion amount is The system calculates the distance in meters; based on this, a local risk sub-envelope is generated around the predicted occupancy area of ​​the obstacle for the next 3 seconds, and the union of the local risk sub-envelopes corresponding to other obstacles is obtained to obtain the risk envelope. If the predicted trajectory corresponding to a candidate control action enters the above risk envelope, it is determined that the candidate control action has a collision risk or dangerous interference risk in the current external environment, and the environmental risk verification unit determines that it has failed; if the predicted trajectory does not intersect with the risk envelope, it is determined that it has passed.

[0036] The driver willingness and takeover capability verification unit is used to determine whether the driver is suitable to accept the execution of the behavior prediction control system and whether they explicitly object. Specifically, if the driver's hands are off the steering wheel for more than 1 second, or their gaze is continuously deviated from the road for more than 1.5 seconds, or the driver monitoring camera detects that their eyes are closed for more than 0.8 seconds, then the driver is determined not to be in a takeover capability state.

[0037] Conversely, if the driver activates the turn signal corresponding to the candidate control action and the steering wheel is held correctly, this can be considered a positive intention signal supporting execution. If, within a preset 600ms confirmation window, a reverse force greater than 2.5 N·m (Newton-meter, a unit of torque), a brake pedal opening greater than 20%, a voice command "Do not change lanes," or cancellation of the turn signal operation is detected, an objection signal is determined to exist. This layer is only considered successful when the driver is in a takeover state and no objection signal exists. After setting, if the driver clearly loses the ability to take over or there is a clear objection signal, the driver's intention and takeover verification unit will determine that the action has failed, thus preventing the corresponding candidate control action from being authorized for execution and avoiding the system executing control commands when the driver lacks the ability to take over or clearly objects.

[0038] The model reliability verification unit is used to determine reliability based on the confidence level of candidate control actions and the historical false trigger rate corresponding to the scene identifier. Specifically, the learning closed-loop module updates the historical false trigger rate for each scene identifier. The historical false trigger rate for a scenario is calculated using the aforementioned formula, based on the number of false triggers determined by control command cancellation events and / or takeover events under that scenario identifier, the number of triggers where candidate control actions are granted execution permission under that scenario identifier, and a smoothing coefficient. Subsequently, the model reliability verification unit uses the historical false trigger rate... The confidence scores of candidate control actions are adaptively verified using a threshold, and the current scene identifier is determined according to the aforementioned confidence score verification threshold formula. Corresponding confidence verification threshold The higher the historical false trigger rate of a scenario, the more prone it is to false triggering in that scenario, and the higher the corresponding confidence verification threshold. Conversely, the lower the historical false trigger rate of a scenario, the lower the corresponding confidence verification threshold. However, the confidence verification threshold is always constrained by the preset minimum confidence threshold and the preset maximum confidence threshold. Then, the confidence level of the current candidate control action is... Confidence verification threshold corresponding to the current scene identifier Compare; when Not less than When the current candidate control action is deemed sufficiently reliable in the scenario, the model reliability verification unit determines it to have passed; when Less than If the current candidate control action is deemed unreliable in the given scenario, the model reliability verification unit will determine that it has failed. Through these settings, the model reliability verification unit can dynamically adjust the confidence threshold for candidate control actions based on historical false triggering patterns in different scenarios. This avoids the problem of misjudging and allowing actions in high-false-trigger scenarios or being overly conservative in low-false-trigger scenarios due to using only a fixed threshold.

[0039] The interpretable consistency verification unit is used to calculate the concept consistency score based on the concept vector and to determine the consistency of the current candidate control action based on the concept consistency score. Since the concept vector is output by the concept bottleneck unit based on the scene embedding vector, the candidate control action is output by the behavior prediction unit based on the same scene embedding vector and the concept vector, and the scene identifier is output by the scene identifier generation unit based on the same scene embedding vector, the concept vector, candidate control action, and scene identifier naturally correspond to the same driving scenario. The role of the interpretable consistency verification unit is not to regenerate candidate control actions, but to verify whether "the current candidate control action has sufficient reasonable interpretable evidence under the current scene identifier," thereby avoiding outputting unreasonable control suggestions based solely on a single high-confidence value. In other words, this layer addresses the problem of "the behavior prediction unit believes that the candidate control action can be executed, but whether the conceptual evidence it provides truly supports the action."

[0040] Specifically, the interpretable consistency verification unit pre-sets a corresponding expectation range for each concept component in the concept vector based on the combination of candidate control actions and scene identifiers. and weight The concept consistency score is calculated based on the degree of deviation of each concept component in the concept vector from its respective expected interval. The concept consistency score satisfies:

[0041]

[0042] in, For the concept vector, the first Each concept component For the first corresponding to the candidate control action and scene identifier The expected range of each concept component is determined by the combination of candidate control actions and scene identifiers; For the first The weight of each concept component, Let be the number of concept components. From this formula, we can see that when the th... Each concept component Falling into the corresponding expected range When the concept component is within the specified time, it does not incur a penalty; when the first... Each concept component Below the lower bound or higher than the upper limit At that time, the consistency verification unit can be interpreted to generate out-of-bounds penalties based on its deviation, and according to weight. This is included in the total penalty item. As the total penalty item increases, the conceptual consistency score... Decrease; Conceptual consistency score decreases when all conceptual components are within a reasonable range. The score is close to 1. Therefore, the concept consistency score can be used to characterize the degree of consistency between the current candidate control action and the interpretable evidence under the current scenario label.

[0043] The optimal concept vector comprises eight conceptual components: target lane forward clearance sufficiency, target lane rear-end collision risk, current lane congestion level, lane line clarity, driver's intention to actively change lanes, driver's ability to take over, road adhesion stability, and lateral maneuvering space sufficiency. All of these conceptual components can be output by the concept bottleneck unit and uniformly normalized to the 0-1 range. These eight conceptual components were chosen because they can jointly characterize whether a candidate control action has a sufficiently reasonable explanatory basis from three aspects: environmental conditions, driver state, and vehicle maneuvering conditions. For example, for the candidate control action "changing lanes to the left," a high action confidence alone is insufficient; it also needs to simultaneously satisfy explanatory conditions such as relatively sufficient forward clearance in the target lane, relatively low rear-end collision risk in the target lane, strong driver's intention to actively change lanes, high driver's ability to take over, and sufficient lateral maneuvering space. Otherwise, although the candidate control action may have a high confidence level, it should not be considered to have sufficient explanatory evidence in the current scenario.

[0044] In a specific example, when the scenario is identified as a "high-speed overtaking scenario" and the candidate control action is "change lanes to the left," the interpretable consistency verification unit pre-sets the expected ranges for each conceptual component as follows: the target lane forward clearance sufficiency is... The risk of a rear-end collision in the target lane is The congestion level of this lane is [missing information]. Lane line clarity is The driver's intention to change lanes was strong. The level of driver takeover is Road adhesion stability is The lateral maneuverability is sufficient. Correspondingly, the weights of each conceptual component can be set to 0.18, 0.20, 0.10, 0.08, 0.18, 0.10, 0.08, and 0.08, respectively. The principle for setting the above expected range and weights is as follows: for conceptual components that are more directly related to the safety and rationality of the current candidate control action, higher weights are assigned; for conceptual components that are supportive but not decisive, relatively lower weights are assigned. For example, in the combination of "high-speed overtaking scenario + left lane change," the risk of rear-end collision in the target lane and the intensity of the driver's intention to actively change lanes are usually more decisive than the clarity of lane lines in determining whether a lane change is allowed, so their weights can be set relatively higher. Through this method of "configuring expected ranges and weights according to the combination of actions and scenarios," the interpretability consistency verification unit no longer remains at the level of abstract judgment, but is concretely grounded in a quantifiable, calculable, and directly implementable conceptual evidence judgment mechanism.

[0045] In a further specific example, if the current concept vector output is: Target lane forward clearance sufficiency 0.82, Target lane rear-end collision risk 0.18, Current lane congestion level 0.66, Lane line clarity 0.87, Driver's intention to actively change lanes 0.74, Driver's ability to take over 0.79, Road adhesion stability 0.91, Lateral maneuvering space sufficiency 0.72, then since each concept component falls within the above expected range, the concept consistency score is [high / high]. A value close to 1 indicates that the interpretable consistency verification unit can determine that the candidate control action has high consistency. Conversely, if the current concept vector output shows a target lane rear-end collision risk of 0.55, a driver's active lane change intention intensity of 0.32, and lateral maneuvering space adequacy of 0.48, then multiple concept components exceed their corresponding expected ranges. Specifically, a high target lane rear-end collision risk, insufficient driver's active lane change intention, and insufficient lateral maneuvering space will each incur significant boundary violations, and under the influence of their weights, significantly lower the concept consistency score. This means that although a candidate control action such as "change lanes to the left" may be output, from the perspective of interpretable evidence, this candidate control action is inconsistent with the current scene identifier and therefore should not be directly verified. Thus, the interpretable consistency verification unit transforms the abstract question of "why the behavior prediction unit outputs the current candidate control action" into a quantifiable judgment question of "which conceptual evidences are satisfied, which conceptual evidences are not satisfied, and how large is the deviation".

[0046] The interpretable consistency verification unit scores the concept consistency. The result is compared with a preset consistency threshold to output a consistency determination. Preferably, the consistency threshold can be set to 0.75. When the concept consistency score... When the score is less than 0.75, the interpretable consistency verification unit is judged as failing; when the concept consistency score is... A score of 0.75 or higher is considered a pass for the interpretable consistency verification unit. Here, the consistency threshold sets a minimum requirement for whether conceptual evidence is sufficient to support the current candidate control action. If the score is too low, it indicates that at least one or more key conceptual components significantly deviate from the reasonable range required by the current scenario identifier and the current candidate control action. In this case, even if the candidate control action has a high confidence level, it cannot be considered to have a sufficiently reliable explanatory basis. The result of the interpretable consistency verification unit will continue to participate in the overall verification process of the dynamic safety verification module, and together with the results of the environmental risk verification unit, driver willingness and takeover capability verification unit, model reliability verification unit, and runtime monitoring verification unit, determine whether the final verification result is execution permission, execution blocking, or only a prompt not to execute.

[0047] The runtime monitoring and verification unit receives the monitoring decision results output by the redundant control module and verifies the candidate control actions and control commands based on these results. Specifically, the backup control unit operates independently of the main control unit, generating monitoring decision results based on candidate control actions, control commands, and the vehicle execution object status, and outputting these results to the runtime monitoring and verification unit. The monitoring decision results characterize whether at least one of the following situations exists: inconsistency between the candidate control action and the control command; the control command exceeding the preset safety constraints of the vehicle execution object; the vehicle execution object status not corresponding to the control command; the main control link timeout; or communication anomalies between the main control unit and the vehicle execution object. The runtime monitoring and verification unit makes a judgment based on the monitoring decision results: if the monitoring decision results indicate that none of the above situations exist, the runtime monitoring and verification unit determines it as passed; if the monitoring decision results indicate that at least one of the above situations exists, the runtime monitoring and verification unit determines it as failed, and accordingly determines that the preset takeover conditions are met. This allows the backup control unit to complete the takeover within a preset takeover time when the preset takeover conditions are met, triggering a takeover event for the control command and blocking or replacing the control command.

[0048] The learning loop module updates the historical false trigger rate for each scene identifier. Corresponding historical false trigger rate satisfy:

[0049] ;

[0050] in, For scene identification The number of false triggers determined by the control command revocation event and / or takeover event. For scene identification The number of times the next candidate control action is granted execution permission. For smoothing coefficients;

[0051] Taking a high-speed overtaking scenario as an example, assuming there were 80 candidate control actions granted execution permission in this scenario, of which 6 were subsequently revoked or taken over events and determined to be false triggers, the smoothing coefficient... If we take 3, then the historical false trigger rate in this scenario is... .

[0052] Furthermore, the model reliability verification unit performs threshold adaptive verification of the confidence level of candidate control actions based on historical false trigger rates, and scene identification. The confidence verification threshold below satisfy:

[0053] ;

[0054] in, Based on the threshold, It is the adjustment coefficient, and ; and These are the preset minimum confidence threshold and the preset maximum confidence threshold, respectively. .

[0055] For example, the base threshold Set the adjustment coefficient to 0.65. Set the minimum confidence threshold to 0.25. Set the maximum confidence threshold to 0.55. If we set it to 0.90, then the confidence verification threshold in this scenario is approximately 0.676. If the confidence level of the current candidate control action... Then the confidence verification score The score can be 1 as described above. If the current concept consistency score... fusion coefficient Taking 0.6, the overall reliability score is... This means that the current behavior prediction model is not only historically stable for this scenario, but also consistent with current explanatory evidence, thus providing strong support for subsequent implementation.

[0056] In summary, the dynamic safety verification module only outputs execution permission when the environmental risk verification unit, driver willingness and takeover capability verification unit, model reliability verification unit, interpretability consistency verification unit, and runtime monitoring verification unit all pass. If the environmental risk verification unit fails, execution is blocked; if the environmental risk verification unit passes but the driver objects significantly or the confidence level is insufficient, only a prompt to not execute is output. For example, in a highway overtaking scenario in the left lane, the vehicle's current speed is 88 km / h, there is a slower vehicle in the lane ahead, the distance to the vehicle in the adjacent lane to the left rear is 45 meters and the approach speed component is 1.2 m / s, the output candidate control action is to change lanes to the left, the candidate control action confidence level is 0.78, the scenario is identified as a highway overtaking scenario, the concept consistency score is 0.89, the driver has activated the left turn signal in advance and is holding the steering wheel with both hands, and the backup control unit has not detected any abnormalities, then the dynamic safety verification module outputs execution permission. Conversely, if the vehicle to the left rear is only 20 meters away and its approach speed component reaches 4 m / s, the risk envelope will be significantly expanded. The predicted lane-changing trajectory of this vehicle intersects with the risk envelope. In this case, even if the model confidence is high, the output should still be to execute the blocking action. This fully reflects the problem that this solution aims to solve, namely, avoiding the mis-execution of dangerous actions due to relying solely on the output of a single model.

[0057] Upon receiving the execution permission, the safety execution module generates control commands and issues them to the vehicle execution objects. In this embodiment, the vehicle execution objects include the electronic power steering system, the electronic stability control system, and the power control unit. For a left lane change maneuver, the safety execution module first generates a fifth-order polynomial lateral trajectory and a corresponding longitudinal speed following curve for the next 3 seconds, then discretizes the trajectory into a target steering angle value, a steering angular velocity limit value, and a target longitudinal acceleration value, and issues these values ​​to the vehicle execution objects. Simultaneously, the haptic feedback component drives the steering wheel vibration motor to output two 200ms haptic prompts at a frequency of 180Hz, reminding the driver that the system has begun executing control actions. The cancellation decision logic continuously monitors the steering wheel reverse force, brake pedal, voice opposition signal, and turn signal status. Once it detects that the reverse force applied by the driver is greater than 2.5 N·m, or that the brake pedal has significantly intervened, or that a voice opposition signal has been detected, a cancellation event is immediately triggered and the control command is cancelled. At the same time, the scene identifier, candidate control actions, and confidence levels of the candidate control actions corresponding to the cancellation event are sent to the learning closed-loop module to form feedback information. This incorporates the driver's genuine rejection into subsequent learning, rather than simply discarding it.

[0058] The redundant control module includes a main control unit and a backup control unit. The main control unit runs a dynamic safety verification module and a safety execution module, and is responsible for the complete execution of the main control link. The backup control unit operates independently of the main control unit, reads candidate control actions, control commands, and vehicle execution object status at least every 20ms, and generates monitoring and decision results. The monitoring rules of the backup control unit include at least: whether the candidate control action is consistent with the control command; whether the control command exceeds the preset safety constraints of the vehicle execution object; whether the status of the vehicle execution object corresponds to the control command; whether the main control link completes the response within 80ms; and whether the communication between the main control unit and the vehicle execution object is abnormal. Taking steering control as an example, if the steering angle velocity issued by the main control unit exceeds the preset safety constraint, or the actuator feedback status shows that the steering execution lag exceeds 100ms, or the message verification between the main and backup fails, the backup control unit determines that the preset takeover conditions are met, completes the takeover within 50ms, triggers the takeover event, and blocks or replaces the original control command. In this embodiment, the replacement method is preferably "aborting the lane change and smoothly returning to the center of the current lane, while applying a slight deceleration of no more than 1.5m / s²". The takeover event, its corresponding scenario identifier, and the monitoring and adjudication results are also sent to the learning closed-loop module to form feedback information.

[0059] The learning closed-loop module converts the execution results of control commands and the revocation or takeover events of control commands into feedback information, and further converts the feedback information into feedback correction samples for the multimodal incremental learning module to perform incremental updates. Specifically, the feedback correction samples include at least: multimodal data from 2 seconds before to 1 second after the triggering time, scene identifier, candidate control action, candidate control action confidence, concept vector, verification result, execution result, whether a revocation event occurred, whether a takeover event occurred, and monitoring decision result. Through the above field settings, the feedback correction samples not only retain the input conditions when the candidate control action is formed, but also retain the subsequent evolution results of the candidate control action in the dynamic safety verification module, safety execution module, and redundant control module. This allows for a complete characterization of the entire process of a candidate control action from proposal, verification to execution, revocation, or takeover, facilitating accurate identification of the causes of false triggering or erroneous execution in subsequent incremental updates.

[0060] The learning loop module assigns weights to the feedback correction samples, and the multimodal incremental learning module performs incremental updates based on the feedback correction samples and their corresponding weights. (Sample weights) satisfy:

[0061] ;

[0062] in, These are the weighting coefficients, and ; An additional coefficient is added to the event, determined based on the event type corresponding to the feedback information; This represents the overall reliability score output by the dynamic safety verification module for candidate control actions. The meaning of the above formula is: when the overall reliability score... A higher overall reliability score indicates that the candidate control action is relatively reliable at the time of triggering, thus the basic amplification of the corresponding feedback correction sample is lower; when the overall reliability score is higher... A lower value indicates that the candidate control action already has significant uncertainty or poor pass rate at the time of triggering. Therefore, the corresponding feedback correction sample needs to receive higher attention in subsequent incremental updates, and thus its sample weight is larger. Further, through the event-added coefficient... The introduction of this can improve the overall reliability score. Under the same conditions, the sample weight corresponding to the control revocation event is greater than the sample weight corresponding to the takeover event, and the sample weight corresponding to the takeover event is greater than the sample weight formed solely based on the control command execution results. This allows the two types of samples, namely driver veto and backup control unit intervention takeover, to better reflect defects and play a stronger corrective role in incremental updates.

[0063] Overall reliability score satisfy:

[0064]

[0065] in, Let be the fusion coefficient, and ; To compare with the confidence level of candidate control actions The corresponding confidence verification score; This is the concept consistency score calculated by the dynamic security verification module based on concept vectors. Confidence verification score. satisfy:

[0066]

[0067] in, Confidence level of candidate control actions. For scene identification The corresponding confidence verification threshold. Therefore... Used to characterize the confidence level of candidate control actions Compared to scene identifiers Corresponding confidence verification threshold The degree of passage; when Not less than hour, A value close to 1 indicates a high degree of confidence; when... Below hour, It decreases as the degree of passage decreases. and Integrate into a comprehensive reliability score This can simultaneously reflect the reliability of candidate control actions in two aspects: "whether the confidence level is sufficient" and "whether the interpretable evidence level is consistent with the current action and scenario," thereby avoiding sample selection based solely on a single confidence level or a single interpretable score.

[0068] In a specific example, if a left lane change maneuver is revoked by a counterforce applied by the driver, the event type of this feedback correction sample is denoted as a revocation event, and the event additional coefficient is... Take 0.40; if an action is taken over by the backup control unit, the event type of the feedback correction sample is recorded as a takeover event, and the event additional coefficient is 0.40. Set the value to 0.25; if the action is executed normally without any abnormality, the event type of this feedback correction sample is recorded as normal execution, and the event additional coefficient is set to 0.25. Set to 0. Weighting coefficient In this embodiment, a value of 1.2 is used. If the current overall reliability score... If the reliability score is 0.956 and the event type is normal execution, then the sample weight is approximately 1.053; if the overall reliability score is... If the value is 0.62 and the event type is a reversal event, then the sample weight is approximately 1.856. This shows that unreliable feedback correction samples rejected by drivers are significantly amplified, thus occupying a higher proportion in subsequent incremental updates. The technical effect of this design is that updates prioritize correcting problematic samples that were historically more likely to cause false triggers, reversals, or takeovers, rather than having critical defects "diluted" by a large number of normal samples.

[0069] The multimodal incremental learning module includes a two-stage distillation update mechanism, comprising behavioral distillation and concept distillation. Both the teacher and student models are complete network models corresponding to the multimodal incremental learning module, each including a spatiotemporal graph neural network unit, a concept bottleneck unit, a behavior prediction unit, and a scene label generation unit. The teacher model preferably uses the previous stable version of the model before the incremental update, and its parameters are frozen during the current update process, only used to provide distillation reference information. The student model preferably uses the current version of the model to be updated, and its initial parameters can be copied from the teacher model parameters or initialized from the parameters of the current online model, participating in backpropagation updates during training. For the same feedback correction sample, the teacher and student models respectively receive the spatiotemporal interaction graph corresponding to the feedback correction sample as input and output scene embedding vector, concept vector, and original score vector corresponding to the candidate control action, respectively. Furthermore, the teacher and student models respectively perform temperature-parameterized processing on the original score vectors output by their respective behavior prediction units. The Softmax transform is used to obtain the candidate control action distributions of the teacher and student models under the temperature parameter. With this setting, behavior distillation is used to constrain the candidate control action distribution output by the student model to be consistent with that of the teacher model, and concept distillation is used to constrain the concept vector output by the student model to be consistent with that of the teacher model. At the same time, the scene embedding vectors of the teacher and student models are also aligned, thereby jointly suppressing the drift caused by incremental updates at the three levels of action output layer, concept interpretation layer and scene representation layer.

[0070] The two-stage distillation update mechanism is preferably executed in the order of behavioral distillation followed by concept distillation. Specifically, in the first stage, behavioral distillation is used to constrain the distribution of candidate control actions output by the student model to be consistent with that of the teacher model. This aims to prevent the student model from drastically shifting the boundaries of the original candidate control actions under the influence of a small number of high-weight feedback correction samples. In the second stage, concept distillation is used to constrain the concept vectors output by the student model to be consistent with those of the teacher model. Furthermore, it combines the scene embedding vectors of the teacher and student models to perform representation alignment, thus maintaining the continuity of the concept interpretation space and the scene representation space. Alternatively, behavioral distillation and concept distillation can be performed jointly in the same training round. However, functionally, the former mainly corresponds to action distribution constraints, while the latter mainly corresponds to concept representation constraints and scene embedding vector constraints.

[0071] The multimodal incremental learning module updates the total loss. For student model parameters Perform incremental update:

[0072]

[0073] in, This is the weighted task loss term determined based on the feedback correction samples and their corresponding sample weights; To match scene identifiers Corresponding historical false trigger rate Relevant distillation weights; and The temperature parameters are respectively for the teacher model and the student model. The distribution of candidate control actions; and These are the concept vectors for the teacher model and the student model, respectively. and These are the scene embedding vectors for the teacher model and the student model, respectively; and These are the loss weighting coefficients; Let represent the Kullback–Leibler divergence. In the total update loss described above, the first part... The student model is directly corrected at the task level based on feedback correction samples; in Part Two... The corresponding behavioral distillation term is used to measure the difference between the candidate control action distributions of the teacher model and the student model under the temperature parameter and to suppress behavioral output drift; Part Three Corresponding concept distillation terms are used to constrain the concept vectors output by the student model to remain consistent with those of the teacher model; Part Four The corresponding scene embedding vector alignment term is used to maintain the stability of the student model's representation of driving scenarios.

[0074] It consists of a weighted cross-entropy loss and a concept supervision loss, wherein the weighted cross-entropy loss uses the sample weights of the feedback correction samples. As a multiplier, it enhances the correction effect of the corresponding samples of revocation and takeover events on the action classification boundary; the concept supervision loss is used to correct the supervision target of the concept vector based on the feedback information. For example, for a left lane change sample revoked by the driver, the target action can be relabeled as "only prompt not to execute" or "slow down and yield"; for samples taken over by the backup control unit and the monitoring decision result shows that the control command exceeds the safety constraints, in addition to correcting the action label, the concept expectation interval corresponding to the sample is also corrected, making the supervision constraints of key concepts such as "lateral maneuvering space sufficiency" and "driver takeover capability" more stringent. On the other hand, distillation weights It can be preferably set as follows: ,in To match scene identifiers The corresponding historical false trigger rate. Therefore, when the historical false trigger rate of a certain scenario is high, it means that false triggers are more likely to occur in that scenario. In this case, the reliance on the stable knowledge of the teacher model is stronger, and the student model updates more cautiously. When the historical false trigger rate of a certain scenario is low, the student model is allowed to make relatively more sufficient parameter adjustments in that scenario.

[0075] The Kullback-Leibler divergence in the behavioral distillation term is used to measure the difference between the teacher and student models in terms of temperature parameters. The difference between the distributions of candidate control actions can be defined as follows:

[0076]

[0077] in, This indicates that the teacher model has temperature parameters. The distribution of candidate control actions, This indicates that the student model has temperature parameters. The distribution of candidate control actions, This indicates the number of categories of candidate control actions. and They represent the first The probability values ​​corresponding to each candidate control action. In this scheme, the technical role of Kullback-Leibler divergence is not simply to calculate mathematical distance, but to transfer the stable action distribution knowledge already possessed by the teacher model to the student model, thereby suppressing drastic fluctuations in the output distribution of the student model due to the limited number of feedback correction samples or unbalanced sample weights. In simple terms, even if some new samples are very important, the student model should not completely "forget" the judgment boundaries in the original stable scene because of these samples. Therefore, Kullback-Leibler divergence is needed to provide flexible constraints on the behavior distribution.

[0078] In a specific parameter setting, the multimodal incremental learning module can use a batch size of 64 and a learning rate of [missing information] when performing incremental updates. Temperature parameters The behavioral distillation coefficient is set to 1.0; the conceptual distillation coefficient is set to... Take 0.3 as the scene embedding vector alignment coefficient. The threshold is set to 0.1. An incremental update is triggered every 500 accumulated feedback correction samples or every 24 hours. After an update is triggered, the learning loop module first organizes the feedback correction samples and assigns sample weights. Then, the teacher model and student model perform forward inference on the same batch of feedback correction samples, calculate the total update loss, and only perform backpropagation updates on the student model parameters. The teacher model parameters remain frozen throughout the process. After the incremental update is completed, the updated student model is promoted to the new online model only when the false trigger rate on the offline validation set decreases and the execution success rate is not lower than that of the previous stable version of the model.

[0079] In summary, the technical problem solved by this invention can be clearly understood as follows: In existing driver behavior prediction and control systems, even if the model outputs a high-confidence action, it may still result in erroneous actions due to insufficient scenario interpretation, a high historical false trigger rate, driver unpreparedness to take over, or abnormal execution links. This embodiment forms a unified scenario representation through a data acquisition module, a graph construction module, and a multimodal incremental learning module. A dynamic safety verification module performs hierarchical review of candidate control actions. A safety execution module enables revocable interaction of control commands. A redundant control module enables independent takeover in case of main control link abnormalities. Furthermore, a learning closed-loop module transforms revocation and takeover events back into high-value feedback correction samples, continuously correcting easily triggered scenarios in subsequent incremental learning. Therefore, its ultimate effect is not only to improve the hit rate of candidate control actions, but more importantly, to reduce the false trigger rate, reduce the probability of dangerous actions, and improve interpretability, verifiability, and operational safety in real-world driving.

[0080] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Those skilled in the art can readily implement the present invention based on the accompanying drawings and the above description. However, any modifications, alterations, or variations made by those skilled in the art without departing from the scope of the present invention, utilizing the disclosed technical content, are equivalent embodiments of the present invention. Furthermore, any modifications, alterations, or variations made to the above embodiments based on the essential technology of the present invention are still within the protection scope of the present invention.

Claims

1. A behavior prediction and control system based on multimodal incremental learning and dynamic safety verification, characterized in that, include: The data acquisition module is used to collect multimodal data, including vehicle operation data, driver interaction data, and external environment perception data. The graph construction module is used to construct a spatiotemporal interaction graph based on the multimodal data; The multimodal incremental learning module is used to encode the spatiotemporal interaction graph to obtain a scene embedding vector, and output candidate control actions, candidate control action confidence, concept vectors and scene identifiers based on the same scene embedding vector; The dynamic security verification module is used to receive the candidate control action, the confidence level of the candidate control action, the concept vector, the scene identifier, and the multimodal data, calculate the concept consistency score based on the concept vector, and perform threshold adaptive verification on the confidence level of the candidate control action based on the historical false trigger rate corresponding to the scene identifier to obtain the verification result. The verification result includes execution permission, execution blocking, or only prompting not to execute. The safety execution module is used to generate control commands and issue them to the vehicle execution object to execute the candidate control actions when the verification result is an execution permission, and to provide the revocable interaction of the control commands to trigger the revocation event of the control commands when the revocation conditions are met; A redundant control module is used to independently monitor the control commands of the safety execution module and take over when preset takeover conditions are met, so as to trigger the takeover event of the control command, block or replace the control command. The learning closed-loop module is used to update the historical false trigger rate based on the execution result of the control command and the cancellation or takeover event of the control command, and to form feedback information. The feedback information is then converted into feedback correction samples for the multimodal incremental learning module to perform incremental updates.

2. The behavior prediction and control system based on multimodal incremental learning and dynamic security verification according to claim 1, characterized in that: The multimodal incremental learning module includes: A spatiotemporal graph neural network unit is used to encode graph features of the spatiotemporal interaction graph to obtain a scene embedding vector; A concept bottleneck unit is used to output the concept vector based on the scene embedding vector; A behavior prediction unit is used to output the candidate control action and the confidence level of the candidate control action based on the scene embedding vector and the concept vector. A scene identifier generation unit is used to output the scene identifier based on the scene embedding vector; The concept bottleneck unit, the behavior prediction unit, and the scene identifier generation unit share the same scene embedding vector and output them in parallel, so that the candidate control action, the confidence of the candidate control action, the concept vector, and the scene identifier correspond to the same driving scene. The concept vector is used to characterize interpretable evidence of the candidate control action, and the scenario identifier is used to characterize the current driving scenario category.

3. The behavior prediction and control system based on multimodal incremental learning and dynamic security verification according to claim 1, characterized in that, The dynamic security verification module includes: An environmental risk verification unit is used to generate a risk envelope based on the vehicle external environment perception data and determine whether the predicted trajectory of the candidate control action intersects with the risk envelope. The driver willingness and takeover capability verification unit is used to determine whether the driver is in a takeover capability state based on the driver interaction data, and to determine whether there is an objection signal within a preset confirmation window. The model reliability verification unit is used to determine the reliability based on the confidence level of the candidate control action and the historical false trigger rate corresponding to the scene identifier. An interpretable consistency verification unit is used to make a consistency determination based on the concept consistency score; The runtime monitoring and verification unit is used to receive the monitoring decision results output by the redundant control module, and to verify the candidate control action and the control command based on the monitoring decision results. When all verification units pass, the verification result is execution permission; when any verification unit fails, the verification result is execution blocking or only prompts not to execute.

4. The behavior prediction and control system based on multimodal incremental learning and dynamic security verification according to claim 3, characterized in that: The interpretable consistency verification unit calculates a concept consistency score based on the concept vector, which is used to characterize the degree of consistency between the current candidate control action and the interpretable evidence under the current scene identifier. satisfy: ; in, For the concept vector, the first Each concept component The first corresponding to the candidate control action and the scene identifier The expected range of each concept component is determined by the combination of candidate control actions and scene identifiers. For the first The weight of each concept component, For the number of conceptual components; When the concept consistency score is less than the preset consistency threshold, the interpretable consistency verification unit determines that it has failed; otherwise, it has passed.

5. The behavior prediction and control system based on multimodal incremental learning and dynamic security verification according to claim 3, characterized in that, The environmental risk verification unit calculates an environmental risk index for each obstacle in the vehicle's external environment perception data, and generates the risk envelope based on the environmental risk index. The obstacle... Environmental risk index satisfy: ; in, For obstacles Distance from this vehicle For safety reference distance, For obstacles The value is relative to the approach speed component of the vehicle, and is positive only when the obstacle is approaching the vehicle, and zero when it is moving away from the vehicle. and For scale parameters; The environmental risk verification unit determines the longitudinal expansion amount and the lateral expansion amount based on the predicted occupied area of ​​each obstacle and the corresponding environmental risk index, generates a corresponding local risk sub-envelope around the predicted occupied area of ​​each obstacle, and obtains the risk envelope by taking the union of each local risk sub-envelope. When the predicted trajectory corresponding to the candidate control action intersects with the risk envelope, the environmental risk verification unit determines that it has failed; otherwise, it has passed.

6. The behavior prediction and control system based on multimodal incremental learning and dynamic security verification according to claim 4, characterized in that, The learning closed-loop module updates the historical false trigger rate for each scene identifier. Corresponding historical false trigger rate satisfy: ; in, For scene identification The number of false triggers determined by the revocation event and / or takeover event of the control command. For scene identification The number of times the next candidate control action is granted execution permission. For smoothing coefficients; The model reliability verification unit performs threshold adaptive verification of the confidence level of the candidate control action based on the historical false trigger rate, and the scene identifier. The confidence verification threshold below satisfy: ; in, Based on the threshold, It is the adjustment coefficient, and ; and These are the preset minimum confidence threshold and the preset maximum confidence threshold, respectively. .

7. The behavior prediction and control system based on multimodal incremental learning and dynamic security verification according to claim 6, characterized in that, The learning closed-loop module assigns sample weights to the feedback correction samples, and the multimodal incremental learning module performs incremental updates based on the feedback correction samples and their corresponding sample weights; the sample weights satisfy: ; in, These are the weighting coefficients, and ; An additional coefficient is added to the event, which is determined according to the event type corresponding to the feedback information. It is used to ensure that, when the comprehensive reliability scores are the same, the sample weight corresponding to the control cancellation event is greater than the sample weight corresponding to the takeover event, and the sample weight corresponding to the takeover event is greater than the sample weight formed solely based on the control instruction execution results. The comprehensive reliability score output by the dynamic safety verification module for the candidate control action satisfies: ,and ; In the formula, Let be the fusion coefficient, and ; To compare with the confidence level of candidate control actions The corresponding confidence verification score is used to characterize the confidence level of the candidate control action. Compared to scene identifiers Corresponding confidence verification threshold The degree of passage.

8. The behavior prediction and control system based on multimodal incremental learning and dynamic security verification according to claim 7, characterized in that: The multimodal incremental learning module includes a two-stage distillation update mechanism, which comprises behavioral distillation and concept distillation. Behavioral distillation is used to ensure that the candidate control action distribution output by the student model remains consistent with that of the teacher model. Concept distillation is used to ensure that the concept vector output by the student model remains consistent with that of the teacher model, and updates the overall loss accordingly. Incremental updates are performed on the student model parameters, and the total update loss is... satisfy: ; in, The weighted task loss term is determined based on the feedback correction samples and their corresponding sample weights. To match scene identifiers Corresponding historical false trigger rate The relevant distillation weights, and The temperature parameters are respectively for the teacher model and the student model. The distribution of candidate control actions, and These are the concept vectors for the teacher model and the student model, respectively. and These are the scene embedding vectors for the teacher model and the student model, respectively. and These are the loss weighting coefficients; This represents the Kullback–Leibler divergence.

9. The behavior prediction and control system based on multimodal incremental learning and dynamic security verification according to claim 1, characterized in that, The safety execution module includes a tactile feedback component and a cancellation determination logic. The tactile feedback component is used to output tactile prompts to the driver when the control command is issued. The cancellation determination logic is used to trigger the cancellation event of the control command and cancel the control command when the reverse force applied by the driver is greater than the preset force threshold or an opposition signal is detected. The scene identifier, candidate control action, and confidence level of the candidate control action corresponding to the cancellation event are sent to the learning closed-loop module to form feedback information.

10. The behavior prediction and control system based on multimodal incremental learning and dynamic security verification according to claim 3, characterized in that, The redundancy control module includes a main control unit and a backup control unit; The main control unit is used to run the dynamic security verification module and the security execution module to execute the main control link when the preset takeover conditions are not met, and to generate and issue the control command based on the verification result; The backup control unit operates independently of the main control unit and is used to generate a monitoring decision result based on the candidate control action, the control command, and the vehicle execution object status, and output it to the runtime monitoring and verification unit. The monitoring decision result is used to characterize whether at least one of the following situations exists: the candidate control action is inconsistent with the control command, the control command exceeds the preset safety constraints of the vehicle execution object, the state of the vehicle execution object does not correspond to the control command, the main control link times out, or the communication between the main control unit and the vehicle execution object is abnormal; when the monitoring decision result indicates that at least one of the above situations exists, it is determined that the preset takeover conditions are met. When the preset takeover conditions are met, the backup control unit completes the takeover within a preset takeover time, triggers the takeover event of the control command, and sends the takeover event, along with the scene identifier and monitoring decision result corresponding to the takeover event, to the learning closed-loop module to form feedback information, and blocks or replaces the control command.